term_5 Request access

Autonomous engineering studio

An engineering studio that keeps working after you close the tab.

term_5 runs on your own machine. It reads the codebase before it edits it, hands each project to a persistent owner, checks its own work in a real browser, and ships through a release that has to pass a health gate — then leaves a record you can audit.

routes mapped
14
renders checked
30
overflow at 390px
0px
5xx after ship
0%
run_7c41a0e5 verified

ship v2 of the client dashboard

  1. map application_map 14 routes · 9 templates · 61 selectors 1.4s verified
  2. plan improvement_plan 6 batches · 14/14 routes accounted for 0.6s verified
  3. delegate agent_delegate → dashboard owner running · one lane per project 3m 12s running
  4. verify browser_audit_pages 30 renders · 0px overflow · 0 console errors 41s verified
  5. ship deployment_deploy health gate passed · 0% 5xx 58s verified

1 objective · 5 steps · 1 project lane · 4m 53s

One durable objective. Every step either has an observation behind it or is labelled as forecast.

Steps accounted for5 / 5

Every step of this objective is accounted for

Release gate

The gate this site went through

A release is a transaction that is allowed to refuse. Each step below either produced evidence or it would have stopped the deploy before anything public changed.

https://term5.syntal.pro · checked in order · any failure stops the release

  1. G01

    clean tree

    a dirty working tree blocks the release

    passed
  2. G02

    image build

    the test suite runs as a build step

    passed
  3. G03

    configuration

    required values present before the container starts

    passed
  4. G04

    proxy config

    validated before the reload

    passed
  5. G05

    certificate

    issued, with HTTP redirected to HTTPS

    passed
  6. G06

    public probe

    the live endpoint answered 200

    passed
  7. G07

    backup

    this app has no database to back up

    n/a
  8. G08

    authorization

    production changes wait for a human

    standing by
  • One writer per project

    Mutations are serialised per repository, so two agents never race the same working tree.

  • Evidence or silence

    A claim about your code cites a tool observation. Forecast work is reported as forecast.

  • Your host, your keys

    It runs on your machine against your repositories. Credentials are held locally and never enter a prompt.

  • It stops when it matters

    Destructive and production actions wait for a human decision instead of being auto-approved.

01The failure modes

Autonomy fails in four predictable ways

Most agent tooling stalls on the same four problems. Each one is a design constraint here rather than a caveat in the documentation.

  • 01

    Context loss

    The tool forgets what your repository looks like between sessions, so every request starts with the same rediscovery.

    Design constraintDurable project memory and a per-project state capsule that is rebuilt from real repository, queue and history state.

  • 02

    Silent breakage

    “Done” means an exit code of zero, and nothing was ever rendered, clicked or measured.

    Design constraintThe real application is rendered in a browser and checked for console errors, failed requests and horizontal overflow.

  • 03

    Scope creep

    A broad request collapses into one-file guesswork, or a refactor lands in files nobody audited.

    Design constraintThe route, template, CSS and script surface is mapped first, and the plan has to account for every entry in it.

  • 04

    No evidence trail

    You cannot tell a healthy release from a build that never passed.

    Design constraintReleases are health-gated transactions, and post-deploy checks are recorded separately from the deploy that triggered them.

02How it works

Five steps, in the order they actually happen

The order matters. Reading comes before editing, and verification comes before shipping — not the other way around.

  1. 01

    Map Read the repository before touching it

    Routes, templates, shared shells, CSS selectors, script hooks and existing tests are enumerated deterministically. No guessing at structure from filenames.

  2. 02

    Plan Account for the whole discovered surface

    A broad request becomes ordered batches with explicit acceptance criteria. Every discovered route ends up changed, verified or deliberately deferred — and it says which.

  3. 03

    Delegate Hand each project to its own owner

    A persistent owner per project keeps its own durable queue. Different projects proceed concurrently; one project mutates one tree at a time.

  4. 04

    Verify Check the running product, not the plan

    The application is rendered at desktop and mobile widths, its console and network activity inspected, and its critical journeys exercised end to end.

  5. 05

    Ship Release only through a gate that can refuse

    Build, back up, validate the proxy configuration, probe over HTTPS and record the release. If a step fails, the release is not recorded.

03Capabilities

What is actually in the box

Eight mechanisms, each of which exists to close one of the failure modes above.

  • Project Owner Agents

    Each project gets a persistent owner with its own durable queue, its own copy of project goals and decisions, and its own compact state capsule.

  • Procedural knowledge engine

    Reusable product, framework and quality playbooks are resolved before work starts, so “build a booking app” arrives with the parts a booking app is expected to have.

  • Durable human tasks

    When a decision genuinely needs you — a credential, a risky approval, a visual direction — it becomes a durable task with an explicit timeout policy instead of a question buried in a transcript.

  • Memory that ages honestly

    Durable notes carry a source and are marked stale when that source changes, so an outdated fact cannot quietly become a plan.

  • Creative studio

    Material visual choices become distinct rendered directions you can compare, and the selected one is translated into an implementable specification before any CSS is written.

  • Production intelligence

    Development, staging and production are separate environments with their own evidence, and incidents keep the deployment and commit they are attributed to.

  • Browser and visual verification

    A real headless browser renders the pages, and screenshots can be inspected for layout, hierarchy and defects rather than inferred from source.

  • Typed infrastructure operations

    Containers, reverse proxies, certificates and releases are driven through typed operations with a validated, reversible path — not generated shell strings.

04A worked example

From one sentence to a recorded release

This is the shape of a real change, including the parts that can fail.

  1. 01You ask

    “The dashboard is unusable on a phone. Fix it and make it feel finished.”

  2. 02It maps

    14 routes, 9 templates, 61 CSS selectors and 4 script hooks are enumerated, and 3 pages are already overflowing at 390px.

    application_map · 3 pages overflowing at 390px

  3. 03It plans

    Six batches across the shared shell and the affected pages, with every route marked changed, verified-acceptable or deferred.

    improvement_plan · 14/14 routes covered

  4. 04It verifies

    30 renders at three viewports, console and network inspected, then the before/after impact compared rather than assumed.

    browser_audit_pages · 0px overflow · 0 console errors

  5. 05It ships

    The image is built with the test suite as a build step, the database is backed up, the proxy configuration is validated, and the public endpoint is probed before the release is recorded.

    deployment_deploy · HTTPS 200 · 0% 5xx

  6. 06It shows its work

    You get the commit, the evidence, the pages it changed and the things it chose not to touch.

    release_verify · 3/3 healthy samples

05Governance

Things this agent will not do

These are enforced by the runtime it runs inside, not by good intentions in a prompt.

  • Ask you to paste a password, key or token into a conversation

    Secrets are requested through a local configuration surface. Their values are never returned to the model, logged or echoed.

  • Delete a persistent database volume while deploying or rolling back

    A container being stopped is not a database being erased. Volumes survive lifecycle operations by design.

  • Auto-approve something destructive because a timer expired

    Timeouts only resolve safe subjective choices. A risky production action with no answer stays blocked.

  • Report a change it did not observe

    Projected and in-flight work is labelled as such. Completed work cites the observation that supports it.

  • Attribute one site's errors to another site's release

    Deployment correlation is evidence, not causation; site-scoped logs stay attached to the site that produced them.

  • Present a mockup as a finished feature

    A design artifact is input to implementation. The rendered, verified application is what counts as done.

06Running it

Three commands and a health check

It needs a Linux host with Docker and a shell. Everything else — including the browser it verifies with — runs from the workspace.

Host
Linux with Docker Engine and the Compose plugin
Optional
An image-generation key, only if you want raster mockups
Not required
A hosted database, an account, or a credit card
Runs offline
Yes, apart from model calls and certificate issuance
getting started
  1. git clone <your-fork> term5 && cd term5 Clone the workspace
  2. docker compose up -d --build Build and start
  3. curl -fsS localhost:8793/healthz Confirm it is alive

07Questions

The ones worth answering

Does it need to reach the internet?

It needs whatever model endpoint and package registries you point it at. Certificate issuance and repository fetches are the only other outbound calls it makes, and both are optional.

Where do my credentials live?

In local configuration on the host, materialised only where a container needs them. They are excluded from model context, which means the agent can confirm a secret exists without being able to read it.

Can it touch a project it was not asked about?

Each project owns its own repository, ports, volumes and release path. Cross-project work is an explicit, recorded handover with a declared dependency.

What happens when it gets something wrong?

The failure is attached to the release that caused it, with evidence. Rollback returns source to a known commit and rebuilds — and it refuses to run with a dirty working tree, so the rollback target is never ambiguous.

How much can it do before it asks me something?

As much as the work allows. Implementation detail — spacing, internal helpers, indexing, routine tests — is decided without interrupting you. It interrupts for credentials, risky approvals and material product choices.

Is a passing build treated as a finished product?

No. Infrastructure running is not completeness. The implementation is compared against the expected feature set for that kind of product, and gaps are implemented rather than declared out of scope.

Bring it your backlog.

Tell us what you are trying to ship and how the current tooling gets in the way. We read every request.

Request access