Obligation
Engineering
A way to build software with coding agents. Each feature is described by a small contract of facts. A policy engine derives everything the feature owes, from its API and MCP tool to its dashboard, alert and runbook. One required check on every commit decides whether each of those things is real. Anything left out needs a named owner, a reason and a review date.
The instructions an agent reads are short, because the rules live in checks. The commands an agent runs are the same ones a person runs. Nothing is accepted because an agent said it was done.
- specwhy it exists
- storywhat a person gets
- contractfacts about one operation
- obligationsderived by policy
- providerswhat the repository contains
- checkgreen on the commit, or not
§1
Problem
Agents produce change faster than people can review it. A merged change proves the code compiles and its tests pass. It says nothing about the API, the documentation, the dashboard or the alert that should have shipped with it.
The usual answer is a longer instruction file. It does not work. A list of things to remember is out of date the day a new surface appears. It costs tokens on every task. And nobody can tell a forgotten MCP tool from one that was left out on purpose. Both look like silence.
So the question changes from was the code written to what does this feature owe, and is each of those things real. Those two questions have to be answered by the repository, not by a prompt and not by the agent.
§2
The idea in sixty seconds
- 2.1
Facts, not checklists.
A feature is described by what it does, who uses it, what data it touches and how risky it is. Those facts fit on one screen.
- 2.2
Obligations are computed.
A policy engine turns the facts into the exact list of things the feature owes: surfaces, access rules, data handling, tests, analytics, operations and docs.
- 2.3
Providers are discovered.
Each obligation is satisfied by something real in the repository, found by reading registries and source. A comment or a claim earns nothing.
- 2.4
Omission is a decision.
A missing obligation fails until a person records why it does not apply, who owns that decision and when to look again.
- 2.5
The commit is the proof.
One required check runs on every commit and judges all of it. No file records that something passed.
- 2.6
One substrate.
People, CI and agents run the same commands and read the same output. Instructions route to them; they do not restate them.
§3
Rules
- 3.1
Describe meaning, not transports.
A contract names a semantic operation and its facts. Routes, MCP tools, CLI verbs and screens are projections of that one identity, never the identity itself.
- 3.2
Stories live in code.
Each user story is a typed identity with acceptance criteria, defined in the codebase. Everything that proves it refers to that identity, and a misspelling fails to compile.
- 3.3
Obligations are derived.
A developer audience implies an API, an MCP tool and a CLI. A delete implies audit, recovery and a destructive preview. An SLO implies a metric, a dashboard, an alert and a runbook. Nobody lists these by hand.
- 3.4
Missing never means not applicable.
Each obligation is met by something real, or carries a decision with a reason, an owner, a review date and an approval. The engine cannot tell the difference between forgotten and omitted, so the repository must.
- 3.5
Debt only goes down.
Features that predate the policy have their gaps recorded exactly, each owned and expiring. The record can shrink. It can never grow, and its history is append-only.
- 3.6
The check is the proof.
A required check passing on the commit is the evidence. No file records that a check passed, because a file can be edited. Agents change code and tests; they never regenerate receipts.
- 3.7
Operations are part of the contract.
Events, logs, metrics, dashboards, alerts, runbooks, smoke tests and rollback are obligations like any other. Promotion needs a receipt from the live system, not a definition in the repository.
- 3.8
Every check earns its place.
A check exists to catch a defect that nothing cheaper catches. It has an owner, a time budget, a planted failure that proves it still rejects bad input, and an expiring exception mechanism. Checks that only verified the verification system were removed.
§4
Stories
Development starts from a specification. The prose spec explains the problem, the reasoning and the acceptance criteria. It is read by people and agents, and the context command quotes exactly the sections that matter for a task.
The stories that are enforced do not stay in prose. Each one is a typed identity in code, with its acceptance criteria beside it. Tests, rendered components, screenshots, docs pages and analytics events refer to a story by that identity. A misspelt reference fails to compile. A reference to a story the registry does not contain fails the build. Only a story written and reviewed by a person can be active; generated placeholders stay in draft and expire.
export const records = defineSpec({ key: 'records', stories: { [SPEC.records.record.trash]: { text: 'As a customer, deleting a record moves it to the trash, ' + 'so a mistake can be undone.', acceptance: [ 'The record leaves the list immediately', 'Restore returns it with its history', 'Only retention destroys the data', ], }, }, })
test('trash, then restore', async ({ page, spec }) => { spec(SPEC.records.record.trash) // throws if no capability owns this story await page.getByRole('button', { name: 'Delete' }).click() await expect(page.getByText('Moved to trash')).toBeVisible() await page.getByRole('button', { name: 'Restore' }).click() await expect(page.getByText('Restored with history')).toBeVisible() })
§5
Contract
Five small artifacts make up one feature. The contract holds facts the tooling can read without running product code. The spec holds behaviour. Everything else points back at an operation or a story by identity.
{ "key": "records", "profile": "api", "risk": "critical", "deliveryStatus": "shipped", "evidenceStatus": "enforced", "spec": "docs/specs/records.md", "rollout": { "killSwitch": "useRecords", "rollbackSignal": "records.record.delete error rate" }, "operations": [ { "id": "records.record.delete", "effects": ["delete"], "audiences": ["customer", "developer", "automation"], "data": { "persisted": true, "personal": true }, "measurement": { "behavior": true, "operational": true, "slo": true }, "access": { "customer": "records.customer.write", "automation": "records.automation.write" } } ] }
The facts are a closed vocabulary. Profile says what shape the feature has: a UI flow, an API, a workflow, a library or a non-functional requirement. Effects say what an operation does to data. Audiences say who it is for. Data says whether it is stored and whether it is personal. Measurement says whether behaviour, health or a service level is tracked. Risk raises the bar. From these, and nothing else, the policy resolves what is owed.
§6
Resolver
The policy engine below is the same code that runs in CI, compiled into this page. Change the facts of one operation and the list of obligations is recomputed. Items highlighted after a change are the ones that change added.
input · one operation
- example
- audience
- effect
- data
- measure
- shape
- risk
output · resolved by the policy engine
91obligations
plan 6authoring 36merge 42promotion 3runtime 4
- surfaces 5
- product-uipublic-apiopenapi-clientmcpcli
- access 18
- access-policyauthenticationauthorizationcredential-permissionapi-key-permissionoauth-app-permissionoauth-consentoauth-discoverypermission-compatibilityerrorsidempotencyrate-limitcompatibilityauditdeletionrecoverydestructive-previewdestructive-approval
- data 7
- data-schemamigrationretentionpii-classificationprivacydata-boundaryresidency
- proof 8
- test-unittest-integrationtest-contracttest-audittest-boundarysecurityperformancereliability
- analytics 4
- analytics-outcomeanalytics-eventanalytics-emitanalytics-insight
- operations 22
- structured-loglog-fieldsredactioncorrelationtracecardinalitytelemetry-costtelemetry-guardrailsmetric-slihealth-dashboardalert-sloevent-healthsignal-livenessalert-livenesssynthetic-livenesssmokerollbackkill-switchrolloutremote-receiptrunbooktroubleshooting
- docs 7
- acceptanceimplementationproduct-guidechangelogrest-referencemcp-referencecli-reference
* generated from another provider. Highlighted items were added by your last change.
The rules are small and legible. Each fact adds a fixed set of obligations, and the sets compose. The table shows the main ones.
| fact | adds |
|---|---|
| audience: customer | a product UI, authentication and authorization, a product guide |
| audience: developer | a public API, a generated client, an MCP tool, a CLI command, their references, a rate limit and a compatibility contract |
| audience: automation | API-key and OAuth-app permissions, consent and discovery, an error contract and a contract test |
| audience: operator | an admin UI, an admin API, an admin MCP tool, support actions |
| effect: delete | audit, idempotency, deletion, recovery, a destructive preview and a destructive approval |
| data: personal | PII classification, privacy, a data boundary, residency, a boundary test |
| measure: slo | an SLI metric, a health dashboard, an SLO alert and a runbook |
| risk: critical | a kill switch, a smoke test, rollback, security, performance and reliability proof, tracing, synthetic liveness |
| risk: regulated | privacy and audit proof, a data boundary and a human approval at promotion |
| profile: ui-flow | empty, loading, error and conflict states, responsive and accessibility proof, an end-to-end test and a rendered story |
Two words carry weight here. An obligation whose state is derived is a projection of another obligation and is satisfied only while its source is. An obligation whose service level is derived may be met by a shared provider, such as one dashboard that covers several operations. A raw escape hatch, such as a generic command that forwards to the API, never earns first-class credit.
§7
Surfaces
A feature is used by people, by developers and by agents, and each of them reaches it through a different surface. The contract does not list those surfaces. It lists audiences, and the policy derives the surfaces. A customer audience owes a product UI on every shell the product has. A developer audience owes an API, a client, an MCP tool and a CLI. An automation audience owes the permissions and consent that let a machine hold a credential. An operator audience owes an admin console and its machine twins.
Each surface then declares the operation it implements, in its own registry, in its own shape. The graph joins them by identity and checks that they agree.
- peopleaudience: customer
- web appdesktop shellmobile shellbrowser extensionspreadsheet add-onadmin console audience: operator
- developersaudience: developer
- REST routeOpenAPI client from the emitted OpenAPISDK from the clientCLI command owned decision: raw API onlywebhook trait: webhook-emitter
- agentsaudience: automation
- MCP toolWebMCP page tool from the MCP declarationcommand palette from the MCP declarationin-app assistant from the MCP declaration
- knowledgederived from each surface
- REST reference from the OpenAPIMCP reference from the served tool listCLI reference from the command manifestproduct guide frontmatter names the operation
● declares the operation itself○ generated from another declaration× omitted by an owned decision
A surface binds to an operation where it is defined. A route carries the operation in its definition and the emitted OpenAPI carries it as an extension. A tool declaration carries it beside its schema and effect. A screen is found by scanning the app for executable calls to the endpoint that implements the operation; a string or a comment does not count.
defineRoute({ method: 'delete', path: '/v1/records/{id}', operation: 'records.record.delete', effects: ['delete'], access: 'records.automation.write', idempotency: 'safe-to-retry', handler: deleteRecord, })
defineCapability({ name: 'delete_record', operation: 'records.record.delete', input: DeleteRecordInput, annotations: { effect: 'delete', runtimes: ['mcp', 'in-app', 'palette'], authz: { minRole: 'editor' }, }, preview: previewDelete, // what the agent shows before acting run: deleteRecord, })
The second declaration is the one that spreads furthest. A tool declared once is served to remote MCP clients, registered on the page for browser agents through WebMCP, listed in the command palette and offered to the in-app assistant, all executing through the same server with the same rules. A parity test pins the page manifest to the live tool list, tool for tool and schema for schema. Every delete asks for typed confirmation on every one of those surfaces, because the confirmation class comes from the declaration, not from the surface.
Native shells, extensions and add-ons are further product UI projections. When one exists it registers the operations it implements and the graph counts it. When it does not, the policy is unchanged and nothing pretends.
| the graph checks | across every projection of one operation |
|---|---|
| effects | every projection declares the same effects, and a delete surface says it is destructive |
| access | the same named policy, the same scopes, the same tenant boundary |
| safety | the same retry contract and the same confirmation class for agents |
| contracts | machine surfaces carry an error contract and a schema contract |
| freshness | generated projections are regenerated in CI and any drift fails |
§8
Stages
Obligations are due in order, and each is judged at the stage where its claim can be true. Plan, authoring and merge are deterministic and make no external writes. Promotion and runtime need evidence from the live system: a dashboard defined in the repository cannot satisfy them on its own.
- planlocal
- authoringlocal, CI
- mergeCI
- promotionlive receipt
- runtimeexpiring receipt
| stage | true when |
|---|---|
| 8.1plan | Acceptance criteria, access policy, retention, PII classification and rollout are decided. |
| 8.2authoring | Implementation, surfaces, unit tests, typed events, logs and docs exist. |
| 8.3merge | Integration, contract and end-to-end tests pass. Dashboards and alerts are defined and consistent with the events the code emits. |
| 8.4promotion | A smoke test passed against the deployed revision. Dashboards and alerts were applied, and a receipt bound to the exact definitions says so. |
| 8.5runtime | Events arrive, metrics have data, alerts are evaluating and synthetic probes pass. The receipt expires within hours and must be renewed. |
§9
Operations
A feature is not finished when its code merges. It also owes the signals that show it working for people, the signals that show it healthy, and the means to act when it is not. Each of these is an obligation with a provider, and each link between them is checked.
- product analyticstyped eventoutcomeinsight
- operational telemetryhealth dashboardSLO alertrunbook
- live proofpromotion receiptruntime receipt
export const recordTrashed = defineEvent({ name: 'records.record_trashed', capability: 'records', operation: 'records.record.delete', story: SPEC.records.record.trash, outcome: 'success', frequency: 'per-action', schema: z.object({ record_kind: RecordKind }), }) export const recordDeleteFailed = defineEvent({ name: 'records.record_delete_failed', operation: 'records.record.delete', outcome: 'failure', schema: z.object({ error_code: ErrorCode }), // required on every failure event })
export const records = defineDashboard({ key: 'records', operationBindings: [bind('records.record.delete')], tiles: [ tile({ key: 'delete-errors', query: ratio(recordDeleteFailed, recordTrashed), alerts: [ ratioCeiling({ name: 'Records: delete failures', maxPercent: 5, minSamples: 20, interval: 'hourly', rationale: 'deletes that fail leave a person unsure whether data is gone', baseline: 'well under one percent across the last quarter', enabled: true, }), ], }), ], })
- 9.1
two streams, kept apart
Product analytics answer whether people get value: use, success and actionable failure per capability. Operational telemetry answers whether the system is healthy: logs, metrics and traces. Logs are never analytics, and an event emitted only to satisfy a count is a defect.
- 9.2
typed events
An event declares its capability, operation, story, outcome and payload schema. A failure event must carry a categorical error code. A per-request event must declare a sample rate or a cost control. Violations throw when the module loads, before any test runs.
- 9.3
structured logs
A log line is a stable message plus typed attributes from one shared vocabulary. Each component declares the fields it may log; an undeclared field is dropped in production and throws in tests. Correlation ids are bound once at the entrypoint and inherited.
- 9.4
metrics with closed labels
Every metric declares its labels and their allowed values. An undeclared label or an id-shaped value is refused, and the series ceiling per metric is asserted.
- 9.5
dashboards and alerts as code
Tiles read named events and metrics. An alert carries an interval, a rationale and a baseline, and binds the operations it watches. A consistency test checks that tiles read only events the code emits, and that every SLO operation has a dashboard and an enabled alert.
- 9.6
receipts and liveness
Applying the definitions to the live analytics system produces a receipt bound to the revision and the definition digest. At runtime a separate receipt proves events arrived, metrics have data and alerts evaluated recently. It expires within hours.
- 9.7
runbooks, kill switches, rollback
Critical paths owe a runbook, a kill switch named in the contract and a rollback signal. Smoke probes against the deployed revision are the promotion proof, not a green deploy job.
§10
Budgets
Fast feedback is a correctness feature. A slow check is skipped or run late, and an agent can only act on the signal it receives. So speed and cost are gates too, with the same shape as every other gate: a declared budget, a check that measures against it, and a failure that names what to fix.
Performance is also an obligation. For a critical capability the policy demands performance and reliability proof, and the provider is a test with a budget in it, not a paragraph in a plan.
| budget | measured against | fails when |
|---|---|---|
| operation latency | a performance test runs the operation against seeded data and asserts its p95; the test is the provider for the performance obligation | the budget is exceeded |
| telemetry cost | metric points, analytics events, log lines and error events per thousand requests, per service | a class exceeds its budget |
| cardinality | closed label sets per metric, a series ceiling, a denylist of id-shaped labels | an undeclared label or value appears |
| event volume | a per-request event must declare a sample rate, anonymity or a cost control | the module throws on load |
| bundle size | compressed JS and CSS per package, compared with the merge base | growth beyond the threshold |
| worker startup | the entry stays small: minified, heavy modules imported lazily, large corpora as separate modules | the shape regresses |
| context packet | the bounded context an agent receives, under a hard token ceiling | the packet cannot fit |
| instructions | bytes, words and estimated tokens per instruction file, and for the whole chain | any file or the chain is over budget |
| time | per-test, per-suite and per-job ceilings; end-to-end artifact size ceilings | a timeout is a failure, never a retry |
| the machine | a dev profile declares its CPU, memory and named resources; heavy commands queue behind an admission broker | a profile that cannot fit the machine refuses to start |
| hot reload | an edit to a shared package becomes visible within its budget, and only in the worktree that made it | the change is late or leaks to a peer |
const BUDGET_MS = { firstPage: 400, deleteWithHistory: 250 } test('lists the first page within budget at a thousand records', async () => { await seedRecords(db, 1_000) const p95 = await percentile(95, times(20, () => listRecords(db, { limit: 50 }))) expect(p95).toBeLessThan(BUDGET_MS.firstPage) })
§11
Inner loop
The method is only as good as the loop a developer or an agent actually runs. Every step is a stable command with bounded output and a JSON form. None of them writes to a live system. The first two answer from the graph in well under a second, so the plan starts from what is owed rather than from what someone remembered.
- contextwhat does this capability owe, and what is blocking?instant
- planwhich obligations are missing, by stage?instant
- buildchange the contract, then the code, then regenerateyou
- verify:packagedoes this one package still pass?seconds
- capability:checkis every obligation real at this stage?seconds
- verify:fastdoes the affected graph still pass?a minute
- pull requestone required check judges the commitminutes
Three loops run at different speeds. The first two block. The third may be observational by suite, but it is never allowed to stay red without an owner.
| loop | trigger | scope | pace |
|---|---|---|---|
| edit | save, pre-commit | changed files and their tests | seconds |
| pull request | push, agent checkpoint | the affected graph, escalated by risk | minutes |
| confidence | merge, nightly, release | the full graph, live providers, slow suites | budgeted per suite |
Agents work in parallel, so each works in its own worktree, and the worktree is a correctness boundary rather than a convenience. Two agents starting the same profile must never share a port, a database, an emulator or a process, and stopping one must leave the other healthy.
| resource | rule |
|---|---|
| identity | derived from the git directory, not the branch name |
| ports | leased by name from a registry, only those the profile needs; an occupied port names its owner and is never killed |
| state | temp data, emulator state, containers and build output are private to the worktree; the package store is shared and content-addressed |
| readiness | a service is ready when its health endpoint answers, not when a port opens |
| preparation | code generation and migrations run only when their inputs change, under a lock |
| capacity | each profile declares weight, CPU, memory and named resources; a broker admits runs and queues heavy ones |
{ "services": ["app", "internal-api", "postgres"], "waves": [["postgres"], ["internal-api"], ["app"]], "ports": { "app": "leased", "internal-api": "leased", "postgres": "leased" }, "readiness": { "app": "GET / 200 <html", "internal-api": "GET /ready 200" }, "resources": { "weight": 3, "cpu": 2.5, "memoryMb": 3072, "named": { "docker-engine": 1 } }, "prepare": ["schema: inputs unchanged, skipped"] }
§12
Gates
Every pull request is judged by one required check. It fails unless each gate below passes. Each gate is deterministic, has an owner, and names exactly what is missing and the smallest command that repairs it.
venture/records stage=merge graph=1f3a…c9 policy=8b2e…04 records.record.delete mcp missing no provider registered records.record.delete mcp-reference missing derived source mcp is missing records.record.delete alert-slo missing no enabled alert binds this operation owner=records-team source=capabilities.json#records.record.delete repair: declare the tool in capabilities/records.ts, then pnpm -w capability:providers venture satisfied=41 missing=3 failed=0 stale=0 not-applicable=1 exit 1
- 12.1
required check
One always-run job aggregates every selected job. A skipped job that had work, a cancelled job or a malformed result fails it. There is no other required check.
- 12.2
selection
Changed files map to affected packages and their dependants. A change outside every package selects everything. The selection and its reasons are an artifact anyone can read.
- 12.3
capability graph
An obligation has no provider at its stage, a provider is below the required service level, two projections of one operation disagree, or an override is invalid or expired.
- 12.4
story proof
An active story names no test, render or event, or evidence names a story that does not exist.
- 12.5
history guard
A debt baseline grew, an expiry moved later, or a prior record was edited. Compared with the merge base, which the change cannot edit.
- 12.6
permission impact
A permission was removed or renamed, or a grant was broadened without an approval bound to the digest of that exact change.
- 12.7
generated freshness
Running every generator changes any tracked file. Generated files are never hand-edited.
- 12.8
analytics
An event name is not a registry literal, a failure event lacks an error code, a tile reads an event nothing emits, or an SLO operation has no dashboard and alert.
- 12.9
instructions
An instruction file is over budget, links to a missing path, names an unregistered command, contradicts a universal rule, or a harness-specific rule file exists at all.
- 12.10
architecture
A dependency points the wrong way, forms a cycle, crosses between ventures, or a route touches storage directly.
- 12.11
ratchets
Dead code, legacy database access and architecture debt each have an owned, expiring baseline. New findings fail; resolved ones must be pruned.
- 12.12
migrations and environment
A migration lacks its manifest or breaks the order graph. A build reads an undeclared variable or a secret enters a cacheable task.
- 12.13
contracts and size
A breaking OpenAPI change without the breaking label. Bundle growth beyond the threshold against the merge base.
- 12.14
quarantine
A quarantined end-to-end test has no issue, no expiry, or an expiry too far out. A weekly run catches expiries that pass while nothing merges.
- 12.15
planted failures
Each policy checker runs against a planted good input and a planted bad input. A checker that stops rejecting its bad input fails the build.
The ratchet. When the policy becomes stricter, older features are neither waved through nor blocked. Each gap is recorded by exact identity with an owner, a reason and an expiry. The record may only fall, a policy change that adds gaps must be appended as a signed entry naming them, and the whole history is compared with the merge base so it cannot be rewritten.
§13
Docs and media
Documentation is an obligation like any other, and so are the images and videos inside it. They are produced by the build from the same components the product uses, so they cannot quietly fall out of date.
- component storyscreenshot, lightscreenshot, darkvisual baseline
- real ui + scenesproduct videodocs embed
- api · mcp · cliREST referenceMCP referenceCLI reference
--- title: Deleting and restoring records specs: ['records'] operationBindings: - { capability: records, operation: records.record.delete } - { capability: records, operation: records.record.restore } --- <Screenshot story="records.record.trash" alt="A record in the trash, with Restore" />
- 13.1
product guide
Every capability, and every operation a person or developer uses, owes a written guide. A docs page declares which operations it documents in its frontmatter; a capability without one fails.
- 13.2
generated references
REST, MCP and CLI references are generated from the emitted OpenAPI, the served tool list and the command manifest. A freshness check fails the build when a reference drifts from the code.
- 13.3
screenshots from code
Component stories tagged with a capability and story are rendered by a headless browser in light and dark, with motion off and the clock frozen. The same image is the documentation asset and the visual-regression baseline.
- 13.4
exact renders
Each render records the source revision and tree it came from. A screenshot counts only when it matches the current source.
- 13.5
videos from code
Product videos are composed from the real interface components with fixtures, then rendered. There is no screen recording, so re-rendering after an interface change refreshes every video.
- 13.6
a decision per capability
Every active capability carries a reviewed video decision: covered, planned or exempt, with a reviewer, a date and a reason. Rendering runs in a gated release workflow, and only when its inputs change.
§14
Commands
The repository is driven by a small set of registered commands with stable names. Each has a description, a JSON form, documented exit codes and bounded output. People and agents run the same ones, and the instruction checker rejects any instruction that names a command the registry does not contain. Select a command to run it.
$ pnpm -w context venture/recordsinstructions: AGENTS.md, venture/AGENTS.mdcapability-contract: exact graph=1f3a…c9 policy=8b2e…04 blockers=2+0docs: coding-practices.md, venture/docs/specs/records.mdverify: pnpm -w verify:package @venture/recordscontext: within budget; sha256:7c1e…--- venture/docs/specs/records.md:41-62 (section; complete; sha256:…) ---## Acceptance criteria
§15
Agents
Instructions are short routers, not manuals. One file at the root, one per venture and, rarely, one per package, each under a checked token budget. They say where to start, how to get context, which commands verify, and the few safety traps that cannot be inferred. They do not contain the surface checklist, because the policy owns it, and they do not restate any rule a gate already enforces.
Every harness reads the same file. The harness-specific file is one line that imports it, and the instruction check rejects any other harness-specific rule file outright. Context is retrieved per task through the context command, not repeated in every prompt.
AGENTS.md universal rules budgeted venture/AGENTS.md venture constraints budgeted package/AGENTS.md local constraints budgeted, rare CLAUDE.md one line: @AGENTS.md .cursorrules rejected copilot-instructions rejected
Skills are verbs. Capabilities, packages and rules are data retrieved from the graph, so there is no skill per feature. Each skill is written once in the open skill format at the repository root, where every harness can discover it, and starts by asking the context command for its workflow and exactly one venture overlay.
- 15.1
plan
Reads the resolved graph as the work list. A missing provider is work, never "not applicable". Reports the graph and policy digests it planned against.
- 15.2
build
Changes the contract first, then the code, then regenerates projections. Generated files are never edited by hand. Verifies one package, then the capability at its stage, then the affected graph.
- 15.3
verify
Runs the cheapest checks that can disprove the change first. Missing, skipped, quarantined, timed-out or stale evidence is an explicit result, never a pass. Read-only.
- 15.4
review
Independent reviewers in parallel, each with the whole diff. Every finding is reproduced against the code before it counts, then fixed with a test that fails without the fix. Discarded findings are recorded with the reason.
- 15.5
migrate
Moves one legacy capability at a time onto the current policy, reducing debt only through the ratchet. No meaningless events, screenshot-only tests or permanent exceptions to satisfy a count.
The agent contract. An implementation agent is a worker inside these rules, not a judge of them.
cannot claim completion
while any in-scope obligation is blocking. The check decides, and the check names what is left.
cannot invent an omission
An agent may implement a not-applicable decision a person approved. It cannot create one.
cannot fabricate later proof
Promotion and runtime evidence come from receipts minted by authorised workflows, never from a local definition or a report.
cannot grant itself authority
Deploys, production data, secrets and destructive actions are privilege classes. A skill or an instruction is never authority.
Authority. Commands resolve to a privilege class from read-only to sensitive-destructive, strictest match wins. A sensitive command needs an explicit invocation and an approval signed outside the repository: a short-lived token bound to the exact command, revision and working directory, verified against a trust root the agent cannot write. Even then the approval authorises; it does not execute. Execution happens in a trusted clean checkout.
{ "capabilityId": "…", "issuer": "release-owner", "revision": "9d4e…", "workingDirectory": "/srv/checkout", "argv": ["pnpm", "venture:db:migrate:prod"], "commandDigest": "sha256:…", "issuedAt": "…", "expiresAt": "…" }
§16
Comparison
Most agent harnesses today are instruction-led: a carefully written instruction file, a set of skills, a few hooks, and a review. Spec-first variants add a written specification before the code. Both improve routing. Neither answers what a feature owes, and neither can tell an agent's report from a fact. Obligation Engineering keeps the routing and moves the truth into the repository.
| question | instruction-led harness | obligation engineering |
|---|---|---|
| where a rule lives | in an instruction file, read every task | in a check, run on every commit |
| scope of a feature | whatever the prompt and the plan remembered | derived from audience, effects, data, measurement and risk |
| a missing surface | silence | fails until someone owns the omission |
| definition of done | tests pass and the agent reports done | every obligation real at its stage: surfaces, access, data, proof, analytics, operations, docs |
| proof | the agent says it ran the checks | one required check computed on the commit |
| legacy gaps | grandfathered, or a growing allowlist | exact, owned, expiring, only shrinking |
| context | the whole instruction file, every time | a bounded packet retrieved per target |
| instructions over time | grow with every incident | stay under a budget; a repeatable rule moves into a gate |
| harness | a rule file per tool, drifting apart | one canonical file; others rejected |
| authority | "be careful with production" | privilege classes and signed, expiring approvals |
| when the agent is wrong | someone notices in review, or later | the check names the capability, surface, stage, owner and repair |
§17
Destination
The north star is a repository that can answer, for any change, two questions without a person remembering and without an agent being believed: what does this owe, and is each of those things real. When a repository can do that, adding an agent adds capacity and nothing else. Adding a surface, a venture or a harness changes the facts, not the method.
It is honest about its limits. A gate proves facts about the repository. Facts about the live system need receipts, and receipts expire. The policy is code that people review, and a stricter policy creates debt that must be owned rather than hidden. Nothing here replaces judgement about what is worth building. It makes sure that whatever is built arrives whole.
Every rule above started as a sentence in an instruction file. Each one was moved into a check the day it was repeated, and the sentence was deleted. That is the whole method, and it is why the instruction files stay short.
end
Questions and corrections: hello@flagbase.com