flagbase
Get in touch
engineering · specification 01rev. 2026-10 · in use

Obligation
Engineering

A way to build software with coding agents. Each feature is described by a small contract of facts. A policy engine derives everything the feature owes, from its API and MCP tool to its dashboard, alert and runbook. One required check on every commit decides whether each of those things is real. Anything left out needs a named owner, a reason and a review date.

The instructions an agent reads are short, because the rules live in checks. The commands an agent runs are the same ones a person runs. Nothing is accepted because an agent said it was done.

  1. specwhy it exists
  2. storywhat a person gets
  3. contractfacts about one operation
  4. obligationsderived by policy
  5. providerswhat the repository contains
  6. checkgreen on the commit, or not
Fig. 1  Six links. The first three are written by people. The last three are computed from them.
Powering apps and services built on the flagbase platform, includingBankSyncUnspendyCommonSchemaand more

§1

Problem

Agents produce change faster than people can review it. A merged change proves the code compiles and its tests pass. It says nothing about the API, the documentation, the dashboard or the alert that should have shipped with it.

The usual answer is a longer instruction file. It does not work. A list of things to remember is out of date the day a new surface appears. It costs tokens on every task. And nobody can tell a forgotten MCP tool from one that was left out on purpose. Both look like silence.

So the question changes from was the code written to what does this feature owe, and is each of those things real. Those two questions have to be answered by the repository, not by a prompt and not by the agent.

§2

The idea in sixty seconds

  1. 2.1

    Facts, not checklists.

    A feature is described by what it does, who uses it, what data it touches and how risky it is. Those facts fit on one screen.

  2. 2.2

    Obligations are computed.

    A policy engine turns the facts into the exact list of things the feature owes: surfaces, access rules, data handling, tests, analytics, operations and docs.

  3. 2.3

    Providers are discovered.

    Each obligation is satisfied by something real in the repository, found by reading registries and source. A comment or a claim earns nothing.

  4. 2.4

    Omission is a decision.

    A missing obligation fails until a person records why it does not apply, who owns that decision and when to look again.

  5. 2.5

    The commit is the proof.

    One required check runs on every commit and judges all of it. No file records that something passed.

  6. 2.6

    One substrate.

    People, CI and agents run the same commands and read the same output. Instructions route to them; they do not restate them.

§3

Rules

  1. 3.1

    Describe meaning, not transports.

    A contract names a semantic operation and its facts. Routes, MCP tools, CLI verbs and screens are projections of that one identity, never the identity itself.

  2. 3.2

    Stories live in code.

    Each user story is a typed identity with acceptance criteria, defined in the codebase. Everything that proves it refers to that identity, and a misspelling fails to compile.

  3. 3.3

    Obligations are derived.

    A developer audience implies an API, an MCP tool and a CLI. A delete implies audit, recovery and a destructive preview. An SLO implies a metric, a dashboard, an alert and a runbook. Nobody lists these by hand.

  4. 3.4

    Missing never means not applicable.

    Each obligation is met by something real, or carries a decision with a reason, an owner, a review date and an approval. The engine cannot tell the difference between forgotten and omitted, so the repository must.

  5. 3.5

    Debt only goes down.

    Features that predate the policy have their gaps recorded exactly, each owned and expiring. The record can shrink. It can never grow, and its history is append-only.

  6. 3.6

    The check is the proof.

    A required check passing on the commit is the evidence. No file records that a check passed, because a file can be edited. Agents change code and tests; they never regenerate receipts.

  7. 3.7

    Operations are part of the contract.

    Events, logs, metrics, dashboards, alerts, runbooks, smoke tests and rollback are obligations like any other. Promotion needs a receipt from the live system, not a definition in the repository.

  8. 3.8

    Every check earns its place.

    A check exists to catch a defect that nothing cheaper catches. It has an owner, a time budget, a planted failure that proves it still rejects bad input, and an expiring exception mechanism. Checks that only verified the verification system were removed.

§4

Stories

Development starts from a specification. The prose spec explains the problem, the reasoning and the acceptance criteria. It is read by people and agents, and the context command quotes exactly the sections that matter for a task.

The stories that are enforced do not stay in prose. Each one is a typed identity in code, with its acceptance criteria beside it. Tests, rendered components, screenshots, docs pages and analytics events refer to a story by that identity. A misspelt reference fails to compile. A reference to a story the registry does not contain fails the build. Only a story written and reviewed by a person can be active; generated placeholders stay in draft and expire.

specs / records.tstypescript, abstracted
export const records = defineSpec({
  key: 'records',
  stories: {
    [SPEC.records.record.trash]: {
      text:
        'As a customer, deleting a record moves it to the trash, ' +
        'so a mistake can be undone.',
      acceptance: [
        'The record leaves the list immediately',
        'Restore returns it with its history',
        'Only retention destroys the data',
      ],
    },
  },
})
tests / trash.e2e.tsthe test names the story it proves
test('trash, then restore', async ({ page, spec }) => {
  spec(SPEC.records.record.trash)   // throws if no capability owns this story
  await page.getByRole('button', { name: 'Delete' }).click()
  await expect(page.getByText('Moved to trash')).toBeVisible()
  await page.getByRole('button', { name: 'Restore' }).click()
  await expect(page.getByText('Restored with history')).toBeVisible()
})
The story records.record.trash is referenced by a test, a rendered component story, screenshots, a docs page, an analytics event and a product video.STORYrecords.record.trashtestcovers(records.record.trash)renderparameters.spec.storyscreenshottrash.png · trash.dark.pngdocs pageoperationBindings: records.record.trasheventrecord_trashed · story: records.record.trashvideodisposition: covered
Fig. 2  One story identity. Everything that proves it points back at that identity, and a reference the registry does not contain fails the build.

§5

Contract

Five small artifacts make up one feature. The contract holds facts the tooling can read without running product code. The spec holds behaviour. Everything else points back at an operation or a story by identity.

capabilities.json
{
  "key": "records",
  "profile": "api",
  "risk": "critical",
  "deliveryStatus": "shipped",
  "evidenceStatus": "enforced",
  "spec": "docs/specs/records.md",
  "rollout": {
    "killSwitch": "useRecords",
    "rollbackSignal": "records.record.delete error rate"
  },
  "operations": [
    {
      "id": "records.record.delete",
      "effects": ["delete"],
      "audiences": ["customer", "developer", "automation"],
      "data": { "persisted": true, "personal": true },
      "measurement": { "behavior": true, "operational": true, "slo": true },
      "access": {
        "customer": "records.customer.write",
        "automation": "records.automation.write"
      }
    }
  ]
}

The facts are a closed vocabulary. Profile says what shape the feature has: a UI flow, an API, a workflow, a library or a non-functional requirement. Effects say what an operation does to data. Audiences say who it is for. Data says whether it is stored and whether it is personal. Measurement says whether behaviour, health or a service level is tracked. Risk raises the bar. From these, and nothing else, the policy resolves what is owed.

§6

Resolver

The policy engine below is the same code that runs in CI, compiled into this page. Change the facts of one operation and the list of obligations is recomputed. Items highlighted after a change are the ones that change added.

input · one operation

example
audience
effect
data
measure
shape
risk

output · resolved by the policy engine

91obligations

plan 6authoring 36merge 42promotion 3runtime 4

surfaces 5
product-uipublic-apiopenapi-clientmcpcli
access 18
access-policyauthenticationauthorizationcredential-permissionapi-key-permissionoauth-app-permissionoauth-consentoauth-discoverypermission-compatibilityerrorsidempotencyrate-limitcompatibilityauditdeletionrecoverydestructive-previewdestructive-approval
data 7
data-schemamigrationretentionpii-classificationprivacydata-boundaryresidency
proof 8
test-unittest-integrationtest-contracttest-audittest-boundarysecurityperformancereliability
analytics 4
analytics-outcomeanalytics-eventanalytics-emitanalytics-insight
operations 22
structured-loglog-fieldsredactioncorrelationtracecardinalitytelemetry-costtelemetry-guardrailsmetric-slihealth-dashboardalert-sloevent-healthsignal-livenessalert-livenesssynthetic-livenesssmokerollbackkill-switchrolloutremote-receiptrunbooktroubleshooting
docs 7
acceptanceimplementationproduct-guidechangelogrest-referencemcp-referencecli-reference

* generated from another provider. Highlighted items were added by your last change.

The rules are small and legible. Each fact adds a fixed set of obligations, and the sets compose. The table shows the main ones.

factadds
audience: customera product UI, authentication and authorization, a product guide
audience: developera public API, a generated client, an MCP tool, a CLI command, their references, a rate limit and a compatibility contract
audience: automationAPI-key and OAuth-app permissions, consent and discovery, an error contract and a contract test
audience: operatoran admin UI, an admin API, an admin MCP tool, support actions
effect: deleteaudit, idempotency, deletion, recovery, a destructive preview and a destructive approval
data: personalPII classification, privacy, a data boundary, residency, a boundary test
measure: sloan SLI metric, a health dashboard, an SLO alert and a runbook
risk: criticala kill switch, a smoke test, rollback, security, performance and reliability proof, tracing, synthetic liveness
risk: regulatedprivacy and audit proof, a data boundary and a human approval at promotion
profile: ui-flowempty, loading, error and conflict states, responsive and accessibility proof, an end-to-end test and a rendered story

Two words carry weight here. An obligation whose state is derived is a projection of another obligation and is satisfied only while its source is. An obligation whose service level is derived may be met by a shared provider, such as one dashboard that covers several operations. A raw escape hatch, such as a generic command that forwards to the API, never earns first-class credit.

§7

Surfaces

A feature is used by people, by developers and by agents, and each of them reaches it through a different surface. The contract does not list those surfaces. It lists audiences, and the policy derives the surfaces. A customer audience owes a product UI on every shell the product has. A developer audience owes an API, a client, an MCP tool and a CLI. An automation audience owes the permissions and consent that let a machine hold a credential. An operator audience owes an admin console and its machine twins.

Each surface then declares the operation it implements, in its own registry, in its own shape. The graph joins them by identity and checks that they agree.

operationrecords.record.deletecustomer · developer · automation · operatoreffect: delete
peopleaudience: customer
web appdesktop shellmobile shellbrowser extensionspreadsheet add-onadmin console audience: operator
developersaudience: developer
REST routeOpenAPI client from the emitted OpenAPISDK from the clientCLI command owned decision: raw API onlywebhook trait: webhook-emitter
agentsaudience: automation
MCP toolWebMCP page tool from the MCP declarationcommand palette from the MCP declarationin-app assistant from the MCP declaration
knowledgederived from each surface
REST reference from the OpenAPIMCP reference from the served tool listCLI reference from the command manifestproduct guide frontmatter names the operation

● declares the operation itself○ generated from another declaration× omitted by an owned decision

Fig. 3  Which surfaces an operation owes is computed from its audiences and effects. Each surface either declares the operation, is generated from one that does, or carries an owned decision to leave it out.

A surface binds to an operation where it is defined. A route carries the operation in its definition and the emitted OpenAPI carries it as an extension. A tool declaration carries it beside its schema and effect. A screen is found by scanning the app for executable calls to the endpoint that implements the operation; a string or a comment does not count.

api / routes / records.tsa route declares its operation
defineRoute({
  method: 'delete',
  path: '/v1/records/{id}',
  operation: 'records.record.delete',
  effects: ['delete'],
  access: 'records.automation.write',
  idempotency: 'safe-to-retry',
  handler: deleteRecord,
})
capabilities / records.tsone declaration, four agent surfaces
defineCapability({
  name: 'delete_record',
  operation: 'records.record.delete',
  input: DeleteRecordInput,
  annotations: {
    effect: 'delete',
    runtimes: ['mcp', 'in-app', 'palette'],
    authz: { minRole: 'editor' },
  },
  preview: previewDelete,   // what the agent shows before acting
  run: deleteRecord,
})

The second declaration is the one that spreads furthest. A tool declared once is served to remote MCP clients, registered on the page for browser agents through WebMCP, listed in the command palette and offered to the in-app assistant, all executing through the same server with the same rules. A parity test pins the page manifest to the live tool list, tool for tool and schema for schema. Every delete asks for typed confirmation on every one of those surfaces, because the confirmation class comes from the declaration, not from the surface.

Native shells, extensions and add-ons are further product UI projections. When one exists it registers the operations it implements and the graph counts it. When it does not, the policy is unchanged and nothing pretends.

the graph checksacross every projection of one operation
effectsevery projection declares the same effects, and a delete surface says it is destructive
accessthe same named policy, the same scopes, the same tenant boundary
safetythe same retry contract and the same confirmation class for agents
contractsmachine surfaces carry an error contract and a schema contract
freshnessgenerated projections are regenerated in CI and any drift fails

§8

Stages

Obligations are due in order, and each is judged at the stage where its claim can be true. Plan, authoring and merge are deterministic and make no external writes. Promotion and runtime need evidence from the live system: a dashboard defined in the repository cannot satisfy them on its own.

  1. planlocal
  2. authoringlocal, CI
  3. mergeCI
  4. promotionlive receipt
  5. runtimeexpiring receipt
Fig. 4  Every obligation is due at exactly one stage. The first three are judged from the repository alone. The last two need a receipt from the live system.
stagetrue when
8.1planAcceptance criteria, access policy, retention, PII classification and rollout are decided.
8.2authoringImplementation, surfaces, unit tests, typed events, logs and docs exist.
8.3mergeIntegration, contract and end-to-end tests pass. Dashboards and alerts are defined and consistent with the events the code emits.
8.4promotionA smoke test passed against the deployed revision. Dashboards and alerts were applied, and a receipt bound to the exact definitions says so.
8.5runtimeEvents arrive, metrics have data, alerts are evaluating and synthetic probes pass. The receipt expires within hours and must be renewed.

§9

Operations

A feature is not finished when its code merges. It also owes the signals that show it working for people, the signals that show it healthy, and the means to act when it is not. Each of these is an obligation with a provider, and each link between them is checked.

  1. product analyticstyped eventoutcomeinsight
  2. operational telemetryhealth dashboardSLO alertrunbook
  3. live proofpromotion receiptruntime receipt
Fig. 5  What a measured operation owes after it ships. The first two rows are checked from the repository. The third needs the live system.
analytics / records.tsan event is bound, typed and budgeted
export const recordTrashed = defineEvent({
  name: 'records.record_trashed',
  capability: 'records',
  operation: 'records.record.delete',
  story: SPEC.records.record.trash,
  outcome: 'success',
  frequency: 'per-action',
  schema: z.object({ record_kind: RecordKind }),
})

export const recordDeleteFailed = defineEvent({
  name: 'records.record_delete_failed',
  operation: 'records.record.delete',
  outcome: 'failure',
  schema: z.object({ error_code: ErrorCode }),   // required on every failure event
})
dashboards / records.tsa dashboard binds the operation and carries its alert
export const records = defineDashboard({
  key: 'records',
  operationBindings: [bind('records.record.delete')],
  tiles: [
    tile({
      key: 'delete-errors',
      query: ratio(recordDeleteFailed, recordTrashed),
      alerts: [
        ratioCeiling({
          name: 'Records: delete failures',
          maxPercent: 5, minSamples: 20, interval: 'hourly',
          rationale: 'deletes that fail leave a person unsure whether data is gone',
          baseline: 'well under one percent across the last quarter',
          enabled: true,
        }),
      ],
    }),
  ],
})
  1. 9.1

    two streams, kept apart

    Product analytics answer whether people get value: use, success and actionable failure per capability. Operational telemetry answers whether the system is healthy: logs, metrics and traces. Logs are never analytics, and an event emitted only to satisfy a count is a defect.

  2. 9.2

    typed events

    An event declares its capability, operation, story, outcome and payload schema. A failure event must carry a categorical error code. A per-request event must declare a sample rate or a cost control. Violations throw when the module loads, before any test runs.

  3. 9.3

    structured logs

    A log line is a stable message plus typed attributes from one shared vocabulary. Each component declares the fields it may log; an undeclared field is dropped in production and throws in tests. Correlation ids are bound once at the entrypoint and inherited.

  4. 9.4

    metrics with closed labels

    Every metric declares its labels and their allowed values. An undeclared label or an id-shaped value is refused, and the series ceiling per metric is asserted.

  5. 9.5

    dashboards and alerts as code

    Tiles read named events and metrics. An alert carries an interval, a rationale and a baseline, and binds the operations it watches. A consistency test checks that tiles read only events the code emits, and that every SLO operation has a dashboard and an enabled alert.

  6. 9.6

    receipts and liveness

    Applying the definitions to the live analytics system produces a receipt bound to the revision and the definition digest. At runtime a separate receipt proves events arrived, metrics have data and alerts evaluated recently. It expires within hours.

  7. 9.7

    runbooks, kill switches, rollback

    Critical paths owe a runbook, a kill switch named in the contract and a rollback signal. Smoke probes against the deployed revision are the promotion proof, not a green deploy job.

§10

Budgets

Fast feedback is a correctness feature. A slow check is skipped or run late, and an agent can only act on the signal it receives. So speed and cost are gates too, with the same shape as every other gate: a declared budget, a check that measures against it, and a failure that names what to fix.

Performance is also an obligation. For a critical capability the policy demands performance and reliability proof, and the provider is a test with a budget in it, not a paragraph in a plan.

budgetmeasured againstfails when
operation latencya performance test runs the operation against seeded data and asserts its p95; the test is the provider for the performance obligationthe budget is exceeded
telemetry costmetric points, analytics events, log lines and error events per thousand requests, per servicea class exceeds its budget
cardinalityclosed label sets per metric, a series ceiling, a denylist of id-shaped labelsan undeclared label or value appears
event volumea per-request event must declare a sample rate, anonymity or a cost controlthe module throws on load
bundle sizecompressed JS and CSS per package, compared with the merge basegrowth beyond the threshold
worker startupthe entry stays small: minified, heavy modules imported lazily, large corpora as separate modulesthe shape regresses
context packetthe bounded context an agent receives, under a hard token ceilingthe packet cannot fit
instructionsbytes, words and estimated tokens per instruction file, and for the whole chainany file or the chain is over budget
timeper-test, per-suite and per-job ceilings; end-to-end artifact size ceilingsa timeout is a failure, never a retry
the machinea dev profile declares its CPU, memory and named resources; heavy commands queue behind an admission brokera profile that cannot fit the machine refuses to start
hot reloadan edit to a shared package becomes visible within its budget, and only in the worktree that made itthe change is late or leaks to a peer
tests / records.performance.test.tsthe performance provider
const BUDGET_MS = { firstPage: 400, deleteWithHistory: 250 }

test('lists the first page within budget at a thousand records', async () => {
  await seedRecords(db, 1_000)
  const p95 = await percentile(95, times(20, () => listRecords(db, { limit: 50 })))
  expect(p95).toBeLessThan(BUDGET_MS.firstPage)
})

§11

Inner loop

The method is only as good as the loop a developer or an agent actually runs. Every step is a stable command with bounded output and a JSON form. None of them writes to a live system. The first two answer from the graph in well under a second, so the plan starts from what is owed rather than from what someone remembered.

  1. contextwhat does this capability owe, and what is blocking?instant
  2. planwhich obligations are missing, by stage?instant
  3. buildchange the contract, then the code, then regenerateyou
  4. verify:packagedoes this one package still pass?seconds
  5. capability:checkis every obligation real at this stage?seconds
  6. verify:fastdoes the affected graph still pass?a minute
  7. pull requestone required check judges the commitminutes
Fig. 6  The loop a person or an agent runs. Each step is one command with a bounded answer. Nothing in it writes to a live system.

Three loops run at different speeds. The first two block. The third may be observational by suite, but it is never allowed to stay red without an owner.

looptriggerscopepace
editsave, pre-commitchanged files and their testsseconds
pull requestpush, agent checkpointthe affected graph, escalated by riskminutes
confidencemerge, nightly, releasethe full graph, live providers, slow suitesbudgeted per suite

Agents work in parallel, so each works in its own worktree, and the worktree is a correctness boundary rather than a convenience. Two agents starting the same profile must never share a port, a database, an emulator or a process, and stopping one must leave the other healthy.

resourcerule
identityderived from the git directory, not the branch name
portsleased by name from a registry, only those the profile needs; an occupied port names its owner and is never killed
statetemp data, emulator state, containers and build output are private to the worktree; the package store is shared and content-addressed
readinessa service is ready when its health endpoint answers, not when a port opens
preparationcode generation and migrations run only when their inputs change, under a lock
capacityeach profile declares weight, CPU, memory and named resources; a broker admits runs and queues heavy ones
$ pnpm -w dev:plan venture --profile appillustrative; nothing is started
{
  "services": ["app", "internal-api", "postgres"],
  "waves": [["postgres"], ["internal-api"], ["app"]],
  "ports": { "app": "leased", "internal-api": "leased", "postgres": "leased" },
  "readiness": { "app": "GET /  200 <html", "internal-api": "GET /ready 200" },
  "resources": { "weight": 3, "cpu": 2.5, "memoryMb": 3072, "named": { "docker-engine": 1 } },
  "prepare": ["schema: inputs unchanged, skipped"]
}

§12

Gates

Every pull request is judged by one required check. It fails unless each gate below passes. Each gate is deterministic, has an owner, and names exactly what is missing and the smallest command that repairs it.

$ pnpm -w capability:check venture/records -- --stage mergeillustrative failure
venture/records stage=merge graph=1f3a…c9 policy=8b2e…04
  records.record.delete  mcp            missing   no provider registered
  records.record.delete  mcp-reference  missing   derived source mcp is missing
  records.record.delete  alert-slo      missing   no enabled alert binds this operation
  owner=records-team  source=capabilities.json#records.record.delete
  repair: declare the tool in capabilities/records.ts, then pnpm -w capability:providers venture
satisfied=41 missing=3 failed=0 stale=0 not-applicable=1
exit 1
  1. 12.1

    required check

    One always-run job aggregates every selected job. A skipped job that had work, a cancelled job or a malformed result fails it. There is no other required check.

  2. 12.2

    selection

    Changed files map to affected packages and their dependants. A change outside every package selects everything. The selection and its reasons are an artifact anyone can read.

  3. 12.3

    capability graph

    An obligation has no provider at its stage, a provider is below the required service level, two projections of one operation disagree, or an override is invalid or expired.

  4. 12.4

    story proof

    An active story names no test, render or event, or evidence names a story that does not exist.

  5. 12.5

    history guard

    A debt baseline grew, an expiry moved later, or a prior record was edited. Compared with the merge base, which the change cannot edit.

  6. 12.6

    permission impact

    A permission was removed or renamed, or a grant was broadened without an approval bound to the digest of that exact change.

  7. 12.7

    generated freshness

    Running every generator changes any tracked file. Generated files are never hand-edited.

  8. 12.8

    analytics

    An event name is not a registry literal, a failure event lacks an error code, a tile reads an event nothing emits, or an SLO operation has no dashboard and alert.

  9. 12.9

    instructions

    An instruction file is over budget, links to a missing path, names an unregistered command, contradicts a universal rule, or a harness-specific rule file exists at all.

  10. 12.10

    architecture

    A dependency points the wrong way, forms a cycle, crosses between ventures, or a route touches storage directly.

  11. 12.11

    ratchets

    Dead code, legacy database access and architecture debt each have an owned, expiring baseline. New findings fail; resolved ones must be pruned.

  12. 12.12

    migrations and environment

    A migration lacks its manifest or breaks the order graph. A build reads an undeclared variable or a secret enters a cacheable task.

  13. 12.13

    contracts and size

    A breaking OpenAPI change without the breaking label. Bundle growth beyond the threshold against the merge base.

  14. 12.14

    quarantine

    A quarantined end-to-end test has no issue, no expiry, or an expiry too far out. A weekly run catches expiries that pass while nothing merges.

  15. 12.15

    planted failures

    Each policy checker runs against a planted good input and a planted bad input. A checker that stops rejecting its bad input fails the build.

The ratchet. When the policy becomes stricter, older features are neither waved through nor blocked. Each gap is recorded by exact identity with an owner, a reason and an expiry. The record may only fall, a policy change that adds gaps must be appended as a signed entry naming them, and the whole history is compared with the merge base so it cannot be rewritten.

Recorded gaps step down over time. An attempt to add a gap is refused.recorded gapstimerefused: debt grew
Fig. 7  The debt record may fall. It may not rise, and its history cannot be rewritten.

§13

Docs and media

Documentation is an obligation like any other, and so are the images and videos inside it. They are produced by the build from the same components the product uses, so they cannot quietly fall out of date.

  1. component storyscreenshot, lightscreenshot, darkvisual baseline
  2. real ui + scenesproduct videodocs embed
  3. api · mcp · cliREST referenceMCP referenceCLI reference
Fig. 8  Screenshots, videos and references are outputs of the build, not uploads.
docs / records / deleting.mdxa page declares what it documents
---
title: Deleting and restoring records
specs: ['records']
operationBindings:
  - { capability: records, operation: records.record.delete }
  - { capability: records, operation: records.record.restore }
---

<Screenshot story="records.record.trash" alt="A record in the trash, with Restore" />
  1. 13.1

    product guide

    Every capability, and every operation a person or developer uses, owes a written guide. A docs page declares which operations it documents in its frontmatter; a capability without one fails.

  2. 13.2

    generated references

    REST, MCP and CLI references are generated from the emitted OpenAPI, the served tool list and the command manifest. A freshness check fails the build when a reference drifts from the code.

  3. 13.3

    screenshots from code

    Component stories tagged with a capability and story are rendered by a headless browser in light and dark, with motion off and the clock frozen. The same image is the documentation asset and the visual-regression baseline.

  4. 13.4

    exact renders

    Each render records the source revision and tree it came from. A screenshot counts only when it matches the current source.

  5. 13.5

    videos from code

    Product videos are composed from the real interface components with fixtures, then rendered. There is no screen recording, so re-rendering after an interface change refreshes every video.

  6. 13.6

    a decision per capability

    Every active capability carries a reviewed video decision: covered, planned or exempt, with a reviewer, a date and a reason. Rendering runs in a gated release workflow, and only when its inputs change.

§14

Commands

The repository is driven by a small set of registered commands with stable names. Each has a description, a JSON form, documented exit codes and bounded output. People and agents run the same ones, and the instruction checker rejects any instruction that names a command the registry does not contain. Select a command to run it.

~/ventureillustrative output for the example capability
$ pnpm -w context venture/recordsinstructions: AGENTS.md, venture/AGENTS.mdcapability-contract: exact graph=1f3a…c9 policy=8b2e…04 blockers=2+0docs: coding-practices.md, venture/docs/specs/records.mdverify: pnpm -w verify:package @venture/recordscontext: within budget; sha256:7c1e…--- venture/docs/specs/records.md:41-62 (section; complete; sha256:…) ---## Acceptance criteria

§15

Agents

Instructions are short routers, not manuals. One file at the root, one per venture and, rarely, one per package, each under a checked token budget. They say where to start, how to get context, which commands verify, and the few safety traps that cannot be inferred. They do not contain the surface checklist, because the policy owns it, and they do not restate any rule a gate already enforces.

Every harness reads the same file. The harness-specific file is one line that imports it, and the instruction check rejects any other harness-specific rule file outright. Context is retrieved per task through the context command, not repeated in every prompt.

instruction chaininstructions:check
AGENTS.md               universal rules        budgeted
  venture/AGENTS.md     venture constraints    budgeted
    package/AGENTS.md   local constraints      budgeted, rare
CLAUDE.md               one line: @AGENTS.md
.cursorrules            rejected
copilot-instructions    rejected

Skills are verbs. Capabilities, packages and rules are data retrieved from the graph, so there is no skill per feature. Each skill is written once in the open skill format at the repository root, where every harness can discover it, and starts by asking the context command for its workflow and exactly one venture overlay.

  1. 15.1

    plan

    Reads the resolved graph as the work list. A missing provider is work, never "not applicable". Reports the graph and policy digests it planned against.

  2. 15.2

    build

    Changes the contract first, then the code, then regenerates projections. Generated files are never edited by hand. Verifies one package, then the capability at its stage, then the affected graph.

  3. 15.3

    verify

    Runs the cheapest checks that can disprove the change first. Missing, skipped, quarantined, timed-out or stale evidence is an explicit result, never a pass. Read-only.

  4. 15.4

    review

    Independent reviewers in parallel, each with the whole diff. Every finding is reproduced against the code before it counts, then fixed with a test that fails without the fix. Discarded findings are recorded with the reason.

  5. 15.5

    migrate

    Moves one legacy capability at a time onto the current policy, reducing debt only through the ratchet. No meaningless events, screenshot-only tests or permanent exceptions to satisfy a count.

The agent contract. An implementation agent is a worker inside these rules, not a judge of them.

  1. cannot claim completion

    while any in-scope obligation is blocking. The check decides, and the check names what is left.

  2. cannot invent an omission

    An agent may implement a not-applicable decision a person approved. It cannot create one.

  3. cannot fabricate later proof

    Promotion and runtime evidence come from receipts minted by authorised workflows, never from a local definition or a report.

  4. cannot grant itself authority

    Deploys, production data, secrets and destructive actions are privilege classes. A skill or an instruction is never authority.

Authority. Commands resolve to a privilege class from read-only to sensitive-destructive, strictest match wins. A sensitive command needs an explicit invocation and an approval signed outside the repository: a short-lived token bound to the exact command, revision and working directory, verified against a trust root the agent cannot write. Even then the approval authorises; it does not execute. Execution happens in a trusted clean checkout.

approval requestsigned outside the repository, expires in minutes
{
  "capabilityId": "…",
  "issuer": "release-owner",
  "revision": "9d4e…",
  "workingDirectory": "/srv/checkout",
  "argv": ["pnpm", "venture:db:migrate:prod"],
  "commandDigest": "sha256:…",
  "issuedAt": "…", "expiresAt": "…"
}

§16

Comparison

Most agent harnesses today are instruction-led: a carefully written instruction file, a set of skills, a few hooks, and a review. Spec-first variants add a written specification before the code. Both improve routing. Neither answers what a feature owes, and neither can tell an agent's report from a fact. Obligation Engineering keeps the routing and moves the truth into the repository.

questioninstruction-led harnessobligation engineering
where a rule livesin an instruction file, read every taskin a check, run on every commit
scope of a featurewhatever the prompt and the plan rememberedderived from audience, effects, data, measurement and risk
a missing surfacesilencefails until someone owns the omission
definition of donetests pass and the agent reports doneevery obligation real at its stage: surfaces, access, data, proof, analytics, operations, docs
proofthe agent says it ran the checksone required check computed on the commit
legacy gapsgrandfathered, or a growing allowlistexact, owned, expiring, only shrinking
contextthe whole instruction file, every timea bounded packet retrieved per target
instructions over timegrow with every incidentstay under a budget; a repeatable rule moves into a gate
harnessa rule file per tool, drifting apartone canonical file; others rejected
authority"be careful with production"privilege classes and signed, expiring approvals
when the agent is wrongsomeone notices in review, or laterthe check names the capability, surface, stage, owner and repair

§17

Destination

The north star is a repository that can answer, for any change, two questions without a person remembering and without an agent being believed: what does this owe, and is each of those things real. When a repository can do that, adding an agent adds capacity and nothing else. Adding a surface, a venture or a harness changes the facts, not the method.

It is honest about its limits. A gate proves facts about the repository. Facts about the live system need receipts, and receipts expire. The policy is code that people review, and a stricter policy creates debt that must be owned rather than hidden. Nothing here replaces judgement about what is worth building. It makes sure that whatever is built arrives whole.

Every rule above started as a sentence in an instruction file. Each one was moved into a check the day it was repeated, and the sentence was deleted. That is the whole method, and it is why the instruction files stay short.

end

Questions and corrections: hello@flagbase.com