---
title: Obligation Engineering
publisher: flagbase
canonical: https://flagbase.com/factory/obligation-engineering
format: agent-readable companion to /factory/obligation-engineering
updated: 2026-10
---

# Obligation Engineering

A way to build software with coding agents. Each feature is described by a
small contract of facts. A policy engine derives everything the feature owes,
from its API and MCP tool to its dashboard, alert and runbook. One required
check on every commit decides whether each of those things is real. Anything
left out needs a named owner, a reason and a review date.

The instructions an agent reads are short, because the rules live in checks.
The commands an agent runs are the same ones a person runs. Nothing is accepted
because an agent said it was done.

All identifiers below belong to one abstracted example capability, `records`.

## The chain

```
spec         why it exists                 written by people
story        what a person gets            written by people
contract     facts about one operation     written by people
obligations  derived by policy             computed
providers    what the repository contains  computed
check        green on the commit, or not   computed
```

## Problem

Agents produce change faster than people can review it. A merged change proves
the code compiles and its tests pass. It says nothing about the API, the
documentation, the dashboard or the alert that should have shipped with it.

A longer instruction file does not fix this. A list of things to remember is
out of date the day a new surface appears, costs tokens on every task, and
cannot tell a forgotten MCP tool from one left out on purpose. Both look like
silence.

So the question changes from "was the code written" to "what does this feature
owe" and "is each of those things real". The repository must answer both, not
a prompt and not the agent.

## The idea in sixty seconds

1. **Facts, not checklists.** A feature is described by what it does, who uses
   it, what data it touches and how risky it is.
2. **Obligations are computed.** A policy engine turns the facts into the exact
   list the feature owes: surfaces, access, data handling, tests, analytics,
   operations and docs.
3. **Providers are discovered.** Each obligation is satisfied by something real
   in the repository, found by reading registries and source. A comment or a
   claim earns nothing.
4. **Omission is a decision.** A missing obligation fails until a person records
   why it does not apply, who owns that decision and when to look again.
5. **The commit is the proof.** One required check runs on every commit. No
   file records that something passed.
6. **One substrate.** People, CI and agents run the same commands and read the
   same output. Instructions route to them; they do not restate them.

## Rules

1. **Describe meaning, not transports.** A contract names a semantic operation
   (for example `records.record.delete`) and its facts. Routes, MCP tools, CLI
   verbs and screens are projections of that identity.
2. **Stories live in code.** Each user story is a typed identity with
   acceptance criteria, defined in the codebase. Everything that proves it
   refers to that identity; a misspelling fails to compile.
3. **Obligations are derived.** A developer audience implies an API, an MCP
   tool and a CLI. A delete implies audit, recovery and a destructive preview.
   An SLO implies a metric, a dashboard, an alert and a runbook.
4. **Missing never means not applicable.** Each obligation is met by something
   real, or carries a decision with a reason, an owner, a review date and an
   approval.
5. **Debt only goes down.** Features that predate the policy have their gaps
   recorded exactly, owned and expiring. The record may shrink, never grow, and
   its history is append-only.
6. **The check is the proof.** A required check passing on the commit is the
   evidence. No file records a pass, because a file can be edited. Agents
   change code and tests; they never regenerate receipts.
7. **Operations are part of the contract.** Events, logs, metrics, dashboards,
   alerts, runbooks, smoke tests and rollback are obligations. Promotion needs
   a receipt from the live system.
8. **Every check earns its place.** A check catches a defect nothing cheaper
   catches, has an owner, a time budget, a planted failure and an expiring
   exception mechanism. Checks that only verified the verification system were
   removed.

## Stories

Development starts from a prose specification that explains the problem, the
reasoning and the acceptance criteria. The stories that are enforced are typed
identities in code. Tests, rendered components, screenshots, docs pages and
analytics events refer to a story by that identity. A reference to a story the
registry does not contain fails the build. Only a story written and reviewed by
a person can be active; generated placeholders stay in draft and expire.

```ts
export const records = defineSpec({
  key: 'records',
  stories: {
    [SPEC.records.record.trash]: {
      text:
        'As a customer, deleting a record moves it to the trash, ' + 'so a mistake can be undone.',
      acceptance: [
        'The record leaves the list immediately',
        'Restore returns it with its history',
        'Only retention destroys the data',
      ],
    },
  },
})
```

```ts
test('trash, then restore', async ({ page, spec }) => {
  spec(SPEC.records.record.trash) // throws if no capability owns this story
  await page.getByRole('button', { name: 'Delete' }).click()
  await expect(page.getByText('Moved to trash')).toBeVisible()
  await page.getByRole('button', { name: 'Restore' }).click()
  await expect(page.getByText('Restored with history')).toBeVisible()
})
```

A story owes a behavioural test, a docs page, a rendered component story and
an analytics outcome. Higher-risk capabilities owe more: security, privacy,
audit and rollback proof.

## Contract

The contract holds facts the tooling can read without running product code.
The vocabulary is closed: profile (ui-flow, api, workflow, library,
non-functional), effects (read, create, update, delete, execute, export,
recover, administer), audiences (public, customer, developer, automation,
operator, internal), data (persisted, personal, exportable), measurement
(behavior, operational, slo) and risk (standard, critical, regulated).

```json
{
  "key": "records",
  "profile": "api",
  "risk": "critical",
  "deliveryStatus": "shipped",
  "evidenceStatus": "enforced",
  "spec": "docs/specs/records.md",
  "rollout": {
    "killSwitch": "useRecords",
    "rollbackSignal": "records.record.delete error rate"
  },
  "operations": [
    {
      "id": "records.record.delete",
      "effects": ["delete"],
      "audiences": ["customer", "developer", "automation"],
      "data": { "persisted": true, "personal": true },
      "measurement": { "behavior": true, "operational": true, "slo": true },
      "access": {
        "customer": "records.customer.write",
        "automation": "records.automation.write"
      }
    }
  ],
  "overrides": [
    {
      "kind": "cli",
      "operation": "records.record.import",
      "state": "not-applicable",
      "reason": "a person stays accountable for each batch",
      "owner": "records-team",
      "reviewAt": "2027-03-01",
      "approval": "docs/decisions/import-is-human.md"
    }
  ]
}
```

## Resolver

The policy engine is a pure function from contract to obligations. Each fact
adds a fixed set, and the sets compose. The interactive resolver on the human
page runs the same code CI runs.

| Fact                 | Adds                                                                                                           |
| -------------------- | -------------------------------------------------------------------------------------------------------------- |
| audience: customer   | product UI, authentication, authorization, product guide                                                       |
| audience: developer  | public API, generated client, MCP tool, CLI command, their references, rate limit, compatibility contract      |
| audience: automation | API-key and OAuth-app permissions, consent, discovery, error contract, contract test                           |
| audience: operator   | admin UI, admin API, admin MCP, support actions                                                                |
| effect: delete       | audit, idempotency, deletion, recovery, destructive preview, destructive approval                              |
| data: personal       | PII classification, privacy, data boundary, residency, boundary test                                           |
| measure: slo         | SLI metric, health dashboard, SLO alert, runbook                                                               |
| risk: critical       | kill switch, smoke test, rollback, security, performance and reliability proof, tracing, synthetic liveness    |
| risk: regulated      | privacy and audit proof, data boundary, human approval at promotion                                            |
| profile: ui-flow     | empty, loading, error and conflict states, responsive and accessibility proof, end-to-end test, rendered story |

Two words carry weight. An obligation whose **state** is derived is a
projection of another obligation (a REST reference from the OpenAPI) and is
satisfied only while its source is. An obligation whose **service level** is
derived may be met by a shared provider, such as one dashboard covering several
operations. Providers are first-class, derived or escape-hatch; an obligation
accepts a provider at its level or above. A raw escape hatch, such as a generic
command that forwards requests to the API, never earns first-class credit.

## Surfaces

The contract lists audiences; the policy derives surfaces. A customer audience
owes a product UI on every shell the product has. A developer audience owes an
API, a client, an MCP tool and a CLI. An automation audience owes the
permissions and consent that let a machine hold a credential. An operator
audience owes an admin console and its machine twins. Each surface declares the
operation it implements in its own registry, and the graph joins them by
identity.

```
operation  records.record.delete

people      ● web app  ● desktop shell  ● mobile shell  ● browser extension
            ● spreadsheet add-on  ● admin console (operator)
developers  ● REST route  ○ OpenAPI client  ○ SDK  × CLI command (owned decision)
            ● webhook (trait: webhook-emitter)
agents      ● MCP tool  ○ WebMCP page tool  ○ command palette  ○ in-app assistant
knowledge   ○ REST reference  ○ MCP reference  ○ CLI reference  ● product guide

●  declares the operation itself
○  generated from another declaration
×  omitted by an owned decision
```

A route declares its operation and the emitted OpenAPI carries it as an
extension. A tool declaration carries it beside its schema and effect. A screen
is found by scanning the app for executable calls to the endpoint; a string or
a comment does not count.

```ts
defineRoute({
  method: 'delete',
  path: '/v1/records/{id}',
  operation: 'records.record.delete',
  effects: ['delete'],
  access: 'records.automation.write',
  idempotency: 'safe-to-retry',
  handler: deleteRecord,
})

defineCapability({
  name: 'delete_record',
  operation: 'records.record.delete',
  input: DeleteRecordInput,
  annotations: {
    effect: 'delete',
    runtimes: ['mcp', 'in-app', 'palette'],
    authz: { minRole: 'editor' },
  },
  preview: previewDelete,
  run: deleteRecord,
})
```

A tool declared once is served to remote MCP clients, registered on the page
for browser agents through WebMCP, listed in the command palette and offered to
the in-app assistant, all executing through the same server with the same
rules. A parity test pins the page manifest to the live tool list. Every delete
asks for typed confirmation on every one of those surfaces, because the
confirmation class comes from the declaration.

Native shells, extensions and add-ons are further product UI projections. When
one exists it registers the operations it implements and the graph counts it.
When it does not, the policy is unchanged and nothing pretends.

Across every projection of one operation the graph checks: effects (a delete
surface says it is destructive), access (same named policy, scopes and tenant
boundary), safety (same retry contract, same confirmation class), contracts
(machine surfaces carry an error and schema contract) and freshness (generated
projections are regenerated in CI; any drift fails).

## Stages

| Stage     | True when                                                                                                  | Judged from       |
| --------- | ---------------------------------------------------------------------------------------------------------- | ----------------- |
| plan      | Acceptance criteria, access policy, retention, PII classification and rollout are decided                  | repository, local |
| authoring | Implementation, surfaces, unit tests, typed events, logs and docs exist                                    | repository, CI    |
| merge     | Integration, contract and end-to-end tests pass; dashboards and alerts are consistent with emitted events  | repository, CI    |
| promotion | A smoke test passed against the deployed revision; dashboards and alerts were applied, with a receipt      | live receipt      |
| runtime   | Events arrive, metrics have data, alerts evaluate, synthetic probes pass; the receipt expires within hours | expiring receipt  |

A receipt names the revision, the graph digest, the policy digest and the
digest of the definitions it applied. A receipt for any other revision or
definition is rejected. Only an authorised workflow can mint one. Authoring and
merge are deterministic and make no external writes.

## Operations

Two streams, kept apart. Product analytics answer whether people get value:
use, success and actionable failure per capability. Operational telemetry
answers whether the system is healthy: logs, metrics and traces. Logs are never
analytics, and an event emitted only to satisfy a count is a defect.

- **Typed events.** An event declares its capability, operation, story,
  outcome and payload schema. A failure event must carry a categorical error
  code. A per-request event must declare a sample rate or a cost control.
  Violations throw when the module loads.
- **Structured logs.** A stable message plus typed attributes from one shared
  vocabulary. Each component declares the fields it may log; an undeclared
  field is dropped in production and throws in tests. Correlation ids are
  bound once at the entrypoint.
- **Metrics with closed labels.** Every metric declares its labels and values.
  An undeclared label or an id-shaped value is refused; the series ceiling is
  asserted.
- **Dashboards and alerts as code.** Tiles read named events and metrics. An
  alert carries an interval, a rationale and a baseline, and binds the
  operations it watches. A consistency test checks that tiles read only events
  the code emits and that every SLO operation has a dashboard and an enabled
  alert.
- **Receipts and liveness.** Applying definitions to the live analytics system
  produces a receipt bound to the revision and definition digest. At runtime a
  separate receipt proves events arrived, metrics have data and alerts
  evaluated recently. It expires within hours.
- **Runbooks, kill switches, rollback.** Critical paths owe a runbook, a kill
  switch named in the contract and a rollback signal. Smoke probes against the
  deployed revision are the promotion proof, not a green deploy job.

```ts
export const recordTrashed = defineEvent({
  name: 'records.record_trashed',
  capability: 'records',
  operation: 'records.record.delete',
  story: SPEC.records.record.trash,
  outcome: 'success',
  frequency: 'per-action',
  schema: z.object({ record_kind: RecordKind }),
})

export const records = defineDashboard({
  key: 'records',
  operationBindings: [bind('records.record.delete')],
  tiles: [
    tile({
      key: 'delete-errors',
      query: ratio(recordDeleteFailed, recordTrashed),
      alerts: [
        ratioCeiling({
          name: 'Records: delete failures',
          maxPercent: 5,
          minSamples: 20,
          interval: 'hourly',
          rationale: 'deletes that fail leave a person unsure whether data is gone',
          baseline: 'well under one percent across the last quarter',
          enabled: true,
        }),
      ],
    }),
  ],
})
```

## Budgets

Fast feedback is a correctness feature. A slow check is skipped or run late,
and an agent can only act on the signal it receives. Speed and cost are gates
with the same shape as every other gate: a declared budget, a check that
measures against it, a failure that names what to fix. Performance is also an
obligation; for a critical capability the provider is a test with a budget in
it.

| Budget            | Measured against                                                                     | Fails when                           |
| ----------------- | ------------------------------------------------------------------------------------ | ------------------------------------ |
| operation latency | a performance test runs the operation against seeded data and asserts its p95        | the budget is exceeded               |
| telemetry cost    | metric points, events, log lines and error events per thousand requests, per service | a class exceeds its budget           |
| cardinality       | closed label sets per metric, a series ceiling, a denylist of id-shaped labels       | an undeclared label or value appears |
| event volume      | a per-request event declares a sample rate, anonymity or a cost control              | the module throws on load            |
| bundle size       | compressed JS and CSS per package, against the merge base                            | growth beyond the threshold          |
| worker startup    | minified entry, heavy modules imported lazily, large corpora as separate modules     | the shape regresses                  |
| context packet    | the bounded context an agent receives, under a hard token ceiling                    | the packet cannot fit                |
| instructions      | bytes, words and estimated tokens per file and for the whole chain                   | over budget                          |
| time              | per-test, per-suite and per-job ceilings; artifact size ceilings                     | a timeout is a failure               |
| the machine       | a dev profile declares CPU, memory and named resources; heavy commands queue         | a profile cannot fit the machine     |
| hot reload        | a shared-package edit is visible within budget, only in its own worktree             | late, or leaks to a peer             |

## Inner loop

Every step is a stable command with bounded output and a JSON form. None writes
to a live system.

```
context            what does this capability owe, and what is blocking?   instant
plan               which obligations are missing, by stage?                instant
build              change the contract, then the code, then regenerate     you
verify:package     does this one package still pass?                       seconds
capability:check   is every obligation real at this stage?                 seconds
verify:fast        does the affected graph still pass?                     a minute
pull request       one required check judges the commit                    minutes
```

Three loops: edit (save, pre-commit; changed files and their tests; seconds),
pull request (push or agent checkpoint; the affected graph, escalated by risk;
minutes) and confidence (merge, nightly, release; the full graph, live
providers, slow suites; budgeted per suite). The first two block. The third may
be observational by suite but never stays red without an owner.

Agents work in parallel, so each works in its own worktree, and the worktree is
a correctness boundary. Identity derives from the git directory. Ports are
leased by name, only those the profile needs; an occupied port names its owner
and is never killed. Temp data, emulator state, containers and build output are
private; the package store is shared and content-addressed. A service is ready
when its health endpoint answers. Preparation runs only when its inputs change.
Each profile declares its weight, CPU, memory and named resources, and a broker
admits runs and queues heavy ones.

A developer machine runs the narrow checks that can disprove a change quickly.
CI produces the proof.

## Gates

Every pull request is judged by one required check. It fails unless each gate
passes. Each gate is deterministic, has an owner, and names what is missing and
the smallest repair.

```
$ pnpm -w capability:check venture/records -- --stage merge
venture/records stage=merge graph=1f3a…c9 policy=8b2e…04
  records.record.delete  mcp            missing   no provider registered
  records.record.delete  mcp-reference  missing   derived source mcp is missing
  records.record.delete  alert-slo      missing   no enabled alert binds this operation
  owner=records-team  source=capabilities.json#records.record.delete
  repair: declare the tool in capabilities/records.ts, then pnpm -w capability:providers venture
satisfied=41 missing=3 failed=0 stale=0 not-applicable=1
exit 1
```

- **Required check.** One always-run job aggregates every selected job. A
  skipped job that had work, a cancelled job or a malformed result fails it.
- **Selection.** Changed files map to affected packages and dependants. A change
  outside every package selects everything. The selection and its reasons are
  an artifact.
- **Capability graph.** An obligation has no provider at its stage, a provider
  is below the required service level, two projections disagree, or an
  override is invalid or expired.
- **Story proof.** An active story names no test, render or event, or evidence
  names a story that does not exist.
- **History guard.** A debt baseline grew, an expiry moved later, or a prior
  record was edited. Compared with the merge base, which the change cannot
  edit.
- **Permission impact.** A permission was removed or renamed, or a grant was
  broadened without an approval bound to the digest of that exact change.
- **Generated freshness.** Running every generator changes any tracked file.
- **Analytics.** An event name is not a registry literal, a failure event lacks
  an error code, a tile reads an event nothing emits, or an SLO operation has
  no dashboard and alert.
- **Instructions.** A file is over budget, links to a missing path, names an
  unregistered command, contradicts a universal rule, or a harness-specific
  rule file exists at all.
- **Architecture.** A dependency points the wrong way, forms a cycle, crosses
  between ventures, or a route touches storage directly.
- **Ratchets.** Dead code, legacy database access and architecture debt each
  have an owned, expiring baseline. New findings fail; resolved ones must be
  pruned.
- **Migrations and environment.** A migration lacks its manifest or breaks the
  order graph. A build reads an undeclared variable or a secret enters a
  cacheable task.
- **Contracts and size.** A breaking OpenAPI change without the breaking label.
  Bundle growth beyond the threshold.
- **Quarantine.** A quarantined end-to-end test has no issue, no expiry, or an
  expiry too far out.
- **Planted failures.** Each policy checker runs against a planted good input
  and a planted bad input. A checker that stops rejecting its bad input fails.

**The ratchet.** When the policy becomes stricter, older features are neither
waved through nor blocked. Each gap is recorded by exact identity with an owner,
a reason and an expiry. The record may only fall. A policy change that adds
gaps must be appended as an entry naming them, and the history is compared with
the merge base so it cannot be rewritten.

**What was removed.** Evidence envelopes, recorded receipts, a registry of
gates and a health check for the gates. None had failed a pull request in a
month; they verified the verification system. A green check on the commit
replaced them.

## Docs and media

Documentation is an obligation, and so are the images and videos inside it.
They are produced by the build from the same components the product uses.

- **Product guide.** Every capability, and every operation a person or
  developer uses, owes a written guide. A docs page declares which operations
  it documents in its frontmatter.
- **Generated references.** REST, MCP and CLI references are generated from the
  emitted OpenAPI, the served tool list and the command manifest. Drift fails
  the build.
- **Screenshots from code.** Component stories tagged with a capability and
  story are rendered by a headless browser in light and dark, with motion off
  and the clock frozen. The same image is the docs asset and the regression
  baseline.
- **Exact renders.** Each render records the source revision and tree it came
  from, and counts only when it matches the current source.
- **Videos from code.** Product videos are composed from the real interface
  components with fixtures, then rendered. No screen recording.
- **A decision per capability.** Every active capability carries a reviewed
  video decision: covered, planned or exempt, with a reviewer, a date and a
  reason.

```mdx
---
title: Deleting and restoring records
specs: ['records']
operationBindings:
  - { capability: records, operation: records.record.delete }
  - { capability: records, operation: records.record.restore }
---
```

## Commands

Registered commands with stable names, a description, a JSON form, documented
exit codes and bounded output. People and agents run the same ones, and the
instruction checker rejects any instruction naming a command the registry does
not contain.

| Command                          | Purpose                                                                  |
| -------------------------------- | ------------------------------------------------------------------------ |
| `context <target>`               | A bounded, hashed context packet: instructions, contract, blockers, docs |
| `capability:new`                 | Scaffold a capability; preview its resolved obligations first            |
| `capability:plan <capability>`   | Resolve the full work matrix without failing                             |
| `capability:check <capability>`  | Fail on any missing, failed or stale obligation at a stage               |
| `capability:graph <capability>`  | The full normalised graph as JSON                                        |
| `capability:providers <venture>` | Discover providers from registries and source                            |
| `capability:baseline <venture>`  | Check, or ratchet down, recorded debt                                    |
| `verify:package <package>`       | Verify one package; empty verification is rejected                       |
| `verify:fast`                    | Lint, typecheck and test the affected graph                              |
| `explain:affected`               | Which packages and CI jobs a change selects, and why                     |
| `instructions:check`             | Validate the instruction hierarchy and budgets                           |
| `dev:plan <venture> --profile`   | The resolved dev profile: services, waves, ports, resources              |
| `generate`                       | Regenerate every projection from canonical sources                       |

Illustrative output for the example capability:

```
$ pnpm -w context venture/records
instructions: AGENTS.md, venture/AGENTS.md
capability-contract: exact graph=1f3a…c9 policy=8b2e…04 blockers=2+0
docs: coding-practices.md, venture/docs/specs/records.md
verify: pnpm -w verify:package @venture/records
context: within budget; sha256:7c1e…
--- venture/docs/specs/records.md:41-62 (section; complete; sha256:…) ---
## Acceptance criteria
```

```
$ pnpm -w capability:plan venture/records
venture/records stage=authoring graph=1f3a…c9 policy=8b2e…04
satisfied=38 missing=2 failed=0 stale=0 pending=0 not-applicable=1
  records.record.delete  mcp          missing  no provider registered
  records.record.delete  product-ui   missing  no executable endpoint reference
handoff authoring={"blocking":2,"pending":19} merge={"blocking":5,"pending":4}
unresolved-owner-decisions=0
```

## Agents

- **Instructions are routers.** One file at the root, one per venture and,
  rarely, one per package, each under a checked token budget. They say where
  to start, how to get context, which commands verify, and the few traps that
  cannot be inferred. They do not contain the surface checklist and do not
  restate any rule a gate enforces.
- **One canonical file.** The harness-specific file is one line that imports
  it (`CLAUDE.md` contains `@AGENTS.md`). Other harness-specific rule files
  are rejected by the instruction check.
- **Context is retrieved, not repeated.** The context command returns a bounded,
  hashed packet per target: instruction chain, contract digests, blocking
  obligations and exact document sections.
- **Skills are verbs.** Plan, build, verify, review and migrate are written
  once in the open skill format at the repository root. There is no skill per
  feature; capabilities and rules are data retrieved from the graph.
- **Plan** reads the resolved graph as the work list; a missing provider is
  work, never not-applicable.
- **Build** changes the contract first, then the code, then regenerates
  projections; generated files are never hand-edited.
- **Verify** runs the cheapest checks that can disprove the change first;
  missing, skipped, quarantined, timed-out or stale evidence is never a pass.
- **Review** runs independent reviewers in parallel, reproduces every finding
  before it counts, and fixes each with a test that fails without the fix.
- **Migrate** moves one legacy capability at a time onto the current policy,
  reducing debt only through the ratchet.

The agent contract: an agent cannot claim completion while an in-scope
obligation is blocking, cannot invent a not-applicable decision, cannot
fabricate promotion or runtime evidence, and cannot grant itself authority.

Authority: commands resolve to a privilege class from read-only to
sensitive-destructive, strictest match wins. A sensitive command needs an
explicit invocation and an approval signed outside the repository: a
short-lived token bound to the exact command, revision and working directory,
verified against a trust root the agent cannot write. The approval authorises;
execution happens in a trusted clean checkout. A skill or an instruction is
never authority.

Hooks are optional. No invariant depends on an editor plugin, a harness hook or
a git hook. CI is the enforcement boundary.

## Comparison

| Question                | Instruction-led harness                 | Obligation Engineering                                           |
| ----------------------- | --------------------------------------- | ---------------------------------------------------------------- |
| where a rule lives      | in an instruction file, read every task | in a check, run on every commit                                  |
| scope of a feature      | whatever the prompt and plan remembered | derived from audience, effects, data, measurement and risk       |
| a missing surface       | silence                                 | fails until someone owns the omission                            |
| definition of done      | tests pass and the agent reports done   | every obligation real at its stage                               |
| proof                   | the agent says it ran the checks        | one required check computed on the commit                        |
| legacy gaps             | grandfathered, or a growing allowlist   | exact, owned, expiring, only shrinking                           |
| context                 | the whole instruction file, every time  | a bounded packet retrieved per target                            |
| instructions over time  | grow with every incident                | stay under a budget; a repeatable rule moves into a gate         |
| harness                 | a rule file per tool, drifting apart    | one canonical file; others rejected                              |
| authority               | "be careful with production"            | privilege classes and signed, expiring approvals                 |
| when the agent is wrong | someone notices in review, or later     | the check names the capability, surface, stage, owner and repair |

## Destination

A repository that can answer, for any change, two questions without a person
remembering and without an agent being believed: what does this owe, and is
each of those things real. When a repository can do that, adding an agent adds
capacity and nothing else. Adding a surface, a venture or a harness changes the
facts, not the method.

Its limits are explicit. A gate proves facts about the repository. Facts about
the live system need receipts, and receipts expire. The policy is code that
people review, and a stricter policy creates debt that must be owned rather
than hidden. Taste, product direction, approvals and every decision to leave
something out stay with people.

Every rule above started as a sentence in an instruction file. Each one was
moved into a check the day it was repeated, and the sentence was deleted. That
is why the instruction files stay short.

## Contact

hello@flagbase.com
