.ai

Evidence and decisions

Method

How the process treats evidence and decisions — claims carry their grounds, every source carries the case against it, losers are kept, corrections are added — and whose shoulders it stands on.

Invent as little as possible

Every concern is assigned to an existing standard. The scheme adopts vocabularies and ideas, not stacks — PROV’s terms, not RDF storage; Git’s model, not a .git folder.

Every source receives one verdict

A standard, a paper, a practice or a family of them

Adopt
taken on, often only in part: an idea, a distinction, a shape, a vocabulary
Adapt
taken with a stated change
Evidence
informs a rule, not a component
Reduce
expressible with relations the scheme already has
Set aside
viable, not chosen — with the reason
Reject
considered and refused — with the reason
  1. CausalityW3C PROVRDF storage
  2. VocabularySKOSOWL inference
  3. Attestationin-toto · SLSA · DSSEthe bundle format
  4. Append-only proofTransparency checkpointscontent addressing alone
  5. RulesDatalog, pure fragmenta second rule language
  6. VersionsGit's object modela .git folder
  7. Tool surfaceMCPgRPC, protobuf
  8. PolicyCedar's disciplineCedar

The rule is to take the smallest part of a standard that does the job. A vocabulary travels: PROV’s relations are already Datalog atoms, so the scheme gets causality without an RDF store. A stack does not travel — it brings its storage, its runtime and its failure modes, and a design built from several stacks inherits all of them.

Take the vocabulary. Leave the stack.

The case against borrowing

Borrowing only the idea means owning the implementation. The scheme’s rule engine is the only maintained, embeddable Go Datalog with stratified negation, and its bus factor is about one. The exit stays cheap because the rules are text the scheme owns and the facts are regenerated, so any lock-in sits in files under its control.

A borrowed name can also overstate itself. The scheme’s identifiers follow the ULID layout — and ULID is a README, not a standard. The standards-track construction is UUIDv7 in RFC 9562, whose monotonicity rules the scheme applies, and the scheme says which is which.

Narrow briefs, drawn boundaries, written as they go

Research runs as parallel agents, one field each. Every brief names the boundary it must not cross, and every report is written to disk section by section.

The process is the same for every agent; only the forms differ. For a person, an interruption is a handoff, a budget is a deadline and a second attempt is new work — for an AI agent, they are the end of a session, a token allotment and a fresh sample from the same model.

Held in memory

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10

A report held only in an agent's memory is lost when the work is interrupted.

Written as it goes

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10

A report written section by section survives; a rerun resumes from the sections already on disk.

Partition check

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10
  11. 11
  12. 12

The briefs are read against the question as a whole. A field that no brief opens gets briefs of its own.

Schematic. The tiles are illustrative.
Objective
One field, stated as the question it must answer.
Output shape
The same deliverables for every brief, the case against each source and a section for gaps included.
Sources and tools
Where to look, and a budget for the shared ones.
Boundary
What it must not research — the neighbouring brief's field.

Every brief asks for the same deliverables in the same shape, the case against each source included, so the reports can be read against each other. The partition is checked too. A set of briefs that all study one side of a question, and none that opens the neighbouring literature, has a hole — and the hole belongs to the partition, not to any agent’s work.

Boundaries, not communication, make parallel work possible.

Why narrow, bounded, on disk and budgeted

Narrow, because each agent then holds a context it can actually use. Bounded, because an explicit edge is what keeps parallel agents apart. Measured on AI agents: they are correlated, not independent — in one published red-team study, 18 of 30 chose the identical branch name. Without an edge, they all research the same easy thing.

On disk, because what is in a file survives an interruption and what is held only in an agent’s memory does not. Each report starts as a skeleton of sections and is appended as each section is finished.

Budgeted, because agents that share a budget exhaust it together, and whatever is done after it runs out is guesswork. Each agent gets its own allotment, or the agents are staggered, and a report that ran short says so in its own gaps section. An agent that hands part of its work on issues a sub-brief of the same shape, written to disk the same way — and accepts it itself only when it is a leaf (The issuer accepts; the owner accepts what branches).

Every claim carries its grade

Sources are saved as bytes, named by their hash; quotations are matched by script. Each claim says how it was verified, and what could not be.

  1. Verified verbatimquotation matched by script against the saved bytes
  2. Paraphraseread in the source, restated
  3. Secondary sourceread through someone who read it
  4. Abstract onlythe full text was not reached
  5. Paywalledread through extracts and open commentary
  6. Not readgeneral knowledge, flagged at the claim

By script

Every quotation is matched against the saved bytes of its source. A quotation that does not match is not a quotation.

By hand

A sample is re-read against the same bytes, because the checker can be wrong too. Dropped symbols or a changed apostrophe are caught here, and the quotation is replaced with the source's exact text.

A page fetched through a summarizer is a paraphrase, never a quotation. What could not be verified is said at the claim, not in a footnote: paywalled standards, login-walled frameworks, unreachable PDFs and image-only scans are each listed by name.

An agent that will guess an answer will also guess a source.

Pin sources by content — and what pinning costs

Web addresses drift. A standards body’s site can start serving a gambling affiliate page; a cited post can quietly resolve somewhere else. A link is not a citation; a digest is.

The case against: pinning makes a poisoned input stable as well as auditable. A bad input that passes review once passes forever, so pinning never replaces review.

The case against every source

No source is admitted on its own say-so. Each comes with who wrote it and why, what it was measured on, and whether it is evidence or a framing.

  1. Q1

    Who wrote it, and why?

    Vendor documentation is graded as a vendor describing its own product.

  2. Q2

    What was it measured on?

    A benchmark built by the authors of the defence it evaluates is noted as such.

  3. Q3

    Is it perishable?

    Benchmark scores age in months; ratios and taxonomies last longer.

  4. Q4

    Do its authors sell the fix?

    Then its framing of the problem is a claim, not a finding.

  5. Q5

    Is it evidence, or a framing?

    A memorable framing is credited for being memorable, not for being a study.

Surveys report the disagreements between sources, not a smoothed consensus — and resolve them by layer where they can. Governance regimes disagree on what scales process weight; the resolution is that what must be produced keys on consequence, while how hard someone checks keys on size (Big work becomes small, checkable tasks).

A survey that reports only support would read the same if its conclusion were wrong.

A disagreement resolved, not averaged

One vendor’s published research practice splits work across parallel agents; another vendor cautions against multi-agent designs. Read closely, it is not a factual dispute: the first explicitly excludes coding and shared-context work, which is where the caution applies. The resolution keeps both — the scheme may be as thorough as it likes, but the context handed to any one agent is minimal.

Correct the record in public

A claim that turns out wrong is marked wrong where it was made, with what replaced it. Summaries are claims too, and are checked like any other.

A correction is a new record that points at the old one — the same rule the scheme applies to everything (Nothing is overwritten. Everything is superseded.). A finding once marked verified and later refuted stays visible as refuted, with both reports linked, so anyone can see what changed and why. Each kind of correction is named, because “wrong” covers very different failures:

  1. refuted

    Was: Governance regimes never scale process weight by size.

    Primary texts contain counter-examples. What must be produced keys on consequence; tailoring and scrutiny key on size.

  2. superseded

    Was: IEEE 1012-2016 is the verification standard.

    Superseded by 1012-2024. Its widely quoted phase names come from the withdrawn 2004 edition.

  3. corrected

    Was: The famous 1012 risk grid is a requirement.

    It is informative, not normative.

  4. misattributed

    Was: “Record the reduction” comes from ISO/IEC/IEEE 29148.

    It comes from ISO/IEC/IEEE 15289 §5.2.

  5. category error

    Was: Organizations mirror their products 70% of the time.

    That figure covers firms. For open collaborative projects, 56% do not support mirroring.

  6. overstated

    Was: ULID is a standard.

    ULID is a README. The standards-track construction is UUIDv7, in RFC 9562.

  7. retracted

    Was: A cited paper supports the claim.

    Read in full, it does not discuss the subject at all. Retracted as a citation; the claim stands or falls on other grounds.

  8. over-claimed

    Was: A summary states more than its sources support.

    A collated review catches it, and the summary is corrected like any other record.

A record that hides its own errors cannot be checked.

One decision at a time, losers kept

Decisions are taken singly and recorded when made, each with the alternatives weighed and why each lost. A reason that only explains the choice is not evidence.

fails the test

“We chose Datalog because it is decidable and can explain every answer.”

True — and it would read exactly the same if Datalog were the wrong choice.

passes

Datalog, over each alternative:

RDF + SPARQL
cannot express general recursion, so rules would need a second language
Cypher, GQL
pattern languages, not rule languages — and a second store beside the record
OPA / Rego
a second rule language, outside the decidable fragment
Our own engine
stratification and negation fail with wrong answers, not crashes

A decision's status

  • settled
  • at risk
  • amended
  • superseded
  • retired

Its number is kept for life. A retired number leaves a gap and is never reused.

Its approval

  • in conversation
  • awaiting signature
  • signed

A draft is never mistaken for an acceptance. Silence is refusal, never assent.

The test is simple to state. A justification that names only the reason for the choice would read exactly the same if the choice were wrong, so it carries no evidence. A sound record says why each alternative lost — and those reasons are the most re-cited part of any decision, because without them every rejected idea comes back and is argued again from scratch.

A reason that only explains the choice isn’t evidence.

Recorded when made, searched before changing

A conversation is not storage: a decision that is not written down when it is made is lost when the session ends. And before anything new is put forward, the record is searched for what already bears on it, so every change reconciles with past decisions instead of quietly contradicting them.

Keeping the losers needs no special machinery. Deciding is an activity that used every option it weighed, so the options that lost stay attached to the decision by the provenance graph itself.

Let worked cases force the decisions

Rules are checked against whole worked cases, record by record. Where a case breaks, the break forces a decision, and the decision names the case.

  1. 01

    Worked case

    A full lifecycle of work, told with a cast and a timeline.

  2. 02

    Transcript

    Every record and every invocation, by actor, machine-readable.

  3. 03

    Check

    Every record run, by tools, against the rules.

  4. 04

    Where it breaks

    A decision is forced — and names what forced it.

  • ✕ The case cannot be expressed.
  • ✕ A trace stops short.
  • ✕ A rule refuses evidence that passed every earlier check.

Many decisions are forced, not chosen in discussion. A rule is checked against complete worked cases — a whole lifecycle of work, told with a cast and written out as every record and every invocation, by actor — not only against the case that prompted it. Rules are also tested by mutation: each test must fail when the rule it guards is changed, so a rule that cannot fail is caught.

A worked case does not care what anyone hoped.

You can play one step of this yourself in Try it: you are the agent.

Review loops that know when to stop

One reviewer per concept, nothing fixed until all are in, findings collated into root causes — and a written rule for telling progress from thrashing.

Schematic. Three finding classes; the numbers are illustrative.
  1. 1

    Findings are classified by someone other than their generator.

  2. 2

    A loop is converging only while open findings strictly decrease, class by class.

  3. 3

    A class that recurs after a structural fix triggers a spike, not another round.

  4. 4

    Reviewer recall is spot-checked with seeded defects, so under-reporting cannot pass for convergence.

Large reviews put one reviewer on each concept, reading the whole corpus, and fix nothing until every reviewer is in: the same defect surfaces from several angles, and fixing early churns lines another reviewer is about to report. The findings are then collated into root causes. A mistake made once and then repeated surfaces as many findings; one fix to the cause clears them all.

A loop without a stopping rule is a budget leak with a progress bar.

An exit criterion that cannot deadlock

“Keep reviewing until it is clean” has no end, and an agent will happily iterate forever. Google’s code-review guidance supplies the form that cannot deadlock: approve a change once it “definitely improves the overall code health”, even if it is not perfect. Monotonic improvement is a test that always has an answer.

State the limits as limits

When a mechanism cannot prove something, the scheme says so plainly. A capability the harness cannot enforce is declared unenforceable, never quietly trusted.

MechanismProvesDoes not prove
A harness signatureThese bytes came out of this run, under this brief, unaltered.That they are true. It attests custody, not content.
VerificationThat the artifact meets its criteria.Anything about the process. Judged work cannot be re-derived: done twice, it does not come out the same.
Two runs that agreeThat one agent, asked twice, answered the same way.Corroboration. One agent asked twice is one opinion, said twice.
The append-only recordThat nothing was removed — to anyone who checks.Anything, against whoever holds the root key.
A gate the agent is asked to respectNothing.It is not a gate. A capability the harness cannot enforce is declared unenforceable.

A boundary that is stated stops people relying on the system beyond it. A boundary that is hidden becomes the place the system fails. So non-repudiation by an agent is called what it is — impossible by construction, since the agent holds no key and the harness signs — and what remains checkable is said instead: whether an agent’s claims match a transcript the harness wrote.

A gate the agent is asked to respect is not a gate.

The one hard no in regulation

Regulation draws the same line, and reading it is a check on the scheme, not a compliance claim. Every regime that addresses it lets a non-human identity act. They diverge on attesting, and one is a flat no: 21 CFR Part 11 requires that “each electronic signature shall be unique to one individual”. Software is not an individual, so under Part 11 software cannot sign at all. In the scheme no agent holds a signing key: the harness signs what a run executed, the owner signs every acceptance that reaches the owner, and every record distinguishes “executed by” from “attested on behalf of”.

Standing on shoulders

Fourteen bodies of prior art: what each is, the part the scheme takes, the part it leaves, the case against each — and what the scheme defines where none reaches.

Each entry links to the source itself. Open The case against on any card: every source carries one.

  1. 01Causality

    W3C PROV

    The W3C's model of provenance: entities, the activities that use and generate them, and the agents responsible.

    We take The backbone. Every artifact is related to another through the activity that consumed one and produced the other.

    We leave RDF as the storage

    The case against

    PROV has no notion of correction or supersession in an append-only store, and its alternateOf is, in its own words, “a necessarily very general relationship”. Having provenance is also not using it: measured on AI agents, those with memory inspected it in about one episode in five.

  2. 02Versions and history

    Git

    The version-control system whose content-addressed objects, commits and branches made history cheap and trustworthy.

    We take Its model: immutable versions, events that move pointers, drafts as branches, acceptance as a merge, corrections as new commits.

    We leave a .git folder, squash and rebase

    The case against

    Git's tooling provides little of what the scheme needs, and a store that looks like a repository risks being treated as one by Git itself — so the store uses none of Git's file or folder names.

  3. 03Attestation

    in-toto, SLSA and DSSE

    The software supply chain's standards for signed statements about what was built, from what, and by whom.

    We take The shape of a brief, a run and a signed result, addressed by the digest of the specification — and the rule that the harness signs, never the agent.

    We leave the in-toto bundle format

    The case against

    SLSA's gold standard is reproducibility, which judged work cannot supply. Per-line signatures in the bundle format leave deletion and replay undetectable. And a signature proves less than people assume: not when, not completeness, not intent.

  4. 04Append-only, proven

    Transparency logs

    Merkle-tree logs — Certificate Transparency, Go's checksum database, witness cosigning — that make deletion detectable.

    We take Signed checkpoints and consistency proofs, so “nothing was removed” is a proof, not a promise. Irreversible acts wait for a witness.

    We leave content addressing as the whole answer

    The case against

    A misbehaving log can show different views to different clients, and the fix is still “an active area of research”. Every transparency design ends in “somebody else must look”; an unwatched log is a cost with no benefit.

  5. 05Rules

    Datalog

    A decades-old logic language that is decidable, terminating and able to explain every answer it gives.

    We take One rule language for every gate, so each “blocked” comes with a proof tree of why.

    We leave engine extensions that forfeit termination

    The case against

    Unfamiliar to most contributors, and a family rather than one language. It cannot count, and some real rules want counting. The chosen engine has a bus factor of about one — a stated risk with a cheap exit, since the rules are text the scheme owns.

  6. 06Vocabulary

    SKOS and the thesaurus standards

    The W3C and ANSI/NISO vocabularies for concepts, labels, hierarchies and scope notes.

    We take One definition per term, alternate labels so things can be found, non-transitive hierarchy, and concept schemes as bounded contexts.

    We leave OWL and RDFS as inference machinery

    The case against

    The field's own verdict on controlled vocabulary is modest: it recovers “a quarter to a third” of records, and “we still do not have definitive proof” that a thesaurus is worth building.

  7. 07Work identity

    Bazel

    Google's build system, which identifies each action by the digest of its full specification and sandboxes it.

    We take Work identified by what it was asked to do, timeout included — and the lesson that a declaration is only true when something enforces it.

    We leave the transparent cache

    The case against

    Bazel's cache assumes actions are reproducible. Judged work is not: reusing a result returns one attempt, not the answer. Reuse is a decision with a recorded reason.

  8. 08Division of labour

    Parasuraman, Sheridan and Wickens

    The 2000 human-factors model that splits automation into four independently set stages.

    We take “Agents make judgements, tools perform the mechanics” as a named, peer-reviewed configuration rather than a slogan.

    The case against

    Designed for human operators of automated systems, which fits an agent who is a person working through tools. The transfer to AI agents is by analogy, not by measurement.

  9. 09Verdicts

    FIPA, A2A and Design by Contract

    Agent-communication and software-contract traditions that separate refusing, failing and not understanding.

    We take “Will not”, “tried and could not” and “the request was wrong” as distinct verdicts, each with its own recovery.

    The case against

    FIPA is effectively defunct; its specification survives only in an archive, because its old address no longer serves it.

  10. 10Rationale

    IBIS, Toulmin and Dung

    Fifty years of work on recording reasons, warrants and attacks as first-class structure.

    We take Claims and defeats as records whose standing is derived, not stored — with an agent as the scribe IBIS always needed.

    We leave participants typing their own map

    The case against

    The scribe can invent: an agent writing twelve claims overnight is not corrected in the room. And a rationale graph that blocks people is one people will route around.

  11. 11Verification

    NASA SE Handbook and IEEE 1012

    Aerospace and software-assurance practice for verification, waivers and independent checking.

    We take Completion as evidence-or-waiver plus every discrepancy closed, and independence three ways: who runs the check, who chooses it, who can cut it short.

    We leave waivers without a clock

    The case against

    No source anywhere bounds a waiver with an expiry. IEEE 1012-2016 is superseded by 1012-2024, its widely quoted phase names come from a withdrawn edition, and its task tables are paywalled.

  12. 12Staleness

    Build systems and suspect links

    Early cutoff from build-system theory; suspect links from requirements-tracing tools such as Doorstop.

    We take Mechanical staleness — every derived item remembers the versions it was built from — and propagation that stops when nothing really changed.

    The case against

    When an agent judges “meaning unchanged”, that judgement is the correctness boundary, and no theory covers a fallible equality.

  13. 13Authority

    Object capabilities

    The security model in which authority is held by reference and can only be narrowed when passed on.

    We take Authority that flows down by attenuation only, recorded so it can be reviewed and revoked.

    We leave ambient authority

    The case against

    Its classic weaknesses are review and revocation. The scheme answers both with an append-only log of grants.

  14. 14Tool surface

    Model Context Protocol

    An open protocol for exposing tools to AI agents over JSON-RPC.

    We take The tool surface: every mechanical operation is a tool any agent can call, and a refusal is a result the agent can read.

    We leave trusting unpinned tool descriptions

    The case against

    In the specification, authorization is optional and stdio transports are outside it. Tool descriptions are read by the agent that calls the tools, not by the owner who approved them — so the tool surface must be pinned by digest.

And what is turned down

Each refusal is kept with its reason:

  • OWL and RDFS as machinery. RDFS domain infers rather than constrains, so a wrong value silently reclassifies a document; OWL’s open world cannot report a missing field. Their property terms are kept as vocabulary.
  • owl:sameAs. In one study only about 51% (±21%) of real-world uses were correct. SKOS’s graded matches are used instead.
  • RRULE and cron. RRULE silently drops invalid instances; cron ORs day-of-month with day-of-week and has no notion of a missed run.
  • Allen’s interval algebra as machinery. Reasoning over it is NP-hard, and parked-and-resumed work is not one interval.
  • Upper ontologies. Their value is shared commitment across institutions; one scheme would pay the whole cost and collect none of it.
  • A second rule language. OPA/Rego, and Cedar as a dependency: two rule languages are worse than one. Cedar’s discipline — decidable, ground, proven — is kept.
  • A forge’s branch protection as the guarantee. It governs branches, while immutable records can live in refs it does not govern.

Where no standard reaches

Six needs fall between the standards. Some are partly covered — FIPA, A2A and Design by Contract already split “will not” from “could not” — but the pieces that make the whole work are defined nowhere else. The scheme defines them.

What no standard covers, the scheme defines.

  1. 01

    Agents as attested producers

    Briefs, runs and results are signed, digest-addressed records. Custody is kept distinct from content.

    Nearest prior art in-toto, SLSA

  2. 02

    The “cannot be done” verdict

    Stated in advance as part of the work's own contract, with a no-progress rule, and reconsidered when the reason for the work is superseded.

    Nearest prior art FIPA, A2A, Design by Contract

  3. 03

    Correction in an append-only store

    A correction is a new record that supersedes the old one. Nothing is edited. None of PROV, CloudEvents, in-toto or SKOS provides this.

    Nearest prior art review dispositions, argumentation, Kanban

  4. 04

    Bounded waivers and parks

    Every deferral carries a clock or a trigger. No source bounds a waiver; the scheme does.

    Nearest prior art triage practice

  5. 05

    Keeping derived knowledge current

    Staleness is mechanical. “Meaning unchanged” is a judgement, made twice and recorded.

    Nearest prior art build systems, suspect links

  6. 06

    Claims admitted only with captured evidence

    A claim enters with a verbatim quote found in bytes the harness read, checked by a different actor from its own reads.

    Nearest prior art none found

No folklore numbers

The process uses no number it cannot trace to a measurement. Most widely repeated process numbers trace to nothing; the one measured sizing number is a ratio of time horizons.

  1. 8/80

    The 8/80 rule for work packages

    No primary source found. The institutional rule is “short or with interim objective milestones” — and every retelling drops the second half.

    no primary source
  2. ≤ 10

    Cyclomatic complexity

    One unquantified ranking of 24 subroutines, by the project's own members.

    one ranking
  3. 200–400

    Lines per code review

    A curve from 300 of an advertised 2,500 reviews, authored by a vendor.

    vendor data
  4. ◁

    The Cone of Uncertainty

    A 1958 chemical-industry shape. The one direct test found a pipe, not a cone.

    contradicted
  5. 150

    Dunbar's number

    Refuted: confidence intervals run from 4 to 520.

    refuted
  6. 3d · 2w · 6w

    Methodology sizing rules

    Three ideal days, two weeks, six weeks, two days and fifteen people — none referenced to any study.

    no study
  7. $ / level

    Cost per assurance level

    Traces to magazine articles from 1994–95.

    magazine
  8. 4–10×

    Time horizons

    Measured on AI agents: the task length they finish 80% of the time is four to ten times shorter than the length they finish half the time. Computed from raw benchmark data, not a summary.

    measured

“Make it small” is a lossy compression of “make its progress objectively measurable” — the half of the institutional rule that every retelling drops. A number with no origin is labelled as convention wherever it appears, so nobody re-imports it as evidence.

Everything else is convention.

The case against the one measured number

Absolute time horizons are perishable within months; the ratio between the 80% and 50% horizons is the durable finding. Nobody has measured a 99% horizon. Even sized against the 80% horizon, one run in five fails — which is why an agent’s “cannot be done” verdict is load-bearing rather than an embarrassment.

Search every section

Commands

  1. zOverview: every section on this pageview
  2. themeSwitch light or darkui

Sections

  1. ixn.aiixn.ai
  2. The schemeThe scheme
  3. The lifecycleThe lifecycle
  4. Agents and trustAgents and trust
  5. NorthstarsNorthstars
  6. Work outgrew the way we organize itixn.ai
  7. Seven storiesSeven stories
  8. Artifacts, not documentsixn.ai
  9. MethodMethod
  10. Agents judge. Tools do the mechanics.ixn.ai
  11. Work goes out as a brief and comes back as evidenceixn.ai
  12. GlossaryGlossary
  13. Nothing is overwritten. Everything is superseded.ixn.ai
  14. Four layers, one truthixn.ai
  15. Big work becomes small, checkable tasksixn.ai
  16. Values, given teethixn.ai
  17. Try it: you are the agentixn.ai
  18. One word, one meaningGlossary
  19. Identity and historyGlossary
  20. Records and relationsGlossary
  21. Work and bundlesGlossary
  22. Trust and authorityGlossary
  23. Process and weightGlossary
  24. Roles and judgementGlossary
  25. Invent as little as possibleMethod
  26. Narrow briefs, drawn boundaries, written as they goMethod
  27. Every claim carries its gradeMethod
  28. The case against every sourceMethod
  29. Correct the record in publicMethod
  30. One decision at a time, losers keptMethod
  31. Let worked cases force the decisionsMethod
  32. Review loops that know when to stopMethod
  33. State the limits as limitsMethod
  34. Standing on shouldersMethod
  35. No folklore numbersMethod
  36. How work moves through the systemSeven stories
  37. The command that lied about successSeven stories
  38. The plan that was wrongSeven stories
  39. No end in sightSeven stories
  40. Rolled back in ninety secondsSeven stories
  41. The rules change mid-flightSeven stories
  42. Most of what arrives is never builtSeven stories
  43. Autonomy is earnedSeven stories
  44. One cycle, at every levelThe lifecycle
  45. Five rules govern the graphThe lifecycle
  46. Accepting the specification is the commitmentThe lifecycle
  47. Gate where a mistake becomes uncorrectableThe lifecycle
  48. Gates are enforced, never requestedThe lifecycle
  49. The issuer accepts; the owner accepts what branchesThe lifecycle
  50. Oversight and capability run beside the workThe lifecycle
  51. Consequence decides what; size decides how hardThe lifecycle
  52. Err smallThe lifecycle
  53. Decompose until every leaf passes four testsThe lifecycle
  54. A budget is a fuse, not a decisionThe lifecycle
  55. The record is the processThe lifecycle
  56. Twelve lines worth keepingSeven stories
  57. Four primitives: source, node, version, actThe scheme
  58. Four classes of nodeThe scheme
  59. A kind is permanent. Nothing is promoted.The scheme
  60. An identity that explains itselfThe scheme
  61. Native names firstThe scheme
  62. Versions are addressed by what they containThe scheme
  63. Paths derive from identity, so nothing movesThe scheme
  64. The card: what an agent reads insteadThe scheme
  65. References point up. Every list is a view.The scheme
  66. Edges are nodes, and staleness is mechanicalThe scheme
  67. Metadata on everything — including the metadataThe scheme
  68. Three written forms. State is a fold.The scheme
  69. The scheme checks itselfThe scheme
  70. The unwritten rules of good work, written downNorthstars
  71. Honest reasoningNorthstars
  72. Trust that survives inspectionNorthstars
  73. Attention to the unexplainedNorthstars
  74. Care for the readerNorthstars
  75. Continuity of attentionNorthstars
  76. The good path is the easy pathNorthstars
  77. The dignity of boundariesNorthstars
  78. Reverence for the irreversibleNorthstars
  79. Minds decide, tools doNorthstars
  80. Intent before effortNorthstars
  81. A record that only growsNorthstars
  82. One answer, one shapeNorthstars
  83. Every mechanism is a value given teethNorthstars
  84. The brief is a sealed contractAgents and trust
  85. Least privilege is absence, enforced three timesAgents and trust
  86. Narrow the task, widen the readingAgents and trust
  87. The run: containment, not trustAgents and trust
  88. One signed manifest, or nothingAgents and trust
  89. Evidence is what the harness sawAgents and trust
  90. Judgements are declared choicesAgents and trust
  91. Nothing checks its own workAgents and trust
  92. Authority is a token the agent never holdsAgents and trust
  93. Sign only what agents must not alterAgents and trust
  94. One owner, one lease, one fencing tokenAgents and trust
  95. Cut at wide dependencies, and own every joinAgents and trust
  96. Drift and done are derived, never declaredAgents and trust
tool · agent · owner — each voice is who acts ↑↓ move ↵ open esc close