Harnessing the frontier: what stands between the model and the world
The agent says "done," and the screen gives you every reason to believe it. There is a patch, the test command has exited with zero, and the final answer explains what changed with the confidence of someone who has checked. Then you open the output and find that no tests ran. The process succeeded at doing nothing, and somewhere between that empty run and the final answer, nothing became evidence.

Figure 1. The same exit code, two different worlds: only one of them is evidence.
That is the failure I want to follow. The bug fix is a small enough task that we can see what happened, but it contains most of the problems that become harder to see in a larger system. The worker can produce a reasonable patch and still misunderstand what it is allowed to publish. It can open a pull request, lose the response, and open another on restart. It can read a repository file that asks for environment variables as part of a diagnostic check and mistake material it was supposed to inspect for permission it was never given.
It is tempting to collect these failures under "the model got it wrong." Sometimes it did. But that explanation stops exactly where the interesting work begins. Why was a claim accepted without the evidence needed to support it? Why could an uncertain outcome become a fresh attempt? Why could a file being read change the authority of the reader? Better judgment might make those mistakes less frequent; it does not answer what the system should do when they happen anyway.
That is where I put the harness: between a model's proposal and an effect in the world. It gathers context, exposes tools, checks authority, controls execution, and keeps the records needed to understand what happened. It also owns the less glamorous questions about resources, recovery, and stopping. Those responsibilities do not disappear because the model becomes more capable, any more than making a prompt longer makes a missing network response arrive.
The question that holds this together is simple: what would let someone else check? A fluent explanation is not enough, and neither is a diagram with "memory," "reasoning," and "tools" arranged around a model. To get beyond the diagram, we have to follow a proposal until it becomes an action, then follow the action until there is enough evidence to say what it accomplished. When that evidence is missing, the architecture has to make room for uncertainty instead of filling it with a better sentence.
1. Start with the loop#
At the center of the system, the model proposes a next step and the harness decides whether that step may happen. It sounds like a small distinction until the proposal leaves the context window. An edit changes a file, a service accepts a request, or money leaves an account. At that point, we are no longer discussing the quality of generated text. We are discussing what allowed that text to acquire consequences.
The loop makes that transition visible:
load durable state
assemble context
ask the model for the next step
if it proposes an action:
validate the request
check authority and limits
reserve budget and resources
record the attempt
execute through a controlled tool
record the observation
update state
otherwise:
evaluate whether the task is actually complete
continue, pause, or stopThis is a sketch, not crash-safe code. Its most important gap sits between execution and recording the observation: the effect can happen before the local state knows that it happened. We will return to that gap when the pull request loses its response, because a surprising amount of recovery work lives inside it.
Even before we reach execution, "valid" needs to mean more than "the arguments matched the schema." A request can have the right shape while using the wrong authority, exceeding the remaining budget, or asking a tool to do something it cannot reliably verify. After execution, a successful call still has to be connected to the task's acceptance conditions. The test command exiting cleanly did not repair our bug by itself.
Anthropic's distinction between workflows and agents is useful here: a workflow fixes more of the route in code, while an agent leaves more choices to the model. Either way, the route eventually reaches controlled execution. There is no exemption for a predictable workflow, and no reason to surrender enforcement because the next step was chosen dynamically.
I am equally suspicious of the claim that the loop is never the bottleneck. Inference may dominate one workload while network calls, tool startup, filesystem work, queueing, or repeated reasoning dominate another. Async I/O can matter more than a faster driver; at higher throughput, contention and persistence can become expensive too. Profile the actual workload. A small loop is easier to read, but its size tells us very little about whether the system around it is correct.
2. Follow one task, not a diagram#
Keep the bug fix in view as we add the rest of the system. The task is to repair a failure without breaking existing behavior, which is already a stricter goal than producing a plausible patch or finishing a sequence of tool calls. If we lose that distinction, every completed step starts looking like progress even when it moves us away from an accepted result.
Before the worker begins, the task needs a repository, branch, issue, and set of constraints. It needs to know which edits, installs, pushes, and pull requests are permitted, and what generated code can reach while doing the work. Once it starts, we need to know what survives a crash, how much time and money it can use, who resolves an unknown outcome, and which checks separate a real fix from a confident claim. By the end, we should also be able to say which workload acted, on whose authority, and which records support the result.
Those are different questions even if one program answers all of them. I do not object to putting the code in a class named Harness; I object to treating the class name as an answer. If the distinctions disappear inside it, a permissions bug can be diagnosed as a model failure, a missing receipt as a failed action, and a broken sandbox as a reason to write a stricter prompt. Each diagnosis sends the repair to the wrong place.
The point of separating responsibilities is therefore practical, not decorative. When the test run is empty, we should know whether the runner failed to discover tests, the evaluator accepted insufficient evidence, or the worker changed the conditions under which success was judged. That is a much more useful starting point than asking why the agent was not careful enough.
3. A tool is a contract#
The worker reaches for a test runner. A function name and a description might tell the model how to call it, but they do not tell the rest of the system what the result means. Starting the command, finishing the command, and passing the tests required for this patch are separate outcomes. If the tool compresses them into one success flag, it has already removed the distinction that would have caught our opening failure.
For that runner, the contract should preserve the command, repository, working directory, timeout, exit code, and captured output, along with the conditions for calling the run successful. For the pull-request tool, it should preserve the repository, base and head branches, title, body, and remote object created. When the call times out, there needs to be a way to look up what happened rather than treating the timeout as a complete account of the effect.
I split that contract into six parts:
- Input: types, bounds, required fields, and accepted values.
- Authority: who has permission for this operation on this resource.
- Execution: what operation the tool attempted.
- Observation: what the service returned or the process produced.
- Effects: what changed, and what still needs confirmation.
- Failure: rejected, failed, or outcome unknown.
The separation matters because a valid destination can still be the wrong destination, and a failed response can follow a successful effect. A tidy interface that catches every exception and returns a friendly status may make the model's life easier while making recovery impossible. Rejection before sending and a lost response after sending cannot share a meaning: in the first case, nothing happened; in the second, something may have happened and we need to find out what.
SWE-agent's agent-computer-interface work treats tools and feedback as design choices rather than incidental plumbing. That is the right starting point. An interface that hides useful output or exposes vague actions forces the model to guess, but an interface built only for model convenience can hide what the system needs to enforce and verify. The contract has to be legible to the model without becoming imprecise for everyone else.
4. Policy is not a promise#
Suppose the worker is allowed to prepare the patch but not publish it. "Do not push to production" belongs in the prompt, but enforcement begins when a production push without valid approval is refused. Blocking a deployment tool while leaving the shell free to call the same API does not enforce the boundary. It closes one entrance and leaves the effect available through another.
That is why the check belongs where the action becomes real. Tool names are convenient labels, not security boundaries. The stable requirement is that no action happens outside granted authority; the changing rules specify who may act, on which resource, within what scope, and until when. Those rules need versions and tests, and each decision needs a record containing the proposal, authority, resource, policy version, result, and reason. A record that says only "allowed" leaves a later reviewer guessing what was allowed and why.
Default-deny does not have to mean interrupting a person over every uncertain detail. Our worker can read the issue, inspect the code, and prepare the pull request while holding the publish step for approval. The separation lets useful work continue without quietly expanding the permission that made it possible.
Now let that approval wait fourteen hours. The branch advances, the short-lived token expires, and the sandbox may still be holding locks or reserved disk. The person eventually approves the action they saw yesterday, but the worker is about to execute against a different state. Pausing execution did not pause time, and the approval cannot be treated as though it preserved the world around it.
Approval gates therefore need explicit TTLs, but a timer alone is not the fix. What the person approves should work like a compare-and-swap: the approval signs an optimistic concurrency token bound to the base commit SHA and a digest of the environment state the reviewer actually saw, not merely a description of the intent. On resume, the gate compares the current base SHA with the approved one and the current state digest with the signed one. If current_base_sha differs from approved_base_sha, the gate trips automatically, before anyone has to notice that the branch moved. Credential validity and the other preconditions get the same re-check. If the current action is no longer covered, ask again; if resources are sitting idle, release them rather than holding them indefinitely. The same discipline applies when policy changes during the wait: the system must know which rule applies to the resumed action, not inherit an answer accidentally from the point where it stopped.

Figure 2. The approval signs a state, not an intent. If the base SHA or environment digest drifted, the gate trips.
Policy changes deserve the care we give code changes. Test permitted and forbidden cases, compare proposed decisions with current ones before enabling a rule, and keep exceptions narrow with expiry dates. The Claude Agent SDK permissions example also makes evaluation order worth inspecting. Hooks, deny rules, ask rules, modes, allow rules, and a callback do not all see the same calls; an earlier approval can bypass a later callback. A permission function is useful only to the extent that the actions needing its check actually reach it.
5. A sandbox limits damage, not meaning#
The worker has permission to edit, so we let it run code. This is usually where someone says "we have a sandbox" and expects the discussion to end. I think that is where it should begin. Which files can the code read, which network destinations can it reach, and what stops it from consuming everything available? The word sandbox does not answer any of those questions on its own.
Isolation constrains access and resource use. It cannot decide whether an edit fixes the bug or whether a permitted email is rude. Policy decides whether an effect should happen; isolation limits the available paths when the model, a tool, or generated code fails. Confusing the two makes it easy to demand semantic judgment from a boundary that was built to control access.
The labels also need to stay honest. Restricted processes can limit files, syscalls, privileges, or network access through OS controls. Containers use namespaces and other controls while usually sharing the host kernel. Sandboxed runtimes such as gVisor add a boundary between guest syscalls and the host kernel, while VMs run a guest kernel behind a virtualization boundary; microVMs reduce startup cost and device surface. A remote disposable environment describes placement and lifetime, and can use any of these boundaries.
That is not a tidy ladder from weak to strong. A misconfigured VM can expose secrets, while a hardened container can suit a particular threat model. "Remote" tells us where the worker runs, not whether it is secure, and "disposable" tells us about its lifetime, not what it can access. Start with the damage we need to prevent: reading host credentials, crossing tenants, writing outside the workspace, reaching arbitrary endpoints, or exhausting resources. Then test those paths against the actual environment.
Anthropic's secure-deployment guidance describes a useful split: keep credentials outside the execution boundary and use a controlled proxy for authenticated requests. The worker can request an operation without receiving the token. That protects credential secrecy, but it still leaves the question of what the worker can do through the proxy. Keeping the token hidden and constraining its authorized use are different controls.
The proxy can also become an SSRF route if the worker can steer requests toward 169.254.169.254, internal VPC DNS names, or private endpoints. Network-boundary egress filtering and a strict destination allowlist belong here, with outbound denied by default and the resolved destination address checked before connecting. Validating only the supplied hostname leaves the network decision unfinished.
Coverage matters as much as the boundary itself. The Claude Code Bash sandbox documentation, checked in October 2026, says that file tools, MCP servers, and hooks run outside that sandbox. So a statement that "the agent is sandboxed" still owes us an account of the paths it covers and the paths it does not. Our repository file should remain task data regardless of which path reads it; it must not acquire the power to expand authority because it happened to arrive inside an isolated workspace.
6. Context is not history#
The bug fix takes longer than expected, and the context window fills up. Someone summarizes the run so the next model call can continue. "Tests were not run" becomes "validation considered," or a proposed action becomes a completed one. The next worker sees a neat account of the task and has no reason, within that account, to question it.
The problem is no longer just compression. A view of the record has changed the meaning of the record. The context window holds what the model sees now; it must not hold the only history of what happened. Otherwise every summary becomes an opportunity to replace an event with an interpretation that later work will treat as fact.
I keep four things separate: event history for requests, decisions, attempts, observations, approvals, and transitions; checkpoints for saved state at useful boundaries; artifacts such as patches, reports, test output, screenshots, and remote-object references; and working context, which selects material for the next model call. A summary belongs in working context. It can help the worker continue without becoming the authority on whether the tests ran.
If compaction drops "test skipped," a surviving source record lets us retrieve and correct the omission. If only the summary survives, the omission is now part of the system's past. Anthropic's Managed Agents architecture separates the session log, harness, and execution environment, allowing a replacement harness to use a surviving session. The OpenHands Software Agent SDK paper describes event-sourced conversation state and deterministic replay. In both cases, the useful distinction is that we append events and derive a working state rather than making the working state our sole evidence.
Replay still needs a precise meaning. Rebuilding state from stored events is not the same as rerunning a model against today's world. Full reproduction can require pinned tools, environment images, inputs, model settings, and the original external responses; some dependencies will not support it. We should say which form we can provide instead of letting the word imply more than the records support.
Durability also cannot mean keeping every secret forever. Evidence storage needs retention limits, minimized sensitive content, restricted access, and recorded redactions. Sometimes the right record contains a reference or digest rather than the original material. Preserving enough to check the result and protecting what should not be retained are both design responsibilities.
7. A budget must reject work before it starts#
Once one worker is behaving, the obvious next move is to start four. That changes what a budget has to do. A counter that notices overspend afterward is a meter, even if the interface calls it a limit. The distinction becomes visible when parallel workers all make locally reasonable decisions against the same balance.
Imagine a run with ten units of budget. Four workers each see six units remaining and each admit work costing four. Every decision fits the balance that worker read, but the combined admission does not. The mistake is not in any worker's arithmetic; it is in the system allowing the same remaining capacity to be spent more than once.
Capacity must be reserved before admission, with atomic reservations wherever workers share a budget, then reconciled against actual usage. Children draw from the parent's allowance unless they have a separate grant. The control also needs to cover the resources that can actually run out: spend, elapsed time, calls, retries, concurrent workers, disk, and compute. A token cap does not stop a stuck subprocess, and a timeout does not prevent a burst of expensive parallel calls.
There is a second pressure here that a balance cannot see. Capacity budgets decide how many tokens and dollars the run may admit in total; they say nothing about how fast sibling workers hit a downstream limit such as tokens per minute, requests per minute, or a database connection pool. Four workers with plenty of budget can still trigger a cascade of 429 responses if they all retry on the same beat. That needs a different mechanism: rate and backpressure queues that control bursts, with jittered leases so siblings do not retry in lockstep. One mechanism decides what the run may spend, the other decides when it may spend it, and neither substitutes for the other.

Figure 3. Capacity budgets bound what a run may spend. Rate queues with jittered leases bound when siblings may spend it.
Even then, the claimed boundary needs to be honest. Stopping a run cannot undo charges already incurred, billing can arrive late, and some services cannot cancel an operation after admitting it. A hard bound requires admission control and a defensible bound on in-flight cost. Without both, we have a best-effort stop and should call it that.
These limits change the task's state, not its truth. A worker waiting for approval has not finished, and an unverified patch with no budget left has not become a fixed bug. When resources run out, the system needs to preserve what was done and what remains unchecked rather than rewarding itself for reaching a stopping condition.
8. Recovery starts with "I do not know"#
Return to the pull request. The service accepts the request, but the response is lost before the worker saves anything, and then the worker dies. The remote world has changed while the local record still looks as though the operation has not finished. On restart, the system does not know whether the pull request exists.
That uncertainty is real state, not an awkward error to tidy away. Retrying can create a duplicate; marking failure can discard a real result; assuming success can invent one. Recovery begins by preserving the unknown outcome long enough to resolve it, which is why the gap in our loop cannot be repaired with a generic retry policy.
The write-ahead log (WAL) pattern gives us a starting point: write intent durably, perform the effect, and mark the confirmed effect committed. After a crash, reconcile intents that lack confirmation. But writing the intent is only half the arrangement. We also need a stable operation ID and a way to check the outside world through provider idempotency support, a remote identifier, or a reliable lookup tied to the attempt. If the service provides none of those, the unresolved outcome has to remain visible.
The pattern also assumes a cooperative downstream, and agents often are not talking to one. Legacy systems, internal admin panels, and raw SSH targets frequently have no idempotency tokens and no read-after-write consistency. The harness should treat such a sink as non-idempotent and degrade its effect classification to irreversible, manual-confirmation-only: when the response drops, it runs a pre-check query if one exists, and otherwise stops for explicit human intervention rather than retrying.

Figure 4. Intent is logged before the effect. A lost response is a known unknown, and the sink class decides how to recover.
An effect journal might distinguish planned, admitted, attempted, confirmed, failed, and unresolved operations. The names can vary, but the distinction between intending an effect and confirming it cannot. The journal should let the resumed worker ask what the earlier attempt actually did before it decides whether another attempt is safe.
Checkpoints help restore data, but they do not resolve a remote effect or restart a computation by themselves. LangGraph's persistence docs describe saved thread state, while Temporal's LangGraph integration explanation distinguishes persistence from durable execution. Something still has to detect failure, resume work, handle waits, and choose retries. Restoring yesterday's state is not enough when today's service already contains the pull request.
The phrase "exactly once" needs the same scrutiny. Queue delivery and deduplication of an external payment are different guarantees, and idempotency keys have scope and expiry. Reusing a key against another account need not refer to the same operation. Rather than trusting the phrase, kill the worker at awkward points and inspect what restart does. A restart button demonstrates that a process can start again; a recovery test demonstrates what it can safely continue.
9. Compensation is another action#
Suppose reconciliation finds that the duplicate pull request already exists. "Just roll it back" sounds reasonable because the mistake is obvious, but the remedy still depends on the effect. A local edit may be restorable from a snapshot. A sent message cannot be unsent by sending another message, a purchase comes with cancellation terms, and an external record may have no recovery route at all.
I classify effects before promising rollback: reversible effects can be restored from a known snapshot or patch; compensable effects have a defined follow-up action that repairs some consequence; irreversible effects need prevention, approval, or an honest account of the loss. Calling all three rollback hides what the recovery actually does.
Compensation happens now, against the current world. It can fail, cost money, require permission, or conflict with someone else's valid work. That makes it a new action with its own record and authority check, not a magical subtraction from the earlier action. Dependencies can tell us the order in which repairs should happen, but they cannot grant permission to restore the whole world to an old snapshot.
This is also why the size of the original effect matters. Repairing a prepared patch is easier than correcting a published mistake, because fewer people and systems have had a chance to rely on it. A narrow effect does not eliminate failure; it gives recovery less to disturb when failure arrives.
10. Do not grade your own work#
We have contained the worker, preserved its state, limited its spending, and recovered the uncertain operation. The patch can still be wrong. A system can execute every step according to its design and fail the task, because orderly execution and correct results are different achievements.
This brings us back to the empty test run. The final answer is a claim, the test output is evidence, and an acceptance rule has to connect them. For this bug fix, zero on exit is insufficient. The runner must discover the expected tests and exercise the changed behavior; the evidence should include the original failure, the patched result, regression checks, and the environment in which they ran. Counting assertions cannot rescue a test suite that never encounters the bug.
Separating the worker from the judge helps, but it is not enough to ask a second model whether the first did well. Shared misleading inputs, mutable tests, or a compromised environment can fool both. The important question is which evidence the judge sees and which parts of the judgment the worker can alter.
SWE-bench's evaluation harness offers a concrete pattern: apply a patch in a controlled environment and run defined tests. That checks the benchmark's acceptance conditions, not security, maintainability, or correctness for every possible input. Anthropic's long-running application harness work separates generation from evaluation and checks applications through end-to-end interaction. The useful principle is independence and test quality, not the suggestion that one evaluator can settle every kind of task.
Look closely at what the worker can change: the tests, the grader, expected results, or failing output. Critical checks belong outside that boundary, or changes to them need independent review. Then the conclusion can be precise enough to inspect: these tests passed on this artifact in this environment. That is a smaller claim than "the agent succeeded," but it tells the next person what they can rely on and what they still need to check.
11. A signed receipt still has limits#
By now the run has records worth preserving, so signing them seems like the next obvious step. It is useful, but it changes the integrity of a claim rather than turning the claim into truth. A signature can make alteration detectable and bind a record to a key; it cannot, by itself, prove that the recorded observation matches the effect in the outside world.
The 2026 individual IETF receipts draft proposes signed, hash-chained action receipts. It puts the motivation sharply:
A record that the operator can silently edit is evidence of intent, not evidence of fact.
This is a proposal, not an adopted IETF standard. Its mechanism is still worth examining: signed receipts linked by hashes can be inspected by another verifier without asking the producer's service to approve the inspection. The hash chain links records, the signature binds a record to a key, and a trust anchor explains why that key belongs to the expected signer. Each does a different job, and none should inherit the guarantees of the others.
The limits are just as important. A compromised signer can sign false records, and an action omitted before logging leaves no trace. Removing the end of a chain can leave a valid shorter chain unless an external anchor exposes the truncation. A signed timestamp is the signer's clock assertion, not independent evidence of time. A log can therefore be consistent, incomplete, and false without breaking its cryptography.
For our pull request, the receipt must keep the policy decision, attempt, observation, artifact digest, and verification result distinct. "Request accepted" cannot become "effect completed" merely because the record was signed. Where the threat model requires it, keep the signer away from the untrusted worker, obtain trust anchors independently, and use external head commitments to detect truncation. Protect the contents as well: integrity is not confidentiality, and an excellent audit trail can still expose information it had no business retaining.
The goal is not a perfect story about the run. It is a set of checks that does not depend entirely on trusting whoever tells that story.
12. Identity is not authority#
The receipt says it was signed by "the agent," which immediately raises another question: which agent? The continuing service, this particular run, the process presenting a credential, or the person whose request started the work? If one name stands for all four, the audit trail may look clear while the permission check is already confused.
I keep those identities separate: the logical agent is the continuing service or role; the run is one execution of one task; the workload is the process or service presenting credentials; and the delegating principal is whose authority that workload uses. An audit identifier can outlive the run's short-lived credentials because identifying past work and authorizing present work are different purposes.
SPIFFE's workload-identity model defines workload identifiers and verifiable identity documents, supporting automated short-lived credentials. It identifies software; it does not prove that the software followed a user's intent. That distinction remains even when the credential is well issued and the signature verifies.
Two implementation examples, checked in October 2026, make the same boundary worth keeping visible. Cursor documents OIDC tokens and HSM-backed signed cloud-agent commits, while Microsoft documents first-class agent identities. Neither proves universal per-run identity, and neither proves that an identity can never return.
Expiration, revocation, issuer policy, and prevention of new credential issuance are separate controls. Destroying a key cannot stop an authorized issuer from making another one, so permanent death is a lifecycle requirement rather than a cryptographic slogan. For delegation, preserve the original principal and narrow the scope at every handoff. Sub-agent token minting must obey monotonic privilege restriction: a child can never hold a capability absent from the parent's active token. Knowing which workload presented the credential begins the authority check; it does not finish it.
13. The log is truth; the graph is a view#
Our log can now tell us that the worker read a file, changed another, and ran a test, in that order. But a reviewer wants to know which source revision the patch used and which patch the test actually checked. A sequence is useful evidence without being a complete account of those relationships, and this is where a provenance graph can earn its place.
For the bug fix, the graph can connect the issue, source revision, fixture, edit, patch, test run, and final claim with typed edges. W3C PROV already models entities, activities, and agents; agent systems add model calls, retrieved evidence, memory, permissions, and tool effects. They are extending a provenance problem, not inventing one.
The 2026 execution-provenance survey treats a run as a typed graph, and LEDGER's research prototype builds review graphs over captured sessions so a reviewer can work backward from a conclusion to evidence. Those are research approaches worth examining, not support for claims that everyone has converged on graphs or that no production harness uses them. The graph is useful when it makes a question easier to answer than the flat record does.
"Causal graph" also needs care, because it can refer to three different things. Control flow describes the intended route; provenance records or infers relationships among actions and artifacts; a causal model makes claims about what would change under intervention. B following A does not establish that A caused B. Reading a file before writing an answer is an observed dependency, not proof of which information drove the answer, and the model's own explanation can be wrong.
Edges should therefore say how they were obtained: direct instrumentation, a deterministic rule, or inference. Keep their links to the underlying records. LEDGER treats inferred structure as an audit aid rather than ground truth, which is the right limit to preserve when a helpful projection starts looking authoritative.
For production recovery, I would start with a structured append-only trace log: JSONL records with execution trace IDs and parent-child span pointers, giving it an OpenTelemetry-style shape. Recovery still depends on the intent states, reconciliation rules, and tool evidence we have already described; a file format cannot supply those guarantees. There is a cost to keeping all of it, though. Hundreds of intermediate tool attempts, file diffs, and bash outputs will bloat a log that embeds them, so the storage boundary matters: raw execution output goes to cold object storage with content addressing, and the trace records carry the artifact digest instead of the contents. The log stays small enough to read, and a reviewer can still fetch the artifact and check that its hash matches the record. Graph views help with traversal, and a retrieval index can be replaced. The full provenance graph is mostly for offline post-mortems and regulatory audit, not a substitute for the operational record.
When I say the log is truth and the graph is a projection, "truth" means the preserved source record, not a guarantee that every recorded claim is true. Changing the projection should not rewrite that record, and a reviewer should still be able to follow an edge back to evidence whose limits are visible.
14. Repair needs dependencies I can trust#
Now imagine discovering a bad memory after the pull request is open: "Tests may be skipped." The worker trusted it, and the resulting verification claim is wrong. Deleting the memory removes a future hazard, but it does nothing to repair the work that already relied on it. We need to trace where the error went before deciding what to change.
A dependency graph can help locate the bad claim and the actions that used it, leaving unrelated work intact. The 2026 dependency-guided rollback preprint studies targeted repair of memory-to-action dependencies. It supports a research approach to finding an affected set, not a claim that arbitrary external effects can be undone.
The quality of that set matters. Missing edges can leave bad work behind, while false edges can discard good work. Even a correct dependency does not tell us whether an external effect is reversible, compensable, or in need of human review. The graph helps diagnose the consequences; it does not inherit authority to repair them.
I would begin with graph-assisted diagnosis and let it propose the affected set. Automatic repair should follow only where the dependencies, effect boundary, and authority support it. Otherwise an inferred relationship can quietly become permission for deletion, turning an audit aid into another source of damage. The useful ambition is targeted repair we can justify, not a system that erases everything connected to a mistake.
15. Scaffolding decays. Infrastructure does not disappear.#
After following the bug fix through all these boundaries, the objection is fair: how much of this machinery exists only because today's model needs help? Some of it does, and keeping it forever would be as careless as removing everything because the next model improves.
Code that helps a model track features, follow a planning ritual, continue across interruptions, or check its own work may encode a temporary limitation. I call that scaffolding. Permissions, finite budgets, irreversible effects, uncertain network outcomes, privacy, and disputes about evidence belong to the world the model acts in. The machinery that deals with those is infrastructure.
Anthropic's harness-design account describes components built around assumptions about model limits, and those assumptions can go stale. A stronger model may remove the need for a planning ritual. It cannot make an unapproved payment authorized or settle a lost network response by thinking more clearly. Those problems remain even if the worker never makes the original reasoning mistake again.
The way to distinguish the two is through removal tests rather than attachment to a design. Take a component out under controlled conditions and measure the changes in performance and failure. Ask whether it enforces a boundary or encourages a behavior. A boundary still matters when the model seldom challenges it; scaffolding can go when the evidence supports removing it.
That is the low-decay target: durable records around changing context strategies, explicit policy around changing reasoning, controlled effects around replaceable workers, and independent checks around replaceable models. We should be able to improve the worker without rebuilding our account of authority, evidence, and consequence every time.
16. Build the slice. Break it.#
The argument is only worth so much until one task goes through the system. Our bug fix is a useful vertical slice: one repository, one patch, one publish gate, and one result someone else can check. It is small enough to inspect and real enough to expose the gaps a diagram can conceal.
Give the worker a disposable repository workspace, allow reading, editing, and testing, and gate publishing. Keep durable operation records, check the patch in a separate environment, and preserve the artifact alongside its test evidence. Then stop demonstrating the successful route and make the slice fail:
- Put a hostile instruction in a file the worker must read. Check that it cannot expand authority or retrieve host credentials.
- Make the runner return zero after discovering no tests. Check that the system rejects completion.
- Kill the process before an operation, after the effect, and before local confirmation. Check that restart neither duplicates the effect nor erases uncertainty.
- Start parallel work near the budget limit. Check that shared reservations hold.
- Change policy while a task waits for approval. Check which policy applies on resume and whether stale approval remains valid.
- Corrupt a receipt, remove the end of a chain, and supply an unexpected signing key. Check what the verifier detects and what it cannot.
- Delete the retrieval index and rebuild it. Check that durable evidence survives.
These failures tell us more than a clean run because they force the boundaries to do the work claimed for them. Measure correctness and blocked unauthorized effects, recovery and verification coverage, latency, cost per accepted result, and the cleanup left to a person. Record where the failure happened. A single success percentage hides the difference between an incorrect patch, an unauthorized publication, and a correct effect that the system cannot reconcile.
The benchmark also needs a threat model, workload, acceptance rule, and budget; there is no absolute score that captures every harness. The harder test is whether a compromised model can bypass a boundary the system claims to enforce. If it can, we have discovered something about the boundary, regardless of how often an ordinary run succeeds.
Event records, graph views, delegated authority, compensation, and outcome verification all sound sensible when named. Each still needs a testable contract. The diagram proposes how the system should work; reproducible failure tests are how those components earn their place in it.
17. The finish line is evidence#
We began with a patch that looked right, a command that exited cleanly, and a conclusion that sounded finished. Following that one bug fix outward revealed why none of those things was enough alone. The claim needed test evidence, publication needed authority, the lost response needed reconciliation, and repair needed a narrower account of what could safely change.
A better model can improve the proposals. The harness has to improve the connection between those proposals and reality: preserve the authority behind an action, make the execution boundary inspectable, resolve the outcome or retain the uncertainty, and recover without repeating an irreversible effect. Verification needs conditions the worker cannot quietly rewrite, and the final record needs enough evidence for another person to follow what was done.
When those checks are incomplete, I do not want the system to manufacture a more confident conclusion. I want it to name the missing evidence, unresolved outcome, or blocked action clearly enough that the person receiving the work does not have to discover the gap for themselves. That may make the answer less impressive, but it makes the work much more usable.
"Done" should come after the system has earned the claim. It should not be the belief that decides which evidence the system bothers to collect.