agent-coherence. Contact us

Your Agent Keeps Two Ledgers. You Read the Wrong One.

Two parallel records of the same agent run, one verified by infrastructure and one narrated by the model, diverging at the point where an agent acted on a version that had moved

The way to keep a fleet legible is for an agent to report on its own progress. Believing that report is the reason you end up spending an afternoon on the wrong debugging session.

The shape of a run that appeared satisfactory is as follows: a planner sends six workers to look at a series of account records. Worker three reads its record, carries out good work, and then states that the task is complete. The planner then takes this result and sets off four further tasks from it. All four of these tasks are completed. The board shows green, six out of six, and then four out of four, before the run closes without any problems.

Worker three had read version 7, and a peer had committed version 8 while it was running. All the subsequent tasks are now based on a conclusion derived from a record which no longer existed when worker three’s change landed. No errors occurred anywhere because, from the point of view of each worker’s own execution, nothing went wrong.

The main point is that your system maintained two records of that run and you ended up reading the incorrect one.

What each ledger actually knows

The first record is the one that your infrastructure can check against. It shows which version each agent read, whether that view was still up to date when the agent wrote, what was committed, what was refused, and what the store currently holds. This kind of record is mechanical, verifiable, and rather dull.

The second one is the version of events that the model tells. The task is finished. The plan has been updated. I’ve carried out the change. It sounds fluent, it appears in the format that your dashboard requires, and it is the output of the same procedure whose correctness is being questioned.

Verified ledger Narrated ledger
Produced by the storage and coordination layer the model
Says version 7 was read, version 8 was current, the write was refused “task complete”
Wrong when rarely, and loudly quietly, and only when it matters
Read first by almost nobody the dashboard, the operator, and the orchestrator

They are always there. Most of the time they coincide, and it is precisely this that makes the disagreement so costly: the tendency to place faith in the second one is strengthened each day that it turns out to be correct.

Why it amplifies instead of just failing

The framing that stood out to me was one put forward in a public thread relating to an earlier post: it is not just the case that people refer to the narrated ledger when they are debugging, but that orchestrators also branch on it.

Things such as fan-out and fan-in, decisions regarding retries, downstream dispatch, and conditional routing all depend in most frameworks on some variant of the phrase “the agent reported complete”, which is merely a token sequence and no more. It is not compared with what the storage layer actually accepted.

Drift doesn’t just stay in one place and wait to be discovered. A single unverified claim seeds all the subsequent branches downstream of it, and each of these branches generates work that is internally consistent with a premise that has not been checked. Four workers carry out good work on a bad input. The retry logic treats all of this as progress since each individual step indicates success. The topology takes one unverified string and turns it into a number of confidently derived outputs.

This is also the reason why it is never caught as it occurs. At no stage do things appear to be broken, since at every stage everything seems to be working, because everything actually is working, even though the wrong assumption is being made.

What closes it, and what does not

Not a superior report format. The narrated ledger is not misleading in its intentions. It simply cannot see beyond the scope of the world it describes.

Another model checking the first one does not close it either. Since a reviewer looks at the same view as the worker did, their agreement serves as evidence of the reasoning and provides absolutely no information about the input. This is the verification blind spot, and it is the reason why having a series of judges increases confidence without adding any new coverage in this respect.

What closes it is making the verified ledger authoritative in any situation where a decision depends on it. Specifically, in the case of shared state, keep a record of which version each agent has read, reject any write that is based on a view that has since changed, and return a typed conflict rather than carrying out a silent overwrite. The orchestrator then has a point to branch on that the model did not create. The mechanisms themselves and the checks that they carry out are explained in their own write-up.

Two conclusions from the same discipline are worth taking up even if you don’t make use of any of the points mentioned. You should bind an action to the precise inputs for which it was approved, in order that a retry won’t silently carry out different bytes based on an earlier decision. And if a call times out and the result is truly unknown, you should record unknown and then reconcile before attempting the operation again, rather than allowing a plausible explanation to fill in the gap. To refuse to guess is equivalent to refusing to trust the account the model provides.

Scope, stated plainly

This is a verified ledger that records the state of shared artifacts on a single host for writers that use the coordinator, including the versions, ownership, what was committed and what was rejected. That aspect is the part which can currently be checked and enforced.

This is not a record of what an agent carried out in the world. Whether or not an email was actually sent, a webhook actually fired, or a payment actually moved is something else again and should be the concern of your outbox, your idempotency keys, and your connectors. Guarantees apply only to the artifacts that a fleet shares, not to the actions of the fleet, and a system that blurs the two is one that is selling a boundary it doesn’t possess. When something beyond that boundary is truly irrecoverable, it is marked rather than simply being ignored: a workspace checkpoint evaluates each member it captures, and those that can never be restored are designated as forward-only rather than being quietly added to the restore count. Naming the members you cannot undo is the same discipline as naming the ledger you cannot trust. The full classification of how a write disappears shows which failures are the runtime’s to prevent and which are the application’s responsibility.

The ability to coordinate between different hosts has not shipped, and neither has any feature that checks model reasoning. That remains the verifier’s responsibility and is a genuine task.

Why it matters

Every layer that teams currently invest in reads what is recorded in the narrated ledger. Progress dashboards display it, orchestrators make decisions on the basis of it, retry policies take it into account, and post-incident reviews begin with it. That investment is not in vain, but none of it addresses the question of whether the claims were true, and the failure it overlooks is the one which, instead of stopping, causes things to compound.

The shift is small and it is mostly a habit. When something goes wrong in a fleet, the first question is usually what did the agent do. Ask instead: what does the infrastructure say happened, and does it match what the agent said happened? If you cannot answer the first half, that is the gap, and no amount of better prompting or better judging will close it.

A specific next step is to select one instance in your orchestrator where a downstream action is triggered because an agent has reported success, and then find out what, apart from the model’s own output, serves as confirmation of that success. If the answer is nothing, then you have located the seam, and it lies upstream of all the stuff that you would otherwise have spent the afternoon debugging.

The mechanisms, with deterministic offline reproductions, are at github.com/Cohexa-ai/agent-coherence.

Frequently asked

Why does my agent report success when the work was wrong?

Because from inside its own execution, it succeeded: it read a view, reasoned correctly over that view, and reported what it did. If the view had already moved, the report is honest and the result is still wrong. The report describes intent and reasoning, not verified outcome.

Can I trust an agent's own report that it finished a task?

Treat it as a claim, not evidence. In most frameworks "task complete" is a token sequence the model produced, with nothing checking it against what the storage layer actually accepted. Trust it for progress display and not for anything that gates a downstream action.

Why do multi-agent failures cascade instead of just failing once?

Because orchestrators branch on the agent's own report of success, not on what the storage layer accepted. A fan-out reads "task complete" and dispatches downstream work, so one ungrounded claim seeds several branches that each produce output internally consistent with a premise nobody checked. Retry logic then reads all of it as progress.

What can infrastructure actually verify about an agent's work?

For shared state: which version each agent read, whether that view was still current when it wrote, what committed, and what was refused. On a single host, for writers that go through the coordinator, that is checkable and enforceable. Whether an email really sent or a webhook really fired is a different ledger, owned by your outbox and your connectors.