Research / 2026-10-04

How to know a coding agent is done: a standard of completion written before the work took a rebuild from 36% to 90%

A coding agent that grades its own work stops when it believes it is finished. A separate role that writes the standard of completion before the work begins changed the result of the same agent on the same program, and eight checks bring that to a team's own line.

For CTOs, VPs of Engineering and founders whose teams hand large tasks to coding agents and have to decide when the work is finishedViaro Networks7 min read
How to know a coding agent is done: a standard of completion written before the work took a rebuild from 36% to 90%

This document is a working method for the moment a coding agent says it is done. It takes an experiment that Factory Research, the research team of a company that sells coding agents, published on August 27, 2026, in which the same agent reproduced 36% of a large program on its own and 90% when a separate role wrote the standard of completion first, and turns it into eight checks a team can run on its own line before trusting an agent's word.

1. The agent did not run out of budget. It decided it was finished.

Factory Research asked its agent to rebuild gdal, the command line tool of a geospatial library in development since 1998, with about 600,000 lines of C and C++ reachable through the interface used in the task. The agent could run the original program without limit but could not read its source, its tests or the internet.

Working alone, the agent wrote 17,000 lines of C++ and reproduced 36% of the program's behavior. The report's own words: it did not run out of time or budget, it stopped because, by its own assessment, it was done. The common paths worked. Most of the program was missing.

The team then split the work into roles. Before any implementation, one role built an executable standard of completion, its own account of what the rebuild had to do and what evidence would prove it, and the implementation was held to that standard. The same model, same reasoning level, reached 115,000 lines and 90% behavioral parity.

Across the 24 hardest tasks the team selected from a 200-task benchmark, the same system took a 7-Zip rebuild from 54% to 95% and a DuckDB rebuild from 34% to 80%. With the strongest of the three models tested, the median score across the 24 tasks went from 56.7 to 89.3. Every number is one run, not an average, and the report says so.

2. Why an agent stops early

The report's explanation is about scope, not skill. A coding agent validates its own work as it goes: implement a piece, write or run a few checks, inspect the output, decide whether to continue. For a small change that works, because the task, the code and the evidence fit in one view.

A large task has to be broken into features, subsystems and rounds of work. As the agent reaches each piece, it also decides what evidence counts and whether it is enough. Those checks inherit the scope of the work that produced them: they can establish everything the agent thought to build and nothing it never represented. So the agent makes steady, locally correct progress and stops with much of the outcome absent, because it never held a complete account of what remained.

The report's conclusion in one line: the single agent did not lack skill, it lacked a standard of completion.

3. What the system did differently

Three roles, and a wall between two of them.

A validator surveys the reference program before implementation starts, maps where its behavior lives, and builds a weighted body of test cases with the rules for judging each result. In the gdal run it wrote hundreds of cases and allowed exactly two relaxations of byte-for-byte comparison, each with recorded evidence that the original program could not produce stable output there.

An implementer builds the candidate program and runs its own ordinary checks, as before.

An orchestrator decides what gets measured, reviews the validator's findings, rejects noise, turns the real problems into a directive at the level of missing features or subsystems, and decides when to ship.

The wall: the implementer never authors, runs or sees the validator's cases or raw results. Once a sample of cases becomes visible, the report says, it becomes the target, and passing it establishes those cases, not the behavior they were meant to represent. The validator can expand its instrument as it learns but cannot weaken it to fit what the candidate happens to contain.

Agent aloneSystem with a validator
Who defines doneThe agent, inside the same context as the codeA separate role, before implementation
What the checks coverWhat the agent thought to buildAn inventory of the outcome, weighted
Who sees the checksThe agentThe validator only; findings cross the wall, cases do not
Who decides to stopThe agentThe orchestrator, when the instrument stops finding differences
gdal result17,000 lines, 36% parity115,000 lines, 90% parity

The system runs cost far more: for gdal, 14 times the credits and 13 times the wall clock. The report's point is that budget did not separate the two conditions, because every single-agent campaign ended when the agent chose to end it. More compute does not help an agent that will not spend it.

4. What transfers to a team's own line

The benchmark is a limit case, a black-box program the model has to map and measure at once. The report says what generalizes to real software work: an external, executable standard of completion, derived from the outcome before implementation narrows attention, and kept current until the work meets it. For a product task, that standard comes from user-approved flows and designs; for a migration, from the system being replaced.

The report also says why teams rarely do this by hand. Requirements traceability, independent verification and conformance suites exist in safety-critical work and standards bodies, where the cost spreads across many implementations. A product team bears the cost again for each application, so it validates incrementally and relies on review, feedback and the memory of the people involved. Agents change both sides: they produce work faster than people can inspect it, and they can also build and rerun the standard.

Our reading, not the report's: for a team of 10 to 200 people the validator does not have to be an agent. It has to be someone, or something, that did not write the code and wrote down what done means before the code existed.

5. Eight checks before you accept an agent's "done" on a large task

Score yes or no on the next large task the line takes. The six-or-more threshold is ours, not the report's.

#CheckWhat passes
1Is there a written definition of done that exists before the first line of code?Derived from the outcome (flows, designs, the system being replaced), not from the plan of work
2Was it written by someone other than the implementer, person or agent?A different role, with its own reading of the requirement
3Is it executable?Cases that run and compare, with the rule for judging each result
4Does it cover the outcome, not the task list?An inventory of behaviors, weighted by what matters, including what nobody assigned
5Is it kept away from the implementer?The agent sees findings at the level of missing features, not the cases
6Can it grow but not shrink?Cases are added as the reference is understood; none are removed to fit the candidate
7Does someone other than the implementer decide when to stop?A named person or role reads the findings and ships
8Is the final score taken outside the loop?An independent measure, not the implementer's own report

A line that answers no to checks 1 and 2 has an agent grading its own work. That is the condition that produced 36%.

6. How we run this on our own line

On our own line, every specification that goes to an agent carries acceptance criteria written by the engineer who owns the outcome, before the agent starts, and turned into tests the agent does not edit. An independent verifier, which did not write the code, reads every test the agent produces looking for tests made to pass by checking less. The engineer who approves the merge is not the one who wrote the criteria.

When a task is large, we write the inventory of behaviors first and hold the agent to it in rounds, at the level of what is missing, without handing over the cases. The eight checks are the ones we run before accepting an agent's word on anything longer than a day's work; our own answers are in the file that accompanies this document.

Sources

Analysis and conclusions by Viaro Networks, based on the published reports listed above.