This document is a working method for the moment a coding agent says it is done. It takes an experiment that Factory Research, the research team of a company that sells coding agents, published on August 27, 2026, in which the same agent reproduced 36% of a large program on its own and 90% when a separate role wrote the standard of completion first, and turns it into eight checks a team can run on its own line before trusting an agent's word.
1. The agent did not run out of budget. It decided it was finished.
Factory Research asked its agent to rebuild gdal, the command line tool of a geospatial library in development since 1998, with about 600,000 lines of C and C++ reachable through the interface used in the task. The agent could run the original program without limit but could not read its source, its tests or the internet.
Working alone, the agent wrote 17,000 lines of C++ and reproduced 36% of the program's behavior. The report's own words: it did not run out of time or budget, it stopped because, by its own assessment, it was done. The common paths worked. Most of the program was missing.
The team then split the work into roles. Before any implementation, one role built an executable standard of completion, its own account of what the rebuild had to do and what evidence would prove it, and the implementation was held to that standard. The same model, same reasoning level, reached 115,000 lines and 90% behavioral parity.
Across the 24 hardest tasks the team selected from a 200-task benchmark, the same system took a 7-Zip rebuild from 54% to 95% and a DuckDB rebuild from 34% to 80%. With the strongest of the three models tested, the median score across the 24 tasks went from 56.7 to 89.3. Every number is one run, not an average, and the report says so.
2. Why an agent stops early
The report's explanation is about scope, not skill. A coding agent validates its own work as it goes: implement a piece, write or run a few checks, inspect the output, decide whether to continue. For a small change that works, because the task, the code and the evidence fit in one view.
A large task has to be broken into features, subsystems and rounds of work. As the agent reaches each piece, it also decides what evidence counts and whether it is enough. Those checks inherit the scope of the work that produced them: they can establish everything the agent thought to build and nothing it never represented. So the agent makes steady, locally correct progress and stops with much of the outcome absent, because it never held a complete account of what remained.
The report's conclusion in one line: the single agent did not lack skill, it lacked a standard of completion.
3. What the system did differently
Three roles, and a wall between two of them.
A validator surveys the reference program before implementation starts, maps where its behavior lives, and builds a weighted body of test cases with the rules for judging each result. In the gdal run it wrote hundreds of cases and allowed exactly two relaxations of byte-for-byte comparison, each with recorded evidence that the original program could not produce stable output there.
An implementer builds the candidate program and runs its own ordinary checks, as before.
An orchestrator decides what gets measured, reviews the validator's findings, rejects noise, turns the real problems into a directive at the level of missing features or subsystems, and decides when to ship.
The wall: the implementer never authors, runs or sees the validator's cases or raw results. Once a sample of cases becomes visible, the report says, it becomes the target, and passing it establishes those cases, not the behavior they were meant to represent. The validator can expand its instrument as it learns but cannot weaken it to fit what the candidate happens to contain.
| Agent alone | System with a validator | |
|---|---|---|
| Who defines done | The agent, inside the same context as the code | A separate role, before implementation |
| What the checks cover | What the agent thought to build | An inventory of the outcome, weighted |
| Who sees the checks | The agent | The validator only; findings cross the wall, cases do not |
| Who decides to stop | The agent | The orchestrator, when the instrument stops finding differences |
| gdal result | 17,000 lines, 36% parity | 115,000 lines, 90% parity |
The system runs cost far more: for gdal, 14 times the credits and 13 times the wall clock. The report's point is that budget did not separate the two conditions, because every single-agent campaign ended when the agent chose to end it. More compute does not help an agent that will not spend it.
4. What transfers to a team's own line
The benchmark is a limit case, a black-box program the model has to map and measure at once. The report says what generalizes to real software work: an external, executable standard of completion, derived from the outcome before implementation narrows attention, and kept current until the work meets it. For a product task, that standard comes from user-approved flows and designs; for a migration, from the system being replaced.
The report also says why teams rarely do this by hand. Requirements traceability, independent verification and conformance suites exist in safety-critical work and standards bodies, where the cost spreads across many implementations. A product team bears the cost again for each application, so it validates incrementally and relies on review, feedback and the memory of the people involved. Agents change both sides: they produce work faster than people can inspect it, and they can also build and rerun the standard.
Our reading, not the report's: for a team of 10 to 200 people the validator does not have to be an agent. It has to be someone, or something, that did not write the code and wrote down what done means before the code existed.
5. Eight checks before you accept an agent's "done" on a large task
Score yes or no on the next large task the line takes. The six-or-more threshold is ours, not the report's.
| # | Check | What passes |
|---|---|---|
| 1 | Is there a written definition of done that exists before the first line of code? | Derived from the outcome (flows, designs, the system being replaced), not from the plan of work |
| 2 | Was it written by someone other than the implementer, person or agent? | A different role, with its own reading of the requirement |
| 3 | Is it executable? | Cases that run and compare, with the rule for judging each result |
| 4 | Does it cover the outcome, not the task list? | An inventory of behaviors, weighted by what matters, including what nobody assigned |
| 5 | Is it kept away from the implementer? | The agent sees findings at the level of missing features, not the cases |
| 6 | Can it grow but not shrink? | Cases are added as the reference is understood; none are removed to fit the candidate |
| 7 | Does someone other than the implementer decide when to stop? | A named person or role reads the findings and ships |
| 8 | Is the final score taken outside the loop? | An independent measure, not the implementer's own report |
A line that answers no to checks 1 and 2 has an agent grading its own work. That is the condition that produced 36%.
6. How we run this on our own line
On our own line, every specification that goes to an agent carries acceptance criteria written by the engineer who owns the outcome, before the agent starts, and turned into tests the agent does not edit. An independent verifier, which did not write the code, reads every test the agent produces looking for tests made to pass by checking less. The engineer who approves the merge is not the one who wrote the criteria.
When a task is large, we write the inventory of behaviors first and hold the agent to it in rounds, at the level of what is missing, without handing over the cases. The eight checks are the ones we run before accepting an agent's word on anything longer than a day's work; our own answers are in the file that accompanies this document.
Sources
- Factory Research (Theo Luan), What it Takes for Coding Agents to Complete Large Software Tasks, August 27, 2026. https://factory.com/news/what-it-takes-for-coding-agents-to-complete-large-software-tasks
Analysis and conclusions by Viaro Networks, based on the published reports listed above.
