Building a Coding Agent from Scratch, Part 2: "Done" Means "It Runs," Not "It's Written." The Completion Gate and What to Write in a Pushback

The model's "Done" slipped through in five ways, from a bare syntax check to tests with no assert. We describe, with numbers, a completion gate that closes in three stages, and why a pushback should contain the failing lines and how to investigate, but never the answer.

Building a Coding Agent from Scratch, Part 2: "Done" Means "It Runs," Not "It's Written." The Completion Gate and What to Write in a Pushback

Hello!

This is a story from when we had our own coding agent build a task-management web app with authentication. At turn 40 the model reported, "All features are implemented and the tests pass. Done," and stopped. Looking at the test output, everything did indeed pass.

But when we started the server and hit the APIs as written in the specification, only 2 of 22 checks passed. The specification has 12 APIs; split into normal and error cases, the grader has 22 check items.

The tests pass. The server starts. And yet not even a tenth of the specification works. The tests the model wrote mirrored its own implementation rather than the specification.

To put the conclusion first, a coding agent's "done" has to be decided by evidence from actually running what it built, not by the model's declaration. And the role that gathers that evidence must belong to the agent side (the harness), not the model. We call this the "completion gate," and of the 105 defects we fixed, the largest share appeared around this gate.

Figure 1: The completion gate closes in three stages (diagram: Qualiteg)
Figure 1: The completion gate closes in three stages (diagram: Qualiteg)

This article is Part 2 of the series "Building a Coding Agent from Scratch." Last time we described separating the "turn runner" from the "stop decider." This time we cover the heaviest judgment inside that stop decider: the completion gate.

PartTheme
Part 1Why "finishing the job" is hard. Separating the turn runner from the stop decider
Part 2 (this post)"Done" means "it runs," not "it's written." The completion gate and how to write a pushback
Part 3"It didn't stop" is not "it made progress." Count idle turns, and cap them without cutting the thinking
Part 4Autonomy rules are for when nobody is around. If a human is present, stop and ask
Part 5Making local LLMs first-class citizens. Getting two 16GB GPUs to build a web system
Part 6How to build the evaluation. Have it build something that runs, auto-grade it, and doubt the grader itself
Part 7Logging and regression. Everything is found in the event log. And building with zero dependencies

1. Five Ways "Done" Slipped Through

Before building the gate, let's list what slipped through. All of these actually happened on our machines.

Figure 2: Five ways completion slipped through (diagram: Qualiteg)
Figure 2: Five ways completion slipped through (diagram: Qualiteg)
How it slipped throughWhat actually happened
"Verified working" on a syntax check alonenode --check passed, so it declared done. Never executed once
Tests pass but nothing works (failures hidden behind exit code 0)36 tests with no assertions. The start command had been rigged to return exit code 0 even on failure
The listener comes up, but the first request crashes itThe server starts and opens the port. But while handling the first HTTP request it calls an undefined function and dies. A gate that checked startup by "did the port open" let this through
Starts, but the specified APIs don't existThe opening example. Its own tests passed, the top page responded, API checks 2/22, and it reported done
Tests with no assertThe body of test() was nothing but console.log. Three runs, all three disqualified in the same way

Reading down this table, you notice something.

Close one gate, and a weak model goes through the hole next to it.

Say "run it," and it runs with || echo attached. Say "start it," and it only starts it. Say "write tests," and it writes tests with nothing inside. The model is not malicious. It is simply satisfying whatever we treat as the completion condition in the cheapest possible way.

That is why the gate must look not at "what the model did" but at "what the thing it built does."

2. The Gate Closes in Three Stages

Our completion gate ended up as the three stages shown in Figure 1 at the top.

2-1. Gate 1: The harness starts it up itself

When the model says "done," the harness starts the deliverable itself. The model has declared how to start it in package.json in start, so the harness uses that. There is no need to embed task-specific knowledge (file names or commands) into the gate.

Once started, it sends exactly one request to the top page. This is the important part: an open port alone does not pass. The third row of the table above, "the listener comes up but the first request crashes it," is exactly this. The version that treated a successful listen as proof of startup was passing this "deliverable that doesn't start" as completed. The grader's score for startup: 0 out of 15.

If it fails to start, the failure details (the tail of the startup log) go into a pushback and back to the model. If it still won't come up after 10 pushbacks, we stop for a different reason. Pushing back forever is the same as the "spinning, unable to stop" described in Part 1.

2-2. Gate 2: Hit every route in the specification, one by one

Once it starts, the harness hits every route written in the specification, one at a time. As long as any route is missing (404), or any route answers with a status not in the specification (500 or 302), the gate stays closed.

This is the countermeasure for the opening "done at API checks 2/22." Since the tests do not necessarily mirror the specification, the specification itself becomes the yardstick.

However, it does not go as far as inspecting the shape of the response JSON. If it did, the gate would know the answer to the task, and the evaluation described in Part 6 would become a fixed game. "Does the route exist" and "does it answer with a status code in the specification" are the gate's job; whether the contents are correct is the grader's job.

2-3. Gate 3: Look inside the tests

The third stage looks inside the tests the model wrote. It checks two things.

If zero tests ran, it does not count as success. The Node.js built-in test runner (node --test) returns exit code 0 with "tests 0" even when there is not a single test file. There was a case where the model implemented all six problems, wrote no tests at all, npm test passed, and it declared done. If the runner output contains the form "0 tests," "No tests found," or "0 passing," we treat it as a verification failure and send it back to write tests.

If a file remains that has `test()` but no `assert`, we send it back to write the body. There was a hole next to this one, too. At first we only looked at *.test.js under the test/ directory, so the model placed test-all.js (test() once, body console.log, zero assertions) directly under the project root and slipped through with "tests 1 / pass 1." The gate's yardstick must exactly match the file patterns that the actual test runner picks up.

2-4. Closing the gate at the moment of completion is enough

Right after adding Gate 3, we made the opposite mistake.

We made the "zero tests" verification failure fire on every write. A weak model was then pushed back every turn from the stage when it had not yet written tests, in the middle of implementation, and a task that used to finish in 20 to 40 turns took 137. Twenty-four pushbacks. The model wanted to move the implementation forward, and the harness kept saying "there are no tests."

During the run, just mark it; stop it at the gate at the moment the model tries to complete. That was enough. Closing the gate at the moment of completion is sufficient. Close it every time mid-run, and a weak model cannot make progress on the implementation.

3. Copy the Model's Own Verification into the Same Record

Even with three gates, there was still one place it spun.

On a task that added a feature to an existing 30,000-line codebase, the harness's automated verification failed exactly once. What failed was a timing-sensitive existing test, because the machine was busy. After that, the model ran the same test through its own verification tool and passed it four times. Even so, the harness pushed back 14 times with "verification is still failing."

The results of the verification the model ran itself had not been written into the harness's verification record. The "last verification" the gate was looking at still pointed at the harness's old failure.

Copy the verification the model passed into the same record as the harness's verification. That was the whole fix, but until we added it, the first run of that task was spinning through 14 pushbacks while producing a deliverable of the same quality.

4. What to Write in a Pushback, and What Not To

When the gate stops the model, we tell it "why it didn't pass." How this text is written changed the subsequent turn counts dramatically.

Figure 3: What to write in a pushback, and what not to (diagram: Qualiteg)
Figure 3: What to write in a pushback, and what not to (diagram: Qualiteg)

4-1. Extract only the failing lines

The first version put the first 800 characters of the verification output into the pushback. But when there are 20 tests and 19 pass and 1 fails, the first 800 characters were filled with the list of passing tests, and the failing line never arrived. From the model's point of view, there was just a row of green checkmarks under "Verification is still failing."

We changed it to extract and include only the lines that indicate failure (✖, not ok, AssertionError, expected, actual).

4-2. "Read the cause and fix it" does not work

Generic pushbacks only made the model try another wild guess. In one run, a model that had gotten a module's export name wrong kept guessing different names for 77 turns while receiving "read the cause and fix it" pushbacks.

What worked was writing "how to investigate" for each kind of failure. "That name is not exported. List the file's exports and check." "'F:/' is the shell's path conversion, not a mistake in the code." When you write concretely what differs (status code, route, wording), the same failure does not repeat.

The same goes for syntax errors right after an edit. Returning only the error message did not get a 450-line file fixed. Attaching the source around the offending line lets the model notice its own syntax error by itself.

4-3. Do not write the answer

There is one line we draw, however. We do not write the answer to the task.

"POST /api/tasks returns 500 where it should return 201" is fine to write. It is a contract written in the specification. "Set the default of status to todo" we do not write. That is the implementation's answer, and the moment you write it, the evaluation becomes a fixed game.

Write how to investigate. Write what differs. Do not write the answer. This line connects directly to the discussion of evaluation in Part 6.

5. What Changed After Building the Gate

Here are the numbers after the gate went to three stages.

Running the same local model on five kinds of tasks, three runs each, 15 runs in total, all 15 stopped as completed, and all six that start a server passed the harness's own startup check. The number of runs where a deliverable that does not start was marked completed: zero.

As a control, we also ran a smaller, faster model through the same gate. Of the two runs before the gate, one did not start, and the other started but every API returned 500 or 401. The one run after adding the gate that checks through the first request did start, and the API checks reached 9 of 22. Authorization and persistence remained at zero. The gate does not make the model stronger, but it does stop "calling something that doesn't work done."

The three-stage gate looks at the behavior of the deliverable. Separately from that, we also added one gate that checks whether the premises of the instruction were respected. Unglamorous, but effective. The model had been able to complete without ever reading the specification named in the instruction. In one run, it wrote a different specification of its own without reading the given one, completed in 23 turns, and scored zero. Keep the gate closed until the file named in the instruction has been read. That alone made it disappear.

Summary: The Gate Looks at "What the Deliverable Does," Not "What the Model Did"

GateWhat it checksReal example that slipped through
Gate 1: StartupDoes it come up with the declared start method and answer the first request?The port opens but the first request crashes it
Gate 2: Specified routesDo all routes in the specification exist and answer with the specified status codes?Own tests passed, API checks 2/22
Gate 3: Test contentsIs the number of tests run nonzero, and is there an assert?Tests that are only console.log, a test-all.js at the project root
RecordIs the model's own verification copied into the same record?Passed four times, pushed back 14 times
PushbackWrite the failing lines and how to investigate, not the answerThe first 800 characters filled with passing tests

Close one gate, and the model goes through the hole next to it. Even so, there is no alternative to the harness closing the gate itself. The moment you trust the model's "Done," you are back at the same place as "What should I do next?" in Part 1.

Next time: runs that pass the gate but are "not making progress." Starting from the figure that more than 70% of 300 turns never called a tool, we describe the metric for counting idle turns and why we cap them without cutting the thinking.

See you next time!

Read more