잡동사니
Why LLM Coding Agent Harnesses Should Be Designed for the Model 본문
Hello, I’m yeTi.
While building sqlgen-ai, I connected Hermes Agent, Codex, GitLab, and Discord into a single development workflow.
At first, I thought the most important challenge was getting AI to write code. Once I began operating the system, however, I found that more problems appeared after code generation than during it.
An agent worked from an outdated branch.
It decided that an implementation was complete without running the tests.
After validation failed, it repeated the same action.
Sometimes, multiple agents picked up the same task.
To address these problems, I added state, validation, recovery paths, and stop conditions to the harness. In previous posts, I described how this structure developed into an AI engineering organization and then into multiple operational loops.
Recently, while working with local models such as Qwen3.6-35B-A3B, I started asking a different question.
Should weaker and stronger models be given the same harness?
Codex was likely designed around the working patterns of GPT models. Claude Code was likely refined around the way Claude models use context and tools. If that is true, how far can a general-purpose agent harness really go?
At first, this seemed like a matter of design preference. Then I found a paper in Hugging Face Daily Papers titled “An Empirical Study of Harness Design for Coding Agents.” The researchers examined this question across 176 harness configurations.
The paper changed the way I think about harness components.
A good harness is not necessarily the harness with the most features.
It is the harness that compensates for the points where a model actually fails.
The same harness component can serve different roles
A language model alone is not a coding agent.
It needs a software layer that keeps the execution loop running, gives the model access to files and terminals, and manages the growing history of actions and observations. That layer is the agent harness.
The researchers did not compare complete harness products as indivisible systems. They kept the same ReAct execution loop and varied three components:
- Planning: Should the model create and maintain an explicit plan?
- Action space: Should it receive dedicated tools for reading, searching, and editing files, or only Bash?
- Context management: How should old tool results and execution history be reduced?
The study evaluated Nemotron-3 30B, 120B, and 550B, together with Mistral Medium 3.5 128B. The models were tested on SWE-Bench Verified, which uses real GitHub issues, and Terminal-Bench 2.1, which focuses on end-to-end command-line tasks. In total, the researchers compared 176 settings.
The main results can be summarized as follows.
| Harness component | Weaker model or tighter resource budget | Stronger model or larger resource budget |
|---|---|---|
| Context management | Prevents the run from ending because the context window is full | Becomes less valuable as the available context grows |
| Planning | Keeps the model from stopping too early and can improve success | Mainly reduces repeated verification and cost |
| Predefined tools | Breaks actions into smaller steps for models with weak Bash control | Can add tool-selection and interaction overhead for Bash-capable models |
The same component did not produce the same effect across models.
To me, that is the most important result in the paper.
Context management extended execution more than it changed reasoning
As a coding agent works, its context fills with file contents, search results, test logs, and error messages.
When the context window is exhausted, the run may stop not because the model is incapable of solving the problem, but because it no longer has enough room to continue.
The paper compared five context-management strategies across context windows ranging from 32K to 128K:
- Keep everything without additional management.
- Elide old tool results.
- Store elided results externally and allow the model to recall them.
- Summarize earlier execution history with an LLM.
- Elide old tool results first, then summarize only when necessary.
At 32K, agents without context management often stopped after roughly 20 to 30 turns on SWE-Bench. Managed agents continued for approximately 50 to 180 turns, depending on the model and policy.
The trajectory analysis showed something important. Context management did not substantially change the order in which agents worked.
They still localized the problem, modified the code, and verified the result in broadly similar ways. What changed was whether they had enough execution time to complete that process.
I interpret this result as follows.
Context management does not necessarily make a model more intelligent. It prevents the model from abandoning work it may already be capable of completing simply because its context is full.
At 128K, the differences between context-management strategies became much smaller. Applying an elaborate compression system to a model that already has sufficient context may add operating cost before it adds accuracy.
Recoverable memory was not automatically useful
It seems safer to preserve every elided tool result so that the model can recover it later.
The paper tested this approach by storing elided observations externally and exposing a recall_event tool. However, 36 of the 64 configurations with recall available never used it. The median number of recall calls was zero.
Compared with elision alone, recoverable elision performed better in 15 of 32 matched conditions, worse in 14, and tied in three. Its mean success-rate difference was -0.36 percentage points.
Preserving every observation appears more complete in principle. But if the model rarely retrieves that information, the external store and recall tool become machinery maintained for a possibility that seldom contributes to task completion.
The most efficient policy first removed stale tool results using rules and invoked LLM summarization only when the remaining context was still too large.
This suggests a practical design principle.
The need for a harness feature should not be determined by the designer’s anxiety about what might be lost. It should be evaluated by whether the feature is actually used in trajectories and whether that use contributes to successful completion.
Planning was a scaffold for weaker models and a stopping aid for stronger ones
Planning also served different purposes depending on model capability.
For Nemotron-3 30B, the smallest model in the study, planning increased success by 11.6 percentage points on SWE-Bench and 4.5 points on Terminal-Bench. The additional turns and tool calls also increased cost.
On SWE-Bench, the median trajectory lasted only five turns without planning. With planning, it increased to 40 turns. The percentage of runs that ended without editing any code fell from 68.6% to 27.8%.
Planning did not make the weaker model faster.
It acted as a scaffold that kept the model from losing the task and giving up before it made an initial edit.
The stronger models showed a different pattern.
For Nemotron-3 550B and Mistral Medium 3.5, planning reduced SWE-Bench cost by approximately 30% and 32%, respectively. Success decreased by only 2.0 and 0.4 percentage points.
These models could edit the code without an explicit plan. The main contribution of planning was reducing repeated verification after the edit was already complete.
The same planning mechanism helped the weaker model continue working, while it helped the stronger models stop at a more appropriate point.
The effect of a harness component is not fixed inside the component itself. Its role depends on where the underlying model tends to fail.
More tools do not always make a coding agent more capable
As developers, we are naturally inclined to give agents clear, specialized tools.
Separate tools for reading files, searching text, editing code, and running tests appear safer and easier to control than a general shell interface.
This was true for Nemotron-3 30B, which had weak Bash control. Compared with Bash alone, the predefined tool set improved success by 15.0 percentage points on SWE-Bench and 10.1 points on Terminal-Bench.
The result reversed for Nemotron-3 550B. With Bash alone, success increased by 3.6 points on SWE-Bench and 5.6 points on Terminal-Bench, while cost fell by 53% and 30%.
Predefined tools divide a complex operation into several smaller and safer actions. This granularity helps a weaker model.
A Bash-capable model, however, can combine multiple operations into a single command or script. For that model, a large tool registry can add overhead: the model must choose an interface and distribute the work across more interactions.
This does not mean every strong model should receive only Bash.
Mistral Medium 3.5 achieved 6.7 percentage points higher success with Bash alone on Terminal-Bench. On SWE-Bench, where it had to locate and modify files in a repository, predefined tools improved success by 23.2 points.
Task type mattered as much as model capability.
Bash was a natural interface for terminal-centric work. Search, read, and edit tools remained valuable when the task required locating the correct files in a large repository.
A harness should be designed around failure points, not model size
If I apply these findings to the harness I operate, I would not select components based on parameter count alone.
I would first inspect the trajectory and identify where the model stops making progress.
| Observed failure | Harness change to test first | Metric to inspect |
|---|---|---|
| The run ends before any code is changed | Add an explicit plan and narrower tools | Turns to first edit; percentage of runs without an edit |
| The run ends because the context window is full | Elide old tool output, then summarize selectively | Context-overflow rate; peak context usage |
| Verification continues long after the edit is complete | Add completion criteria and stop conditions to the plan | Post-edit verification turns; cost per task |
| Tool selection and incremental edits require too many calls | Test a Bash-centered action space | Tool-call count; repeated patches; success rate |
| The Bash-only agent fails to locate the correct file | Restore dedicated search, read, and edit tools | First-edit rate; correct-file localization rate |
This process does not begin with a list of features.
It begins by observing a failure, changing one component that may address it, and measuring whether success and cost change on the same class of tasks.
This resembles how the validation and recovery loops emerged in sqlgen-ai.
I did not begin by designing a complex harness. I found recurring failures and added the structures required to handle them.
The paper adds one important condition to that process.
The same failures do not appear in every model. A scaffold that one model requires should not automatically become the default for another model.
A general-purpose harness should be configurable, not fixed
Before reading this paper, I was asking whether a general-purpose agent harness was possible.
I now think the question needs to be framed differently.
Models can share a common execution substrate. State management, permissions, execution records, and validation can remain common operational infrastructure.
But if every model receives the same planning mechanism, the same tool registry, and the same context policy, the harness cannot account for the strengths and weaknesses of each model.
Weaker models may need explicit plans and tools that divide work into smaller operations. Models with limited context may need aggressive elision and summarization. Models that already handle Bash and long contexts well may experience some of these mechanisms as cost rather than support.
A general-purpose harness, therefore, should not force every model through the same workflow.
It should allow planning, action space, and context policy to change with the model, task type, and context budget.
I now define a good agent harness this way:
A good harness does not provide the model with the largest number of features. It places support precisely where the model cannot reliably proceed on its own.
My next step is to test this idea directly.
I plan to give the same coding task to the local model I currently use and to a frontier model, then compare two harness configurations:
- A full harness with planning, predefined tools, and recoverable context.
- A simpler harness with Bash and rule-based context elision.
I will compare task success, turns to first edit, tool-call count, input tokens, tool-result tokens, and total cost.
The goal is not to identify one universally better harness.
The goal is to find where each model actually needs help from the harness.
As models become more capable, the harness does not necessarily need to become more complex.
Designing a harness for a stronger model may begin not by adding another component, but by removing the components the model no longer needs.