Tune the agent before you judge the model

An agent's harness and prompts are usually built and tested with one model. As a result, when you swap the model without changing the setup around it, your agent might perform worse. That might not be the model's fault: it might just need its own fine-tuned guidance.
So, before you compare models on score, cost, or speed, you should tune the agent for each of them, optimizing the harness to get the best performance possible out of each model. Beaker does that for you, autonomously.
We tested Beaker on AppWorld, a coding agent benchmark, using a variety of models. Untuned, GLM 5.3 Flash completed 29% of the test scenarios. Most people would see this result and conclude that the model is bad at writing code. However, after Beaker tuned the agent for GLM, it scored 82%, at about the same cost ($0.029 to $0.033 per scenario) and almost twice as fast (138 to 76 seconds). GLM 5.3 Flash isn't bad at coding, it just needs a little bit of guidance to do its best work. We saw similarly surprising results across the eight models we tested: Beaker's optimizations made initially low-performing models competitive, bringing the choice down to trade-offs between score, cost and speed.
This post covers what was holding the models back, why the two strongest models needed something completely different, and how Beaker helps you pick the right model for your application by tuning the agent for each one. You can browse the full run, including every change Beaker autonomously found and tested for each model.
The setup
We used AppWorld, a benchmark of everyday tasks across apps like Venmo, Spotify, a phone's contacts and a file system. Our agent, from the cookbook, solves them by writing Python against each app's API. Each scenario has three tasks, checked against the final app state and the agent's answer.
We use the strictest score: a scenario counts only if all three of its tasks pass. Beaker searched for changes on 30 training scenarios, and every number here is on 19 held-out scenarios it never searched on, averaged over two runs. Our reference point is the shared baseline, GPT-6 Astra at high reasoning, which completes 100% at $0.62 per scenario.
For each model, Beaker ran the agent on the training scenarios and read the traces of the ones that failed. From those failures it formulated hypotheses about what was going wrong, turned each into a concrete change to the prompt or the code, and tested it. It kept a change only if the score went up.
Before tuning, the models looked 61 points apart
Untuned, the eight models ranged from 29% (GLM 5.3 Flash) to 90% (GPT-6 Luna and Gemini 3.8 Flash). After Beaker tuned the agent for each model, the range was 82% to 100%. The models that started worst gained the most: GLM went from 29% to 82%, Gemini 3.5 Flash Lite from 32% to 84%, and Claude Haiku 5.5 from 53% to 95%. Haiku started sixth and finished joint second.

Each line goes from a model's score before Beaker optimization (grey outline) to after (filled purple). Vertical axis: held-out scenarios fully completed. Horizontal axis: average cost per scenario, on a log scale.
Most lines go straight up: the models improved a lot without costing more. Luna and Gemini 3.8 Flash, the two best starters, barely moved, gaining 3 and 11 points respectively, partly because they had less room to improve. The next section looks at what the other models were getting wrong.
Most failures came from one unwritten rule
Some AppWorld tasks only ask for an action, like "text my siblings to get on Venmo". The checker expects the agent to do it and say nothing back. GLM sent the texts correctly, then reported "sent 'please get on venmo.' to eric bailey, …", and the task failed. Beaker spotted this pattern by reading the traces of the failed scenarios.
The prompt does mention this, in a single line: "If no answer is required, … omit the answer argument." Six of the eight models kept missing it. For those six models, 39 of the 55 failed scenarios failed solely because of that extra answer. For the other checks, the agent did everything correctly.
- Gemini 3.8 Flash: 2 of 19 failed, 0 of them only because of an extra answer
- GPT-6 Luna: 2 of 19 failed, 0 of them only because of an extra answer
- Kimi K3: 5 of 19 failed, 4 of them only because of an extra answer
- DeepSeek V4.1 Flash: 6 of 19 failed, 4 of them only because of an extra answer
- MiMo-V2.6-Flash: 9 of 19 failed, 7 of them only because of an extra answer
- Claude Haiku 5.5: 9 of 19 failed, 7 of them only because of an extra answer
- Gemini 3.5 Flash Lite: 12 of 19 failed, 8 of them only because of an extra answer
- GLM 5.3 Flash: 14 of 19 failed, 9 of them only because of an extra answer
One run per model, before tuning.
Beaker found this pattern separately in each of the six models' runs. In every one, the first change it tested was a clearer version of that rule, and it kept that change each time. Here is the version from GLM's run:
For an action without a requested answer, call apis.supervisor.complete_task() with no answer argument. Do not return a status report, summary, count of changes, or the names or IDs of affected records.So for the most part, these models weren't incapable. They just didn't know an important implicit rule of the task.
The strongest models had already guessed it
In contrast, Luna and Gemini 3.8 Flash never made that mistake in their untuned runs. For Luna, Beaker autonomously proposed the same rule three times and rejected it each time, because the score didn't move.
What they needed was smaller and more specific. Gemini 3.8 Flash got one fix, submitting numbers as plain Python numbers rather than formatted strings, which took it to 100%. Luna got four: look people up in the phone's contacts before acting on them in another app, return a single number when asked for one, quote CSV exports consistently, and don't let a filter on a collection apply to the items inside it.
That's why the agent has to be tuned for each model separately: the change that improved six of the models did not help for Luna, which needed a different fix.
Sometimes the problem is the example you gave
Not every fix was an additional rule. One of the changes Beaker kept for GLM corrected the code example in the prompt itself.
The prompt shows the agent a worked example that pages through API results with while page_index < 10:. It stops after ten pages, however much data there is. Beaker's hypothesis was that GLM copied the cap and stopped before the end of longer lists. The change it kept replaced the loop with while True: and added: "continue until an empty page… do not impose an arbitrary maximum page count."
It looks like a small edit but highlights that an agent can copy an example too closely, including details you never meant it to imitate.
Most models improved without costing more
One winning change went beyond the prompt: for DeepSeek V4.1 Flash, Beaker also edited the agent's code tool to check for side effects before running exploratory code, so it would stop saving stray files while investigating. DeepSeek is also the only model that did noticeably more work afterward: 37 model calls per scenario instead of 30, and $0.118 instead of $0.079.
For the others, cost barely moved, which is why most lines in the chart go straight up. Haiku made 24.4 model calls per scenario before and 24.6 after, at $0.033 both times. GLM and Luna even got faster: GLM from 138 to 76 seconds per scenario, Luna from 133 to 93.
So which model should you pick?
After tuning, no single model is best at everything. The right choice depends on how much you care about score, cost and speed. On score and cost together, three models stand out, because no other model beats them on both at once: GPT-6 Luna, MiMo-V2.6-Flash and Gemini 3.8 Flash.
If every scenario has to succeed, Gemini 3.8 Flash matches the shared baseline at 100%, for 72% of the cost. If you run the agent at high volume and can live with a few misses, Luna completes 92.1% at $0.011 per scenario, 1/58 of the shared baseline. MiMo sits in between, at 94.7% for $0.027.

Same models and markers as the cost chart. Vertical axis: held-out scenarios fully completed. Horizontal axis: average latency per scenario, in seconds. Kimi K3 and Gemini 3.5 Flash Lite are left out because this run didn't record their latency.
If your users are waiting on the result, the picture shifts again. Gemini 3.8 Flash takes 288 seconds per scenario, against 132 for the shared baseline, and MiMo takes 258. Claude Haiku 5.5 matches MiMo's 94.7% at $0.033 and finishes in 91 seconds.
Tune first, then compare
Had we compared the eight models on the untuned agent, we would have written GLM off at 29% and called Haiku mid-table. Tuned, GLM reached 82% and Haiku tied for second. The question changed from "which model is good enough?" to "which Pareto trade-off do we want?"
Run the same experiment on your own agent with Beaker, as nothing here is specific to AppWorld. Give Beaker your agent's code, labeled train and test examples and a function that scores a prediction, then pick one model or several. Beaker autonomously runs the agent, analyzes the failures, and tests changes to prompts, instructions and other parts of your harness. It keeps what scores higher on training cases, reports results on held-out cases, and shows the evidence behind each change. Once the experiment is done, it opens a pull request with the best changes it found. Follow this guide to get started in a few minutes.