Changing everything except the model

Salesforce AI Research and Salesforce Agentforce published a framework this week called DarwinX, which improves an AI agent by evolving the scaffolding around it — the prompts, tools, skills and workflows that make up the harness — while leaving the model weights untouched.

On WebArena-Infinity, a browser task benchmark, Salesforce reports the agent going from completing 43.5% of tasks to 93%. On Terminal-Bench 2.1 it reports 75.5% to 83.2%, and on SWE-bench Verified 80.8% to 84.2%. The primary base model was GPT-5.5, with Claude Opus 4.8 also tested.

These are the developer’s own figures for its own system, and no independent evaluation has yet reproduced them. The browser number in particular is large enough that it invites scrutiny of how much headroom the starting harness left on the table.

The two problems it claims to solve

Harness tuning is not new. What DarwinX targets are two failure modes that make it hard to do by hand.

A software engineer at a monitor displaying source code
Salesforce reports gains on browser, terminal and software engineering benchmarks. Illustrative image. ThisIsEngineering · pexels · Pexels License

The first is path dependence: an early choice in how the scaffolding is written narrows what can be explored later, and the process gets stuck in a local optimum nobody chose. The second is cross-task interference, where a change that improves one kind of task quietly degrades another.

DarwinX handles both by maintaining several harness variants at once and letting improvements compete, rather than repeatedly rewriting one version. A change only advances if it improves capability without regressing on the other tasks in the suite.

Why this matters to people who do not train models

Most teams building on AI do not have a fine-tuning pipeline. They call a hosted model and write everything around it. For them, the harness is the only thing they can change, and it is usually tuned by hand, by intuition, until it seems good enough.

The branches of a tree seen from below against a pale sky
The framework keeps several harness variants alive and lets them compete. Illustrative image. Ishtiak Ahamed · pexels · Pexels License

A systematic search over that space is therefore directly usable in a way that a training result is not. Salesforce has released an open-source framework called Beagle under the Apache 2.0 licence so teams can run harness evolution against their own agents. The senior author on the research is Ran Xu.

What would confirm it

Two things. The first is someone outside Salesforce reproducing the WebArena-Infinity jump on a harness they did not write, since the size of an improvement depends heavily on how weak the starting point was. The second is evidence that an evolved harness holds up on tasks that were not in the evaluation suite — the whole method is a search against a fixed set of benchmarks, and searches against fixed sets have a well-known habit of finding the set rather than the skill.