A cost nobody had measured
Text watermarking has been sold as a provenance feature with no downside: it biases token selection in a way that leaves a detectable signature and, the argument goes, does not affect the quality of what comes out. Lasso Security published research this week arguing that for agents, it does.
The firm tested Google DeepMind’s SynthID-Text, the scheme adopted by Anthropic and OpenAI among others, across seven open-weight models: phi-4, Llama-3.1-8B, Qwen3-32B, Qwen3-4B, gemma-3-12b, gemma-3-27b and Granite-3.2-8B. Tool-calling accuracy was measured on the BFCL v4 single-turn AST benchmark; refusal behaviour on HarmBench and JailbreakBench.
Tools first
Watermarking reduced tool-calling accuracy on six of the seven models. The failures were of a specific kind: the agent selected the wrong tool, or passed incorrect arguments, producing malformed input and parsing failures downstream.

That pattern is what you would expect if the mechanism is doing what it says. A tool call is a highly constrained output — the function name and the argument structure have to be exactly right. Nudging token probabilities is cheap in prose, where many continuations are acceptable. In a JSON function call there is one acceptable continuation, and a nudge is a defect.
Refusals second
The safety result is the one worth arguing about. On bare harmful requests, Lasso found the effect on refusals was small. Under prompt injection, attack success rates rose significantly.

“Watermarking changes refusal behavior on bare harmful requests, but the effect becomes more pronounced under prompt injection,” the researchers wrote. “At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection.”
A refusal under injection is already a marginal decision — the model is being pushed towards compliance by text in its context, and the refusal wins narrowly or not at all. Anything that perturbs the distribution at that margin will move some cases across the line.
The practical consequence
Lasso’s recommendation is procedural rather than dramatic: security evaluations of agents should be run on watermarked outputs, because that is what the deployment will actually produce. Most published agent safety evaluations are not.
Two caveats belong with this. The seven models tested are all small open-weight models, not the frontier systems where watermarking is most often deployed, and the effect could be smaller on larger models with more headroom. And the research comes from a company selling AI security tooling, which is a reason to want independent replication rather than a reason to dismiss it. Nobody has published a contrary result, but nobody had published this one either until now.