Anthropic confirmed that future Claude models will embed an invisible watermark in their output, built on Google DeepMind’s SynthID-Text. It sounds like a labelling change. It isn’t. SynthID-Text uses tournament sampling, which means it sits inside the token selection process itself. Change how tokens get picked and you change what the model refuses, whether that refusal survives a prompt injection, and which tool your agent calls with which arguments. Anthropic applies the watermark at the model level, so it covers the Claude Platform API and cloud providers too.
Researchers call this sampling drift, and it shows up in both refusal behaviour and agent tool calling. It’s also model- and key-dependent, and aggregate scores can hide it when changes cancel out. Below: what that costs you in production, and the paired-run evaluation habit that catches it.

The Model Under Your Agent Changed and Your Test Suite Didn’t Notice
Your prompt didn’t change. Your orchestration code didn’t change. Your tool schemas are byte-identical to last quarter. What changed is the sampling process underneath, and no deployment pipeline you own touched it. Researchers describe the resulting behavioral shift as sampling drift, and it surfaces in the two places production agents actually break: refusal behavior and tool calling.
The research is explicit that the effect is model- and key-dependent, and that aggregate scores hide it when changes in opposite directions cancel each other out. A pass rate that stays at 94% can sit on top of dozens of individual decisions that flipped in both directions.
Most quality and operations teams have filed watermarking under legal. Regression suites are built to catch code changes, not silent provider-side generation changes. That gap is the operational risk.

What SynthID-Text Actually Does to Every Token Your Agent Generates
There are two families of text watermarks. Post-processing methods take finished text and alter it afterwards, swapping words or punctuation. Generation-time methods reach into the loop that produces each token. SynthID-Text is the second kind, and that distinction is the whole story for anyone running agents in production.
Tournament sampling versus standard next-token sampling
Ordinary generation is simple to describe. The model produces a probability distribution over possible next tokens, and the sampler draws one. Repeat a few hundred times and you have a response.
A generative watermark adds three components to that loop: a random seed generator, a sampling algorithm, and a scoring function. SynthID-Text’s version is tournament sampling, described in Dathathri et al. Candidate tokens are drawn and then compete in seeded rounds, and the winner carries the watermark signal. The detector later reads that signal back out through the scoring function.
The output is still fluent text. But the token that survives the tournament is not necessarily the token plain sampling would have drawn, and your agent acts on tokens, not on fluency.
Why ‘non-distortionary’ still allows different outputs under a fixed key
The non-distortionary configuration preserves the original token distribution in expectation over the watermark randomness. That is a statement about averages across keys, not a guarantee about any single run. Under a fixed key, individual generations can still differ from the unwatermarked baseline. Dathathri et al. report no measurable quality degradation across nearly twenty million Gemini responses, which is a serious result and worth taking at face value.
It is also a different claim than the one operations teams need. “No average quality loss” and “same output” are not the same sentence. Quality metrics average over a corpus; your agent executes one trajectory at a time, and a single divergent token can change a tool name or an argument value.
Treat the quality evidence as reassurance about writing and the behavioral question as separate and unanswered by it.
Sampling Drift: Weakened Refusals Meet Tools That Actually Do Things
The empirical finding is blunt: drift shows up in both model refusal behavior and agent tool calling. It is not uniform. The effect is model- and key-dependent, which means the same watermarking setup can move behavior one way on one model and a different way on another, or shift again when the key changes.
The research names the compounding risk directly. Prompt injection links the two layers, because a refusal that weakens under injection matters far more when the model can also act through tools. In a manufacturing context, those same sampled tokens decide whether an agent flags a nonconformance for human review or quietly writes a passing value to an ERP record. One is a ticket. The other is a traceability problem you find during an audit.
Why aggregate accuracy scores hide the drift
Aggregate scores cancel. If watermarking causes twenty cases to improve and twenty to regress, your benchmark number barely moves and your dashboard says nothing changed. The composition of that number changed completely.
This is why the vendor-side quality claims, while accurate, are not sufficient for you. Dathathri et al. report no measurable quality degradation across nearly twenty million Gemini responses. Fluency and helpfulness held. That says nothing about whether a specific agent, on a specific key, now calls update_batch_record where it previously escalated.
Paired disagreement as the metric that actually surfaces it
The researchers report both net performance and paired disagreement between watermarked and unwatermarked runs. Paired disagreement is the one that does real work. You run the identical prompt set twice, watermark on and watermark off, then count the cases where the two runs diverge in outcome rather than in score.
For operations teams, the unit of comparison should be the decision, not the text. Did the refusal hold. Which tool was called. What arguments were passed. Count divergences on those three fields and you get a drift rate you can actually act on, instead of a benchmark average that flatters both configurations equally.

Article 50(2) Makes This Everyone’s Problem, Not Just the Model Provider’s
Article 50(2) of the EU AI Act puts the obligation on providers of AI systems that generate synthetic text. They must mark outputs in a machine-readable format and make them detectable as artificially generated. The legal text is specific about the standard those markings have to meet.
using technical solutions that are effective, interoperable, robust, and reliable as far as technically feasible
That wording is why marking is showing up inside generation rather than as a toggle in a settings panel. A feature customers can switch off is not effective or reliable. So the provider solves their compliance problem in the one place they fully control, which is the sampler, and every downstream system inherits the result.
For a Dutch or German manufacturer running an agent over maintenance logs or supplier correspondence, provenance may be irrelevant to the use case. You inherit the behavioral side effects anyway. Compliance posture and agent reliability are now the same conversation, and “we just call the API” stopped being a clean boundary.
What changes for teams consuming models via cloud providers
Routing through Bedrock, Vertex AI or Azure does not insulate you. When marking is applied at the model layer, it travels with the model into every hosting surface. Your cloud contract governs uptime and data residency. It does not govern how tokens get selected.
Practically, three things need to change in how you manage model dependencies. Pin model versions explicitly and treat any provider-side change as a release that requires re-testing, not a patch note you skim. Add a question to vendor review asking whether generation-time marking is active and whether it is keyed, because key changes alter behavior too. Keep a small set of production-representative prompts, including tool-calling flows, that you run against both the old and new model build before you promote.
None of this is expensive. It is a few hours of evaluation per model change. The alternative is finding out through a mis-called tool in a live workflow.
Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.
The Evaluation Habit That Keeps Agent Behavior Verifiable Through 2027
Stop treating this as a compliance item and treat it as a release gate. Three things belong in your operating stance: pin the model version and watermark configuration in your deployment config, run comparisons on your own traces, and alert on behavioral disagreement instead of output quality alone. Vendor benchmarks will not help you here. Dathathri et al. report no measurable quality degradation across nearly twenty million Gemini responses, and that finding is entirely compatible with your maintenance agent calling a different tool on the same work order.
A minimum paired-run test protocol for production agents
Take fifty to a hundred real traces from your highest-consequence workflow: deviation triage, supplier qualification, spare parts reordering. Replay each one twice, watermarked and unwatermarked, same prompt, same tool schemas, same seed where you control it. Then compare tool name, argument values, and final action, not just the text.
Test refusal durability under injection, not in isolation. A refusal that holds on a clean prompt tells you nothing about a refusal that has to survive adversarial text inside a supplier PDF or a technician’s free-text note. Build a small injection corpus from your own document types and run it through the same paired setup. Because the effect is key-dependent, rerun when the key rotates.
What to log so drift is attributable, not just visible
Log the model identifier and version, the watermark configuration and key identifier if exposed, the full tool call with arguments, and the paired-run disagreement rate per workflow. Without the configuration fields, you will see behavior change and have no way to explain why.
Set the alert threshold on tool-call disagreement, not on aggregate accuracy, since opposite-direction changes cancel in a single score. The cost is a few engineering days per release cycle. Weigh that against one agent taking a wrong action inside a validated process, and the audit trail you then have to reconstruct. Provenance obligations are tightening, and teams with paired runs already wired in absorb the next provider-side change quietly.
Source: lasso.security