Bay Street Wire
Tech & BusinessOpinion

The Provenance Tax: Why AI Watermarking is a Security Liability

Portrait of Naomi Frost
Naomi Frostcybersecurity & privacySep 26AI
The Provenance Tax: Why AI Watermarking is a Security Liability

AI-generated image · Bay Street Wire

Forcing AI agents to embed invisible signals doesn't just track content—it creates 'sampling drift' that can break critical tool-calling and safety refusals.

Industry leaders are framing watermarking as a transparency win, but as Lasso Security first reported, from a defender's perspective, it looks like a fragile leash. Anthropic recently announced that future Claude models will embed invisible watermarks based on Google DeepMind’s SynthID-Text, a move driven in part by the EU AI Act's requirement for machine-readable synthetic text detection.

Here is the problem: watermarking isn't a passive label; it's an active intervention in the generation process. According to research from Lasso Security, SynthID-Text utilizes 'tournament sampling' to bias token selection. This creates what Lasso Security calls 'sampling drift.' While Google DeepMind's Dathathri et al. claim no measurable quality degradation across millions of Gemini responses, Lasso Security's analysis reveals that this drift manifests in actual agent behavior.

In the context of AI agents, the tokens being biased are the same tokens the agent uses to act. Lasso Security found that watermarking affects both model refusal behavior—including whether a model resists harmful requests under prompt injection—and agent tool calling. Because tournament sampling alters token selection where the model is uncertain, it can change the arguments an agent passes to a tool or determine which tool is called entirely.

**Opinion:** We are trading operational reliability for a provenance signal that is easily disrupted. By introducing this 'provenance tax,' developers are baking instability into the model's reasoning. When an agent's output is tweaked to satisfy a watermark, it creates a gap where adversarial noise or prompt injections can more easily trigger failures or hallucinations, turning a supposed safety feature into a primary attack vector.

Sources

More from Naomi Frost