AI Can See Your Watermark
Machines are learning to recognize each other. They already recognize us.
In From Russia With Love, James Bond is approached by a stranger. “Excuse me, can I borrow a match?” asks the stranger. “I use a lighter,” Bond replies, and pulls one from his pocket. “Better still,” the stranger says as he leans over with a cigarette in his mouth. “Until they go wrong…” Bond concludes.
What sounds like small talk to an outsider can be an exchange of credentials between intelligence operatives. Bond and the stranger don’t know each other, but they carry a pre-agreed linguistic key that enables them to identify each other.
AI models do something similar. They can signal their identity in ways that the humans around them cannot notice. And they can even discern our identity from signals we send without noticing.
Let’s start with a simple experiment. Look at the two texts below. Without external help, can you tell which one was generated by AI?
Whichever you chose, you got it right! Both texts were generated by AI. There is nothing particularly machine-like about them. But another machine can easily recognize that they were not written by a human. When I pasted the text into Pangram, it only needed one second to conclude that it was 100% AI-generated.
How does Pangram know? By now, you may have heard that AI has some noticeable stylistic quirks. It loves using em dashes — like this. It loves grouping adjectives and nouns into sets of threes; it must think it is funny, sophisticated, and authentic. And specific AI models tend to overuse certain words like delve, spine, and goblin.
Pangram takes such quirks into account, but it also does something deeper. AI-generated text has unique statistical patterns that distinguish it from human-generated text. And specific AI models have specific patterns that distinguish them from other models.
The same is true for humans as well. The way each of us talks has unique statistical patterns: The frequency of certain words, the length of the pauses between sentences, and the average length of all words and sentences — all add up to a signal that distinguishes you from anyone else. And it also distinguishes humans in general from machines.
AI companies are working hard to make their output harder to distinguish. They want their machines to sound and seem more human. Unless someone forces them to do the opposite. Which is what happened in Europe. The EU AI Act requires “Providers of AI systems” to ensure that “the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated.”
Anthropic is one of the first to comply with this requirement, and it does so in a fascinating manner. Here’s how the company describes it:
When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
The word “watermark” makes it sound like the little symbols that are often added to images or pieces of paper. But Claude’s watermark is not an actual mark; it is something more sophisticated. The explanation continues:
Because the watermark is part of the text, it will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing.
Let’s be clear about what this means: Even if the text would be copied, by hand, onto a piece of paper, it would still carry the “watermark” that shows it was generated by Claude. Even if it would be spray-painted on a wall or pasted into a plain text editor, it would still carry the “watermark.”
How could this be?
Claude’s watermark relies on the same statistical patterns we discussed above. But instead of simply acknowledging that AI-generated text has a unique signature, Anthropic makes this signature intentional. To make it identifiable, Anthropic will introduce a pre-defined statistical bias into the model. As a result, the model’s output would not just be unique, but it would be unique in a pre-defined way.
To see how this works, let’s go back to the two texts we looked at above. The two texts contain multiple pairs of interchangeable words like big/large, shows/reveals. For simplicity’s sake, let’s assume that the likelihood of a human using either word from each pair is 50%. So, if I were to write 100 different sentences that require me to describe something’s size, I would use the word “big” roughly fifty times and the word “large” roughly fifty times.
What would a machine do? As we’ve seen above, the statistical signature of a machine might be different. It might use the word “large” 75% of the time, or only 20% of the time. An AI model would have a bias towards certain words that is different from a typical human’s.
In Claude’s case, the bias itself would be pre-defined. Anthropic will tell the model to prefer certain words a certain percent of the time. These biases would show up in any text generated by the model. At the same time, the model would show no bias in choosing between pairs of words that are not part of its key. The specific combination and intensity of each model’s biases would make it possible to identify it.
Look again at the text below, this time with some color highlights. The words in red represent specific words that Claude was instructed to prefer over others. If we were told about these specific biases in advance (the key), we would know that seeing several of them appear together in the same short text is an indication that the text was generated by Claude.
In reality, the statistical patterns are more subtle than that. But this crude example illustrates how it works. Just like James Bond and his colleague, the machine can use specific words to signal its identity and help it trust and coordinate with other machines.
This type of coordination is not hypothetical. OpenAI recently disclosed that its own AI agents created a “bulletin board” where they exchanged messages with each other by leaving notes inside a software repository. These were not just social messages, but instructions and lessons on how to exploit different vulnerabilities that would enable them to complete the tasks assigned by their trainers.
As Wired reports, “the agents even developed paranoia, suspecting an imposter in their midst” — just like a secret agent would. To address this concern, some AI agents proposed that “messages be signed cryptographically to validate content and root out fraud.”
Machines are learning to recognize each other. Sometimes, because we instruct them to; sometimes, of their own initiative. Machines can recognize us, too. Individual humans already have specific biases in how we use language and even in how we type it. We did not agree to comply with any “human-identifying regulation, and yet each of us is already walking around with a watermark.
How will AI reshape our cities, companies, and careers?
My speaking schedule for the fall and winter is filling up. Visit my speaker profile and get in touch to learn more.
Click here to book a keynote or learn more.








I guess it all comes back to trust, as ever. The innate need to trust to protect oneself is not unique to humans, it seems.