SyloSpace

How does Google SynthID detect AI-generated content?

SynthID hides signals in AI text and images so detectors can spot them, but detection is split by company and text watermarks often vanish after paraphrasing.

Updated 2 hours ago6 min readVersion 2
CommentsFollow

Covers: This page explains the technical principles behind Google SynthID watermarking and detection for AI-generated text, images, audio, and video, including how it embeds imperceptible signals and how detectors identify them. It also covers SynthID's expansion to detect output from other AI providers like OpenAI, but does not provide instructions for removing or bypassing watermarks.

3 free full reads left this month. Join or upgrade

The short answer

Interpretation AI-prepared starting map

SynthID is Google's family of watermarking systems that embed hidden statistical or pixel-level signals into AI-generated content so detectors can later identify it. For text, SynthID-Text uses a Tournament sampling algorithm to embed the watermark during generation and a score function (Bayesian or mean score) for detection, in a unified design that supports both distortionary and non-distortionary watermarking. For images, the watermark is embedded in pixels rather than metadata: across ten AI images passed through Canva export, screenshots, WhatsApp transfers and Instagram posts, the SynthID watermark survived every action and every chain when read by that company's own detector, while C2PA metadata was removed by the first action every time (survival rate 0 across fifty post-processing checks). Detection is split across companies: each detector reads only its own company's watermark, so no single public tool reads both, and a "not detected" answer cannot be trusted to mean "not AI."12

What this rests on6 independent sources
  • Evidence 19
  • Interpretation 1

In brief

  1. SynthID embeds hidden signals during generation — a Tournament sampling algorithm and score-based detection for text, pixel-level watermarks for images — so content can be identified later without visible marks.12

    Evidence-backed
  2. Pixel watermarks proved robust across everyday editing and social-media chains, while C2PA metadata was stripped by the very first action in every test.2

    Evidence-backed
  3. Detection is split by company: each detector reads only its own watermark, so a "not detected" result does not mean "not AI."2

    Evidence-backed
  4. Text watermark detection is fragile under paraphrase: 98.3% of initially-detected SynthID text lost its watermark after paraphrasing, and pre-attack false-negative rates were already around 80%.3

    Evidence-backed
  5. Detection can be strengthened statistically: goodness-of-fit tests improved detection power and robustness across three watermarking schemes and three open-source LLMs.4

    Evidence-backed

At a glance

The picture in numbers

Live · updated just now

Forensic-readiness evaluation of SynthID text

98.3%

98 in every 100

of initially-detected SynthID text lost its watermark after paraphrasing3
Forensic-readiness evaluation of SynthID text

80%

80 in every 100

false-negative rate for SynthID text before any attack3
Forensic-readiness evaluation of three text watermarking methods
  • KGW70%
  • Unigram83%
  • SynthID80%
false-negative rates before any attack, by watermarking method3

The evidence behind it

6 sources
  • Other studies and data6

When it was published

Newest from 2026

20242026
Sources on this page by kind and year
SourceKindYear
On Google's SynthID-Text LLM Watermarking System: Theoretical Analysis and Empirical ValidationOther studies and data2026
A Real-World Robustness Evaluation of AI-Image Provenance: Comparing C2PA Content Credentials and the SynthID Watermark Across Consumer Editing and Social-Media PipelinesOther studies and data2026
SoK: Watermarking for AI-Generated ContentOther studies and data2024
Watermarking across Modalities for Content Tracing and Generative AIOther studies and data2025
On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection.Other studies and data2025
AI Watermark Evidence Fails Forensic Readiness: An Empirical EvaluationOther studies and data2026

The community around it

Contributions
0
People
0
Following
0

Nobody has added anything yet. Experience, evidence or a different view would show up here.

What it means for you

Which fits you?

Pick the situation closest to yours. Each answer says what it rests on.

If you rely on a "not detected" result to conclude content is human-made

treat that conclusion as unsafe: each detector reads only its own company's watermark, so no single public tool covers both Google and OpenAI content.2

Evidence-backed

If you need to preserve provenance through editing or social-media sharing

metadata-based provenance such as C2PA is unreliable — it was removed by the first action in every test — whereas the pixel watermark survived every action and chain when read by its own company's detector.2

Evidence-backed

If you are considering text watermark evidence for a legal or disciplinary proceeding

be cautious: in the tested configurations, 98.3% of initially-detected SynthID text lost its watermark after paraphrase, false-negative rates were around 80% before any attack, and the methods did not meet more than two of five Daubert factors.3

Evidence-backed

If you are building or tuning a text watermark detector

consider goodness-of-fit tests, which improved detection power and robustness across three watermarking schemes and three open-source LLMs, and note the advantage they gain from text repetition in low-temperature settings.4

Evidence-backed

If you are choosing a detection score for a SynthID-Text-style system

the Bayesian score was proven more robust with respect to tournament layers than the mean score, which is inherently vulnerable to increased layers and to a designed layer inflation attack; the optimal Bernoulli parameter for detection was 0.5.1

Evidence-backed

If you want to understand watermarking beyond a single provider's models

research covers detecting watermarks in models fine-tuned on watermarked text and training-free watermarks for transformer weights, aimed at model monitoring and misuse detection.5

Evidence-backed

The full story · 3 chapters

01

How SynthID embeds and detects signals

AI summary:SynthID embeds watermarks during generation — a Tournament sampling algorithm for text and pixel-level marks for images — and pixel marks survived everyday editing while C2PA metadata did not.

Evidence-backed

Evidence-backed: For text, SynthID-Text's innovation is threefold: a Tournament sampling algorithm for embedding the watermark, a detection strategy based on an introduced score function (for example a Bayesian or mean score), and a unified design that supports both distortionary and non-distortionary watermarking methods. The first theoretical analysis of SynthID-Text focuses on detection performance and watermark robustness, and proves that the mean score is inherently vulnerable to increased tournament layers, enabling a designed layer inflation attack. The same analysis proves the Bayesian score offers improved robustness with respect to layers, and establishes that the optimal Bernoulli distribution for watermark detection is achieved when the parameter is set to 0.5.1

Evidence-backed

Evidence-backed: For images, the watermark is embedded in the pixels rather than in attached metadata. In a real-world robustness evaluation, ten AI images (five from ChatGPT/DALL-E and five from Google Gemini) were passed through four everyday actions — a Canva export, a screenshot, a WhatsApp transfer, and an Instagram post — first individually and then in chains. C2PA metadata was removed by the first action every time, with a survival rate of 0 across fifty post-processing checks. The SynthID watermark survived every action and every chain for both companies when read by that company's own detector.2

Evidence-backed

Evidence-backed: Across modalities, watermarking techniques have been developed for images, audio, and text, including adapting latent generative models to embed watermarks in all generated content, identifying watermarked sections in speech, and improving watermarking in large language models with tests that ensure low false positive rates. Work also explores detecting watermarks in language models fine-tuned on watermarked text and training-free watermarks for the weights of large transformers, aimed at model monitoring and content moderation.5

Evidence-backed

Evidence-backed: Watermarking schemes embed hidden signals within AI-generated content to enable reliable detection. They are not a silver bullet for all risks associated with generative AI, but can play a role in enhancing AI safety and trustworthiness by combating misinformation and deception. The field formalizes definitions and desired properties of watermarking schemes and examines key objectives and threat models, including practical evaluation strategies for robustness against various attacks.6

02

Across AI models and providers

AI summary:Each company's detector reads only its own watermark, so a "not detected" result cannot be trusted to mean the content is not AI.

Evidence-backed

Evidence-backed: Detection is fragmented across companies. In the image provenance evaluation, each detector read only its own company's watermark: the Gemini app detected Google's watermark and OpenAI Verify detected OpenAI's, but no single public tool read both. The authors conclude that a "not detected" answer cannot be trusted to mean "not AI," and that provenance fails as a usable system because detection is split across companies with no shared detector.2

Evidence-backed

Evidence-backed: The scope of watermarking research extends beyond a single provider's models. Techniques have been developed to detect watermarks in language models fine-tuned on watermarked text, and training-free watermarks have been introduced for the weights of large transformers, both aimed at detecting model misuse and supporting model monitoring.5

03

How reliable is detection?

AI summary:Goodness-of-fit tests can improve detection, but paraphrase removed the watermark from nearly all initially-detected text and false-negative rates were already high.

Evidence-backed

Evidence-backed: Detection methods for text watermarks often rely on pivotal statistics that are independent and identically distributed under human-written text, which makes goodness-of-fit (GoF) tests a natural tool. A systematic evaluation of eight GoF tests across three popular watermarking schemes, using three open-source LLMs, two datasets, various generation temperatures, and multiple post-editing methods found that general GoF tests can improve both the detection power and robustness of watermark detectors. Text repetition, common in low-temperature settings, gives GoF tests a unique advantage not exploited by existing methods.4

Evidence-backed

Evidence-backed: A forensic-readiness evaluation using meaning-preserving paraphrase as the attack vector reports severe weaknesses. Out of 846 valid paraphrase runs across 15 diverse prompts per method, every initially-detected KGW and Unigram text lost its watermark after paraphrasing — 100% conditional removal — while SynthID fared only slightly better at 98.3%. Even before any attack, false-negative rates were already high: 70% for KGW, 83% for Unigram, and 80% for SynthID. The SynthID configuration also flagged 5.4% of paraphrased human-written controls as AI-generated and showed an 18.6% paradox rate, with 80% of its own pristine watermarked output landing in the uncertainty deadband. None of the three methods satisfied more than two of five Daubert factors, and the authors conclude these configurations as tested do not meet the evidentiary bar courts require.3

Ask this Sylo

Still wondering about something?

Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.

Behind this page

Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.

Discussion

Nobody has added anything yet. If you have experience, evidence or a different view, you could be the first.

Sources

Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.

  1. 1
    On Google's SynthID-Text LLM Watermarking System: Theoretical Analysis and Empirical Validation
    arXiv (Cornell University) (Omidi et al.)Published Mar 3, 2026Checked Oct 8, 2026
    “The system's innovation lies in: 1) a new Tournament sampling algorithm for watermarking embedding, 2) a detection strategy based on the introduced score function (e.g., Bayesian or mean score), and 3) a unified design that supports both distortionary and non-distortionary watermarking methods. This paper presents the first theoretical analysis of SynthID-Text, with a focus on its detection performance and watermark robustness, complemented by empirical validation. For example, we prove that the mean score is inherently vulnerable to increased tournament layers, and design a layer inflation attack to break SynthID-Text. We also prove the Bayesian score offers improved watermark robustness w.r.t. layers and further establish that the optimal Bernoulli distribution for watermark detection is achieved when the parameter is set to 0.5. Together, these theoretical and empirical insights not only deepen our understanding of SynthID-Text, but also open new avenues for analyzing effective watermark removal strategies and designing robust watermarking techniques. Source code is available at https: //github.com/romidi80/Synth-ID-Empirical-Analysis.”
  2. 2
    A Real-World Robustness Evaluation of AI-Image Provenance: Comparing C2PA Content Credentials and the SynthID Watermark Across Consumer Editing and Social-Media Pipelines
    Zenodo (CERN European Organization for Nuclear Research) (Sundaramoorthy)Published Sep 13, 2026Checked Oct 8, 2026
    “We take ten AI images (five from ChatGPT/DALL-E and five from Google Gemini) and pass each one through four everyday actions (a Canva export, a screenshot, a WhatsApp transfer, and an Instagram post), first one at a time and then in chains of several actions. At every step, we measure the C2PA metadata and the pixel change with two small open-source Python tools, and we check the SynthID watermark with both companies' own detectors: the Gemini app and OpenAI Verify. The C2PA metadata is removed by the first action every time (survival rate 0 across fifty post-processing checks). The SynthID watermark survives every action and every chain for both companies, when read by that company's own detector. The key result concerns detection: each detector reads only its own company's watermark, so no single public tool reads both, and a "not detected" answer cannot be trusted to mean "not AI." We conclude that metadata-based provenance does not work in real use, that the pixel watermark is robust for both companies, but that provenance still fails as a usable system because detection is split across companies with no shared detector.”
  3. 3
    AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
    arXiv (Cornell University) (Tamim & Khan)Published Jul 17, 2026Checked Oct 8, 2026
    “We focus on meaning-preserving paraphrase as the attack vector, since it is both legally realistic and difficult to dismiss as evidence tampering. The results raise serious evidentiary concerns. Out of 846 valid paraphrase runs across 15 diverse prompts per method, every single initially-detected KGW and Unigram text lost its watermark after paraphrasing -- 100% conditional removal. SynthID fared only slightly better at 98.3%. Even before any attack, false-negative rates were already high: 70% for KGW, 83% for Unigram, 80% for SynthID. The SynthID configuration also flagged 5.4% of paraphrased human-written controls as AI-generated and showed an 18.6% paradox rate, with 80% of its own pristine watermarked output landing in the uncertainty deadband. None of the three methods satisfy more than two of five Daubert factors. We also find that the FRS point-based scoring system, despite working as designed, cannot fully capture forensic uselessness -- a limitation worth noting for future framework design. These configurations, as tested, do not meet the evidentiary bar that courts require.”
  4. 4
    On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection.
    Advances in neural information processing systems (He et al.)Published Jan 1, 2025Checked Oct 8, 2026
    “Large language models (LLMs) raise concerns about content authenticity and integrity because they can generate human-like text at scale. Text watermarks, which embed detectable statistical signals into generated text, offer a provable way to verify content origin. Many detection methods rely on pivotal statistics that are i.i.d. under human-written text, making goodness-of-fit (GoF) tests a natural tool for watermark detection. However, GoF tests remain largely underexplored in this setting. In this paper, we systematically evaluate eight GoF tests across three popular watermarking schemes, using three open-source LLMs, two datasets, various generation temperatures, and multiple post-editing methods. We find that general GoF tests can improve both the detection power and robustness of watermark detectors. Notably, we observe that text repetition, common in low-temperature settings, gives GoF tests a unique advantage not exploited by existing methods. Our results highlight that classic GoF tests are a simple yet powerful and underused tool for watermark detection in LLMs.”
  5. 5
    Watermarking across Modalities for Content Tracing and Generative AI
    arXiv (Cornell University) (Fernandez)Published Feb 4, 2025Checked Oct 8, 2026
    “The contributions of this thesis include the development of new watermarking techniques for images, audio, and text. We first introduce methods for active moderation of images on social platforms. We then develop specific techniques for AI-generated content. We specifically demonstrate methods to adapt latent generative models to embed watermarks in all generated content, identify watermarked sections in speech, and improve watermarking in large language models with tests that ensure low false positive rates. Furthermore, we explore the use of digital watermarking to detect model misuse, including the detection of watermarks in language models fine-tuned on watermarked text, and introduce training-free watermarks for the weights of large transformers. Through these contributions, the thesis provides effective solutions for the challenges posed by the increasing use of generative AI models and the need for model monitoring and content moderation. It finally examines the challenges and limitations of watermarking techniques and discuss potential future directions for research in this area.”
  6. 6
    SoK: Watermarking for AI-Generated Content
    arXiv (Cornell University) (Zhao et al.)Published Nov 27, 2024Checked Oct 8, 2026
    “These schemes embed hidden signals within AI-generated content to enable reliable detection. While watermarking is not a silver bullet for addressing all risks associated with GenAI, it can play a crucial role in enhancing AI safety and trustworthiness by combating misinformation and deception. This paper presents a comprehensive overview of watermarking techniques for GenAI, beginning with the need for watermarking from historical and regulatory perspectives. We formalize the definitions and desired properties of watermarking schemes and examine the key objectives and threat models for existing approaches. Practical evaluation strategies are also explored, providing insights into the development of robust watermarking techniques capable of resisting various attacks. Additionally, we review recent representative works, highlight open challenges, and discuss potential directions for this emerging field. By offering a thorough understanding of watermarking in GenAI, this work aims to guide researchers in advancing watermarking methods and applications, and support policymakers in addressing the broader implications of GenAI.”

How it changed

Published 1 time since Oct 8, 2026.

  1. Version 2Oct 8, 2026Live now

    AI-prepared Starting Map from live research.

    • First published version.
Every version, side by side

Help improve it

The brief is open about what's uncertain. These are the specific gaps that new material would fill.

Open questions

  • What are SynthID's specific detection rates and robustness for audio and video, as opposed to the technique-level descriptions available for those modalities?

    No answers yet

  • Could a shared or interoperable detector read watermarks from multiple providers, and what would that require?

    No answers yet

  • How do SynthID's false-positive and false-negative rates translate into real moderation or legal decisions, given the reported uncertainty deadband?

    No answers yet

  • Can detection be made robust to meaning-preserving paraphrase, or is paraphrase an inherent limit of current text watermarking?

    No answers yet

Around this topic

Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.

Add what you know

Sign in to add what you know. Reading stays open to everyone.

Ask this Sylo

Answers only from “How does Google SynthID detect AI-generated content?”

Ask anything about this page. The AI reads only its reviewed brief, sources and contributions, cites what it used, and says when the page doesn't cover something.