How are AI-generated images detected in photography competitions?
No single screening method — AI detector, metadata analysis, or human expert — is reliable enough on its own to catch AI-generated competition entries.
Covers: The technical, procedural and metadata-based methods competitions use to screen entries for AI generation or AI editing, and how contest rules define AI use. Does not cover how to create AI images or how to evade detection.
Also answers: How do photo contests catch AI-generated images? · Can photography competitions detect AI images? · How are AI photos detected in contests? · What tools detect AI-generated photos in competitions?
- One page for this question6 other ways of asking lead here
- 6 independent sourcesEvery claim links to what supports it
- 2 connected pages2 changed this week
- Clean discussionScreened before anything appears
Before you read, make a guess
Fill in the blank: ?% of expert reviewers correctly identifying manipulated images
Drag the slider to fill in the blank
The short answer
Interpretation AI-prepared starting mapCompetitions and forensic examiners have three broad tools for screening entries: automated AI-image detectors, metadata and file-structure analysis, and human expert review. None is currently reliable enough to stand alone. A diagnostic study of 104 manipulated research images found 24 PhD-level expert reviewers performed at chance level (mean accuracy 50.5%, SD 6.9%), while the best commercial AI detector reached only moderate discrimination (AUC 0.790, 95% CI 0.695–0.885). A separate forensic study found an AI expert system scored 91% versus a 69% human mean across 20 examiners on a mixed set of original, transmitted, edited and synthetic videos. Detector research is advancing — one framework reported a 19% relative improvement over prior state-of-the-art — but a 400-plus-study review concluded that inconsistent evaluation datasets make methods hard to compare and that generalizing to unseen generators remains the central unsolved problem.1234
- Evidence 16
- Interpretation 5
Did this answer your question?
Be the first to voteIn brief
In one diagnostic study, 24 PhD-level experts identified manipulated images at chance level (50.5% mean accuracy), while the best commercial detector reached only moderate discrimination (AUC 0.790).1
Evidence-backedIn a separate preliminary forensic study, an AI expert system scored 91% against a 69% human mean across 20 examiners on a mixed video dataset.2
Evidence-backedAt least one major contest, Nikon's Small World in Motion, discovered an AI-generated prize-winning image and is re-evaluating its rules and procedures as a result.5
Evidence-backed
At a glance
The picture in numbers
Live · updated just now
50.5%
51 in every 100
- AI expert system91%
- Human examiners (mean)69%
19%
19 in every 100
400 studies
The evidence behind it
6 sources- Reviews of many studies1
- Other studies and data4
- Background1
Published in 2026
| Source | Kind | Year |
|---|---|---|
| Interpol review of detection of AI-generated image and video deepfakes, 2022-2025. | Reviews of many studies | 2026 |
| Biomedical Research Images Manipulated by Generative AI to Alter Scientific Outcomes: Diagnostic Study of Human and Automated Detection. | Other studies and data | 2026 |
| Penny-Wise and Pound-Foolish in AI-Generated Image Detection. | Other studies and data | 2026 |
| CoDA: Color Distribution Probing for Efficient and Generalizable AI-Generated Image Detection. | Other studies and data | 2026 |
| Evaluating the use of a novel expert system for interpretation of media authentication results. | Other studies and data | 2026 |
| Prize-winning image which sparked backlash was AI-generated, Nikon rules | Background | 2026 |
The community around it
No one has added to this page yet. Firsthand experience, a newer study or a different reading of the numbers would show up here, credited to you.
What it means for you
Which fits you?
Pick the situation closest to yours. Each answer says what it rests on.
If you are running a competition and deciding how to screen entries
treat any single method as a flag rather than a verdict: automated detectors, metadata and file-structure checks, and human review each have documented failure modes, and the research review recommends embedding detection in a broader evidence-evaluation framework rather than relying on one score.41
InterpretationIf you are relying on human expert reviewers to spot AI manipulation
expect weak results on hard cases: 24 PhD-level reviewers scored at chance level (50.5% mean accuracy) on high-fidelity manipulated research images, so expert review alone is not a dependable screen.1
Evidence-backedIf you are considering an automated expert system for forensic or contest screening
a preliminary study found an AI expert system scored 91% versus a 69% human mean across 20 examiners on mixed video, but the authors call it preliminary and frame it as a basis for evaluating deployment alongside human interpretation, not as a replacement.2
Evidence-backedIf you are choosing a detector and comparing published accuracy claims
be cautious about cross-paper comparisons: the field's own review found inconsistent use of evaluation datasets makes methods difficult to compare, and generalization to unseen generators is the main open challenge.4
Evidence-backedIf you are writing or revising contest rules on AI use
note that Nikon responded to an AI-generated prize-winning image by re-evaluating the rules and procedures of its Small World in Motion contest, which suggests rules may need to specify permitted AI operations rather than only banning AI outright.5
Evidence-backedIf you are an entrant worried your legitimate photograph could be flagged
the sources here do not describe any contest's false-positive rate or appeal process, so the practical risk and any route to challenge a ruling are not established by the available evidence.5
InterpretationThe full story · 3 chapters
01
How detection methods work
AI summary:Detectors read statistical traces in pixels, newer frameworks adapt vision-language models, and forensic workflows add file structure and metadata.
Evidence-backed: Automated detectors look for statistical traces left by generative models. One approach probes color distribution: real photographs tend to show smoother, more stable color patterns, while synthetic images often carry characteristic color imbalances introduced by neural generation. A detector built on this cue uses only 1.48 million parameters and reported state-of-the-art results on standard benchmarks and the best results on the cross-domain FakeForm evaluation, while staying competitive in cross-model photorealistic settings.6
Evidence-backed: Another line of work adapts large pre-trained vision-language models to the detection task. One framework (PoundNet) uses a learnable prompt design and a balanced training objective to keep broad knowledge from upstream object-classification tasks while improving detection. Trained on a single standard AI image dataset and tested across 10 large-scale public detection datasets with 5 evaluation metrics, it reported a 19% relative improvement in detection performance over prior state-of-the-art methods while retaining 63% performance on object classification.3
Evidence-backed: A review of more than 400 studies published between 2022 and 2025 mapped the wider field. It covers generative architectures including GANs, diffusion models and autoregressive models, and notes progress in visual fidelity, user control, content consistency and computational efficiency. On the detection side, it found that research effort is concentrated on generalizing to unseen generators and staying robust against common perturbations and adversarial attacks.4
Evidence-backed: Beyond pixels, forensic workflows use file structure data, attribute similarity analyses, proprietary structural data and metadata values. In one study, both an AI expert system and trained human examiners from diverse digital forensic backgrounds were given the same dataset of original, transmitted, edited and synthetic videos along with this supporting file and metadata information.2
02
How reliable are these methods?
AI summary:Experts scored at chance on manipulated images while an AI system beat humans on video, but the studies differ too much to compare.
Evidence-backed: Human expert review is the weakest link in the evidence. In a diagnostic study of 104 western blot and subcutaneous xenograft tumor images, a high-fidelity generative model produced forgeries capable of altering a study's conclusions. Twenty-four PhD-level expert reviewers could not reliably distinguish the forgeries from authentic figures, with mean accuracy of 50.5% (SD 6.9%) — essentially chance. The best-performing commercial AI detector managed only moderate discrimination, with an area under the curve of 0.790 (95% CI 0.695–0.885).1
Evidence-backed: A separate forensic study pointed the other way for automated systems. Comparing an AI expert system against trained human examiners on the same video dataset, across 20 respondents the human mean score was 69% while the expert system scored 91%. The authors present this as a preliminary result and frame it as a basis for evaluating whether AI expert systems are appropriate for forensic examinations alongside human interpretation.2
Evidence-backed: The research review identifies the core reliability problem: methods struggle to generalize to generators they were not trained on and to remain robust against perturbations and adversarial attacks. It also notes that inconsistent use of evaluation datasets makes it difficult to compare methods, and recommends forensically relevant benchmark datasets, more attention to explainability, and embedding detection methods in an evidence-evaluation framework.4
Interpretation: The two reliability findings point in opposite directions — chance-level human accuracy in one study, a large automated advantage in another — but they tested different media (still biomedical images versus video), different tasks and different sample sizes. They should not be read as a single verdict on whether machines or humans are better screeners.12
03
What competitions actually do
AI summary:Nikon ruled a prize-winning Small World in Motion image AI-generated and is re-evaluating its contest rules and procedures.
Evidence-backed: On 9 October 2026, BBC News reported that a prize-winning image which sparked backlash was ruled AI-generated by Nikon, the camera-maker, which said it is now re-evaluating the rules and procedures of its Small World in Motion contest. This is the only source here describing a real competition outcome, and it shows a contest discovering an AI entry after the fact and responding by revisiting its rules rather than by describing a screening method.5
Your turn
Have your say
See where others stand. Join free to add your perspective. One answer per account.
How do you feel about this?
No votes yetQuick questions from connected pages
Before you go
What to remember
Try to recall each hidden figure before you reveal it. Remembering, not rereading, is what makes it stick.
In one diagnostic study, PhD-level experts identified manipulated images at chance level (50.5% mean accuracy), while the best commercial detector reached only moderate discrimination (AUC 0.790).
In a separate preliminary forensic study, an AI expert system scored against a 69% human mean across 20 examiners on a mixed video dataset.
Detector research is improving — one framework reported a relative gain over prior state-of-the-art — but generalizing to unseen generators and resisting adversarial attacks remain unsolved.
Your reading
0 of 3 chaptersThis answer keeps changing
When new evidence or a better source comes in, this page is updated (it's on version 2, last changed 42 minutes ago). Follow it to be told when that happens.
Ask this Sylo
Still wondering about something?
Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.
More on AI
Everything on AI ›How does AI-generated imagery affect photography competitions and trust?
Is AI really using up our drinking water?
Is artificial intelligence really using up our drinking water, and how much water do data centres actually consume?
How do AI companies handle employees who raise safety concerns?
How do AI companies handle employees who raise safety concerns about their products or research?
What are the risks of AI agents taking autonomous actions in the real world?
Why are data centre developments facing local protests?
How do AI agents give false tips to police and how are they audited?
How can AI agents generate false tips to police, and how are these systems audited?
Behind this page
Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.
Discussion
Sources
Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.
- 1Biomedical Research Images Manipulated by Generative AI to Alter Scientific Outcomes: Diagnostic Study of Human and Automated Detection.Journal of medical Internet research (Luo et al.)Published Oct 2, 2026Checked Oct 11, 2026
“UnlabelledIn a diagnostic study of 104 western blot and subcutaneous xenograft tumor images, a high-fidelity generative model produced forgeries that could alter the conclusions of a study; 24 PhD-level expert reviewers could not reliably distinguish the forgeries from authentic figures (mean accuracy 50.5%, SD 6.9%), while the best-performing commercial AI detector achieved only moderate discrimination (area under the curve 0.790, 95% CI 0.695-0.885), revealing critical vulnerabilities in current research-integrity safeguards.”
- 2Evaluating the use of a novel expert system for interpretation of media authentication results.Journal of forensic sciences (Epstein et al.)Published Jun 11, 2026Checked Oct 11, 2026
“These tests were organized into distinct authentication pathways to produce concise, natural language conclusions comparable to opinions rendered by human examiners. The automated nature of an expert system may allow for demonstrably accurate results at scale without extensive human resources. This preliminary study compared the accuracy of the artificial intelligence expert system's opinions with those of trained human examiners from diverse digital forensic backgrounds. Both the expert system and the examiners were provided the same dataset, consisting of original, transmitted, edited, and synthetic videos. Both groups received file structure data, attribute similarity analyses, proprietary structural data, and metadata values. Using a 14-question survey, responses were collected and evaluated for accuracy, and error rates were calculated. Across 20 respondents, the human mean score was 69% as compared to the expert system score of 91%. These results can be used to evaluate the appropriateness of deploying artificial intelligence expert systems for forensic examinations, alongside human interpretations.”
- 3Penny-Wise and Pound-Foolish in AI-Generated Image Detection.IEEE transactions on pattern analysis and machine intelligence (Wang et al.)Published Jul 1, 2026Checked Oct 11, 2026
“To address this trade-off issue, we propose a novel learning framework (PoundNet) for the generalization of AI-generated image detection on a pre-trained vision-language model. PoundNet incorporates a learnable prompt design and a balanced objective to preserve broad knowledge from upstream tasks (object classification) while enhancing generalization for downstream tasks (AI-generated image detection). We train PoundNet on a single standard AI image dataset, following common practice in the literature. We then evaluate its performance across 10 large-scale public AI-generated image detection datasets with 5 main evaluation metrics, forming the largest benchmark test set for assessing the generalization ability of AI-generated image detection models, to our knowledge. The comprehensive benchmark evaluation demonstrates that PoundNet successfully balances generalization with knowledge retention, achieving a remarkable relative improvement of 19% in AI-generated image detection performance compared to state-of-the-art methods, while maintaining a strong performance of 63% on object classification tasks.”
- 4Interpol review of detection of AI-generated image and video deepfakes, 2022-2025.Forensic science international. Synergy (van et al.)Published Jun 13, 2026Checked Oct 11, 2026
“We analyzed more than 400 studies published between 2022 and 2025, specifically addressing the generation and detection of AI-generated deepfakes as opposed to manipulation of existing media. We provide an overview of the developments in various generative architectures such as GANs, diffusion models, and autoregressive models, highlighting progress in visual fidelity, user control, content consistency, and computational efficiency. In addition, we outline the most common designs of detection methods, looking at various types of features that are used for detection. We conclude that innovations are mainly centered on addressing the challenges of generalizability to unseen generators and robustness against common perturbations and adversarial attacks. However, the inconsistent use of the evaluation datasets makes it difficult to compare the methods. Our review has identified several directions for future research, such as making methods more directly applicable to forensic practice by incorporating forensically relevant benchmark datasets, paying more attention to explainability, and embedding detection methods in an evidence evaluation framework.”
- 5Prize-winning image which sparked backlash was AI-generated, Nikon rulesBBC NewsPublished Oct 9, 2026Checked Oct 11, 2026
“The camera-maker says it is now re-evaluating the rules and procedures of its Small World in Motion contest.”
- 6CoDA: Color Distribution Probing for Efficient and Generalizable AI-Generated Image Detection.IEEE transactions on pattern analysis and machine intelligence (Jia et al.)Published Oct 1, 2026Checked Oct 11, 2026
“Motivated by this broader setting, we revisit color-distribution probing as an efficient complementary cue for AI-generated image detection. We observe that, especially for photographic content, real photographs tend to exhibit smoother and more stable color patterns, whereas synthetic images often show characteristic color imbalances introduced by neural generation. Based on this observation, we propose CoDA, a compact 1.48M-parameter detector built on a Noise-Quantization Probe, together with a theoretical analysis linking probe responses to color non-uniformity. Experiments show that CoDA achieves state-of-the-art performance on standard benchmarks and the best results on the challenging cross-domain evaluation of FakeForm, while remaining highly competitive in cross-model photorealistic settings. These results suggest that persistent generative artifacts can provide a practical foundation for efficient and robust AI-generated image detection. The models and FakeForm benchmark will be made publicly available.”
How it changed
Published 1 time since Oct 11, 2026.
- Version 2Oct 11, 2026Live now
AI-prepared Starting Map from live research.
- First published version.
Help improve it
The brief is open about what's uncertain. These are the specific gaps that new material would fill.
“What competitions actually do” rests on one independent source
A second, independent source that confirms or challenges it would make this part more reliable.
Open questions
How many competitions screen entries with automated detectors, metadata checks or human review, and at what stage — before judging, after a complaint, or only after a prize is awarded?
No answers yet
What happens when a detector flags a legitimate photograph, and do contests offer a route to contest a ruling?
No answers yet
How do contest rules define permitted AI use — for example, denoising, upscaling or generative fill — and are those definitions consistent across competitions?
No answers yet
How do detector accuracy figures measured on benchmark datasets translate to the mixed, compressed and resized files that competitions actually receive?
No answers yet
What have entrants and judges experienced when AI screening was applied to their work?
No answers yet
Around this topic
Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.