← Back to Issues

AI Disinformation:  Securing the Instruments of Truth

The Problem

Adversaries are poisoning the AI models that defense analysts, intelligence platforms, journalists, and policymakers rely on to sort real from fake. [3]  Not generating fakes for human audiences; that threat is real, but it is the one we already see.  The deeper threat is upstream: corrupting the systems we trust to tell us what is true.  If the instruments we use to evaluate truth have themselves been compromised, then every detection tool built on top of them is a gauge attached to a broken sensor.  It may give confident readings, but those readings cannot be trusted.

This is not theoretical.  In 2024, Nicholas Carlini and colleagues at Google demonstrated that poisoning web-scale training datasets is immediately practical and inexpensive.  Their research showed that for as little as $60, an attacker could have poisoned 0.01% of LAION-400M or COYO-700M, two widely used open training datasets, by purchasing expired domains that host dataset shards or by injecting edits into sources like Wikipedia just before scheduled crawls. [4]  By 2025, training data poisoning had moved from academic demonstration to real-world deployment: researchers documented cases where hidden prompts in GitHub code comments poisoned a fine-tuned model (DeepSeek’s DeepThink-R1), web-seeded text manipulated Qwen 2.5’s search-augmented outputs, and Grok 4’s training data was saturated with jailbreak prompts posted on X to the point where a single short command could strip all its safety guardrails. [5]

The Intelligence Advanced Research Projects Activity (IARPA) has studied this class of threat since 2019 through its Trojans in Artificial Intelligence (TrojAI) program, producing over 150 publications on detecting poisoned models before deployment. [6][7]  The research is there.  What is missing is the translation of that research into operational practice, procurement requirements, and red-team exercises across the broader defense and intelligence ecosystem.

What We’re Doing Instead

The United States is getting better at catching deepfakes, and that matters, but it is not the same problem.  DARPA’s Semantic Forensics (SemaFor) program developed hundreds of detection analytics across multiple media types before concluding and transitioning to the Digital Safety Research Institute in late 2024. [1]  In December 2024, the Defense Innovation Unit awarded a $2.4 million, two-year prototype contract to Hive AI for deepfake detection of video, image, and audio content. [2]  These investments are necessary.  They are also aimed at the downstream symptom, finding fakes after they are generated, rather than the upstream cause: ensuring the models we rely on to make those determinations have not themselves been compromised.

We Already Saw What Happens When These Systems Fail

During the 2026 Iran conflict, AI-generated disinformation was deployed at scale during active combat operations, and it worked.  An AI-generated video depicting the aircraft carrier USS Abraham Lincoln on fire was circulated online by an account claiming to be part of the Iranian press.  President Trump was initially taken in by the video before being informed by a U.S. general that the carrier was intact and undamaged. [8]  An AI-generated video falsely depicting the Burj Khalifa in Dubai on fire was viewed tens of millions of times. [9]  The state-linked Tehran Times published an AI-generated photograph of damage to the U.S. Navy’s Fifth Fleet headquarters in Bahrain that depicted a greater degree of destruction than had actually occurred. [9][10]

The platform-level response was reactive and limited.  X identified a Pakistani individual who had hacked 31 accounts and repurposed them to distribute AI-generated war videos under names like “Iran War Monitor.” [11]  One AI-generated video of an Iranian aircraft confronting a U.S. naval vessel was viewed over seven million times and received more than 15,000 likes before being flagged. [12]  X announced a 90-day demonetization policy for accounts sharing undisclosed AI-generated conflict imagery, but this was a response to the flood, not a defense against it. [13]

The most alarming failures, however, were not in human judgment.  They were in the AI verification tools themselves.  BBC Verify reported that X’s own AI chatbot, Grok, was falsely identifying AI-generated videos and photos as genuine. [9]  The Guardian found that Google’s Gemini was doing the same thing: when presented with an authentic photograph of a mass burial of victims of the Minab school attack, Gemini falsely described it as depicting the burial of earthquake victims in Turkey in 2023; Grok claimed it showed a mass burial of COVID-19 victims in Jakarta in 2021. [10]  The AI systems people turned to for verification were themselves producing false confirmations.

That is the upstream corruption problem made concrete.  It is no longer sufficient to ask “is this image real?”  We must also ask: “can we trust the system telling us whether it’s real?”

The Gap in Current Policy

On June 5, 2026, the White House issued National Security Presidential Memorandum NSPM-11, directing accelerated AI adoption across the national security enterprise. [14]  Its four pillars, Adoption, Adaptation, Assurance, and Accountability, establish a significant framework, and its Section 4(c) specifically names distillation attacks as a security concern.  That is a meaningful acknowledgment that AI systems are not only tools but targets.

But NSPM-11 does not address training data integrity for the broader commercial model ecosystem that defense analysts actually use every day.  In March 2026, the Pentagon disclosed plans to set up secure environments where AI companies could train military-specific model versions on classified data. [15]  That effort covers government-controlled models in government-controlled environments.  It does not cover Claude, GPT, Gemini, Grok, or any of the other commercial models that analysts, intelligence officers, journalists, and policymakers routinely use to process information, evaluate sources, and support decisions.  The broader ecosystem remains unaddressed.

Current AI red-team exercises compound the gap.  They test whether models produce harmful content when prompted: can you make it say something dangerous?  They do not test whether models have already been influenced by adversarial training data: has someone already made it believe something false?  The first question is about misuse.  The second is about corruption.  We are testing extensively for one and barely at all for the other.

What I Propose

The detection work the Pentagon has funded is valuable and should continue.  What I am proposing is to extend the defense to the layer where the real vulnerability now lives: the integrity of the models themselves.

Mandatory training data provenance documentation for any AI model used in federal procurement.  If we cannot verify what a model was trained on, we should not trust its outputs where accuracy matters.  I propose extending NSPM-11’s Assurance pillar to require, as a condition of federal procurement, that any AI model supplier document the provenance of its training data: what sources were used, how they were curated, what adversarial content screening was performed, and what ongoing monitoring is in place for post-training data integrity.  This does not require suppliers to disclose proprietary model architectures or weights.  It requires them to account for the data their models learned from, which is a fundamentally different question.

Red-team protocols that specifically test for training data poisoning.  IARPA’s TrojAI program has spent six years developing detection methods for exactly this class of attack and has produced over 150 publications. [6][7]  That research should be translated into operational red-team protocols required for any AI system deployed in a defense or intelligence context.  Current red-team exercises ask: “can we make this model produce harmful outputs?”  The new protocols should also ask: “has this model already been trained on adversarial data that biases its outputs in ways its operators cannot see?”

A public model-integrity reporting mechanism analogous to the CVE database for software vulnerabilities.  The cybersecurity community already maintains the Common Vulnerabilities and Exposures (CVE) system, a standardized, publicly accessible database that identifies and catalogs known vulnerabilities in software so that organizations can assess their exposure and prioritize remediation.  No equivalent exists for AI model integrity.  I propose creating one: a federally maintained, publicly accessible registry where documented instances of training data poisoning, adversarial manipulation, or systematic output bias in widely used AI models can be reported, cataloged, and tracked.  This would give every organization using AI, not just the Pentagon, but newsrooms, courts, hospitals, financial institutions, and election security offices, a standardized way to evaluate whether the models they depend on have known integrity issues.

Allied coordination on training data integrity standards.  This is not a problem the United States can solve alone.  The commercial AI models in question are trained on globally sourced data and deployed worldwide.  I propose establishing bilateral and multilateral working groups with allied nations to develop shared training data integrity standards for AI models used in defense and intelligence contexts.  Several allies are already moving in this direction: Israel’s AI defense infrastructure buildout, including joint projects with U.S. defense firms, has been characterized as taking a provenance-first approach to model development, [3] and the research base for this work is international.  Coordination, rather than redundant national efforts, will produce better standards faster.

The Democratic Dimension

This is not only a Pentagon problem.  If adversaries can poison the models that journalists use to verify sources, that analysts use to evaluate intelligence, that courts use to process evidence, and that citizens use to evaluate the information on which they cast their votes, then the entire epistemic infrastructure of democratic decision-making is compromised.  The 2026 Iran conflict demonstrated this at wartime speed, but the same vulnerability exists in peacetime: a subtly poisoned model does not announce itself with a dramatic failure.  It simply nudges outputs in a direction its operators never chose and may never notice.

Detection tools are the fire alarm.  Training data integrity is the building code.  We need both, and right now we are investing heavily in the first while barely acknowledging the second.

What I Don’t Know

I don’t know the right institutional home for the model-integrity registry I’m proposing.  NIST maintains the National Vulnerability Database and has the institutional culture for this work, but the scope here extends into national security territory that may require a joint structure with the intelligence community.  I’d rather get the function right than the org chart.

I don’t know how far training data provenance requirements can practically extend into the commercial model ecosystem without creating compliance burdens that drive smaller AI companies out of federal procurement entirely, concentrating the market further among a handful of large providers.  That tradeoff needs to be studied, not assumed away.

I don’t know whether the allied coordination I’m proposing is better pursued through existing structures (NATO, Five Eyes, bilateral defense agreements) or through a new purpose-built arrangement.  The answer likely depends on classification levels and on which allies are willing to move first.

I use AI extensively, at work, for research, for this campaign, and in building this very page.  I am not approaching this from fear of the technology.  I am approaching it from an engineering perspective: any system whose inputs can be corrupted cannot be trusted, no matter how sophisticated its outputs appear.  The answer is not to stop using AI.  The answer is to secure the foundation it runs on.

References

[1] DARPA, “Semantic Forensics (SemaFor)” program page, darpa.mil; DARPA, “Furthering Deepfake Defenses” (Mar. 2025), darpa.mil.  Program concluded Sept. 2024; technology transition to Digital Safety Research Institute (DSRI) ongoing.

[2] “Hive Secures DoD Contract for Deepfake Detection, Pioneering AI Defense Against Emerging Threats,” Business Wire (Dec. 5, 2024); “The US Department of Defense is investing in deepfake detection,” MIT Technology Review (Dec. 5, 2024), technologyreview.com.

[3] “The AI disinformation gap the Pentagon may be missing,” Breaking Defense (July 2026), breakingdefense.com.

[4] N. Carlini, M. Jagielski, C. A. Choquette-Choo, D. Paleka, W. Pearce, H. Anderson, A. Terzis, K. Thomas, and F. Tramèr, “Poisoning Web-Scale Training Datasets Is Practical,” 2024 IEEE Symposium on Security and Privacy (SP), pp. 407–425, 2024.

[5] Lakera, “Introduction to Data Poisoning: A 2026 Perspective,” lakera.ai (2026).  Documents real-world training data poisoning incidents in 2025 including DeepSeek DeepThink-R1, Qwen 2.5, and Grok 4.

[6] IARPA, “Trojans in Artificial Intelligence (TrojAI)” program page, iarpa.gov.

[7] “Intelligence Community AI Cybersecurity Program Achieves ‘Massive Scientific Impact,’” AFCEA SIGNAL Media (Mar. 2025), afcea.org.  Reports 150+ publications from the TrojAI program.

[8] “The deepfake video that fooled Donald Trump,” The Independent (Mar. 18, 2026).  Cited via Wikipedia, “Misinformation during the 2026 Iran war.”

[9] T. Copeland, “AI-generated Iran war videos surge as creators use new tech to cash in,” BBC Verify (Mar. 7, 2026), bbc.co.uk.

[10] T. McClure, “A photo of Iran’s bombed schoolgirl graveyard went viral.  Why did AI say it wasn’t real?,” The Guardian (Mar. 17, 2026), theguardian.com.

[11] A. Singh, “X Exposes Pakistan Man Using 31 Accounts To Post AI Videos Amid US-Israel-Iran Conflict,” NDTV World (Mar. 5, 2026), ndtv.com.

[12] L. Carter and M. Workman, “Can you believe your eyes?  War in the age of AI,” ABC News (Mar. 4, 2026), abc.net.au.

[13] L. Tress, “X cracks down on AI-generated war footage as Iran misinformation runs rampant,” The Times of Israel (Mar. 4, 2026), timesofisrael.com.

[14] The White House, National Security Presidential Memorandum/NSPM-11, “Artificial Intelligence in the National Security Enterprise” (June 5, 2026), whitehouse.gov.

[15] “The Pentagon is making plans for AI companies to train on classified data, defense official says,” MIT Technology Review (Mar. 17, 2026), technologyreview.com.

How We’ll Know It’s Working

Goals

A federal procurement requirement for training data provenance documentation from any AI model supplier; operational red-team protocols that test for training data poisoning alongside existing safety and bias testing; and a publicly accessible model-integrity registry modeled on the CVE database.

Metrics

Percentage of federally procured AI models for which training data provenance documentation has been submitted and reviewed; number of adversarial training data tests conducted in defense and intelligence red-team exercises annually; number of model-integrity reports filed, cataloged, and resolved in the public registry.

Review

Annual public reporting to Congress on the state of AI model integrity across the federal procurement ecosystem, including identified vulnerabilities, remediation actions, and unresolved gaps.

Severability

Each proposal, provenance requirements, red-team protocols, the model-integrity registry, and allied coordination, can be implemented independently.  None depends on the others to function.

Sunset

Subject to periodic review and revision as the threat landscape and AI technology evolve.  The model-integrity registry, like the CVE database, should be designed as permanent infrastructure, not a time-limited program.

Share Your Voice

Tell me what you think.  Tell me what I’m missing.  Tell me where I’m wrong.  I will listen and update the website accordingly, and add my thoughts on those ideas.

All fields optional except your comment.  Anonymous submissions welcome.

✓ Thank you. Your comment has been received.
11

Moving this slider doesn’t submit anything by itself. Click Submit Comment below to send it.

Average ranking so far: 11 (starting point -- no rankings yet)

May we post your comment for others to see and review?
CONTACT THE CAMPAIGN