SecurityTechInsider AI security & governance
EN/ NL
Governance

Checking AI citations in three steps

Studies show AI fabricates both non-existent sources and wrongly linked metadata. Here is how to set up source checking as a separate, verifiable workflow.

22 September 2026 4 min
Illustration for this article: Checking AI citations in three steps. Oxidised copper and patinated brass sheet, corrosion blooming across the surface.
Organisations must verify citations independently rather than trust model output, with documented checks recorded per source. Image: SecurityTechInsider — original editorial illustration

You must establish source checking as a separate, documented workflow that runs independently of the model generating the text. Do not rely on citation formatting or the model's own assessment of its sources. Every reference in high-stakes work requires three distinct checks: existence, link validity, and claim support.

The prompt is an analysis of 22 September 2026 of citation hallucination and fabrication in AI systems, which argues that models produce citations that are formatted correctly but factually wrong or non-existent at rates between 3% and 18% depending on the system type. A study examined over 220,000 URLs from deep-research and search-augmented systems across 32 fields and found that deep-research agents, despite producing more citations, hallucinated URLs more frequently than simpler search-augmented models. In our assessment, this means that citation volume is not a proxy for reliability, and that any organisation handling sensitive or high-trust information must treat source verification as a mandatory step rather than an optional refinement.

Why does citation volume not guarantee accuracy?

Deep-research agents are built to search, retrieve and synthesise independently, which means they generate substantially more references than conventional search-augmented models. The research shows this productivity comes with a cost: a higher proportion of hallucinated or non-resolving URLs. The counterintuitive finding is that an agent supplying dozens of sources may introduce more errors than one naming only a handful. The correction does not lie in changing how the model generates citations; it lies in building a separate verification step after generation is complete.

What are the failure modes in citation checking?

  • Fabricated URLs — links that do not exist anywhere, invented by the model to fill a citation slot.
  • Link rot and non-resolution — URLs that once existed but no longer resolve, or that never pointed to the claimed source.
  • Metadata mismatch — a URL that resolves but points to a different publication than the one cited, or to a version that does not contain the claimed passage.
  • Unsupported claims — sources that exist and are correctly linked but do not actually substantiate the specific claim attributed to them.
  • Plausible fabrication — citations that are specific in content, correctly formatted, attributed to real researchers, and entirely invented.

Can the model check its own citations?

No. When tested on their own output, language models achieved only 38% accuracy in judging whether their citations were valid. If the checker is as fallible as the generator, the result is apparent certainty rather than control. Self-verification by a single model is therefore unsuitable as a sole safeguard and creates a false sense of security. External verification of both source existence and claim support is not additional diligence; it is a minimum requirement for any work where the citation matters to the decision.

What concrete controls must you be able to demonstrate?

  1. Record the source check as a separate workflow — document which sources were checked, by whom, when, and using which method, distinct from the generation step.
  2. Verify URL existence and resolution independently — test each link outside the model environment to confirm it resolves and points to a real publication.
  3. Confirm metadata alignment — open the resolved source and verify that the title, author, date and publication match what the citation claims.
  4. Read the cited passage and assess claim support — locate the specific claim in the source text and record whether the source actually supports it or merely touches on a related topic.
  5. Document exceptions and human judgement — record any sources you accept despite incomplete verification, with the reason and the person who made that decision.

What does this mean for high-stakes work?

A faulty citation in consequential decisions undermines not only the text but the knowledge base beneath it. Research into biomedical publications found that the share of papers containing at least one fabricated reference rose sharply between 2023 and early 2026, and that these citations were often specific, well-formatted and attributed to real researchers. It is precisely that plausibility that makes surface-level checking inadequate. For anyone delivering work on which others will act, source checking must be an explicit step in the verification layer, with recorded exceptions and a documented human final judgement.

Automated tools can flag suspicious patterns and test link resolution at scale, but they cannot read context or judge whether a source truly supports a claim. What tooling can do is make the checking process traceable and repeatable. What remains yours is the decision to accept or reject each source on its merits, and the responsibility for that choice.

Sources: This article draws on reporting and guidance from arXiv, Journal of Science Communication and CIDRAP, University of Minnesota.

Marit Halversen

Written by

Marit Halversen

Covers AI governance and regulatory design, with a focus on how compliance obligations land on architecture rather than on paperwork.