Why Does ChatGPT Make Up Citations?

ChatGPT predicts plausible text, and a citation is just text. The measured fabrication rates, why niche topics fare worse, and how to protect your draft.

The short answer

ChatGPT makes up citations because it generates plausible text, and a citation is just very structured text.

The model predicts one token at a time based on patterns in its training data. It learned what citations look like: author names that fit a field, journal titles that sound established, years that match a topic's research wave, DOIs with the right shape. When you ask for sources, it produces text with exactly those properties.

Nothing in that process checks whether the produced reference exists. There is no built-in lookup against Crossref or PubMed. Existence was never part of the objective. Plausibility was.

So the model does not lie, and it does not retrieve. It composes. Sometimes the composition lands on a real paper, because real papers dominated the training data. Often it lands on a reference that never existed, assembled from real fragments.

What the research measures

Fabrication rates are measured, not anecdotal. The peer-reviewed numbers:

StudyModel testedFabricated citations
Walters & Wilder, Scientific Reports, 2023GPT-3.5 / GPT-455% / 18%
Bhattacharyya et al., Cureus, 2023GPT-3.547% (only 7% were real and fully accurate)
Chelli et al., JMIR, 2024GPT-4 / Bard28.6% / 91.4%
Linardon et al., JMIR Mental Health, 2025GPT-4o19.9%

Walters and Wilder collected 636 citations across 42 topics and found 55% of GPT-3.5's citations and 18% of GPT-4's were fabricated, with substantive errors in many of the real ones. (Scientific Reports)

In medicine, Bhattacharyya and colleagues found that of 115 references GPT-3.5 generated, 47% were fabricated, 46% were real but inaccurate, and only 7% were both authentic and accurate. (Cureus)

For systematic-review prompts, Chelli and colleagues measured hallucination rates of 28.6% for GPT-4 and 91.4% for Google's Bard. (JMIR, 2024)

And in mid-2025, Linardon and colleagues tested GPT-4o: 19.9% of 176 generated citations were fabricated, and 45.4% of the real ones contained errors, most often broken DOIs. (JMIR Mental Health, 2025)

Newer models improved. None got close to safe. Roughly one in five citations from a current model is invented, before counting the merely wrong ones.

Niche topics make it worse, and your thesis is a niche topic

The 2025 study found something that matters for thesis writers specifically.

Fabrication tracked topic familiarity. On a heavily-published topic, major depressive disorder, GPT-4o fabricated 6% of citations. On less-published topics, binge eating disorder and body dysmorphic disorder, fabrication jumped to 28% and 29%. (JMIR Mental Health, 2025)

The mechanism explains why. Dense literature gives the model strong memorized patterns, so its plausible-text generator reproduces real references more often. Thin literature forces it to compose, and composition invents.

A thesis lives in thin literature by design. You write about the narrow intersection nobody covered, which is exactly where fabrication rates triple. The general-purpose fabrication statistics understate your risk.

Why the fakes look so convincing

Fabricated citations survive skimming because every part is drawn from reality.

The model blends. A real researcher's name, because that name appeared near your topic in training data. A real journal, because that journal publishes this field. A title assembled from the field's phrasing. A page range and volume with the right format. Each fragment passes a plausibility check, and your eye checks fragments.

The tell sits in the combination: that author never wrote that paper, that journal never published that title. Only a database lookup or a search exposes it, which is why our guide on how to check if a citation is real starts with the DOI and the exact-title search rather than with reading the reference list harder.

One more convincing trick: invented titles often restate your claim almost word for word. The model generates a source that fits your sentence perfectly because your sentence shaped the generation. Real literature never fits that well.

The consequences are no longer hypothetical

Fabricated citations stopped being a curiosity in 2023 and started ending up in sanctions and retractions.

A federal judge fined two lawyers and their firm $5,000 after they filed a brief citing six court cases ChatGPT had invented, and stood by the fakes when challenged. (CBS News)

Springer Nature retracted a $169 machine-learning textbook in 2025 after Retraction Watch checked 18 of its citations and found two-thirds either did not exist or were badly wrong. (Retraction Watch)

Senior Australian academics apologised to a parliamentary inquiry after AI-generated case studies in their submission, produced with Google Bard, turned out to be fiction. (Information Age)

Universities responded. Libraries now publish dedicated guides on the problem. UNC Charlotte's guide warns that a hallucinated citation "may look real and mix together a combination of real and made-up elements." (UNC Charlotte Library) When librarians write warning pages, markers read them too, which connects to what we cover in can professors tell if you used ChatGPT: a fabricated reference is the one AI trace a marker can prove.

Students sit right in the blast radius

AI-assisted writing became the norm fast. In the UK, the 2026 HEPI student survey found 94% of full-time undergraduates use generative AI to help with assessed work, and 12% submit AI-generated text directly. (HEPI, 2026)

Combine the numbers. Most students use these tools, current models invent roughly one citation in five, and the invention rate climbs on exactly the niche topics theses cover. Unchecked reference lists now carry fabricated sources as a statistical default, not as bad luck.

Similarity software will not save you either, because a fabricated reference is original text. Our post on whether Turnitin checks citations explains that gap.

How to protect your draft

Treat every AI-suggested reference as unverified until a database says otherwise.

For a few sources, verify by hand: resolve the DOI at doi.org, search the exact title on Google Scholar, match every field. For a full reference list, batch the first pass with the free Check My Thesis citation checker, which cross-checks each entry against Semantic Scholar, OpenAlex, arXiv, PubMed, CrossRef, Google Books, DBLP, and Open Library and flags verified, hallucinated, and outdated references. For one suspicious source, the hallucination detector shows the strongest database match with the evidence behind it.

If fabricated sources already reached your draft, follow the cleanup workflow in ChatGPT fake citations: what to do. And when you choose tooling, our best citation verification tools guide compares the options.

FAQ

Does ChatGPT know when it invents a citation?

No. The model has no internal flag separating retrieved facts from composed text. Confidence in the output tells you nothing, which also means asking it "are these sources real?" often returns a confident yes for invented references.

Did newer models fix citation hallucination?

They reduced it. GPT-4 fabricated 18% in the 2023 Scientific Reports study, and GPT-4o fabricated 19.9% in the 2025 JMIR Mental Health study, with far higher rates on niche topics. Reduced is not fixed. Verify regardless of model.

Do browsing or research modes solve this?

Retrieval-backed modes ground answers in fetched pages, which lowers fabrication when they actually retrieve. You still verify, because models mix retrieved and composed content, and metadata errors survive retrieval. The checking workflow stays the same either way.

Is one fabricated citation really a big deal in a thesis?

Yes. A marker who finds one invented source rereads everything else with new eyes. Universities treat fabricated references as an integrity issue because you attested to sources you never read. The cost of checking is one evening. The cost of one fake is your credibility.

Practical takeaway

ChatGPT composes citations the way it composes everything: by plausibility, with no existence check. The measured fabrication rate on current models sits near one in five and climbs on niche topics like yours.

So never let an AI-suggested reference into your bibliography unverified. Run the free citation checker over your list, open what it flags, and read what you cite.

Related reading