
I know something about AI hallucinations. I know how convincing they can look, how easily an invented citation can pass as real, and how much damage can be done before anyone discovers the truth.
That is why a new investigation by Guardian Australia got my attention. The Guardian found that Australia’s parliamentary inquiry system is being “flooded” with submissions containing invented studies, nonexistent academic papers, distorted research findings, and citations that look real but lead nowhere.
At least 39 submissions to parliamentary inquiries contained apparent hallucinated references. More than 100 papers included ChatGPT tags in their links, evidence that the platform had been used somewhere in the research process. Some documents contained only a few incorrect citations. In others, every cited reference appeared to be invented.
These submissions were sent to parliament to influence government policy.
Christian Downie, a professor at the Australian National University, described the central danger: Parliament could end up making decisions based on “evidence that doesn’t exist.”
One submission about family violence and suicide cited a nonexistent paper attributed to Divna Haslam, a clinical psychologist and associate professor at the University of Queensland. It also misstated her team’s actual findings. Haslam called the fabricated reference “scary,” in part because it looked so convincing.
Then something even stranger happened. Google’s AI summary described the imaginary paper as if it were real.
Think about that process for a moment. An AI system invents a source. Someone places that source in an official parliamentary submission. Google finds the submission and treats it as evidence that the source exists. ChatGPT or another AI system can then find the same document and repeat the claim. The Guardian calls it an “ongoing cycle” of misinformation.
A hallucination now has a source. The source is wrong, but it looks official. The next person searching for it may have no idea where the fiction began.
The organization responsible for the family-violence submission acknowledged using AI but argued that the incident was ultimately a human error because the wrong version of the document was uploaded. Of course humans are responsible for what they submit and publish. I believe that deeply and have had to confront it publicly in my own work.
But saying a human should have caught the error does not explain where the error came from. A person did not mistakenly type the wrong page number or misspell an author’s name. AI invented a study, created a plausible citation, attributed it to a real academic, and misstated her research. The human failure was not catching what the machine had fabricated.
That distinction matters. Technology companies often describe hallucinations as if they are minor imperfections that become consequential only when careless people fail to check the results. But these systems are remarkably good at creating false information that looks authoritative. Names, titles, links, journals, dates, quotations, and academic citations can all be manufactured in seconds.
I have seen this firsthand. When AI-generated citations and false quotations entered my own work, they did not look imaginary. They arrived with the surface details of legitimate research. I was responsible for failing to catch them. But accepting that responsibility does not require me to pretend that AI played no role in creating them.
The Australian findings also show that this is not confined to one ideology or political group. The questionable submissions came from people and organizations across the political spectrum. This is not fundamentally a partisan problem. It is a breakdown in how we establish whether evidence is real.
The scale is new. The Guardian reports that large language models have allowed false citations to spread at “unprecedented levels.” Producing a polished, research-heavy policy document once required research. Now anyone can ask an AI system to construct an argument and provide the supporting evidence. The bibliography may look impressive even when none of it exists.
The problem is already moving beyond parliamentary submissions. Deloitte reportedly refunded part of its fee for a $440,000 report commissioned by the Australian government after the document was found to contain AI-generated errors, including fake academic references and an invented court citation.
These mistakes can affect decisions about domestic violence, housing, healthcare, climate, immigration, and public safety. As Haslam told the Guardian, “We just can’t risk that.”
OpenAI’s response was that people should use ChatGPT as “a first draft, not a final source.” That is good advice, but it lets the technology off too easily. A first draft can be rough. It can be incomplete. It should not invent its evidence and then present that evidence with the confidence and formatting of genuine scholarship.
We do not need to ban AI from government, research, or public life. But we do need rules that recognize what these systems can do. If generative AI is used to prepare a parliamentary or government submission, that use should be disclosed. Every citation needs to be checked against the original source. Parliamentary committees and government agencies need a process for identifying fabricated references before they enter official reports.
Technology companies also have a responsibility here. They cannot continue building products that create plausible falsehoods and then place the full burden of detection on the user. If a system has retrieved and verified a source, it should say so. If it is generating an answer based on probability, that should be unmistakably clear.
A citation has always been more than decoration. It is a promise that the evidence exists and that someone else can find it, read it, and challenge it. AI hallucinations preserve the appearance of that promise while quietly breaking it.
Australia’s parliament has discovered that fabricated evidence is entering the machinery of government. It will not be the last institution to face this problem. Courts, universities, newsrooms, publishers, corporations, and legislatures everywhere are using the same tools and confronting the same risk.
The question is no longer whether AI hallucinates. We know that it does. The question is how much invented evidence has already entered the public record, and whether our institutions are prepared to find it before decisions are made.