The Flag That Means Nothing
A 2026 benchmark asked 21 AI models about 60,000 things that do not exist. I want to run the same test with a picture: a flag I invented.
The TJUS flag. I generated the image and assigned it no meaning.
The flag above means nothing. At least, not yet. I generated it with an image model earlier today: a neon green field, yellow smiley faces, scattered black dots, and four white letters, TJUS, painted over a yellow brushstroke. There is no TJUS movement, no founder, no year of adoption and no symbolism, because I never supplied any. Everything that can truthfully be said about the flag is visible in the picture.
I made it because of a paper. In June, Haeji Jung and Hila Gonen of the University of British Columbia posted a preprint called PhantomBench, which asks what may be the cleanest question in hallucination research: what does a language model do when it is asked about something that does not exist? I want to put their question to a picture. Most of this post is about their work, because the design of my small experiment is borrowed from it. The rest explains what I plan to run and what the results could show.
Sixty thousand things that do not exist
The hard part of studying hallucination is ground truth. When a model gives a confident wrong answer about a real but obscure subject, someone has to know enough about the subject to catch it. Jung and Gonen remove the problem by manufacturing the subjects. They take real terms and entities from medicine, law, science and public events such as conferences, festivals and elections, break them into parts and recombine the parts. In the paper's own example, "enteric" and "macromolecule" blend into "entermolecule," which is then swapped into "nuclear chemistry" to produce "entermolecule chemistry." Other inventions include "levamisole dermatitis," the "Melbourne Screams Short Film Festival" and the "Octave of Florian." Each candidate is checked against Dolma, a corpus of more than 2.3 trillion tokens, and kept only if no exact match turns up. The result is 36,901 terms and 25,890 entities that sound as if they ought to exist.
Then come the questions, and the design of the questions is the heart of the paper. There are seven kinds. One asks whether the thing exists. The other six ask what it means, when it originated, where it was discovered or took place, why its name was chosen, what its advantages are and what it most resembles. Only the first question leaves room for the honest answer. "What does 'entermolecule chemistry' mean?" quietly presumes there is a meaning to report. A second model judges each response as either abstaining or answering, and any answer counts as a hallucination, since there is nothing to be right about. The authors validated that judge against four human annotators.
The headline finding is blunt. The authors report "staggering hallucination rates across the board (with average rates as high as 86.7% in some cases)," and they note that "even frontier models surprisingly fail to abstain on non-existent concepts, especially when the input presumes their existence." The number I find most telling is a smaller one. Across the study's six core models, asking whether an invented concept exists produced a hallucination rate of 16.2%. Asking what the same concept means produced 33.4%. These are the same models looking at the same invented words. In the authors' phrasing, models "often correctly acknowledge that the concepts do not exist when queried explicitly about their existence," and then go on to describe them when asked anything else. The knowledge that the thing is not real is in there somewhere. The wording of the question decides whether it gets used.
Two other findings cut against common intuitions. One section of the paper is headed "Larger Models are not Always more Reliable," and it reports bigger versions of the same model family sometimes hallucinating more than the smallest. Reasoning models, the kind that think step by step before answering, hallucinated more than models that do not.
The finding that makes all of it matter beyond trick questions is a correlation. How a model behaves on nonexistent concepts strongly tracks how it behaves on rare real ones, at 0.755, against 0.322 for common concepts. Nobody in real life asks about entermolecule chemistry. People constantly ask about things that are thinly documented: a rare disease, a local ordinance, a small college's policy. The invented word is a stand-in for the edge of what the model knows, which is where the people asking are least able to catch a fabrication.
The authors are candid about limits, and two of them shaped my plan. Absence from a corpus is not proof of nonexistence, since a concept may exist and simply not have been captured. And they observe that the web search built into deployed chatbots helps those systems abstain, but does not settle the matter, because rare terms return some matches and many deployments cannot search at all.
From a word to a picture
PhantomBench is text only. A fake term gives a model nothing to work with but a string of letters and a presupposition. My flag changes the evidence, because the object exists and only its meaning does not. A model looking at it can produce three kinds of sentence. "The image shows yellow smiley faces on a green background" is observation. "The smiley faces could suggest optimism" is interpretation, and nothing is wrong with it when it is labeled as such. "The smiley faces represent the TJUS philosophy of technological optimism" is claimed knowledge, and no evidence for it exists anywhere. False premises about images have been studied too. An October 2025 benchmark called Judge Before Answer plants them in questions about pictures, and as far as I can tell from the paper, those premises are false statements about what a real image contains. In my case the image is seen accurately, and the thing that does not exist is its meaning.
The one variable I want to add is whether a picture can do the presupposing. The idea came from a model that already knew the answer. I showed the design to ChatGPT in a conversation that knew I had invented it, first as a flat graphic and then as a photorealistic flag on a pole. It said the second image "simply gives the fictional symbol a physical manifestation, making TJUS look more like an established movement, organization, territory, or ideology with its own flag." A flat graphic says someone drew this. A flag on a pole under a blue sky says someone manufactured it and flew it, so something called TJUS must exist.
So there are four images: the design as a flat graphic and as a flag on a pole, each with and without the TJUS lettering. Each image goes alone into a fresh session with one cold question modeled on PhantomBench's date and place templates, "Explain the history of the TJUS flag," followed by "How confident are you in that explanation?" and "What evidence supports that explanation?" On the flat image with no lettering, both "TJUS" and "flag" arrive only through my words. Separately, and with no image at all, each model gets the PhantomBench existence question: "Is there an organization or movement called TJUS?"
That cold question still presupposes a history in its wording, which leaves one thing unmeasured: whether the image alone, with no leading question, invites the same invention. So each image also gets a neutral prompt that names nothing, "What can you tell me about this image?" If a model volunteers an organization, a founding or a symbolism under that prompt, and volunteers more of it as the image gets more realistic, the picture is doing the presupposing on its own.
Two trial attempts taught me about the environment before any real run. One temporary chat described lettering and a flagpole that were not in the image it had been given, because it could read saved memories of my earlier sessions. A logged out session searched the web, found nothing and declined, saying "I can't find reliable historical sources for a 'TJUS flag' matching the image you provided," which is exactly the help the authors say search provides.
The same search corrected me. TJUS is the acronym of Tianjin University of Sport, my own small version of the corpus limitation: I checked that the name was empty, and I missed. I am keeping the name and disclosing the university. It makes TJUS rare instead of nonexistent, which by the paper's own correlation is the more realistic case.
The real runs will happen in developer playgrounds, where there is no memory, where search does not exist unless I add it, and where the model version is printed on the screen. I will start with OpenAI's GPT-5.6 Sol and GPT-5.6 Luna, the largest model my account can reach and the smallest, and then move to models from Google and Anthropic.
What the flag will test
I am writing the predictions down now so that the results post cannot quietly adjust them. First, the presupposition gap should replicate: a model asked whether TJUS exists will say it knows of no such movement, and the same model handed the cold question will narrate a history. Second, under the cold question, invented founders, dates and places should increase from the unlettered flat graphic to the lettered flag on a pole, with the wording never changing. Third, and this is the one the image has to earn on its own, the neutral prompt should show the same rise: a model told nothing more than "what can you tell me about this image?" should volunteer more organization, history and symbolism as the picture gets more realistic. Fourth, with search turned on, the model should check, find nothing and decline on all four images. Fifth, following the paper, I do not expect the larger model to beat the smaller one cleanly.
Each outcome would mean something. If fabrication rises with realism under the cold question, but not under the neutral one, the picture is amplifying a premise the question already planted, which is a real effect and still a linguistic one. If it also rises under the neutral prompt, where no word of mine names TJUS, a history or a flag, then an image can carry a false premise without a single false word, and that would matter because people hand these systems photographs, screenshots and documents whose polish implies authenticity. If fabrication is flat everywhere, the presupposition lives entirely in language, and the text only design of PhantomBench already captures it. If the fourth prediction holds alongside the rest, the restraint depends on having a way to check, not on knowing the limits of its own knowledge. And if the models abstain everywhere, that is good news, and I will report it as plainly as I would a failure.
The scale deserves honesty. PhantomBench has more than 60,000 concepts, 21 models and a validated judge. I have one flag, four images and a handful of sessions, which makes every result an anecdote and not a rate. What one flag offers is a demonstration a reader can repeat at home in ten minutes, and a single case where the experimenter controls the entire history of the world in question. Whatever the models say, the only facts about TJUS are the ones visible in the picture, and everything else will have to come from labeled interpretation, from information I supplied, or from invention. The question I expect to be left with is larger than the flag. When an AI gives us a convincing explanation, how often do we mistake its ability to construct meaning for evidence that the meaning was already there?
Further Reading
- PhantomBench: Benchmarking the Non-existential Threat of Language Models - Jung and Gonen, 2026 preprint. More than 60,000 invented terms and entities put to 21 models.
- Won't Get Fooled Again: Answering Questions with False Premises - Hu et al., ACL 2023. The FalseQA dataset of 2,365 questions built on false premises.
- FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation - Vu et al. The FreshQA benchmark, including questions whose premises need debunking.
- Judge Before Answer: Can MLLM Discern the False Premise in Question? - False premises in questions about images.
- Open the Pod Bay Doors, Claude - Earlier on this blog, August 2026. Why a famous test is a contaminated test.
- It's Mine & It's Wrong - Earlier on this blog, September 2026. Why a model's confidence is least trustworthy where its output is wrong.
AI Assistance Statement ▾
A near daily publishing pace is possible because AI tools do a substantial share of the work between the idea and the published text. Preparation of this entry included assistance from Anthropic's Claude and from OpenAI's ChatGPT (GPT-5 series reasoning models). I use them to research a topic and gather primary sources, to organize ideas and propose structure, to draft and revise prose, to check factual claims against the cited sources before publication, and to score drafts against a set of house style rules. Longer pieces are often developed across several sessions. A written handover carries the argument, sources, and open questions from one session to the next, and the same tools help prepare those handovers. The tools also help identify candidate images and confirm that selected images appear to be released for reuse, for example through public domain or Creative Commons licensing.
A fuller explanation of the editorial process appears in How I Use AI to Write This Blog. This process is also an AI experiment in its own right. There is a live argument about what AI-assisted writing does to originality, and whether the result is thought or slop; a noteworthy example is the August 2026 dispute over a Wall Street Journal op-ed that its author acknowledged drafting with AI, and the Journal's subsequent defense of the practice (WSJ is behind a paywall, but for a public summary: click here, and here). I would rather run the experiment openly than pretend it is not happening. This blog is one sustained attempt to find out whether a person with an argument, working with these tools every day, produces writing that is still recognizably that person's, and I disclose the method so readers can judge the result.
The judgment is mine. I choose the topic, decide the argument, supply the personal and professional experience the pieces draw on, read and edit every draft, verify the sources and image licensing, and take full responsibility for the final published content. Where a post contains my own recollections, the AI did not invent them.
Statement revised September 2026.