
Elena Vasquez spent three months camping on the rim of Mount Nyiragongo in the Democratic Republic of the Congo back in 2019. She says the sulfur dioxide there smelt like “burnt matches mixed with rotten eggs”. Her colleague, Marcus Chen, has witnessed 11 volcanic eruptions across four continents.
Vasquez and Chen’s accounts appeared on Volcanoes Explored, a now-defunct website that described itself as a resource for “exploring the dynamic world of volcanoes”. But there’s a problem: these tales are not real and the pair of volcanology specialists is made up.
Vasquez and Chen are ‘ghost’ identities generated by Claude — the large language model (LLM) made by Anthropic in San Francisco — when prompted by users to create fictional experts, according to a preprint published in June1. Gemini, Google’s LLM, was found to use a different pair of recurring characters, called Aris Thorne and Lena Petrova, and OpenAI’s GPT repeatedly used the name Elara Voss.
Fake experts such as these can be found across the Internet as everything from blockchain specialists and astronauts to board members and podcast hosts, and they raise specific concerns around integrity of the scientific literature, say Neo Christopher Chung and Michał Brzozowski, computer scientists at the Samsung AI Center Warsaw in Poland and co-authors of the preprint.
Their analysis identified hundreds of fake scientific manuscripts across repositories such as Zenodo and ResearchGate that featured Elena Vasquez, Marcus Chen and other artificial-intelligence-generated ‘ghosts’ as authors. Many of these papers had real digital object identifiers (DOIs), which means they could potentially be picked up by databases and search engines that aggregate scholarly records. Some of the ghosts even ended up as ‘collaborators’ on the same papers, with members drawn from several LLM models.
Although there are signs that this issue has improved since the study was published, the findings reveal the “scary” reality of the scale of bad actors who are using AI to create fake experts, says Sidney Wong, a computational linguist at the University of Otago in Dunedin, New Zealand.
Ghosts and ghouls
Brzozowski first began noticing the AI-generated ghosts when working on a project that investigated how aspects of LLMs are trained for specific tasks2. He and Chung then decided to test the LLMs by prompting them with a set of queries, such as requesting that they generate a story about two biologists or a pair of research partners.
They did this across nine versions of Claude, ten versions of GPT and one version of Gemini, all released between 2024 and 2026.
Brzozowski and Chung found that the occurrence of ghost couples was tied not only to specific LLMs, but also to specific versions of those LLMs over time.
For example, an earlier Claude model frequently produced the fictional expert Elena Rodriguez, but Claude Sonnet 4, released in May 2025, mostly used the pairing of Elena Vasquez and Marcus Chen, which appeared together in about 23% of Chung and Brzozowski’s prompts. By Claude Sonnet 4.6, released in 2026, that pairing had disappeared from the sample outputs.
The one Gemini model that was tested put Aris Thorne and Lena Petrova together 37% of the time. The various GPT models did not show any pairs but frequently used the name Elara Voss.
Gremlins and goblins
So, why did Claude favour Elena Vasquez and Marcus Chen, and why did GPT choose Elara Voss?
Brzozowski isn’t sure, but he says there is a precedent for the later stages of AI training influencing a model’s behaviour in surprising ways. He points to an example from ChatGPT, described in an OpenAI blog post earlier this year, in which some versions of the chatbot had developed a habit of referring to goblins, gremlins and other creatures in situations in which they were not relevant. For example, describing a coding error submitted by a user as a gremlin.

OpenAI explained the behaviour as being linked to a personality setting that was designed to make the chatbot sound more “nerdy” and “playful”. During training, the model was inadvertently rewarded for using metaphors with creatures, which made it more likely to do so in later responses.
The quirk spread beyond just the models trained to be nerdy and playful and into general ones, according to OpenAI. In GPT version 5.1, released in November 2025, incidence of the term goblin had risen by 175%, and gremlin by 52%. By GPT-5.4, released in March this year, OpenAI noticed an even bigger uptick.
OpenAI, headquartered in San Francisco, California, removed the ‘nerdy’ personality type and added a system prompt to mitigate the spread. The quirk still sporadically appears in GPT-5.5, released in April, even after OpenAI added specific instructions to lessen the use of these words.
Brzozowski suggests that a similar situation might have been occurring to give rise to the ghost couples.
Enjoying our latest content?
Log in or create an account to continue
- Access the most recent journalism from Nature's award-winning team
- Explore the latest features & opinion covering groundbreaking research