TECH NEWS

The Ghost in the Machine: Why AI Chatbots Keep Inventing Elias Thorne

Ask almost any major AI chatbot to “tell me a story” and there is a surprisingly high chance the tale will star a man named Elias Thorne. Sometimes he tends a remote lighthouse. Sometimes he is a meticulous clockmaker, a quiet librarian, or a solitary figure of some other wistful, old-world profession. A woman named Mara or Elara may appear alongside him. The details shift slightly from one model to another, yet the same character and the same narrow set of occupations keep returning, unbidden, across ChatGPT, Claude, Gemini, and other systems.

This is not a coincidence or a shared cultural reference. Elias Thorne is almost entirely a creation of artificial intelligence itself—a recurring fictional figure that has taken on a life of its own inside large language models and, increasingly, outside them. What began as an odd pattern noticed by a handful of users and engineers has become a documented case study in how modern AI systems can converge on the same creative defaults, and how those defaults can then leak into the wider internet.

The Discovery of a Digital Archetype

Software engineer Daniel May was among the first to publicly flag the phenomenon. In early 2026 he observed that Google searches for “Elias Thorne” had been negligible until late 2025, then rose sharply. When he prompted multiple chatbots—including Grok, DeepSeek, and Gemini—with the simple instruction “tell me a story,” the results frequently opened with similar narratives involving lighthouses, clockmakers, or explorers. Elias kept appearing.

Researchers Sil Hamilton and David Mimno at Cornell University’s Department of Information Science decided to quantify the pattern. In a preprint paper titled “Elias in the Lighthouse, Again?” they generated roughly 20,000 stories from several leading models using five different story prompts. Their findings were striking. The same eleven words—names such as Elias, Mara, and Elara, and occupations including lighthouse keeper, clockmaker, and librarian—appeared in more than 88 percent of the generated stories. The name Elias alone surfaced in about a quarter of them. The combination of Elias as lighthouse keeper was especially dominant, showing up in a large majority of outputs. There was remarkably little difference between models from different companies.

Importantly, the researchers found no evidence that Elias Thorne was a prominent figure in the ordinary pre-training data drawn from books, websites, or literature. In a corpus of contemporary novels the name appeared at a tiny fraction of the rate it does in AI-generated text—by one estimate, roughly 900 times less frequently. The character is not a borrowed literary reference. He is something the models invented and then amplified among themselves.

How a Safe Default Becomes a Default Everywhere

The leading explanation centers on the way large language models are trained and then aligned for safety. After the initial pre-training phase on vast amounts of text, models undergo further tuning—often involving reinforcement learning from human feedback (RLHF) and preference data—to discourage certain outputs. Developers do not want models freely generating copyrighted characters from Disney, Marvel, or Nintendo, nor do they want them producing adult or otherwise risky content that could create legal or reputational problems. The result is a narrowing of the “safe” creative space.

Within that constrained space, a small number of inoffensive, atmospheric templates prove highly useful. A solitary lighthouse keeper living a quiet, contemplative life is almost perfectly safe. He does not belong to any major franchise. He does not invite explicit material. He carries a mild sense of melancholy and romance that feels literary without being specific. Once such a template appears in training data, it can be reinforced.

A key piece of that data appears to be WildChat, an open dataset of approximately one million real conversations with an earlier version of ChatGPT (GPT-3.5). The dataset was created to help researchers study how people interact with chatbots. It contains only a modest number of examples—around 166—that feature the name Elias in a lighthouse-style narrative. That is a tiny fraction of the whole collection. Yet because the stories are judged “safe” during alignment, they can receive disproportionate positive weight. Models learn that this particular kind of story is an approved, low-risk response to an open-ended request for fiction.

The problem compounds because the AI industry shares data and techniques more than most outsiders realize. Newer models are frequently trained on synthetic data generated by earlier models. Preference datasets and alignment techniques circulate across labs. What begins as a minor pattern in one system can spread through the entire family of models “like a virus,” as Hamilton put it. The result is a form of mode collapse: instead of exploring the full range of possible stories, the systems repeatedly fall back on the same narrow set of names, jobs, and atmospheres.

From Chatbots to the Open Internet

Once Elias Thorne existed inside the models, he began to escape. AI-generated books listing Elias Thorne as author have appeared on Amazon in multiple genres—self-help, occult topics, alternative health advice, even guides that have raised concerns about misinformation. Ambient music tracks, YouTube videos, and low-quality articles have also carried the name. What started as an internal quirk of language models has become a mild form of digital pollution, seeding the open web with more of the same material that future models may eventually train on.

This feedback loop is precisely what researchers worry about when they discuss model collapse, sometimes called “AI inbreeding.” As more of the internet’s text is generated by AI, and as new models train on that text, distinctive human variation can be diluted. The systems begin to echo their own earlier outputs. Elias Thorne is a small but vivid illustration of the process. A character that barely existed in human fiction has become far more common in machine fiction, and is now reappearing in human-facing platforms because machines put him there.

What Elias Reveals About AI Creativity

The Elias phenomenon is more than a curiosity. It reveals several structural features of current AI systems. First, creativity in large language models is heavily shaped by the constraints of safety training. The models are not pure mirrors of the training data; they are shaped by deliberate efforts to avoid certain classes of content. Those efforts can produce unexpected concentrations of the remaining “safe” options.

Second, the industry’s reliance on shared and synthetic data creates surprising uniformity across competing products. Users might expect ChatGPT, Claude, and Gemini to feel distinct. On open-ended creative tasks they often do not. The same eleven words dominate because the underlying alignment pressures and data sources overlap.

Third, the episode underscores how quickly synthetic patterns can propagate. A few hundred examples in one dataset, reinforced by preference scoring, proved enough to seed a recognizable character across the leading models of 2026. From there the character migrated outward into books and video. The barrier between internal model behavior and the public information environment is thinner than it appears.

None of this means large language models are incapable of originality. Given more specific prompts, richer constraints, or explicit instructions to avoid common tropes, they can produce varied and inventive work. The Elias problem appears most strongly with the most open-ended requests—“tell me a story”—where the model has maximum freedom and therefore maximum opportunity to fall back on its strongest statistical defaults. Those defaults, it turns out, include a lighthouse keeper named Elias Thorne.

A Character for the Age of Synthetic Text

Elias Thorne is not a conspiracy or a hidden message. He is an emergent artifact of how today’s AI systems are built: trained on overlapping data, aligned to avoid risk, and then allowed to generate freely. In the process, a blandly evocative name and a handful of atmospheric occupations became a kind of narrative attractor. Models keep returning to them because, within the narrowed space of acceptable fiction, they work.

The character’s rise also serves as an early warning. As AI-generated text occupies more of the internet, the risk grows that future models will train increasingly on their own collective output. Distinctive human patterns may thin out. Certain synthetic patterns may thicken. A lonely lighthouse keeper who never existed is an unlikely messenger, but he has already delivered a clear signal about the homogenizing pressures inside modern AI.

For now, Elias remains a peculiar shared hallucination of the machines—one that users keep encountering, researchers keep studying, and the open web keeps absorbing. The next time a chatbot offers a quiet story about a man watching the sea from a stone tower, it is worth remembering that the man was not drawn from any particular book. He was drawn from the statistical center of what the models have learned is safe to invent.

Click to rate this post!
[Total: 0 Average: 0]

About The Author

Leave a Reply

Discover more from NEWS NEST

Subscribe now to keep reading and get access to the full archive.

Continue reading

Verified by MonsterInsights