Wednesday, September 9, 2026

It's like something from a science fiction story. (No, it's literally* something from a science fiction novel)

[* Can you use literally to describe something being derived from something even if the result is not a direct quote? Well, I just did so apparently the answer is yes.]

This makes perfect sense when you think about it. Large language models are designed to generate plausible strings of text with respect to their training data. Think about all the times you've heard chatbots like Claude or ChatGPT compared to science fiction or described in sci-fi terms. If you ask an LLM what it's "thinking," it is not at all surprising that it would reference the sci-fi books and plot descriptions in its training data when coming up with a reasonable-sounding answer, particularly when you take into account the noise that comes with the feedback loops of agentic systems. 

From Cal Newport: [emphasis added]

OpenAI was looking at these reasoning traces to find unnerving examples of plotting behavior from its “swarm.” This seems like a natural thing to do, but there are two problems with this approach:

  1. These chain-of-thought traces don’t necessarily reflect the actual logic behind an LLM’s ultimate answer or suggestion. Multiple studies have shown that these models sometimes invent reasoning that sounds plausible, but may be completely unrelated to how they arrived at the response. (See, for example, ​this paper​ from ICML 2026, or ​this paper​ from NeurIPS 2023.)
  2. Research has also shown that referencing the fact that an LLM is an AI system in a prompt increases the chances that the LLM’s output will reflect sci-fi style narratives about AI running amok. Because it was trained on many such stories, the model assumes that this is the type of output it’s supposed to produce. If you take sci-fi tales out of a model’s training set, it’s less likely to talk in terms of AI running amok. (See, for example, ​this study​.)

Put these two observations together, and it’s clear that it borders on research malpractice to soberly report on carefully curated clips from these traces to imply that somehow the combination of these prompt loops and the LLM they are prompting is a unified sentient entity with malicious intent. It’s more likely that the LLM in question is simply post-hoc rationalizing its outputs with well-worn tropes it encountered during training.

Though much more difficult to study experimentally, I wonder if this might explain how frequently sci-fi and fantasy elements show up in cases of AI psychosis. 

No comments:

Post a Comment