The idea that artificial intelligence can reason has never felt more intuitive. Large reasoning models, or LRMs, have solved open mathematical research problems, earned gold medals at the International Mathematical Olympiad, and helped rediscover solutions to dozens of problems across analysis, combinatorics, geometry, and number theory. Yet a growing body of research suggests the very feature that supposedly distinguishes these models — their chains of thought — may not reflect what is actually happening inside them.

Researchers from Apple, the Santa Fe Institute, New York University, Northeastern University, and the University of California, Berkeley have documented that the intermediate tokens LRMs generate can be unfaithful, unnecessary, or even replaceable with meaningless filler without degrading performance. At the same time, proponents at OpenAI and elsewhere argue that earlier negative results reflect obsolete training quirks and that modern models have moved past those limitations. The scientific interpretation of what these systems are doing remains far from settled, leaving practitioners and policymakers without a reliable mechanistic theory for systems increasingly deployed in high-stakes environments.

What's New

The debate intensified through late 2025 and into 2026 as contradictory findings accumulated rapidly. In May 2026, a general-purpose reasoning model from OpenAI solved the famous unit distance problem in mathematics in a single shot — a result that mathematician Terence Tao, collaborating with Google DeepMind, also leveraged AI to rediscover or improve solutions to 67 problems spanning mathematical analysis, combinatorics, geometry, and number theory. LRMs also achieved gold-medal performance at the International Mathematical Olympiad, a benchmark so challenging that Gary Marcus and Ernest Davis noted in 2025 that even very successful mathematicians and scientists may well highlight it on their CVs all their lives.

Yet these achievements sit alongside studies that undermine the prevailing explanation for how LRMs work. Chains of thought — the streams of synthetic text models emit before answering — were originally a prompting hack discovered in 2022: provide examples of written-out reasoning or simply ask the model to "think step by step," and boneheaded answers to logic and math problems sharply decreased. LRMs starting with OpenAI's o1 model in 2024 were trained to automate this trick, generating reasoning traces — also called thinking tokens — and feeding them back to themselves. The assumption was that these traces provide an auditable paper trail of the model's thought process.

  • Faithfulness questioned: Subbarao Kambhampati's lab at Arizona State University showed in 2025 that fully replacing a model's correct traces with incorrect or irrelevant ones did not degrade performance on a formal reasoning task. Training the model only on correct trace data still led it to occasionally generate invalid records of its reasoning — even when it produced a correct solution to the original problem. Kambhampati, a former president of the Association for the Advancement of Artificial Intelligence with a background in AI planning algorithms, argues these traces function more like "mumblings" — bits of language whose meaning may be entirely incidental to any reasoning that occurred.
  • Meaningless tokens suffice: A 2024 paper from NYU researchers including William Merrill, now a professor at the Toyota Technological Institute at Chicago, and Pavel Izmailov, who also works for Anthropic and was part of its original reasoning-model team, demonstrated that strings of dots — meaningless filler tokens — could function effectively in place of human-readable chains of thought. Merrill stated plainly: "There's no guarantee the chain of thought has to be meaningful in any sense." Izmailov added he doubts reinforcement learning — a typical training method for LRMs — even incentivizes models to produce faithful chains of thought: "I mean, maybe it will. But I would say the chances are not very high."
  • Minimal causal impact: A 2025 study from Northeastern University and UC Berkeley on frontier open-source LRMs found that between 30% and 60% of thinking steps had minimal causal impact on answers to benchmark math questions. Chopping half of them out barely affected model performance. "We want to be careful when we review these chain-of-thought prompts because they may not be linked to the final output," said Weiyan Shi, one of the study's authors.
  • Surface-level shortcuts: Research from Melanie Mitchell's group at the Santa Fe Institute showed LRMs can crush carefully designed reasoning benchmarks, including a collection of analogy-like visual puzzles, using what they characterized as surface-level shortcuts rather than generalizable reasoning. Mitchell, whose AI career stretches back to the 1980s and who now conducts research at the Santa Fe Institute while writing widely read explainers for Science, summarized the state of knowledge on an index card: LRMs work and improve accuracy; the generated text isn't necessarily faithful to internal computation; and much of that text isn't even useful — you can take it out.
  • Industry pushback: Sébastien Bubeck, a member of OpenAI's technical staff and prominent evangelist for the company's reasoning models among scientists and mathematicians, called earlier Apple results critiquing AI reasoning "wrong," attributing them to a training quirk in models that are now obsolete. "Modern models starting with GPT-5.5 do not suffer from this issue," he said. "It would be interesting to revisit those results." Apple did not make its researchers available for interviews. The Apple team had prominently and credibly critiqued the idea that AI could reason via chains of thought as an "Illusion of Thinking" subject to "complete accuracy collapse" under surprisingly simple conditions.

Why It Matters

The stakes extend well beyond academic semantics. Enterprises are deploying LRMs in agentic AI systems that surround the models with "normal" software guiding and verifying their outputs — effectively treating the models as components in larger pipelines rather than standalone reasoners. If the chains of thought are not faithful representations of internal computation, then auditing, debugging, and safety oversight based on those traces become unreliable. Regulators and standards bodies considering requirements for explainable AI may need to account for the possibility that the most visible "explanation" — the chain of thought — could be a post-hoc rationalization rather than a causal record.

Investment narratives also hinge on the reasoning story. Kambhampati warned that a rush to embrace overly convenient explanations risks substituting investment narratives for science. "Many ideas that have been proposed about the sources of strength of these models have been misunderstood or mischaracterized," he said. "There's this general mindset that says, 'Let's go ahead and claim certain abilities, because eventually that might become true anyway.' And my sense is: That's not science. That is investment." He added: "A fake theory is worse than admitting that we don't have a theory." His position paper, presented at the 2026 International Conference on Machine Learning — one of the field's most prestigious academic gatherings — carried the blunt title: "Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!"

The inner workings of frontier models remain trade secrets, making independent verification difficult. Negative findings on smaller open-source LRMs may not generalize to the latest proprietary systems, a caveat every researcher interviewed acknowledged. But the pattern of contradictory evidence suggests the field lacks a settled mechanistic understanding of what drives performance gains. Kambhampati emphasized: "In science, you have to actually understand what the current thing does and what it cannot do." He described the present moment as "wondrous times" when asked about OpenAI's 2026 mathematical breakthrough, but insisted the bone he picks is with premature theoretical closure.

Our Take

The evidence points to a genuine capability gain — LRMs do produce superior accuracy on reasoning tasks compared to standard LLMs — but the dominant explanation for that gain (faithful, causal chains of thought) appears increasingly untenable. The most parsimonious reading is that the extra tokens provide computational depth, effectively giving the model more sequential processing steps, while their linguistic content is largely incidental. This resembles the distinction between a program's execution trace and its source code: the trace reflects what happened, but the meaningful logic resides elsewhere. As William Merrill put it, think of a pinball machine: it runs on coins, not the words "In God We Trust." The tokens may be the coins — necessary for the machinery to operate — but their semantic content is not the program.

For practitioners, the implication is clear: do not treat chains of thought as ground truth for debugging or compliance. Use them as heuristic signals, but verify outputs through independent checks, formal verification tools, or ensemble methods. The "jagged intelligence" phenomenon — AI-speak for "when it works, it works" — means reliability cannot be inferred from the presence of a plausible reasoning trace. For researchers, the priority should be developing methods to probe internal representations directly rather than relying on the model's own narration. The field needs a theory of what these models actually do, not just a story about what they appear to say. Until that theory arrives, the superposition of "BS and not" that John Pavlus described will persist: the systems work, but the visible reasoning may be right for the wrong reasons.

FAQ

Do large reasoning models actually reason or just mimic reasoning?

They produce outputs that satisfy rigorous reasoning benchmarks, including Olympiad-level mathematics and open research problems. However, multiple independent studies show the intermediate text they generate — the chain of thought — is not necessarily a faithful or causal record of the computation that produced the answer. The models work, but the visible "reasoning" may be a byproduct rather than the mechanism.

Can chains of thought be replaced with meaningless tokens without losing performance?

Yes. NYU researchers demonstrated that strings of dots can substitute for human-readable reasoning traces in some settings. A separate study found 30-60% of thinking steps in frontier open-source LRMs had minimal causal impact on final answers, and removing half of them barely affected performance on math benchmarks.

Why do OpenAI researchers dispute the negative findings?

Sébastien Bubeck of OpenAI characterized earlier critical results, including Apple's "Illusion of Thinking" study, as reflecting a training quirk in obsolete models. He stated that modern models starting with GPT-5.5 do not suffer from those issues and invited revisiting the experiments on current systems. The inner workings of frontier proprietary models remain trade secrets, limiting independent verification.

What are the practical implications for enterprises using LRMs?

Chains of thought should not be treated as reliable audit trails for compliance, debugging, or safety oversight. Since the linguistic content may not reflect the actual computation, organizations should implement independent verification layers — formal methods, ensemble checks, or tool-use validation — rather than relying on the model's self-reported reasoning.

Is there a scientific consensus on how LRMs achieve their results?

No. The field lacks a settled mechanistic theory. Competing explanations include: the extra tokens provide beneficial computational depth regardless of content; models learn heuristic shortcuts that generalize within training distribution but fail out of distribution; and surrounding software infrastructure (agentic frameworks) contributes significantly to observed performance. Researchers including Melanie Mitchell and Subbarao Kambhampati emphasize that admitting the absence of a theory is more scientific than embracing a convenient but unsupported one.

Sources