Picture Albert Einstein alone with a thought experiment. He imagines himself in a sealed elevator, free-falling through space, and from that quiet mental image he teases out the equations that rewrite physics. That moment—pure intuition, sensory imagination, a leap with almost no data—is the kind of creative jump current language models struggle to mimic.
Why a sea of text is not the same as a lived world
DeepMind scientist Tom Zahavy put this bluntly in a recent analysis titled 'language models cannot jump'. His argument slices thinking into three modes: induction, deduction, and hypothesis generation. Modern models excel at the first. They compress mountains of patterns and can predict the next sentence with uncanny accuracy. Advanced systems also handle deduction decently; automated reasoning tools, including projects like AlphaProof, show machines can follow logical chains when rules are given.
But when it comes to hypothesis generation—the act of imagining a new framework from almost no empirical input—machines stumble. Why? Because they lack a body in the world, and they lack the sensory intuition that humans use to forge causal links. You cannot derive Einstein's initial premises purely by analyzing existing texts. His insight emerged from embodied metaphors and sensory thought experiments, not from statistical pattern-matching across literature.

Some thinkers, such as Jürgen Schmidhuber, once framed scientific discovery as an extreme form of data compression. DeepMind challenges that position. The paper points to how revolutionary theory often arises in data-poor contexts, reliant on counterfactual scenarios and vivid imaginative synthesis. Language models, trained on retrospective patterns, rarely invent new ontologies that reframe how the world is understood.
There are practical consequences. Current AI coding and automated discovery tools remain bounded by predefined frameworks. They can optimize inside a box. They struggle to redraw the box itself. And no amount of extra compute, bigger data centers, or denser training will automatically confer the kind of causal intuition that underpins radical scientific leaps.
So what does this mean for research and industry? It reframes expectations. Language models are powerful collaborators for pattern-heavy tasks, drafting, hypothesis refinement, and exploring permutations. They augment, rather than replace, the human spark that invents the questions and the conceptual tools. In other words: great for scaffolding. Not yet primed to originate revolutions.
Computational scale is not a substitute for embodied insight—human imagination still makes the leap.
There is an open path forward. Hybrid approaches that tie models to sensors, experiments, and interactive environments could shrink the gap. But until machines acquire forms of situated understanding and causal intuition beyond text, the boldest scientific ideas will remain the province of human minds.




Discussion
Leave a Comment
Comments
No comments yet. Be the first.