LLMs can’t jump

The fundamental limitation of modern artificial intelligence lies not in its lack of processing power or data access, but in its lack of a physical presence within the world. In a comprehensive research paper titled “LLMs can’t jump,” Tom Zahavy of Google DeepMind explores the cognitive barriers inherent in software-based intelligence that operates without a biological or mechanical body. The research suggests that while Large Language Models (LLMs) have achieved unprecedented success in mimicking human language and logic, they remain fundamentally incapable of replacing human creative genius. This deficiency stems from the fact that current AI architectures only master specific types of inference, leaving them unable to bridge the gap between statistical prediction and true scientific invention.

The Triad of Inference: Induction, Deduction, and Abduction

To understand the limitations identified by DeepMind, one must look at the three primary modes of logical reasoning: induction, deduction, and abduction. The paper argues that for AI to achieve a level of intelligence comparable to human genius, it must master all three. However, current systems are heavily skewed toward the first two.

Induction involves identifying patterns within large datasets to make generalizations. Thanks to the massive scale of modern internet data and sophisticated statistical pattern recognition, LLMs have arguably mastered induction. They can predict the next word in a sequence or identify stylistic trends across millions of documents with remarkable accuracy. This is the foundation of generative AI’s ability to write essays, compose poetry, and generate code.

Deduction, the process of reaching a logical conclusion from established rules or premises, has also seen significant progress. DeepMind’s own AlphaProof and AlphaGeometry systems have demonstrated that AI can solve complex mathematical proofs by strictly adhering to logical frameworks. In these controlled environments, where the "rules of the game" are clearly defined, AI can outperform human experts by processing permutations at a speed impossible for the human brain.

The "impenetrable ceiling," according to Zahavy, is abduction. Abductive reasoning is the ability to generate the most likely explanatory hypothesis when faced with incomplete or scarce data. It is the "leap" of intuition that allows a scientist or thinker to propose a new law of nature when existing data is insufficient to prove it. While AI can interpolate within a dataset (induction) and follow rules to a conclusion (deduction), it struggles to extrapolate entirely new frameworks from a vacuum of information.

The Einstein Challenge: Beyond Data Compression

The DeepMind paper enters a long-standing debate in the field of computer science regarding the nature of intelligence. Jürgen Schmidhuber, a pioneer in the field of neural networks, has famously argued that scientific discovery is essentially an advanced form of data compression. According to this view, the more efficiently a system can compress information, the better it understands the underlying laws of the universe.

Zahavy refutes this assessment by using Albert Einstein’s development of General Relativity as a primary counterexample. When Einstein formulated his breakthroughs in the early 20th century, there was virtually no visual or empirical data regarding the warping of spacetime. He did not arrive at his conclusions by analyzing vast datasets or compressing existing astronomical observations. Instead, Einstein relied on what Zahavy calls “embodied thought experiments.”

Einstein’s most famous mental exercise involved imagining a person falling inside a sealed elevator. By visualizing the physical sensations of weightlessness and acceleration, he developed the Equivalence Principle—the realization that gravity and acceleration are indistinguishable. This was not a result of statistical pattern matching; it was the result of physical intuition derived from sensory experience.

The paper notes that while a modern LLM could easily handle the complex mathematical deductions required to calculate planetary orbits once provided with Einstein’s equations, the AI could never have invented the initial premise. The AI lacks the "physical grounding" necessary to imagine the sensation of falling, which was the catalyst for the entire mathematical framework.

The Role of Physical Grounding in Cognitive Development

The lack of a physical body—or "embodiment"—means that AI software exists in a purely symbolic world. It processes tokens (words, pixels, or numbers) without any direct connection to the physical realities those tokens represent. This creates what researchers call a "symbol grounding problem."

For a human, the word "heavy" is not just a statistical neighbor to the word "weight" or "gravity." It is linked to the physiological memory of muscle strain, the visual of a falling object, and the tactile sensation of pressure. These sensory inputs form a "world model" that allows humans to predict cause and effect in the physical realm instinctively.

Google DeepMind Paper Says LLMs Will Never Replace Human Genius Because They Lack The Creative Leap Necessary To Make New Scientific Theories

Current AI discovery frameworks, including automated coding tools and scientific assistants, remain tethered to the "rulebooks" provided by their training data. They can optimize existing systems but rarely create entirely new ones. The DeepMind paper suggests that simply expanding the infrastructure of AI—adding more data centers, more GPUs, and more parameters—will not solve this. A larger model might become a more efficient "stochastic parrot," but it will not magically develop the genius-level intellect required for revolutionary scientific shifts because it lacks the intuitive foundation provided by physical existence.

A Chronology of AI Reasoning Milestones

The path to the current limitations identified by Zahavy can be traced through several key milestones in AI development:

  • 1950s – 1990s: Symbolic AI. Early AI focused on deduction, using "if-then" statements to solve logic puzzles. While precise, these systems were brittle and could not handle the messiness of real-world data.
  • 2010s: The Deep Learning Revolution. The rise of neural networks allowed AI to master induction. Systems like AlexNet (image recognition) proved that machines could identify patterns in data without explicit programming.
  • 2016: AlphaGo. DeepMind’s AlphaGo combined induction (learning from games) with deduction (searching through possible moves) to defeat world champion Lee Sedol. However, it still operated within the closed world of a game board.
  • 2022 – 2024: The LLM Explosion. Models like GPT-4 and Gemini demonstrated that massive scale could mimic human-like reasoning. Yet, as these models reached the limits of available text data, the "hallucination" problem highlighted their lack of a factual world model.
  • 2024: AlphaProof and Zahavy’s Paper. While DeepMind announced breakthroughs in mathematical deduction, Zahavy’s paper serves as a sober reminder that the "abductive" leap remains elusive.

Official Perspectives and Industry Reactions

While Google DeepMind has not issued a formal corporate rebuttal to Zahavy’s paper, the research reflects a growing internal consensus among many AI researchers that "scaling laws" alone may be reaching a point of diminishing returns.

In contrast, leaders at organizations like OpenAI have historically leaned into the "scaling hypothesis," suggesting that with enough compute and data, emergent properties will eventually bridge these cognitive gaps. However, even within those organizations, there is a shifting focus toward "Reasoning" models (such as the o1 series) that attempt to simulate the "System 2" thinking described by psychologists—a slower, more deliberate form of deduction.

Independent critics of AI, such as Gary Marcus, have long argued that LLMs lack "common sense" and a "world model." Zahavy’s paper provides a more technical and philosophically grounded version of this critique from within one of the world’s most advanced AI labs. It suggests that the industry may need to pivot from purely digital architectures to systems that can interact with and learn from the physical world.

The Future: Physical Multimodal World Models

The conclusion of the DeepMind research offers a roadmap for the next generation of artificial intelligence. To bridge the gap between logical calculation and true scientific invention, AI must evolve beyond text and image processing.

The paper advocates for the development of "physical multimodal world models." This would involve giving AI systems the ability to run embodied simulations in virtual environments—often referred to as "sim-to-real" training. In these environments, an AI would not just read about gravity; it would "experience" it through a simulated body. By interacting with objects, observing collisions, and navigating three-dimensional space, the AI could build the same kind of physical intuition that Einstein used in his thought experiments.

This shift would represent a move away from "Big Data" toward "Rich Data." Instead of reading the entire internet, the AI would focus on high-quality sensory-motor data. This could lead to AI that can not only solve equations but also propose the "abductive" hypotheses that lead to those equations in the first place.

Broader Implications for Science and Society

If Zahavy’s assessment is correct, the immediate threat of AI "replacing" human scientists and creative thinkers is overstated. Instead, the role of AI will likely remain that of a powerful "copilot"—an engine for induction and deduction that requires a human "abductor" to provide the creative spark.

In the medical field, an AI might be able to analyze millions of patient records to find correlations (induction) and suggest dosages based on medical guidelines (deduction). However, the "leap" to a new theory of cellular aging or a revolutionary drug delivery mechanism may still require a human who understands the physical nuances of biology in a way a software model cannot.

The "LLMs can’t jump" paper serves as a vital check on the hyperbole surrounding Artificial General Intelligence (AGI). It reminds the scientific community that intelligence is not merely the processing of information, but the ability to relate that information to the physical world we inhabit. Until AI can "jump"—or at least understand the physical sensation of jumping—it will remain a reflection of human knowledge rather than a source of original genius.

Related Posts

Xbox Strategic Realignment and the Potential Departure from Steam in the Next Generation of Gaming Hardware

The Microsoft gaming ecosystem is currently navigating a period of profound structural transformation, marked by a departure from long-standing distribution strategies and a re-evaluation of its presence on third-party platforms.…

Lenovo Silently Rolls Out Wildcat Lake-Powered ThinkCentre Neo 50a 24 Gen 7 AIO PCs With 120 Hz Display

The headline feature of the Gen 7 refresh is the transition to Intel’s "Wildcat Lake" processor family, specifically the Intel Core 5 320 and the Intel Core 7 350. These…

Leave a Reply

Your email address will not be published. Required fields are marked *

You Missed

A 30-Year-Old Engineer’s Reddit Post Sparks Wide Debate on the Scarcity of Ambitious, Childfree Men in Modern Dating.

A 30-Year-Old Engineer’s Reddit Post Sparks Wide Debate on the Scarcity of Ambitious, Childfree Men in Modern Dating.

EA Sports FC 26 Reports Record Daily Active Users as Series Momentum Surges Amid Potential Industry Acquisition

EA Sports FC 26 Reports Record Daily Active Users as Series Momentum Surges Amid Potential Industry Acquisition

Xbox Strategic Realignment and the Potential Departure from Steam in the Next Generation of Gaming Hardware

  • By admin
  • August 1, 2026
  • 3 views
Xbox Strategic Realignment and the Potential Departure from Steam in the Next Generation of Gaming Hardware

OpenAI CEO Sam Altman’s "Cool Use Case" for AI in Family Life Sparks Viral Debate Over Technology’s Role in Human Connection

OpenAI CEO Sam Altman’s "Cool Use Case" for AI in Family Life Sparks Viral Debate Over Technology’s Role in Human Connection

The AI Industry Faces a Call for Paused Progress Amidst Growing Concerns

The AI Industry Faces a Call for Paused Progress Amidst Growing Concerns

Rails Patches Critical Active Storage Flaw with Remote Code Execution Potential

Rails Patches Critical Active Storage Flaw with Remote Code Execution Potential