The history of scientific progress has rarely been a matter of simply collecting more observations and arranging them into neat patterns.
Time and again the decisive steps have come from a quieter, less predictable movement of the mind, one that begins with the world as it is experienced and somehow arrives at a new set of foundational principles from which the rest can be derived.
Albert Einstein once tried to describe this movement in a letter to his friend Maurice Solovine.
He sketched a cycle that began with raw sensory experience, passed through an intuitive leap to abstract axioms, and only then proceeded by strict logical deduction to testable consequences.
The middle step, the leap itself, remained mysterious even to him. It was not induction in the ordinary sense, nor was it deduction. It was something closer to abduction, the generation of a novel explanatory hypothesis when the available data did not yet force the conclusion.
That old diagram has returned to public attention through a position paper written by Tom Zahavy of Google DeepMind.
The paper, titled “LLMs can’t jump,” argues that large language models (LLMs) have become extraordinarily proficient at two portions of Einstein's cycle while remaining structurally closed off from the third.
Through statistical pattern matching across vast corpora they perform a powerful form of induction.
Through formal systems and specialized solvers they can carry out rigorous deduction, as demonstrated by systems that generate mathematical proofs at the level of international olympiads.
What they cannot do, according to the argument, is execute the abductive jump that takes incomplete or even contradictory experience and produces an entirely new set of axioms.
Axioms can be described as statements accepted as true, acting as starting points and rules for further logic, math, and philosophy.
Key aspects include assumptions, first principles, and unprovable foundations.
Humans have that. In fact, this is one of the traits that make us humans.
Computers however, don't have this.
The example chosen is Einstein's own formulation of general relativity. Classical mechanics was still highly successful. The observational anomalies were real but modest. Nothing in the data compelled a wholesale revision of the geometry of space and time.
The breakthrough required a conceptual rupture, an intuitive translation of physical reality into a fresh formal framework that could not have been reached by interpolating within the existing theoretical vocabulary.
Once those new axioms were stated, the mathematical consequences could be worked out with ordinary logic.
The difficulty lay in inventing the premises themselves.
What has always made such ruptures possible is a distinctly human restlessness, a curiosity that refuses to treat the current account of the world as final.
Progress in understanding has depended less on the accumulation of confirmed results than on the persistent urge to ask what still lies beyond them, even when the prevailing theory continues to work well enough for practical purposes.
That urge has driven individuals to abandon comfortable frameworks, to risk professional isolation, and in some cases to exhaust their lives in pursuit of a clearer picture.
The willingness to keep pressing against the limits of the known, to invent new axioms when none are yet required by the evidence, and to die trying if necessary, has been the quiet engine of the discontinuous advances that mark the history of science.
The paper treats the difficulty of replicating this capacity as more than a temporary shortcoming of current architectures.
It suggests that the autoregressive, text-trained nature of LLMs ties them to the space of already articulated human thought.
They can recombine, refine, and even surprise within that space, yet they lack the grounded, multimodal contact with the physical world that would allow them to propose premises the human literature has never contained.
Scaling parameters and compute may enlarge the calculator, but it does not supply the capacity to step outside the system of axioms already given, nor does it generate the internal drive that has repeatedly pushed human inquiry past the point where existing explanations seemed sufficient.
This claim has circulated widely because it touches a deeper tension in contemporary discussions of AI.
One widespread expectation has been that sufficient data and sufficient scale would eventually produce genuine scientific invention as an emergent property of compression.
Einstein's account of discovery, and the DeepMind paper's reading of it, presents a counter-image: invention as a discontinuous act that begins in experience and only afterward becomes formal.
The human contribution has never been merely computational.
It has included the restless curiosity that keeps asking for more, the readiness to dismantle workable theories in search of deeper ones, and the acceptance that the search itself may consume a lifetime without guarantee of success.
In other words, AIs may perform well above human capacity. Artificial General Intelligence, and probably beyond that.
But the gap is real.
Computers cannot emulate the strange human impulse to wonder about something simply because it is there, to pursue an idea with no immediate practical value, or to keep searching after the original question has already been answered.
They can process, reason, predict, and discover patterns at extraordinary speed, but the desire to ask why in the first place is something different.
Perhaps that is where the real difference lies.
Not in how much intelligence a machine can possess, but in what makes intelligence reach beyond what it already knows.
















































































































































































































































































































































































