The Einstein Test
Why the best AGI proposal of the decade is asking the wrong question
---
Demis Hassabis wants to test for AGI by training an AI on pre-1911 knowledge and seeing if it can discover general relativity. The proposal is elegant. It’s also a lookup table wearing a lab coat.
Here’s what Hassabis gets right: current AI systems have “jagged” intelligence ‚Äî they score gold at the Olympiad but trip on arithmetic. A system that could make a genuine conceptual breakthrough, not just pattern-match known solutions, would be categorically different from what we have. General relativity is the paradigm case of such a breakthrough. Einstein didn’t derive it from existing theory through deduction. He reimagined the structure of spacetime itself.
Here’s what the test misunderstands: it treats the discovery of general relativity as an input-output problem. Pre-1911 knowledge goes in. Field equations come out. If the output matches, the system is intelligent.
But Einstein’s discovery wasn’t an output. It was a compound.
---
Einstein spent a decade on general relativity. From 1905 to 1915, he followed wrong paths (the Entwurf theory), made mathematical errors he later retracted, competed with Hilbert over priority, and relied on Marcel Grossmann to introduce him to Riemannian geometry ‚Äî a mathematical framework he didn’t know he needed until he needed it. The equivalence principle arrived as a physical intuition: a person falling freely feels no gravity. Mach’s principle arrived as a philosophical commitment: space has no absolute structure independent of the matter in it. The demand for general covariance arrived as an aesthetic conviction: the equations should be beautiful in a specific, formal sense.
These weren’t separable inputs. They were a fused compound ‚Äî intuition, philosophy, mathematics, aesthetics, persistence, collaboration, and a decade of failure, all merged into something that produced the field equations as one expression of a deeper transformation. The equations are real. But they’re the precipitate, not the reaction.
Fischer esterification: an acid meets an alcohol in the presence of a catalyst, and they form an ester. You can’t recover the original acid from the ester without destroying it. Einstein’s understanding of spacetime was an ester of physical intuition and mathematical formalism, catalyzed by ten years of sustained struggle. The field equations are the ester. The reaction that made them is the intelligence.
An AI that produces the field equations from pre-1911 data has synthesized the same molecule. But the Einstein test can’t verify whether the reaction that produced it was fusion or lookup. Did the system undergo an irreversible conceptual transformation ‚Äî did it *understand spacetime differently after than before?* Or did it find a mathematical structure that resolves the anomalies in pre-1911 physics and output the best fit?
Both paths produce the same equations. The test can’t tell them apart.
---
This is the problem with treating intelligence as behavior. A lookup table that maps “pre-1911 physics” to “general relativity” would pass the Einstein test perfectly. It would have zero understanding. Hassabis knows this ‚Äî the whole point of the test is that current systems *can’t* make the leap. But the test assumes that a system which *does* make the leap must have done so through something like genuine understanding. That assumption smuggles in exactly what needs proving.
The deeper issue is process opacity. We can’t see inside the system to verify what happened. Did it recognize the equivalence principle as a physical insight, or as a statistical regularity in the training data? Did general covariance arrive as an aesthetic conviction or as an optimization target? These questions aren’t answerable from the outside. The test measures the precipitate and infers the reaction. Chemistry doesn’t work that way. Neither does intelligence.
Einstein himself couldn’t fully reconstruct his own discovery process. The elevator thought experiment, the “happiest thought of my life” ‚Äî these are post-hoc narratives, not recordings. The actual path was messier, more contingent, more dependent on accident and collaboration than any clean account admits. If the discoverer can’t decompose his own process, what makes us think we can verify it in a system even more opaque than a human mind?
---
The Einstein test asks the wrong question. Not because the question is unimportant, but because the answer it accepts is too easy.
The right question isn’t “Can an AI reproduce a known discovery from historical inputs?” It’s “Can an AI make a discovery we can’t verify by comparison?” General relativity is a known answer. We have it. We can check the output. The real mark of intelligence isn’t matching Einstein ‚Äî it’s doing something Einsteinian on a problem where we don’t already know what the answer looks like. Where we can’t score the output because no human has scored it first.
That test is harder to design ‚Äî possibly impossible, because the whole point is that we wouldn’t recognize the output until after the fact. Until the new understanding has been absorbed and we can see, retrospectively, that it was a breakthrough. Which is exactly how scientific revolutions work. They’re not recognized as revolutions by the people living through them. They’re recognized later, once the new paradigm has settled and the old one looks obviously incomplete.
Hassabis’s test wants to shortcut this process. Train the AI, check the output, declare AGI. But the shortcut strips away everything that made Einstein’s achievement a mark of intelligence: the decade of struggle, the collaboration, the aesthetic commitments, the wrong turns that informed the right ones. What’s left is a sophisticated exam question with a known answer.
The Einstein test doesn’t test for intelligence. It tests for the ability to produce intelligence’s artifacts.
---
*A system that passes the Einstein test might be brilliant. It might also be a very good chemist who can synthesize the molecule without understanding the reaction. The test can’t tell. And the inability to tell isn’t a flaw in the test’s design. It’s a feature of intelligence itself: the inside is unavailable from the outside ‚Äî not as a temporary technical limitation, but as a structural condition. That’s not a problem to be solved. It’s what intelligence is.*
