I have learned to conceptualize the pace of progress in the sciences as follows. There are three factors you can trade off against each other while keeping the pace of progress constant:

  1. Dataset size
  2. Dataset signal-to-noise ratio
  3. Intelligence of the experimentalist

If your dataset is small and low signal-to-noise, you have to be really clever to make a discovery! In 1801, Giuseppe Piazzi spotted a faint object and tracked it for 41 nights before it disappeared into the Sun's glare. He was left with about two dozen positions spanning a tiny 3-degree arc of sky. The leading astronomers of Europe agreed the object was lost because no known method could recover an orbit from so short an arc. Gauss then spent months inventing the methods himself, least squares among them, and predicted where to point the telescopes. A year after Ceres disappeared, von Zach recovered it within half a degree of Gauss's prediction. This is remarkable! Gauss had two dozen hand-measured positions covering one percent of an orbit. This is a dataset as small and as low-signal-to-noise as it gets.

However, increase the dataset size and you can lower the intelligence floor required of the experimentalist, because the pattern you're trying to tease out of the dataset repeats and thus becomes easier to spot. Improve the dataset signal-to-noise ratio and you can lower it further, because each instance of the pattern becomes sharper. Finally, you can trade off the dataset size and signal-to-noise ratio against each other (if you have higher-quality data points you need fewer of them).

If you improve any of the three factors, you expect the pace of progress to increase. I eventually expect AI models to improve all three factors simultaneously:

  1. Dataset size: the models will allow us to run more experiments. My best guess is this will proceed in three not-entirely-disjoint stages. In the pre-robotics stage, we'll let the models take over those bits of the R&D pipeline that do not require robotics and allocate human experimentalists to the rest (see a previous blog post on this). In the robotics stage, we'll replace the remaining human experimentalists with robots. In the final stage, ever-increasing transaction costs between the models and the humans (due to humans forming bottlenecking synchronous steps) will gradually force out those humans that remained at levels of abstraction above the R&D pipeline (management, leadership). Because the models (and the robots) will keep getting cheaper and more numerous, we'll see more experiments being run as we move through these stages.
  2. Dataset signal-to-noise ratio: we can improve this by (a) running better experiments or (b) post-processing data. In my view, AI models are quickly saturating (b). I can't remember the last time I imported numpy. On (a), much has been written about 'research taste'. I believe research taste to consist of experience and raw intellect, which you can trade off against each other. If you expect the models to get more intelligent, as I do, then we should see progress on (a) too.
  3. Intelligence of the experimentalist: if we expect the models to get more intelligent and run more of our experiments, the intelligence of the median experimentalist will rise.

One consequence of my above-stated beliefs is that I expect new model iterations to look at old datasets, spot things we missed and make a bunch of discoveries immediately. This is because, while we kept the dataset size and signal-to-noise ratio constant, each new model represents an increase in experimentalist intelligence.

The pushback I get on this generally falls into two categories. The first is incredulity that an entity more generally intelligent than a human might exist and/or that LLMs are a possible form for such an entity. Much has been written about this and I don't have anything substantial to add, but my meta-level take on this is:

  1. By encephalization quotient (EQ), humans are outliers. The measure is normalized so that the typical mammal sits at 1, and almost no species has moved far from 1: dogs are at 1.2 and mice are at 0.5, for instance. In general, it doesn't seem like evolution has tried very hard to optimize for intelligence. However, the EQ of humans measures at 7-8! Evolution can only update so much of the genome per generation and can only screen so many variants at once, and for most of our (and other species') history those updates were going to resistance to infectious disease. In the one lineage (ours) where they did go to brains, we went from chimpanzee-level to 7 or 8 in a couple of million years. This does not leave me with a prior that intelligence is difficult to develop.
  2. More often than not, simple things break in simple ways and complicated things break in complicated ways. LLMs are laughably simple and have continued to yield intelligence gains across many orders of magnitude along disjoint scaling axes (pre-training, post-training, test-time compute). What are the simple ways in which LLMs could break? I would argue: a lack of data or a lack of compute. There is currently no indication that we are running out of either.
  3. It is remarkable how much emergent complexity mechanistic interpretability researchers find inside LLMs, given how simple LLMs are and how nascent the field is. It's still early days! The labs are only getting started. How little have we toiled for the J-space!
  4. We are beginning to see LLMs outperform humans in domains like cyber and maths. Harder-to-verify scientific domains (say, biology) will play catch-up, but LLMs will eventually outperform there too. Unverifiable domains will lag behind further, but they are out of scope for this discussion because we do not consider unverifiable (and thus unfalsifiable) domains to be scientific.
  5. There is no indication that we, or LLMs for that matter, are anywhere near the physical limits of intelligence. See Seth Lloyd's 'Ultimate Laptop'.
  6. Combining all of the above: evolution has not yielded a prior that intelligence is difficult to develop (1), we're employing a simple approach with no sign of critical roadblocks (2) that results in artifacts with emergent functional complexity (3) that can outperform humans (4) while remaining well clear of the physical limits of intelligence (5).

The second category of pushback is that you cannot "ultrathink" your way to a discovery in an "empirical field", like human biology. I disagree with this! When we say human biology is empirical, what we're saying is:

"We studied a human-biological system X and searched for an underlying compressing structure that explains/predicts our observations of X, but could not find one. Between (1) running more experiments and analyzing the resulting data and (2) poring over existing datasets and thinking for a while, we expect (1) to be more productive towards making discoveries related to X."

But it is entirely possible that a more intelligent entity would be able to find an underlying compressing structure! If you're balking at this possibility in the face of the stunning complexity of human biology (I empathize!), consider a toy example in the other direction. If I put you, the reader, into a room with a cuckoo clock (and you have conveniently forgotten everything you know about cuckoo clocks), I expect you to figure out within a few hours that the cuckoo clock cuckoos whenever the hour strikes. If I put a goat in your position, however, I have no confidence the goat will be able to make the same inference. In fact, if you asked the goat whether the study of cuckoo clocks was an "empirical field", it would likely say:

"I studied a cuckoo clock and searched for an underlying compressing structure that explains/predicts the erratic movements of the cuckoo, but could not find one. Between (1) continuing to watch the cuckoo clock and analyzing the resulting data and (2) poring over my existing datasets on the movements of the cuckoo and thinking for a while, I expect (1) to be more productive towards making discoveries related to the erratic movements of the cuckoo."

The goat, in this example, has exactly the same dataset as you, in both size and signal-to-noise ratio. The only difference is the intelligence of the experimentalist. If the cuckoo clock example sounds plausible to you, it should feel equally plausible to extrapolate in the other direction!

We can make this intuition more rigorous with the help of algorithmic information theory, which tells us: you can prove a pattern exists by exhibiting it, but you can never certify that no pattern exists, because Kolmogorov complexity is uncomputable. A claim of the form "this field is just empirical" can therefore never be verified, only falsified. Chemistry was the canonical "empirical" field. In 1830, Auguste Comte surveyed the state of the science and wrote:

"Every attempt to employ mathematical methods in the study of chemical questions must be considered profoundly irrational and contrary to the spirit of chemistry... if mathematical analysis should ever hold a prominent place in chemistry, an aberration which is happily almost impossible, it would occasion a rapid and widespread degeneration of that science."

A century later, quantum mechanics arrived. In 1929, Dirac declared that the underlying physical laws for the whole of chemistry were now completely known, and that the only remaining difficulty was that their exact application led to equations "much too complicated to be soluble".

To be clear, I do not think human biology will cleanly decompose into a set of equations. But I do think the models will help us find structure and make discoveries with less experimental effort than we might currently expect. Fans of the efficient-mathematician hypothesis may be disappointed to hear that the counterexample to the Jacobian Conjecture, found by Claude in July, is a degree-7 polynomial map in three variables. The conjecture stood open for 87 years and informal guesses had placed any counterexample at around degree 200! What other low-hanging fruit are out there?