Protein Folding

For half a century the central practical problem in structural biology was that proteins are chains of amino acids that fold into precise three-dimensional shapes, that the shape determines what the protein does, and that predicting the shape from the sequence was hard to the point of hopelessness. The way you found out a protein’s structure was to crystallise it and fire X-rays at it, which took a good postdoc somewhere between many months and several years per protein, and which failed outright for large classes of protein that would not crystallise. Over five decades of this the field accumulated an archive of perhaps two hundred thousand experimentally determined structures, and the acquisition of each one was a small career event.

In late 2020 a system called AlphaFold entered the discipline’s biennial blind assessment and produced predictions accurate enough that the organisers said, in effect, that the problem was solved. By the middle of 2022 its authors had released a database of over two hundred million predicted structures — approximately every protein known to science, which is to say a thousand times the entire accumulated experimental archive, produced in about eighteen months and given away. In 2024 the work took a share of the Nobel Prize in Chemistry.

If you are looking for evidence that machines can do research rather than assist it, this is the exhibit. It is not automation of the laboratory technician. It is the automation of a genuinely hard scientific inference, performed better than the humans who had spent their lives on it, at a scale that makes the previous method look like a hobby.

And now the deflating part, which the people who cite this case rarely include.

Protein folding had three properties that almost nothing else in science has. It was a clean, well-posed question with a single unambiguous right answer — the structure is what it is, and you can check. It came with an enormous, high-quality, decades-in-the-making labelled dataset, assembled by exactly the laborious human process the machine then replaced. And it had a metric everyone agreed on, published biennially, which meant progress could be optimised against. Take away any one of those and the method does not obviously work; most scientific questions have none of the three, because most scientific questions are not ‘what is the shape of this thing’ but ‘does this intervention help, and compared to what, and in whom’.

Then notice the second thing, which is what happened next. Knowing the shape of a protein is one step, quite an early one, in a chain that runs through target validation, medicinal chemistry, animal models, three phases of clinical trial and a regulator. Solving a fifty-year bottleneck did not make drugs appear. It moved the bottleneck to the next station down the line, where it now sits, in a queue measured in years and governed by human physiology and the Food and Drug Administration.

Which is the pattern, stated once more in a form I hope is now becoming familiar. The intelligence was real; the acceleration was real; and the aggregate rate of progress in medicine went on being set by something else entirely. There is a version of the coming decades in which this happens twenty times over in twenty fields, and each time it is a genuine triumph, and the growth rate barely moves.

from Everywhere But the Statistics: The Macroeconomics of Artificial Intelligence, Told Through Its Bottlenecks (2026)


GPT-5.6 Sol:

The passage is persuasive because it makes a distinction that is routinely blurred in arguments about AI: a spectacular increase in the productivity of one cognitive task is not the same thing as a spectacular increase in the productivity of the whole system containing that task. AlphaFold is close to an ideal case for demonstrating both sides of that proposition.

The historical core is right. AlphaFold2’s CASP14 performance really was extraordinary: a median GDT_TS of 92.4, with structures in many cases approaching experimental accuracy and a huge gap over competing computational methods. The AlphaFold Database then expanded in July 2022 to more than 200 million predicted structures, covering nearly all catalogued proteins, and Hassabis and Jumper received half of the 2024 Chemistry Nobel for protein-structure prediction.

But I think the passage contains two different arguments, one much stronger than the other.

AlphaFold really is evidence for machine performance on a scientific inference task

The phrase ‘automation of a genuinely hard scientific inference’ is basically fair. This wasn’t merely automating pipetting, image segmentation, literature search, or some clerical component around science. Given sequence information and associated evolutionary data, AlphaFold inferred a latent physical structure that scientists cared about and that could subsequently be checked against experimental observations. CASP was particularly convincing precisely because the targets were blinded: the relevant experimental structures existed but had not been disclosed.

That matters for the broader AI debate. One weak argument sometimes made against AI science is that machines can only manipulate representations supplied by human scientists and therefore cannot be doing anything scientifically substantive. AlphaFold makes that position difficult to maintain. Scientific work contains enormous amounts of inference from observations to otherwise unknown properties. There is no principled reason that inference ceases to be intellectually important merely because an algorithm performs it.

But ‘machines can do research’ is somewhat stronger than the example establishes. AlphaFold did not autonomously decide that protein structure prediction was an important problem, invent CASP, accumulate the Protein Data Bank, choose what experiments needed doing, determine what biological questions mattered, or decide what its predictions meant. Humans built the epistemic infrastructure within which AlphaFold could achieve its startling result. Calling the resulting computation research is defensible, but calling AlphaFold an autonomous researcher would smuggle in considerably more.

There is another slightly misleading sentence: it supposedly performed the task ‘better than the humans who had spent their lives on it’. Structural biologists were mostly measuring structures experimentally, not trying to eyeball amino-acid sequences and guess coordinates. AlphaFold substituted computation for a large amount of experimental structure determination in cases where prediction suffices; it did not demonstrate that neural networks are intrinsically superior to biochemists at doing the same activity.

The account makes protein structure prediction too tidy

This is the main scientific weakness.

The passage needs protein folding to be a pristine example of a problem with ‘a single unambiguous right answer’. CASP structure prediction was unusually close to that ideal, but proteins themselves are not.

Proteins are dynamic objects. They can occupy multiple conformations; ligand binding, interactions with other molecules, post-translational modifications and environmental conditions can change their structures; intrinsically disordered regions do not possess one well-defined folded structure at all. Even AlphaFold3 generally produces static structures rather than the physical ensemble present in solution. Recent work explicitly treats predicting those ensembles as an unsolved problem.

So I would rewrite the claim as:

CASP created an unusually clean supervised prediction problem by asking for a structure corresponding to a hidden experimental target.

That is subtler than saying nature itself supplied one unambiguous answer.

Likewise, ‘protein folding’ is being used somewhat loosely. AlphaFold2 largely solved a major version of the protein structure prediction problem. It did not explain the physical folding pathway by which an unfolded chain traverses an energy landscape and arrives at its functional ensemble. The AlphaFold paper itself carefully describes its target as the ‘structure prediction component’ of the protein-folding problem.

And the opening history is compressed for dramatic effect. X-ray crystallography was enormously important, but experimental structural biology was not simply ‘crystallise it and fire X-rays’ until AlphaFold arrived; NMR and, increasingly, cryo-electron microscopy were other major routes. AlphaFold has also turned out to be extremely useful inside experimental structure determination rather than simply replacing it. Experimentalists now use predicted models as priors or starting points, and 2026 work is explicitly combining AlphaFold with cryo-EM, crystallography and other experimental measurements.

Those qualifications actually strengthen the economic argument. AlphaFold didn’t abolish structural biology. It changed the production function of structural biology.

The bottleneck argument is excellent

The strongest sentence in the extract is essentially:

‘The intelligence was real; the acceleration was real; and the aggregate rate of progress … went on being set by something else.’

That is a much better framework for thinking about AI economics than asking whether AI is ‘transformative’ in the abstract.

Suppose producing a useful medicine requires:

target discovery → target validation → structural understanding → hit discovery → lead optimisation → toxicology → trials → manufacturing → regulatory approval.

If structure determination takes a year and you reduce it to a minute, you have achieved something remarkable. But once structure determination approaches zero cost, its contribution to total cycle time approaches zero too. The shadow price moves elsewhere.

This is essentially Amdahl’s law applied to economic production. If a process contains complementary serial stages, making one stage arbitrarily fast eventually produces diminishing end-to-end returns because the unchanged stages dominate total time.

A toy example makes it obvious. Imagine drug development requires:

  • 2 years structural work

  • 3 years medicinal chemistry and preclinical work

  • 7 years clinical/regulatory work.

Total: 12 years.

Make the first stage 1,000× faster and the process takes about 10 years rather than 12.

The local productivity statistic is astonishing: roughly 1,000×. The system-level productivity improvement is about 20%.

That gap between the two numbers is precisely what a lot of AI economic discussion misses.

And drug development remains a good illustration. A 2026 Nature Reviews Drug Discovery review says explicitly that impressive validation on protein folding does not automatically constitute validation of an end-to-end drug-discovery process; applications of AlphaFold2 structures to virtual screening have produced mixed results, while later systems improve some downstream steps but don’t erase the remainder of the development pipeline.

The passage therefore offers a good antidote to a common extrapolation:

Task benchmark → occupational automation → sector productivity → GDP growth

Those arrows are not identities. Every transition depends on complements.

But the macroeconomic conclusion can also be pushed too far

‘The growth rate barely moves’ is possible. It is not something AlphaFold demonstrates.

There are at least three reasons.

First, bottlenecks are endogenous. Once one input becomes dramatically cheaper, people reorganise production around it. Scientists attempt experiments that previously made no economic sense. Entirely new research programmes become feasible. New complementary technologies get developed because their returns have increased. AlphaFold has already enabled proteome-scale structural analyses that simply weren’t practical when structures were scarce one-offs.

That is why treating a production pipeline as permanently fixed can badly underestimate a general-purpose technology. A factory doesn’t merely continue making exactly the same quantity of goods after electricity arrives, substituting an electric motor for the steam engine while everything else remains unchanged. Eventually its architecture changes.

Second, the right outcome variable isn’t necessarily time to one FDA-approved drug. Cheap structure prediction can produce more targets investigated, more negative hypotheses killed early, better experimental design, broader basic biological knowledge and a larger portfolio of shots on goal. If each medicine still takes ten years but society can investigate ten times as many plausible programmes simultaneously, output may rise without much reduction in latency.

In other words, there are two margins:

speed per project and number of projects.

The excerpt focuses almost entirely on the first.

Third, twenty AlphaFolds in twenty fields are not necessarily equivalent to twenty isolated bottleneck removals. If the improvements occur at successive complementary stages of the same production chains, the effects multiply. Automating protein structure alone may save relatively little total drug-development time. Automating structure, assay design, medicinal chemistry, toxicology prediction, trial recruitment and regulatory documentation would be a fundamentally different proposition.

This is where the metaphor of ‘the bottleneck moving downstream’ has an important corollary:

If AI keeps following the bottleneck downstream, eventually there may not be much bottleneck left.

Whether that happens is almost the whole economic question.

There is also a deeper lesson about datasets

The paragraph about AlphaFold’s unusually favourable conditions is important, but I’d modify it.

Its three properties — clean objective, huge dataset, accepted benchmark — certainly helped. AlphaFold benefited enormously from decades of experimentally determined protein structures and enormous sequence databases. The breakthrough depended on precisely the scientific infrastructure that preceded it.

But saying ‘take away any one of those and the method does not obviously work’ risks overfitting our explanation of AI success to the supervised-learning paradigm of the 2010s.

One of the central questions now is whether AI systems can manufacture substitutes for those missing ingredients:

weak or self-supervised signals instead of labels; simulators instead of experimentally labelled examples; active learning to decide which experiment to perform next; multimodal models combining literature and measurements; automated labs that generate their own training data; learned reward models where no simple benchmark exists.

If those work, AlphaFold’s conditions cease to be a hard boundary. They become an especially favourable point on a continuum.

And conversely, if they do not work reliably, the extract’s pessimism becomes much more powerful. Then AlphaFold looks less like the first example of a general phenomenon and more like the exploitation of an unusually machine-legible scientific niche.

That is probably the most important unresolved issue in the entire analogy.

So I think the passage is directionally right, but for a slightly different reason

AlphaFold is not quite the immaculate example presented here. Protein structure is less well posed than the prose suggests; predicted structures aren’t equivalent to 200 million experimentally observed structures; experimental methods did not become obsolete; and the model automated one extremely important inferential mapping rather than the complete activity called structural biology.

Yet none of those corrections damages the central thesis.

The useful lesson is that AI capability and economic impact are separated by a network of complements.

The naïve AI optimist observes a 1,000-fold improvement in one input and mentally carries the 1,000× through to final output. The naïve pessimist notices that final output barely moves initially and concludes that the technological advance wasn’t important. Both are wrong.

The interesting question is what happens to the next bottleneck.

If AI repeatedly encounters things it cannot accelerate — human subjects, physical construction, scarce energy, regulatory deliberation, longitudinal observation, biological gestation, political coordination — then we could indeed get a world full of astonishing demonstrations accompanied by surprisingly ordinary GDP statistics.

If, instead, AI progressively attacks the complementary intellectual stages, while automation, institutional change and capital accumulation attack the physical ones, AlphaFold may look in retrospect not like an exception but like an early example of a much broader process.

So the decisive macroeconomic variable isn’t really ‘How intelligent does AI become?’

It is closer to:

‘What fraction of the binding constraints on useful production are themselves functions of intelligence?’

That question is much harder—and much more consequential—than whether a machine can ace another benchmark.

Previous
Previous

Genetic Criticism

Next
Next

Two Wolves