Skip to content

What “quantum advantage” and “AGI” claims have in common

Both are defined by benchmarks that the claimant frequently selects. A short list of questions retires most of what circulates.

These two fields make announcements in a similar shape: a capability threshold is named, a benchmark is chosen, a result is reported against it, and the threshold is understood to have been crossed. The weakness is the same in both cases — the benchmark is doing the definitional work, and the party announcing usually chose it.

For a quantum result the load-bearing question is what the best classical algorithm does on the same task. The history here is unambiguous: several headline advantage claims were substantially narrowed within months by improved classical simulation. A result that has survived serious classical attack is worth considerably more than a fresh one, and the gap between the two is invisible in the announcement.

For a claim about general capability the load-bearing question is contamination and construct validity — whether the evaluation was in the training data, and whether it measures the general ability it is named after or a narrower proxy that correlates with it under test conditions and stops correlating in deployment.

A compact filter covers most of it. What is the strongest baseline, and did the claimant try to make it strong? Is the task chosen because it is important, or because the system is good at it? What would falsify this claim, and has anyone attempted it? Who reproduced it independently? None of these require domain expertise to ask, and they retire a surprising proportion of what circulates at this intersection.