The debate over whether machines can “do” fundamental science is now essentially over - the answer is yes, and they will get even better with time. The models’ design was never that of a scientist — and that may be the distracting point. Frontier models were built to predict the next token. In doing so, they have compressed the scientific literature into a statistical map of how arguments, equations and experimental narratives cohere. That is not insight in the traditional sense, but it is just as important in the scientific discovery process. It is a machine that has internalised the nature of discovery. Couple it to question-asker at the start, and a tester at the end, the generation in the middle as the heavy-lifter then gives us the scientific method at industrial speed.
Humans remain slow generators. Agents are not.
When a claim survives independent check, the loop is science, whether or not anyone calls it that.
What the 2025-2026 record now shows:
DeepMind’s AlphaEvolve improved known constructions on about a fifth of the open problems it attempted;
A Stanford “virtual biotech” of tens of thousands of agents designed a lung-cancer strategy later independently pursued by industry;
Google’s Co-Scientist produced hypotheses on leukaemia, fibrosis and antimicrobial resistance that held up in the laboratory;
An OpenAI and Molecule.one loop improved a stubborn medicinal-chemistry reaction across more than 10,000 wet-lab runs;
Earlier this month, Open-AI shared an AI-generated solution to the Navier–Stokes Millennium Prize Problem. Treat it as a test of method, proof awaits independent verification;
Over the next 18 months, AI co-scientists will become ordinary infrastructure. The scarce resource will no longer be injection of fresh ideas, but testing/verification, and the judgement of which questions are worth asking.

