Large language models are slowly but steadily mastering the linguistics of the universe— mathematics. China’s DeepSeek has upped the AI game with a mathematical model that can identify and correct its own errors. In one of the world’s most prestigious undergraduate maths competitions, not only did the model top the charts, but it also broke the highest human score attained in the competition.
DeepSeekMath-V2 scored 118 out of 120 points on questions from the 2024 William Lowell Putnam Mathematical Competition, 28 points clear of the top human score, at 90 points. The model’s performance in the International Mathematical Olympiad (IMO) 2025 and the 2024 China Mathematical Olympiad was at par with gold-medalists.
“We are at a point where AI is about as good at maths as a smart undergraduate student,” says Kevin Buzzard, a mathematician at Imperial College London. “It is very exciting.”
In February, Google DeepMind developed AlphaGeometry 2, an AI problem solver, which also achieved a similar feat. Gemini’s Deep Think reproduced the results in July.
The initial approach focused on getting the correct answer to a problem; however, the flaws were apparent. A correct answer can arise from fortunate errors, and the approach impeded the ability to prove axioms and formulae, which require logical reasoning rather than merely refining the ability to derive the final answer.
Tong Xie, a chemist specializing in AI-driven discovery at UNSW Sydney, Australia, notes that the significant shift in prioritizing reasoning over the final answer led to the development of DeepSeek and Deep Think.
DeepSeekMath-V2 introduces us to self-verifiable mathematical reasoning. The model has a built-in verifier that is trained to evaluate mathematical proofs via step-by-step deduction, to uncover logical flaws, and to assign scores based on the rigor of the proofs. The verifier’s critiques are evaluated through meta-verification, reducing hallucinations. These features work with a proof generator that builds solutions and self-evaluates, refining arguments until no further errors are found.
The system operates in a feedback loop: the verifier polishes the generator, and the generator produces more challenging proofs that serve as training data for the verifier. The system solved 5 of 6 2025 IMO problems, but it didn’t perform well on the most difficult problems from other IMOs.
Math-V2 is an improvement over Deep Think, as it requires minimal human involvement and is cost-effective and scalable due to its self-verification feature. Deep Think uses an external symbolic language called Lean, which requires heavy expert input.
LLMs have started to resemble human mathematicians, as they can reason in plain language. Moreover, Math-V2 is open-source, allowing it to be downloaded online for free and used by mathematicians to reason and train. Although the technology is far from being of pivotal help, it can still aid many amateur mathematicians in determining how a problem should be solved or how a proof should be derived. In devising ways to solve problems, this approach opens avenues of procedures and methods for a student new to the domain.
Copyright @smorescience. All rights reserved. Do not copy, cite, publish, or distribute this content without permission.
SUBSCRIBE TO OUR NEWSLETTER
.......... ..........Subscribe to our mailing list to get updates to your email inbox.
Monthly Newsletter














