A Major Milestone
The International Math Olympiad is the world’s most prestigious math competition for pre-college students. Held annually since 1959, elite competitors from each of 176 countries solve six exceptionally challenging problems in algebra, combinatorics, geometry, and number theory. Medals are awarded based on percentages, and approximately eight percent receive a prestigious gold medal.
When the popularity of AI first dramatically rose in 2022, chatbots struggled with math and code. Recently, companies like Google and OpenAI have developed AI systems that are better equipped to solve complex problems that the average person cannot solve.
The accomplishment is impressive enough for a scientist to leave it on their CV for the rest of their career. Google and OpenAI’s result this month marks the first time a machine reached this level of success. In last year’s contest, DeepMind used systems like AlphaGeometry and AlphaProof -- both designed for math -- to answer questions. However, these systems were not chatbots; they were able to answer questions only after mathematicians translated them into Lean: a computer language designed to solve math problems. Even then, the process took two to three days and resulted in a silver medal. Deep Think and OpenAI’s unreleased model were able to achieve gold fully in natural language -- without any human intervention -- and within the competition time limit.
IMO President Prof Dr. Gregor Dolinar called the achievement “astonishing,” and confirmed that IMO judges verified Deep Think’s answers — OpenAI hired independent reviewers instead. He also said graders found Deep Think’s answers to be “clear, precise, and most of them easy to follow.”
Reasoning Systems
Deep Think’s new unreleased model was built using a “reasoning” system that can “reason” through tasks involving math, science, and computer programming. These systems, like any other AI, initially learn through vast numbers of datasets. Then, through a procedure known as reinforcement learning, AI systems go through an extensive trial-and-error process to learn additional behavior.
Google describes the setup as one that enables models to “simultaneously explore and combine multiple possible solutions before giving a final answer, rather than pursuing a single, linear chain of thought.” Additionally, they provided their latest version of Gemini Deep Think with access to a curated collection of high-quality solutions to complex math reasoning problems, as well as some general tips on how to approach IMO-style problems.
OpenAI used similar methods for its breakthrough moment. Noam Brown, a prominent researcher at OpenAI, confirmed the new experimental model was focused on scaling up “test-time compute.” By allowing the model to reason, or ‘think’, for longer periods and deploying parallel computing power to reason through numerous paths simultaneously, their model performed substantially better than last year.
Reasoning systems and reinforcement capabilities helped the models achieve their remarkable jump in score, but a more dramatic evolution may be necessary on the path to artificial general intelligence (AGI).
Drawbacks
Competition answer quality differed between the two companies. IMO questions are proof-based, meaning two different solutions to the same problem can have extremely different qualities but still be correct. While DeepThink had good answers, OpenAI’s answers were a lot messier and less well produced. And while its solutions were technically correct, they were not very well written and followed a solution path that most humans would not.
Furthermore, neither model was able to get any points on problem number 6. With so little information revealed about the models and the data on which they were trained, it is difficult to guess where the models failed.
In December, an OpenAI system surpassed human performance on a reasoning test called ARC-AGI, but the company spent nearly $1.5 million in electricity and computing costs to complete the test, against the competition rules. Google and OpenAI have not yet revealed the electricity and computing costs associated with their model’s IMO performance, but Brown called it “very expensive”.
In general, reasoning systems like the ones used in the competition can be extremely costly, because they spend an immense amount of additional time thinking about a response. The amount of commute power and monetary capital needed to keep this level of intelligence alive raises questions about the feasibility of AGI and superintelligence in general.
Can Money Buy Intelligence?
Everything comes at a cost. Whether humanity will reach superintelligence through AI or not remains to be solved, but we can be sure it will cost us at a scale never seen before.
Junehyuk Jung, a math professor at Brown University and visiting researcher in Google’s DeepMind AI unit, says that Deep Think’s performance “suggests AI is less than a year away from being used by mathematicians to crack unsolved research problems at the frontier of the field.”
Jung, who won IMO gold in 2003, believes that there is potential for collaboration between AI and mathematicians when AI can solve difficult reasoning problems in natural language. He also says that both Google and OpenAI believe AI models will soon be capable of applying to research questions in other fields, such as physics or computer science.
In a CBS 60 Minutes interview with DeepMind CEO Demis Hassabis, Hassabis described AI as a machine lacking imagination. In other words, we can still think of it, at its best, as an average of all the human knowledge in the world. Until it can think on its known — novel ideas, conjectures, and thoughts altogether — AI will be a step behind human ingenuity and creativity.
Experts do not yet agree on when, if at all, AGI will arrive. While some estimates believe it is only a few years away, others see it taking centuries. One survey of thousands of recent AI publication authors forecasted the arrival of “high-level machine intelligence,” when AI will accomplish every task better or more cheaply than humans. The median estimate showed a 25% chance in the 2030s and a 50% chance by 2047.
Expected feasibility of many AI milestones moved substantially earlier in the course of one year (between 2022 and 2023)
Another collection of surveys from over 5000 researchers and experts indicated a 50% probability of achieving AGI between 2040 and 2061. A survey from the AAAI 2025 Presidential Panel on the Future of AI Research also suggested that the current approach to AI will be unlikely to lead to Artificial General Intelligence (76% of respondents), making the predictions harder to support.
The culmination of AGI is also known as Singularity. It describes a time when systems can combine human-level thinking with rapid and perfect memory. Many fear its implications of machine consciousness, because a machine that can self-improve and recognize its own flaws may easily surpass human capabilities.
Others see potential in AGI as an ally to aid us rather than harm us. They argue that intelligence is multi-dimensional, and Artificial General Intelligence will be different — not superior — to human intelligence.
There are also limitations regarding the amount of compute-power our world can provide; humans may not have the capacity to accelerate AGI. The human brain, which the technology seeks to surpass, has never been fully modeled, and the impossibility of it makes it hard to envision Artificial General Intelligence coming to fruition.
Despite the uncertainty, the mean estimate for AGI’s arrival has been decreasing rapidly in the past few years. As new advancements shake headlines nearly every week, humanity is inching its way closer and closer to a form of intelligence that has the potential to surpass its own: the future of AI will either be humanity’s liberation or its doom.