Chinese companies Huawei and Xiaohongshu announced that their AI systems achieved full marks on the 2026 International Mathematical Olympiad problems, after each provided solutions to the six problems faced by human competitors in this year's edition.

According to a statement by the two companies, Huawei's 'Seelia' system and Xiaohongshu's 'dots-note-3.0' model scored 42 out of 42 points, seven points per problem. The two systems did not participate in the official competition for students; the companies received the problems after the human competitors finished the test, and then the solutions were submitted in a separate evaluation process.

Six Complete Problems

The International Mathematical Olympiad requires participants to solve six problems spread over two days, and the final answer alone is not enough; a complete proof must be presented showing the logical steps leading to the result. Hence, a perfect score signifies more than the ability to perform calculations; it means the solutions met the proof requirements according to the grading process.

Xiaohongshu said its model received full seven points on each of the six problems after grading organized by the Olympiad committee. It clarified that any human intervention during the test was prohibited, including giving hints, modifying answers, or choosing among multiple solutions generated by the system.

Huawei also announced that the 'Seelia' system demonstrated comprehensive abilities in solving problems from various mathematical areas. However, the announced results pertain to the AI's answers to the competition paper and do not mean that the two systems officially won the Olympiad or competed with the students in a single ranking.

Seven human competitors also recorded perfect scores, so the result does not mean the two models outperformed all participants (Adobe).

Only Seven Students

Shanghai hosted the 67th International Mathematical Olympiad from July 10 to 21, 2026, with 666 students from 117 countries. Official results showed that only seven human competitors achieved the perfect score of 42 points, while the gold medal threshold was set at 29 points.

Fifty-five students won gold medals, 105 silver, and 189 bronze, along with 141 honorable mentions. These figures place the announced performance of the two models at the highest possible result on the contest paper, but it remains the outcome of a separate evaluation conducted after the problems were delivered to the technical labs.

Therefore, it is not correct to say that the AI 'outperformed all humans,' because seven students achieved the same result. Additionally, other labs may later announce results for their models on the same problems.

Jump in One Year

The 2026 result represents progress compared to what AI labs announced during the previous edition. In 2025, an advanced version of Google DeepMind's 'Gemini Deep Think' scored 35 points, after solving five of the six problems completely, a level equivalent to a gold medal in that edition.

The shift from five solved problems to six complete problems in one year reveals rapid progress in models' ability to handle long chains of reasoning, revise hypotheses, and produce proofs amenable to human evaluation.

This type of performance does not necessarily rely on a single answer generated directly by a language model. Advanced problem-solving systems may use multiple search processes, try different solution paths, and then verify logical consistency before presenting the final proof. However, the two companies have not published in the available materials all the technical details needed to independently compare model sizes, computational resources, or number of attempts.

The ability to solve Olympiad problems does not mean mastery of open mathematical research, which remains more complex and less amenable to direct evaluation (Adobe).

Limitations of the Test

Despite the difficulty of Olympiad problems, success in them does not necessarily equate to the ability to conduct open mathematical research. Competition problems are designed to have specific solutions that can be evaluated within a limited time, while scientific research problems may require weeks or months, and sometimes lack a known path or ready answer.

A recent research study testing advanced models on 25 private mathematical research problems showed that all evaluated models scored less than 10 percent, indicating a persistent gap between excelling in math competitions and solving new research problems.

The result of Huawei and Xiaohongshu remains an indicator of AI progress in producing mathematical proofs within specific, constrained problems. But it is not sufficient alone to judge its ability for general mathematical innovation or to replace human verification, especially in problems without clear reference answers.

"); googletag.cmd.push(function() { onDvtagReady(function () { googletag.display('div-gpt-ad-3341368-4'); }); }); }