Skip to main content

Xiaohongshu Large Model Scores Perfect IMO Gold: Breakthrough in Mathematical Reasoning

Introduction: Xiaohongshu Quietly Achieves Great Things

Wait, Xiaohongshu, you really have been quietly accomplishing amazing things!

Just now, IMO 2026 results were freshly released, and large model submissions have reached new heights this year. And the first one to get a perfect score from official grading was actually Xiaohongshu?

According to official IMO assessment, Xiaohongshu's large model dots-note-3.0 answered all six questions correctly, winning the gold medal with a perfect score.

How significant is this achievement? Let's put it this way: this is the first time a Chinese large model has received official IMO gold-level certification, and it's also only the second large model globally to achieve this feat after Google's Gemini. And the key point! It's a perfect score! Google only got 5/6 correct back then...

This is shocking. Is this still the Xiaohongshu I remember? Making its debut in a world-class competition and getting off to a flying start!

Digging deeper, the problem-solving approach of Xiaohongshu's large model dots-note-3.0 is also eye-opening: it completed the entire process from reading questions, analysis, to writing proofs using only natural language. The final answers were not only all correct but also innovatively proposed unconventional solutions for certain problems.

Double CMO gold medalist Liu Hanzuo commented after reviewing: "The model ultimately presented a correct and beautiful proof idea. This solution has a compact structure and natural logic, and even among IMO solutions, it belongs to the relatively concise and elegant category."

CMO gold medalist Wang Qiantong reached a similar conclusion: "From a mathematical aesthetics perspective, Xiaohongshu's large model dots-note-3.0's solution is very beautiful and elegant. The conventional solution is already quite clever, but dots directly tackled the hardest and most essential part of the problem."

In his view, dots solves problems concisely and directly hits the essence. After reading it, you'll feel each step follows logically, but when actually facing the problem, human contestants would rarely think of approaching it from this angle. In other words, the IMO full-score gold medal may be far from the upper limit of Xiaohongshu's large model dots-note-3.0 capabilities.

Why Large Models Need IMO Today

Let's start with a quick explanation of what IMO is and why companies are rushing to send their large models to take this exam.

IMO stands for International Mathematical Olympiad, where the world's top high school students gather every year. Terence Tao and nearly half of contemporary Fields Medal winners came out of IMO, and Google co-founder Sergey Brin also won an IMO gold medal.

The competition is typically held over two days, with contestants needing to complete 3 questions in 4.5 hours each day, totaling 6 questions covering algebra, combinatorics, geometry, and number theory. Each question is worth 7 points, with a full score of 42 points. But just having the answer isn't enough; to score points, contestants must write logically complete proof processes that can withstand step-by-step scrutiny.

For large models, this is extremely difficult, yet also the most intuitive way to demonstrate capabilities.

Take Gemini as an example. In 2024, Google first challenged IMO and ultimately only achieved a silver-level score of 28 points. Back then, the system still needed to translate natural language questions into formal languages like Lean first, with some questions taking days to complete. The following year, Gemini Deep Think returned and directly completed proofs end-to-end in natural language, solving 5 problems in 4.5 hours and winning the gold medal.

Behind this is a generational leap in capabilities. Deep down, mathematics is the touchstone for large model intelligence and reasoning ability—the farther a large model goes in mathematics competitions like IMO, the more it demonstrates its underlying advantages.

Another recent example: just this week, Fable 5 participated in constructing a counterexample to the Jacobian Conjecture. This conjecture has remained unresolved for 87 years since being proposed in 1939, with even top mathematicians like Zhang Yitang having attempted and failed. The counterexample found this time, though not yet formally peer-reviewed, also marks that current large model capabilities are sufficient to participate in genuine mathematical discovery. And mathematics competitions are the hardest capability threshold before models enter deep waters of research.

Additionally, IMO has a unique advantage over benchmarks—unknown.

Each competition's questions are crafted by experts and strictly confidential before the official competition. This means models cannot get questions in advance, and real-time participation ensures models cannot conduct targeted training beforehand.

A leading model researcher gave this judgment: "In today's model training, you can basically assume that once certain data appears on the internet, it will quickly be used for model training and the model will know it. So if you want to test whether a model is actually intelligent, you can only test things that haven't been made public."

This is precisely the most irreplaceable value of synchronously participating in official IMO assessment for large models. It tests not how many problems and mathematical knowledge the model has memorized, but when a truly unknown problem first appears, whether the model can independently understand, explore, and ultimately complete rigorous reasoning.

And having only new questions isn't enough—official assessment uses strict competition-day processes. After students finish exams each competition day, the model team only then receives that day's questions and must submit PDF answers within the specified deadline. The model directly reads English questions and generates proofs; staff only ensure system operation. Final answers are reviewed by IMO judges from different countries according to official standards.

When new questions for this edition, time limits, complete proofs, and third-party professional reviews all come together, IMO becomes the capability exam that's currently hard to bypass through memorization and packaging.

Losing No Points is Harder Than Winning Gold

For this very reason, Xiaohongshu's large model dots-note-3.0 winning gold—and with a perfect score—is particularly striking.

Getting all six questions correct first proves the stability of the Agentic reasoning system. The model needs to understand natural language conditions, judge which information is truly critical, propose potentially valid intermediate conclusions, and then organize them into a complete proof. Throughout this process, any misreading of a condition or any skipped step in deduction leaves loopholes that judges can directly point out.

Xiaohongshu's large model dots-note-3.0 achieving full marks across six different types of problems means its performance doesn't rely on accidental inspiration on any single question.

Secondly, Xiaohongshu's large model dots-note-3.0 directly handles everything end-to-end using natural language in one go. This also means the model can directly understand questions like human contestants and independently find paths, possessing a complete closed loop from problem understanding to answer expression. It represents the model having transferable general reasoning capabilities.

Specifically, during this problem-solving process, Xiaohongshu's large model dots-note-3.0 used an agent approach, allowing the model to reason and solve problems using natural language combined with Python code execution. During reasoning, it also leverages recursive self-criticism capabilities to examine its own arguments, thereby solving more complex real-world tasks.

However, what's truly surprising is the model's innovative thinking demonstrated in the third problem.

This is a combinatorial game problem. Common solutions discussed online for this year's IMO questions first transform the original problem into a connectivity problem in graph theory, then establish proofs using graph theory language. This approach itself is already quite clever, but Xiaohongshu's large model dots-note-3.0 didn't follow it—instead, it used induction.

Xiaohongshu's large model dots-note-3.0's solution grasped the most essential structure of the problem and designed perfectly appropriate induction objects and induction propositions. This explains why CMO winner Wang Qiantong described this solution as beautiful and elegant. The depth of Xiaohongshu's large model dots-note-3.0's analysis of the problem structure naturally produces a feeling of sudden enlightenment.

Returning to the result itself, the model achieving a perfect score in this edition is actually quite difficult. We can look at it from several levels.

First level: the scoring difficulty of this entire set of competition questions is significantly higher than last year. The IMO gold medal cut-off for 2026 was 29 points, compared to 35 points last year. In just one year, the gold medal line dropped by 6 points—nearly the score for an entire question. So the gold medal this year carries more weight than Gemini's last year.

Second level: there's also a gap between perfect score and gold medal, with a full 13-point difference this year. 29 points means a contestant can enter the gold medal range by fully solving 4 questions and getting 1 point from the remaining two. But Xiaohongshu's large model dots-note-3.0 went all in, directly reaching the highest point this exam can measure.

After all, as the old saying goes, a perfect score is only the upper limit of the test paper, not the upper limit of ability. Winning gold indeed represents entering the world's most elite group, but a perfect score represents completing another extremely rare screening among that elite group.

Xiaohongshu's Large Model dots-note-3.0 Will Be Open-Sourced

Choosing IMO as the starting point was also a very smart move by Xiaohongshu.

In today's industry, when promoting large model underlying technical capabilities, there are generally two directions: Coding and Mathematics. The reason is simple: results for both types of tasks are hard to gloss over with language packaging. Compared to knowledge Q&A and subjective evaluations, they more directly examine whether the model can truly understand complex problems, decompose tasks, and continuously make progress.

IMO is particularly outstanding among these—high visibility, sufficient industry recognition, and it directly demonstrates deep reasoning capabilities when models face highly difficult unknown problems. So for Xiaohongshu, this was an important technical debut.

In the past, external understanding of Xiaohongshu's technology focused more on content communities, recommendation algorithms, and content understanding. In the foundation model field, there has been a lack of a sufficiently public and authoritative scorecard that doesn't require complex explanation.

Parameter scale is difficult to directly equate with capability, and a few percentage points on conventional benchmarks hardly form a clear industry perception. IMO perfect score is different—42 points is right there, enough to prove that Xiaohongshu already possesses foundation model capabilities worthy of serious industry examination.

With one IMO perfect score gold medal, it let the industry clearly see its technical quality for the first time—a noteworthy new player has emerged in the foundation model field.

P.S. The corresponding model will soon be open-sourced~ Stay tuned

Source: QbitAI · July 22, 2026