Over the winter break, OpenAI announced their new model, o3. Much of the media attention rightly focused on its impressive results on the ARC-AGI benchmark, but for us at Galois, the most significant result was something else—the model’s 25% score on a benchmark called Frontier Math. Frontier Math is designed to measure human-like mathematical abilities in an AI. Before o3, generative AIs were scoring just 2% on this benchmark. The fact that OpenAI was able to achieve a leap to 25% suggests that AI mathematics is improving very rapidly.1 In fact, we may soon see AI mathematicians that rival or surpass humans on certain kinds of problems. Let's break this down: What is Frontier Math? What did o3 do? And what does it mean for the future of mathematics? A long-term problem in AI is figuring out whether the AI is truly capable at a particular task—or just regurgitating answers it has memorized. Modern AIs like GPT-4 are trained on gargantuan quantities of data—essentially every piece of text available on the Internet. That makes it hard to devise a fair test: if you ask an AI a question, maybe it just saw the answer somewhere before! This is a particular problem for advanced mathematics. Producing the answer to a specific math puzzle is often easy once you know the “trick.” It's finding the trick that’s the real challenge. To test whether AIs are truly able to reason mathematically, we need a set of questions that are both extremely challenging and brand new—questions that have never been posted online before. That’s exactly what Frontier Math is: a set of very challenging math questions written specifically to test AIs. The whole benchmark set is private, so no AI can simply memorize the answers. Another challenge in testing AI’s ma

o3, Frontier Math, and the Future of Mathematics
Mike Dodds
2 min read


