Search
Events

Benchmarks in Leipzig

Published June 8, 2026

Designed by mathematics researchers for mathematics researchers, this hands-on event brought together participants to collaboratively develop research-level benchmarks aimed at exploring the limitations and capabilities of rapidly advancing AI models for mathematical reasoning.

The program combined practical sessions on prompting and benchmark design with collaborative working groups. Participants formulated, refined, and tested research-level mathematical problems. They also discussed the current landscape of AI in mathematics and reflected on what makes a problem genuinely challenging for today’s models.

The focus of the project was the interactive platform Science Bench, which was used to compile a dataset of research-level mathematics questions with known solutions. 

The researchers evaluated 100 benchmark questions in three stages. After five state-of-the-art models attempted the questions in stage one, 41 questions remained unsolved. After additional 20 runs of the best three of those models in stage two, 16 questions remained unsolved. And after the final three runs of two "deep think" models, all except two questions were solved.

To share the outcomes with a broader community, the project leaders report the results - which are largely based on the contributions of the 35 participants of the Benchmark Conference, as well as the work of 14 external researchers – in a recent arXive article. 

The evaluation results demonstrate that the mathematical reasoning capabilities of large language models are becoming increasingly impressive.