LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches
Abstract
Mathematical reasoning is widely regarded as a hallmark of human intelligence, and determining whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligence and cognitive science. Moreover, the growing integration of LLMs into scientific workflows creates a practical need for rigorous evaluation of their mathematical capabilities. Existing mathematical reasoning benchmarks, however, are often limited by synthetic problem settings and growing contamination from widely circulated datasets. We present LiveMathematicianBench, a dynamic multiple-choice benchmark for evaluating research-level mathematical reasoning using recent arXiv papers published after model training cutoffs. By grounding evaluation in newly published theorem statements, the benchmark provides a more realistic testbed for assessing whether models can reason about natural mathematical claims beyond memorized benchmark patterns. LiveMathematicianBench introduces a taxonomy of thirteen problem categories based on the logical form of theorem statements, enabling fine-grained evaluation across reasoning types such as implication, equivalence, existence, and uniqueness. It further introduces a proof-sketch-guided distractor generation pipeline, in which proof sketches are used to construct plausible but invalid answer choices that reflect misleading proof directions. This design makes the benchmark more sensitive to genuine mathematical understanding rather than surface-level answer matching. We evaluate models under both selection and sketch-aware selection settings to distinguish answer recognition from reasoning supported by proof-level cues. Overall, LiveMathematicianBench offers a scalable and continuously updated benchmark for studying research-level mathematical reasoning in large language models.