arXiv:2607.07391cs.AI2026-07

测试模型如何用最少信息问出关键问题并解数学题。

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

论文配图:MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning
图 1 · 摘自论文原文
  • 设计难题只缺一个必要事实,需自然语言提问获取
  • 2310个题目覆盖代数到信号处理等22类数学问题
  • 能准确提问但算不对,或根本提不出问题,可分诊模型短板

现有数学推理评测通常提供全部解题所需信息,而交互式评测又混入工具、检索和长对话。我们提出MIRA-Math,聚焦一种更窄的能力:解决隐含状态唯一确定但求解者视角缺失一个必要原子事实的数学问题。求解者必须在严格预算下以自然语言请求缺失信息,并将返回的事实整合为精确答案。固定约束的LLM响应者仅知晓数据集提供的原子事实,若请求匹配则原样返回,否则拒绝。因此实例生成、提示类型、验证与最终答案判定均为确定性过程,请求指标在固定LLM中介通道下测量。MIRA-Math包含从22类数学主题生成的2,310个实例,涵盖代数、概率、线性系统、离散结构、信号处理、马尔可夫链、电路、插值及数值边值问题。前沿与小型模型实验表明,提问成功率与最终答案准确率可分离:模型可能问对但计算失败,或未获标准提示即已失败。我们开放生成器、验证器、提示、运行元数据与文档,支持可复现的最小信息请求评估。

原文摘要 · Abstract (English)

Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and long-horizon dialogue. We introduce MIRA-Math, a benchmark for a narrower diagnostic capability: solving mathematical problems whose full latent state has a unique answer, but whose solver-facing view is missing exactly one necessary atomic fact. The solver must request the missing information in natural language under a strict budget and then integrate the returned fact into an exact final answer. A fixed constrained LLM responder sees only the dataset-provided atomic fact and must either offer the quoted fact when the request matches it, or decline otherwise. Thus, instance generation, typed hint specifications, validation, and final-answer verification are deterministic, while request metrics are measured under a fixed LLM-mediated responder channel. MIRA-Math contains 2{,}310 generated instances from 22 typed mathematical families spanning algebra, probability, linear systems, discrete structures, signal processing, Markov chains, circuits, interpolation, and numerical boundary-value problems. Experiments across frontier and small models show that request success and final-answer accuracy are separable: models may ask for the right fact yet fail the downstream computation, or fail before obtaining the canonical hint. We release generators, verifiers, prompts, run metadata, and dataset documentation to support reproducible evaluation of minimal information requesting in mathematical reasoning.

数学推理信息请求评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。