arXiv:2512.05943cs.AI2025-12Conference of the …被引 2

让视觉语言模型的推理过程可追踪,发现隐藏错误。

TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models

  • 用子问题对分解复杂任务,逐步检验推理一致性
  • 推理路径一致性与最终答案正确性相关,能定位错误环节
  • 生成可信区间,帮助筛选可靠结果和改进模型

可靠的数学与科学推理仍是大型视觉语言模型面临的开放挑战。标准的最终答案评估常掩盖推理错误,导致隐性失败持续存在。为此,我们提出TRACE框架——透明推理与一致性评估,旨在诊断推理轨迹而非仅关注最终结果。其核心是辅助推理集(Auxiliary Reasoning Sets),即紧凑的子问题-答案对,用于分解复杂问题,通过基于一致性的度量评估中间步骤,揭示传统评估忽略的失败。实验表明,辅助推理集间的连贯性与最终答案正确性相关,并能精确定位失败发生的推理环节,为模型改进提供可操作信号。此外,TRACE定义了置信区域,区分可靠与不可靠的推理路径,支持有效过滤、调试与模型优化。

原文摘要 · Abstract (English)

Reliable mathematical and scientific reasoning remains an open challenge for large vision-language models. Standard final-answer evaluation often masks reasoning errors, allowing silent failures to persist. To address this gap, we introduce TRACE, a framework for Transparent Reasoning And Consistency Evaluation that diagnoses reasoning trajectories rather than only end results. At its core, TRACE leverages Auxiliary Reasoning Sets, compact sub question answer pairs that decompose complex problems, evaluate intermediate steps through consistency-based metrics, and expose failures overlooked by standard evaluation. Our experiments show that consistency across ARS correlates with final-answer correctness and helps pinpoint the reasoning steps where failures arise, offering actionable signals for model improvement. Furthermore, TRACE defines confidence regions that distinguish reliable from unreliable reasoning paths, supporting effective filtering, debugging, and model refinement.

视觉语言模型推理分析一致性评估模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。