新基准评估多模态推理过程,揭示模型隐藏缺陷。
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
- 设计可验证中间步骤的诊断性评测集,支持细粒度分析。
- 发现模型普遍存在选择性输出(精度远高于召回)和推理顺序混乱问题。
- 提出因果过程奖励机制,无需人工标注即可提升推理能力。
我们提出CRYSTAL(Clear Reasoning via Yielded Steps, Traceability, and Logic)基准,包含6,372个实例,通过可验证的中间步骤评估多模态推理能力。提出两个互补指标:Match F1基于语义相似性匹配计算步骤级精确率与召回率,Ordered Match F1进一步惩罚推理链顺序错误。参考答案通过类德尔菲流程生成:四个独立的多模态大语言模型生成推理轨迹,经语义聚类整合并由人工质量审核。对20个MLLMs(包括未参与构建的商业前沿系统)的评估揭示了答案准确率无法察觉的系统性失败:普遍存在的选择性输出(精度远高于召回)、非单调缩放权衡、以及推理顺序错乱——无一模型在正确顺序中保留超过60%的匹配步骤。除评估外,提出因果过程奖励(CPR),一种耦合答案正确性与步骤对齐的乘法奖励;CPR-Curriculum则逐步提升训练难度。在GRPO下,该方法使Match F1提升32%,而加法奖励策略失效。
原文摘要 · Abstract (English)
We introduce CRYSTAL (Clear Reasoning via Yielded Steps, Traceability, and Logic), a diagnostic benchmark with 6,372 instances that evaluates multimodal reasoning through verifiable intermediate steps. We propose two complementary metrics: Match F1, which scores step-level precision and recall via semantic similarity matching, and Ordered Match F1, which further penalizes disordered reasoning chains. References are constructed through a Delphi-inspired pipeline in which four independent MLLMs generate trajectories, which are then aggregated via semantic clustering and validated through human quality gates. Evaluation of 20 MLLMs, including commercial frontier systems not used during benchmark construction, reveals systematic failures that are invisible to answer accuracy: universal cherry-picking (precision far exceeds recall), non-monotonic scaling trade-offs, and disordered reasoning in which no competitive model preserves more than 60% of matched steps in the correct order. Beyond evaluation, we propose the Causal Process Reward (CPR), a multiplicative reward that couples answer correctness with step-level alignment, and CPR-Curriculum, which progressively increases reasoning difficulty during training. CPR-Curriculum achieves a 32% improvement in Match F1 via GRPO where additive reward strategies fail, improving reasoning without manual step annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。