用自动生成的评分标准提升多模态模型推理可靠性
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
- 通过自聚合方法自动提取有效推理路径,构建无须人工标注的评分规则
- 在6个基准上达到顶尖性能,推理过程更符合逻辑
- 适合需要可信推理的AI系统开发者和评测研究者
多模态大模型已从感知任务发展到复杂多步推理,但基于最终答案正确性的强化学习常导致虚假推理。为此,我们提出AutoRubric框架,将强化学习与过程级监督结合,通过自动收集的评分标准生成奖励。核心创新在于可扩展的自聚合方法,能从成功轨迹中提炼一致的推理节点,实现无需人工标注或强教师模型的问题特定评分规则构建。联合使用评分标准奖励与结果奖励后,AutoRubric在六个多模态推理基准上取得当前最佳表现,并在专项评估中显著提升推理忠实度。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) often leads to spurious reasoning since only the final-answer correctness is rewarded. To address this limitation, we propose AutoRubric, a framework that integrates RLVR with process-level supervision through automatically collected rubric-based generative rewards. Our key innovation lies in a scalable self-aggregation method that distills consistent reasoning checkpoints from successful trajectories, enabling problem-specific rubric construction without human annotation or stronger teacher models. By jointly leveraging rubric-based and outcome rewards, AutoRubric achieves state-of-the-art performance on six multimodal reasoning benchmarks and substantially improves reasoning faithfulness in dedicated evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。