arXiv:2510.14738cs.CL2025-10ACL被引 9

用自动生成的评分标准提升多模态模型推理可靠性

AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning

  • 通过自聚合方法自动提取有效推理路径,构建无须人工标注的评分规则
  • 在6个基准上达到顶尖性能,推理过程更符合逻辑
  • 适合需要可信推理的AI系统开发者和评测研究者

多模态大模型已从感知任务发展到复杂多步推理,但基于最终答案正确性的强化学习常导致虚假推理。为此,我们提出AutoRubric框架,将强化学习与过程级监督结合,通过自动收集的评分标准生成奖励。核心创新在于可扩展的自聚合方法,能从成功轨迹中提炼一致的推理节点,实现无需人工标注或强教师模型的问题特定评分规则构建。联合使用评分标准奖励与结果奖励后,AutoRubric在六个多模态推理基准上取得当前最佳表现,并在专项评估中显著提升推理忠实度。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) often leads to spurious reasoning since only the final-answer correctness is rewarded. To address this limitation, we propose AutoRubric, a framework that integrates RLVR with process-level supervision through automatically collected rubric-based generative rewards. Our key innovation lies in a scalable self-aggregation method that distills consistent reasoning checkpoints from successful trajectories, enabling problem-specific rubric construction without human annotation or stronger teacher models. By jointly leveraging rubric-based and outcome rewards, AutoRubric achieves state-of-the-art performance on six multimodal reasoning benchmarks and substantially improves reasoning faithfulness in dedicated evaluations.

多模态推理强化学习评分标准生成奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。