用细粒度评分提升多模态模型的复杂推理能力
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
- 通过子问题级反馈实现精准打分,避免粗略二元评价
- 在12个基准上6项领先,新数据集STEM-Bench表现突出
- 适合需要严谨逻辑验证的AI推理研究者使用
现有视觉-语言模型在涉及多个问题的复杂推理任务中表现不佳,传统奖励机制仅提供单一二值评分,难以有效指导模型处理多步骤问题。为此,我们提出StructVRM,一种将多模态推理与结构化可验证奖励模型对齐的方法。核心是训练一个基于模型的验证器,可在子问题层面提供细粒度反馈,评估语义和数学等价性,而非依赖严格字符串匹配,从而实现此前难以实现的局部得分机制。大量实验表明,所训练的Seed-StructVRM模型在十二个公开多模态基准中的六个及新构建的高难度STEM-Bench数据集上达到当前最优性能。结果验证了使用结构化可验证奖励进行训练是提升多模态模型在复杂现实推理领域能力的有效路径。
原文摘要 · Abstract (English)
Existing Vision-Language Models often struggle with complex, multi-question reasoning tasks where partial correctness is crucial for effective learning. Traditional reward mechanisms, which provide a single binary score for an entire response, are too coarse to guide models through intricate problems with multiple sub-parts. To address this, we introduce StructVRM, a method that aligns multimodal reasoning with Structured and Verifiable Reward Models. At its core is a model-based verifier trained to provide fine-grained, sub-question-level feedback, assessing semantic and mathematical equivalence rather than relying on rigid string matching. This allows for nuanced, partial credit scoring in previously intractable problem formats. Extensive experiments demonstrate the effectiveness of StructVRM. Our trained model, Seed-StructVRM, achieves state-of-the-art performance on six out of twelve public multimodal benchmarks and our newly curated, high-difficulty STEM-Bench. The success of StructVRM validates that training with structured, verifiable rewards is a highly effective approach for advancing the capabilities of multimodal models in complex, real-world reasoning domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。