arXiv:2606.07801cs.AI2026-06

通过优化最弱模态维度,提升多模态推理的可靠性

Improving Multimodal Reasoning via Worst Dimension Optimization

论文配图:Improving Multimodal Reasoning via Worst Dimension Optimization
图 1 · 摘自论文原文
  • 针对多模态推理中各维度表现不均的问题,提出最弱维度优化方法
  • 在MMMLU和VQAv2数据集上,推理准确率提升3.2%和2.8%
  • 适合需要鲁棒多模态推理的系统开发者与研究者

多模态推理需要在从视觉定位到逻辑一致性等多重约束下保持路径完整性。然而,当前过程奖励模型依赖启发式定义的奖励,并对各项因素同等加权,可能导致某些维度的失败被主导因素掩盖,无法保证推理过程在一般情况下的有效性。

原文摘要 · Abstract (English)

Multimodal reasoning requires a path that retains integrity over a wide range of constraints, from visual grounding to logic consistency. However, the current Process Reward Models focus on heuristically defined rewards that equally weigh these factors, which may lead to the concealment of individual dimension failures by the dominating factors, without guaranteeing the validity of the reasoning process in general.

多模态推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。