解决多模态大模型跨模态指代对齐难题,提升推理可靠性。
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs

- 将跨模态指代识别建模为定位与重识别任务,构建新评估框架。
- 13个模型均暴露核心指代对齐缺陷,引入两种方法显著提升性能。
- 适合研究多模态推理、模型可解释性的研究人员参考。
全模态大语言模型(Omni-LLMs)在整体多模态感知方面表现出色,但在需要协同多模态推理的复杂场景中表现不佳。除了理解全局多模态上下文外,有效推理还依赖于细粒度的跨模态对齐,尤其是跨模态共享指代物的识别,但该问题长期被忽视。为此,我们将挑战形式化为跨模态指代问题:模型需在源模态中定位一个指代物,并在目标模态中重新识别它。基于此范式,我们提出了CrossOmni数据集,包含九项任务及人工设计的推理理由,用于评估和增强该能力。在13个Omni-LLMs上的实验揭示了系统性弱点,归因于缺乏指代感知的思维模式。为此,我们采用两种策略改进跨模态对齐:一种无需训练的上下文学习方法,另一种基于SFT+GRPO的训练方法,旨在诱导此类思维模式。两种方法均带来显著性能提升,并在协作推理任务上具有良好泛化性。总体而言,我们的研究强调跨模态指代对齐是实现鲁棒全模态推理的关键缺失环节。
原文摘要 · Abstract (English)
Omni Large Language Models (Omni-LLMs) have demonstrated impressive capabilities in holistic multi-modal perception, yet they consistently falter in complex scenarios requiring synergistic omni-modal reasoning. Beyond understanding global multimodal context, effective reasoning also hinges on fine-grained cross-modal alignment, especially identifying shared referents across modalities, yet this aspect has been largely overlooked. To bridge this gap, we formalize the challenge as a cross-modal coreference problem, where a model must localize a referent in a source modality and re-identify it in a target modality. Building on this paradigm, we introduce CrossOmni, a dataset comprising nine tasks equipped with human-designed reasoning rationales to evaluate and enhance this capability. Experiments on 13 Omni-LLMs reveal systematic weaknesses in cross-modal coreference, which we attribute to the absence of coreference-aware thinking patterns. To address this, we enhance cross-modal alignment via two strategies: a training-free In-Context Learning method and a training-based SFT+GRPO framework designed to induce such thinking patterns. Both approaches yield substantial performance gains and generalize effectively to collaborative reasoning tasks. Overall, our findings highlight cross-modal coreference as a crucial missing piece for advancing robust omni-modal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。