arXiv:2506.01275cs.AI2025-06EMNLP

测试模型在四种模态间对比推理选信息的能力,发现当前模型仍有明显短板。

Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D

  • 构建四模态对比推理数据集,让模型选最符合问题的模态
  • 顶尖模型在四模态下准确率仅42%,整体56%
  • 适合研究多模态推理、信息筛选与评估方法的学者

现实决策常始于判断哪个模态包含最相关信息。尽管近年多模态模型在处理多样输入上取得进展,但其能否在多个模态间进行对比推理以选出最符合自然语言提示的模态仍不明确。我们认为这一能力至关重要,尤其在检索增强和决策时场景中,系统需评估多个信号并识别相关者。为此,我们提出Contra4,一个涵盖图像、音频、视频和3D的跨模态对比推理数据集。每个样本包含一个自然语言问题和多个候选模态实例,模型需选择语义最匹配的。Contra4结合人工标注描述与混合模型往返一致性过滤,确保高质量监督,生成17.4万条训练样本和2.3千条人工验证的测试集。任务微调使性能相对基线提升56%,但顶级模型在整体任务中仅达56%准确率,四模态设置下为42%,凸显当前多模态模型在对比推理上的显著局限。

原文摘要 · Abstract (English)

Real-world decision-making often begins with identifying which modality contains the most relevant information for a given query. While recent multimodal models have made impressive progress in processing diverse inputs, it remains unclear whether they can reason contrastively across multiple modalities to select the one that best satisfies a natural language prompt. We argue this capability is foundational, especially in retrieval-augmented and decision-time contexts, where systems must evaluate multiple signals and identify which one conveys the relevant information. To evaluate this skill, we introduce Contra4, a dataset for contrastive cross-modal reasoning across four modalities: image, audio, video, and 3D. Each example presents a natural language question alongside multiple candidate modality instances, and the model must select the one that semantically aligns with the prompt. Contra4 combines human-annotated captions with a mixture-of-models round-trip-consistency filter to ensure high-quality supervision, resulting in 174k training examples and a manually verified test set of 2.3k samples. While task-specific fine-tuning helps improve performance by 56% relative to baseline, state-of-the-art models still achieve only an absolute of 56% accuracy overall and 42% in four-modality settings, underscoring a significant limitation in current multimodal models.

多模态对比推理信息筛选数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。