首个评估多模态模型带引用生成能力的基准,解决幻觉问题。
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
- 构建多模态引用生成评测集,涵盖学术论文与评审互动数据。
- 模型在多模态输入下引用生成可靠性差,存在系统性模态偏见。
- 揭示模型依赖不同信息源的内部机制,指导未来研究方向。
多模态大语言模型(MLLMs)在融合多种模态方面取得进展,但常出现幻觉。一种有前景的解决方案是生成带引用的文本,提供可验证的透明链条。然而,现有工作主要聚焦于纯文本的引用生成,对多模态场景的挑战关注不足。本文提出MCiteBench,首个用于评估MLLM在多模态上下文中生成带引用文本能力的基准。该基准数据源自学术论文及评审-反驳互动,包含多样信息源和多模态内容。实验结果表明,当处理多模态输入时,MLLM难以可靠地进行输出溯源。进一步分析揭示了系统性的模态偏见,并发现模型在生成引用时对不同来源的内部依赖模式,为理解模型行为提供了洞见,也为未来多模态引用任务指明了方向。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have advanced in integrating diverse modalities but frequently suffer from hallucination. A promising solution to mitigate this issue is to generate text with citations, providing a transparent chain for verification. However, existing work primarily focuses on generating citations for text-only content, leaving the challenges of multimodal scenarios largely unexplored. In this paper, we introduce MCiteBench, the first benchmark designed to assess the ability of MLLMs to generate text with citations in multimodal contexts. Our benchmark comprises data derived from academic papers and review-rebuttal interactions, featuring diverse information sources and multimodal content. Experimental results reveal that MLLMs struggle to ground their outputs reliably when handling multimodal input. Further analysis uncovers a systematic modality bias and reveals how models internally rely on different sources when generating citations, offering insights into model behavior and guiding future directions for multimodal citation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。