arXiv:2510.26160cs.CV2025-10KDD被引 11

首个面向可穿戴设备的多模态多轮问答基准,填补领域空白。

CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark

  • 构建6.5K图像-问题-答案三元组与2K多轮对话数据集。
  • 单轮与多轮问答真实度仅32%和43%,显示现有方法严重不足。
  • 适合研究多模态RAG、可穿戴交互及对话系统优化者使用。

智能眼镜等可穿戴设备正在改变人们与环境互动的方式,使用户能够针对视野中的实体获取信息。多模态检索增强生成(MM-RAG)在支持此类问题中起关键作用,但目前尚无针对可穿戴场景的综合性基准。为此,我们提出CRAG-MM——一个面向多模态多轮对话的综合检索增强生成基准。该基准包含6.5K个(图像,问题,答案)三元组和2000个跨13个领域的视觉驱动多轮对话,其中6.2K张为模拟可穿戴设备拍摄的第一人称图像。问题设计涵盖五类图像质量缺陷、六种问题类型、不同实体流行度、信息动态性差异以及不同对话轮次。设置三项任务:单源增强、多源增强和多轮对话,每项均配备图像-知识图谱与网页检索的专用语料库及API。评估显示,简单RAG方法在单轮和多轮问答上的真实度仅为32%和43%,行业领先方案也仅达32%/45%,表明仍有巨大提升空间。该基准已举办KDD Cup 2025,吸引约1000名参赛者提交5000份提交,优胜方案相较基线提升28%,凸显其对领域发展的早期推动作用。

原文摘要 · Abstract (English)

Wearable devices such as smart glasses are transforming the way people interact with their surroundings, enabling users to seek information regarding entities in their view. Multi-Modal Retrieval-Augmented Generation (MM-RAG) plays a key role in supporting such questions, yet there is still no comprehensive benchmark for this task, especially regarding wearables scenarios. To fill this gap, we present CRAG-MM -- a Comprehensive RAG benchmark for Multi-modal Multi-turn conversations. CRAG-MM contains a diverse set of 6.5K (image, question, answer) triplets and 2K visual-based multi-turn conversations across 13 domains, including 6.2K egocentric images designed to mimic captures from wearable devices. We carefully constructed the questions to reflect real-world scenarios and challenges, including five types of image-quality issues, six question types, varying entity popularity, differing information dynamism, and different conversation turns. We design three tasks: single-source augmentation, multi-source augmentation, and multi-turn conversations -- each paired with an associated retrieval corpus and APIs for both image-KG retrieval and webpage retrieval. Our evaluation shows that straightforward RAG approaches achieve only 32% and 43% truthfulness on CRAG-MM single- and multi-turn QA, respectively, whereas state-of-the-art industry solutions have similar quality (32%/45%), underscoring ample room for improvement. The benchmark has hosted KDD Cup 2025, attracting about 1K participants and 5K submissions, with winning solutions improving baseline performance by 28%, highlighting its early impact on advancing the field.

多模态RAG可穿戴基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。