多模态检索增强生成系统,让机器人更懂复杂场景下的实时信息。
Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding
- 融合视频、音频、文本多源信息,动态推理决策
- 在MMBench-Video上性能超越开源模型,用更少资源完成任务
- 适合需要实时理解与辅助的智能机器人应用
为有效融入人类社会,智能体需具备适应变化、过滤信息并做出明智决策的能力。随着机器人日益融入日常生活,将认知负担转移给系统变得愈发重要,尤其在信息密集且动态的场景中。为此,我们提出Multi-RAG——一种多模态检索增强生成系统,旨在通过整合与推理多源信息流(包括视频、音频和文本),提升情境理解能力并减轻人类认知负荷。作为实现长期人机协作的关键一步,Multi-RAG探索了多模态理解如何成为动态、以人为中心场景下自适应机器人辅助的基础。我们在挑战性的多模态视频理解基准MMBench-Video上评估其性能,结果表明,相较于现有开源视频大语言模型(Video-LLMs)和大视觉语言模型(LVLMs),Multi-RAG在更少资源与输入数据条件下实现了更优表现,展现了其在真实世界动态场景中构建高效人机自适应辅助系统的潜力。
原文摘要 · Abstract (English)
To effectively engage in human society, the ability to adapt, filter information, and make informed decisions in ever-changing situations is critical. As robots and intelligent agents become more integrated into human life, there is a growing opportunity-and need-to offload the cognitive burden on humans to these systems, particularly in dynamic, information-rich scenarios. To fill this critical need, we present Multi-RAG, a multimodal retrieval-augmented generation system designed to provide adaptive assistance to humans in information-intensive circumstances. Our system aims to improve situational understanding and reduce cognitive load by integrating and reasoning over multi-source information streams, including video, audio, and text. As an enabling step toward long-term human-robot partnerships, Multi-RAG explores how multimodal information understanding can serve as a foundation for adaptive robotic assistance in dynamic, human-centered situations. To evaluate its capability in a realistic human-assistance proxy task, we benchmarked Multi-RAG on the MMBench-Video dataset, a challenging multimodal video understanding benchmark. Our system achieves superior performance compared to existing open-source video large language models (Video-LLMs) and large vision-language models (LVLMs), while utilizing fewer resources and less input data. The results demonstrate Multi- RAG's potential as a practical and efficient foundation for future human-robot adaptive assistance systems in dynamic, real-world contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。