arXiv:2510.20193cs.IRcs.CL2025-10综述被引 3

综述多模态问答的检索与跨模态推理架构,助你快速掌握前沿方法。

Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures

  • 按检索方式、融合策略、生成路径分类多模态问答架构
  • 分析多个基准数据集上的性能差异与精度-延迟权衡
  • 适合关注跨模态对齐与上下文感知系统的研究者

问答系统传统上依赖结构化文本数据,但多媒体内容(图像、音频、视频及结构化元数据)的快速增长为检索增强型问答带来了新挑战与机遇。本文综述了集成多媒体检索流程的问答系统最新进展,重点关注视觉、语言和音频模态与用户查询对齐的架构。我们基于检索方法、融合技术与答案生成策略对方法进行分类,并分析了基准数据集、评估协议及性能权衡。此外,我们指出关键挑战,如跨模态对齐、延迟-精度权衡与语义定位问题,并提出开放问题与未来研究方向,以构建更鲁棒、更具上下文感知能力的多媒体问答系统。

原文摘要 · Abstract (English)

Question Answering (QA) systems have traditionally relied on structured text data, but the rapid growth of multimedia content (images, audio, video, and structured metadata) has introduced new challenges and opportunities for retrieval-augmented QA. In this survey, we review recent advancements in QA systems that integrate multimedia retrieval pipelines, focusing on architectures that align vision, language, and audio modalities with user queries. We categorize approaches based on retrieval methods, fusion techniques, and answer generation strategies, and analyze benchmark datasets, evaluation protocols, and performance tradeoffs. Furthermore, we highlight key challenges such as cross-modal alignment, latency-accuracy tradeoffs, and semantic grounding, and outline open problems and future research directions for building more robust and context-aware QA systems leveraging multimedia data.

多模态问答跨模态对齐检索增强评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。