从分阶段处理到统一建模,推动多模态问答性能提升
Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language Representation
- 通过对比三种框架,分析多模态问答的演进路径
- 统一文本中心架构使EM和F1得分显著提高
- 适合关注多模态推理系统设计的研究者
多模态数据的快速增长催生了对跨异构数据源(如文本、表格、图像)进行推理的问答系统需求。本文系统比较了三种代表性框架——多模态自适应提取(MAE)、Solar与UniMMQA,梳理了多模态问答从模态自适应流水线到统一架构的演进历程。研究聚焦各方法如何建模跨模态交互、转换异构输入及执行推理,揭示其在模态表示、推理机制与答案生成上的关键差异。分析表明,由预训练语言模型(PLMs)驱动的统一文本中心范式正取代显式的模态特异性处理,带来显著性能提升:在基准数据集上,该转型显著提高精确匹配(EM)与F1分数,其中UniMMQA表现最稳定且可扩展。然而,仍存在模态转换中的信息丢失、多阶段流水线中的误差传播及细粒度跨模态依赖捕捉不足等挑战。本研究深化了对当前设计趋势的理解,并为统一多模态推理系统的发展方向提供洞见。
原文摘要 · Abstract (English)
The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of reasoning across heterogeneous sources such as text, tables, and images. In this paper, we present a comprehensive methodological comparison of three influential frameworks, namely Multimodal Adaptive Extraction (MAE), Solar, and UniMMQA, tracing the evolution of multimodal question answering from modality-adaptive pipelines to fully unified architectures. We examine how each approach models cross-modal interactions, transforms heterogeneous inputs, and performs reasoning, highlighting key design differences in modality representation, reasoning, and answer generation. Our analysis demonstrates a clear shift from explicit modality-specific processing toward unified text-centric formulations enabled by pre-trained language models (PLMs). Empirical comparisons across benchmark datasets show that this transition leads to substantial improvements in both Exact Match (EM) and F1-Scores, with UniMMQA achieving the most consistent and scalable performance. Despite these advances, we identify persistent challenges, including information loss during modality transformation, error propagation in multi-stage pipelines, and limitations in capturing fine-grained cross-modal dependencies. Overall, this study provides a deeper understanding of current design trends and offers insights into the future direction of unified multimodal reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。