arXiv:2503.09081cs.CVcs.AI2025-03被引 3

将视觉音频统一转为结构化文本,提升长视频问答准确率

Everything Can Be Described in Words: A Simple Unified Multi-Modal Framework with Semantic and Temporal Alignment

  • 用结构化文本统一处理多模态数据
  • 长视频问答准确率提升最高达16.9%
  • 适合需要统一多模态推理的场景

尽管多模态学习已取得显著进展,但现有方法在不同模态的表征与推理上常存在不一致。本文提出UMaT,一种理论基础坚实的框架,将视觉和听觉输入统一为大语言模型可处理的结构化文本,解决了语义对齐、时间同步及高效稀疏信息检索问题。通过减少冗余并采用结构化文本表示,实现了统一的多模态推理,使长视频问答准确率提升高达13.7%(最长视频提升16.9%)。

原文摘要 · Abstract (English)

While multi-modal learning has advanced significantly, current approaches often create inconsistencies in representation and reasoning of different modalities. We propose UMaT, a theoretically-grounded framework that unifies visual and auditory inputs as structured text for large language models, addressing semantic alignment, temporal synchronization, and efficient sparse information retrieval. It significantly improves state-of-the-art Long Video Question Answering accuracy (up to 13.7%, and 16.9% on long videos) via redundancy minimization and structured textual representation for unified multi-modal reasoning

多模态视频问答语言模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。