将视觉音频统一到文本空间,提升长视频理解能力
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding
- 用信息论优化跨模态语义对齐,统一视觉与听觉输入
- 长视频问答准确率提升22.6%,超30分钟视频增益达27.3%
- 适合需要长序列推理和多模态融合的研究者
尽管多模态学习取得显著进展,现有方法常独立处理不同模态,导致表征与推理不一致。我们提出MANTA(通过文本对齐实现多模态抽象与归一化),一个理论驱动的框架,将视觉与听觉输入统一至结构化文本空间,以无缝对接大语言模型。MANTA解决四大挑战:(1) 基于信息论优化的跨模态语义对齐;(2) 针对信息密度差异的自适应时间同步;(3) 多尺度内容层级表示;(4) 长序列中稀疏信息的上下文感知检索。我们在长视频问答任务上进行大量实验,结果显示MANTA使最先进模型整体准确率提升22.6%,尤其在超过30分钟的视频上提升达27.3%。此外,在时间推理任务上提升23.8%,跨模态理解提升25.1%。框架引入新颖的密度估计技术,有效减少冗余同时保留罕见信号,为通过结构化文本统一多模态表征奠定新基础。
原文摘要 · Abstract (English)
While multi-modal learning has advanced significantly, current approaches often treat modalities separately, creating inconsistencies in representation and reasoning. We introduce MANTA (Multi-modal Abstraction and Normalization via Textual Alignment), a theoretically-grounded framework that unifies visual and auditory inputs into a structured textual space for seamless processing with large language models. MANTA addresses four key challenges: (1) semantic alignment across modalities with information-theoretic optimization, (2) adaptive temporal synchronization for varying information densities, (3) hierarchical content representation for multi-scale understanding, and (4) context-aware retrieval of sparse information from long sequences. We formalize our approach within a rigorous mathematical framework, proving its optimality for context selection under token constraints. Extensive experiments on the challenging task of Long Video Question Answering show that MANTA improves state-of-the-art models by up to 22.6% in overall accuracy, with particularly significant gains (27.3%) on videos exceeding 30 minutes. Additionally, we demonstrate MANTA's superiority on temporal reasoning tasks (23.8% improvement) and cross-modal understanding (25.1% improvement). Our framework introduces novel density estimation techniques for redundancy minimization while preserving rare signals, establishing new foundations for unifying multimodal representations through structured text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。