提出O-MARC压缩框架,让视频理解模型更高效准确。
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

- 无需训练的压缩方法保留关键视觉记忆和音频锚点
- 在4个基准上平均得分45.8,优于全序列推理
- 适合追求高效视频理解的开发者和研究者
多模态大语言模型实现统一音视频理解,但长联合标记序列导致推理成本高,现有基准未能充分隔离用户生成视频中的音视频关联。我们提出UGC-AVQA,一个包含1,000个视频和4,816个问答对的公开基准,通过音频移除测试确保问题需同时依赖声学与视觉证据。为降低推理开销,我们提出OMAC,一种无需训练的即插即用压缩方法,保留显著视觉记忆和时序对齐的音频锚点。为进一步提升紧凑模型对压缩输入的鲁棒性,引入O-MARC压缩蒸馏框架,用于学习记忆压缩的多模态上下文。在Qwen2.5-Omni-3B上,O-MARC将四个基准平均得分提升至45.8,优于全标记序列推理的44.1和OmniZip的41.0。OMAC还保持高效推理,相比全标记推理降低34.6%延迟(1.53倍加速)和34.7%内存占用。
原文摘要 · Abstract (English)
Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio visual association in noisy user generated videos. We introduce UGC-AVQA, a public UGC benchmark with 1,000 videos and 4,816 QA pairs, where an audio removal test ensures that benchmark questions require both acoustic and visual evidence. To reduce inference cost, we propose OMAC, a training free plug in compression method that preserves salient visual memory and temporally grounded audio anchors. To further make compact models robust to compressed inputs, we introduce O-MARC, a compression distillation framework for learning with memory compressed multimodal contexts. On Qwen2.5-Omni-3B, O-MARC improves the average score across four benchmarks to 45.8, outperforming full token inference at 44.1 and OmniZip at 41.0. OMAC also keeps inference efficient, reducing latency by 34.6\% (1.53$\times$ speedup) and memory by 34.7\% compared with full token inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。