arXiv:2601.10323cs.CVcs.CL2026-01ACL被引 3

实时处理音视频流的全能助手,能主动预警也能即时响应。

ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding

论文配图:ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding
图 1 · 摘自论文原文
  • 将音频与视频帧同步对齐,解决模态粒度不匹配问题。
  • 引入轻量级发声头,实现响应触发与生成解耦,精准控制时机。
  • 在12个基准上验证,主动任务表现领先,兼顾实时交互能力。

近期的全模态大语言模型在统一建模音频、视觉和文本方面展现出潜力,但流式音视频理解仍具挑战:现有方法普遍存在模态支持不全或缺乏自主主动监控能力。为此,我们提出ROMA,一个用于统一反应式与主动式交互的实时全模态助手。ROMA将连续输入处理为同步的多模态单元,通过将密集音频与离散视频帧对齐,缓解粒度差异。针对在线决策,我们设计了一个轻量级发声头,将响应触发与生成过程解耦,确保精确触发且无任务冲突。通过自建流式数据集及两阶段课程训练,逐步优化模型对流式格式的适应性与主动性响应能力。为统一评估标准,我们将多样基准整合为涵盖主动(警报、叙述)与反应(问答)场景的统一评测套件。在12个基准上的大量实验表明,ROMA在主动任务上达到当前最优性能,同时在反应任务中保持竞争力,验证了其在统一实时全模态理解中的鲁棒性。

原文摘要 · Abstract (English)

Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing approaches suffer from disjointed capabilities: they typically exhibit incomplete modality support or lack autonomous proactive monitoring. To address this, we present ROMA, a real-time omni-multimodal assistant for unified reactive and proactive interaction. ROMA processes continuous inputs as synchronized multimodal units, aligning dense audio with discrete video frames to handle granularity mismatches. For online decision-making, we introduce a lightweight speak head that decouples response initiation from generation to ensure precise triggering without task conflict. We train ROMA with a curated streaming dataset and a two-stage curriculum that progressively optimizes for streaming format adaptation and proactive responsiveness. To standardize the fragmented evaluation landscape, we reorganize diverse benchmarks into a unified suite covering both proactive (alert, narration) and reactive (QA) settings. Extensive experiments across 12 benchmarks demonstrate ROMA achieves state-of-the-art performance on proactive tasks while competitive in reactive settings, validating its robustness in unified real-time omni-multimodal understanding.

多模态实时系统主动交互语音视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。