arXiv:2605.17360cs.CV2026-05

首个实时双向多模态交互评估基准,解决AI响应时机与内容质量难题。

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

论文配图:Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
图 1 · 摘自论文原文
  • 构建双场景评估框架:持续描述与主动提醒,测试模型实时响应能力。
  • 最先进模型仅39.6%得分,主动提醒任务仅20.0%,暴露时机与内容双重短板。
  • 引入基于大模型的自动评测,精准衡量响应时间与内容一致性。

实时双向交互对真实场景下的多模态AI系统至关重要,要求模型持续处理流式输入并适时响应。然而,现有多模态大模型(MLLMs)大多在离线环境下评估,即全部视频输入完成后才生成回应。尽管近期研究开始探索实时双向MLLMs,但尚无系统性基准或自动化评估方法。为此,我们提出Omni-DuplexEval,一个用于系统评估实时双向交互的基准。该基准包含两个互补场景:(1) 实时描述,评估模型生成随多模态输入动态演进的时间对齐连续响应的能力;(2) 主动提醒,评估模型识别关键事件并在恰当时刻响应的能力。Omni-DuplexEval包含660个视频,附有细粒度人工标注标签和精确时间元数据,覆盖9个基于真实场景的任务,所有问题均为开放式提问。我们进一步提出基于LLM-as-a-Judge的自动化评估框架,通过时间感知与序列推理联合评估响应内容一致性与响应时机,与人类判断高度一致。在主流双向MLLM上的实验揭示显著局限:最优模型整体得分仅39.6%,主动提醒任务仅20.0%。分析指出两大挑战:模型难以平衡及时响应与连贯整体内容生成,且常无法同时确定何时响应及何内容回应。我们期望本工作推动MLLMs在实时交互方向的发展。

原文摘要 · Abstract (English)

Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming inputs and respond at appropriate moments. However, most existing multimodal large language models (MLLMs) are evaluated in offline settings, where the entire video input is processed before any response is generated. While recent work has started to explore real-time duplex MLLMs, there is still no comprehensive benchmark or automatic evaluation method for this setting. To address this gap, we propose Omni-DuplexEval, a benchmark for systematically evaluating real-time duplex interaction. The benchmark consists of two complementary scenarios: (1) Real-Time Description, which evaluates the ability to generate continuous, time-aligned responses that track evolving multimodal inputs, and (2) Proactive Reminder, which evaluates the ability to identify salient events and respond at appropriate moments. Omni-DuplexEval contains 660 videos with fine-grained, human-annotated labels and precise temporal metadata, spanning 9 tasks grounded in real-world scenarios, where all questions are formulated as open-ended queries. We further introduce an automatic evaluation framework based on LLM-as-a-Judge, which enables systematic assessment by jointly evaluating response-content alignment and response timing through timestamp-aware and sequential reasoning, achieving strong alignment with human judgments. Experiments on state-of-the-art duplex MLLMs reveal substantial limitations. The best-performing model achieves only 39.6% overall, while scoring only 20.0% on Proactive Reminder. Our analysis identifies two key challenges: models struggle to balance timely responses with coherent, holistic content generation, and they often fail to determine both when to respond and what to produce. We hope our work facilitates further progress in MLLMs.

多模态实时交互评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。