arXiv:2603.07966cs.CV2026-03

提出视觉语音对齐新基准,测试模型理解手势指向的实时能力。

Listening with the Eyes: Benchmarking Egocentric Co-Speech Grounding across Space and Time

  • 构建时空同步的双语视频评测集,需同时预测对象、位置和时间。
  • 人类准确率达96.9%,顶尖模型仅17.0%,暴露多模态对齐缺陷。
  • 移除视觉输入后模型性能提升至42.9%,揭示视频界面干扰问题。

在情境协作中,说话者常使用不明确的指示性指令(如“把那个给我”),其指代对象需结合语音与短暂的手势动作才能确定。然而,许多具身任务基准允许语言单一路径,使多模态大模型无需学习音视对齐即可表现良好。为填补这一差距,我们提出「自视角语音-手势对齐(EcoG)」任务,要求模型联合预测‘什么’、‘哪里’和‘何时’。为此,我们构建了EcoG-Bench,一个包含811段自视角视频的双语(英/中)诊断评测集,具备密集空间标注与毫秒级动作监督,并采用渐进式认知评估协议。测试表明:人类在该任务上达到96.9%的严格准确率,而最佳原生视频-音频模型(Gemini-3-Pro)仅为17.0%。进一步消融实验显示,若以带时间戳的帧样本与外部验证的语音识别(词级时间)替代原生多模态接口,同一模型性能提升至42.9%。结果表明,多模态接口可能阻碍时间对齐线索的可观察性,独立于模型推理能力。

原文摘要 · Abstract (English)

In situated collaboration, speakers often use intentionally underspecified deictic commands (e.g., ``pass me \textit{that}''), whose referent becomes identifiable only by aligning speech with a brief co-speech pointing \emph{stroke}. However, many embodied benchmarks admit language-only shortcuts, allowing MLLMs to perform well without learning the \emph{audio--visual alignment} required by deictic interaction. To bridge this gap, we introduce \textbf{Egocentric Co-Speech Grounding (EcoG)}, where grounding is executable only if an agent jointly predicts \textit{What}, \textit{Where}, and \textit{When}. To operationalize this, we present \textbf{EcoG-Bench}, an evaluation-only bilingual (EN/ZH) diagnostic benchmark of \textbf{811} egocentric clips with dense spatial annotations and millisecond-level stroke supervision. It is organized under a \textbf{Progressive Cognitive Evaluation} protocol. Benchmarking state-of-the-art MLLMs reveals a severe executability gap: while human subjects achieve near-ceiling performance on EcoG-Bench (\textbf{96.9\%} strict Eco-Accuracy), the best native video-audio setting remains low (Gemini-3-Pro: \textbf{17.0\%}). Moreover, in a diagnostic ablation, replacing the native video--audio interface with timestamped frame samples and externally verified ASR (with word-level timing) substantially improves the same model (\textbf{17.0\%}$\to$\textbf{42.9\%}). Overall, EcoG-Bench provides a strict, executable testbed for event-level speech--gesture binding, and suggests that multimodal interfaces may bottleneck the observability of temporal alignment cues, independently of model reasoning.

多模态语音对齐动作识别评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。