arXiv:2603.21664cs.CV2026-03

让AI精准识别对话中谁在何时说了什么,突破多模态模型的视觉偏见陷阱。

HumanOmni-Speaker: Identifying Who said What and When

  • 用25帧/秒原始视频采样,通过视觉差分编码压缩每帧为6个令牌。
  • 在新基准上实现高精度说话人定位与唇语识别,超越现有模型。
  • 适合研究多说话人交互、跨模态对齐和端到端语音定位的学者。

尽管通用多模态大模型在联合感知方面取得进展,但在理解复杂多人对话中的“谁说了什么、何时说的”这一核心人类交互能力上仍存在根本性缺陷。当前模型存在‘能力幻觉’——依赖传统基准中的视觉偏差来绕过真正的跨模态对齐,并依赖稀疏、低帧率的视觉采样,破坏了唇部运动等关键高频动态。为打破此幻觉,我们提出视觉注册说话人分离与识别(VR-SDR)方法及HumanOmni-Speaker基准。通过严格消除视觉捷径,该范式要求仅凭自然语言查询实现端到端时空身份绑定。针对底层架构感知差距,我们提出HumanOmni-Speaker,其核心是视觉差分编码器。该模型以25 fps采样原始视频,将帧间运动残差显式压缩为每帧仅6个令牌,有效捕捉精细视觉音素(visemes)与说话人轨迹,避免令牌爆炸。最终,HumanOmni-Speaker展现出强大多模态协同能力,原生支持端到端唇读与高精度空间定位,无需侵入式裁剪,在多种说话人中心任务中表现卓越。

原文摘要 · Abstract (English)

While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, multi-person conversational dynamics to accurately answer ``Who said what and when.'' Current models suffer from an ``illusion of competence'' -- they exploit visual biases in conventional benchmarks to bypass genuine cross-modal alignment, while relying on sparse, low-frame-rate visual sampling that destroys crucial high-frequency dynamics like lip movements. To shatter this illusion, we introduce Visual-Registered Speaker Diarization and Recognition (VR-SDR) and the HumanOmni-Speaker Benchmark. By strictly eliminating visual shortcuts, this rigorous paradigm demands true end-to-end spatio-temporal identity binding using only natural language queries. To overcome the underlying architectural perception gap, we propose HumanOmni-Speaker, powered by a Visual Delta Encoder. By sampling raw video at 25 fps and explicitly compressing inter-frame motion residuals into just 6 tokens per frame, it captures fine-grained visemes and speaker trajectories without triggering a catastrophic token explosion. Ultimately, HumanOmni-Speaker demonstrates strong multimodal synergy, natively enabling end-to-end lip-reading and high-precision spatial localization without intrusive cropping, and achieving superior performance across a wide spectrum of speaker-centric tasks.

说话人分离多模态唇语识别视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。