arXiv:2604.10367cs.AIcs.SD2026-04被引 1

让虚拟人能像真人一样边说边听,实时互动

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels

论文配图:Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels
图 1 · 摘自论文原文
  • 用多头高斯核建模说话与倾听的时序差异,引入渐进式时间先验
  • 实现双流音频输入下自然对话与精准口型同步,超越现有方法
  • 适合做虚拟助手、数字人直播的开发者和研究者参考

音频驱动的人类视频生成在独白场景已取得显著进展,主要得益于强大视频生成基础模型的发展。然而真实人类交流是全双工交互过程,虚拟代理不仅需表达自身话语,还需自然响应对方的对话音频。现有方法多简单将传统音频驱动范式扩展至倾听场景,依赖严格的帧对齐导致对长程对话动态响应僵化,而直接引入全局注意力则严重损害口型同步。针对说话与倾听行为间独特的时序尺度差异,我们提出多头高斯核,将这一物理直觉作为渐进式时间归纳偏置显式注入模型。基于此,构建了可同时处理双流音频输入的全双工交互虚拟代理。此外,我们推出了严格清洗的说话-倾听数据集VoxHear,包含完全解耦的语音与背景音频轨道。大量实验表明,该方法成功融合强时序对齐与深层语境语义,为生成高度自然且响应灵敏的全双工交互数字人树立了新基准。

原文摘要 · Abstract (English)

Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation foundation models. Moving beyond monologues, authentic human communication is inherently a full-duplex interactive process, requiring virtual agents not only to articulate their own speech but also to react naturally to incoming conversational audio. Most existing methods simply extend conventional audio-driven paradigms to listening scenarios. However, relying on strict frame-to-frame alignment renders the model's response to long-range conversational dynamics rigid, whereas directly introducing global attention catastrophically degrades lip synchronization. Recognizing the unique temporal Scale Discrepancy between talking and listening behaviors, we introduce a multi-head Gaussian kernel to explicitly inject this physical intuition into the model as a progressive temporal inductive bias. Building upon this, we construct a full-duplex interactive virtual agent capable of simultaneously processing dual-stream audio inputs for both talking and listening. Furthermore, we introduce a rigorously cleaned Talking-Listening dataset VoxHear featuring perfectly decoupled speech and background audio tracks. Extensive experiments demonstrate that our approach successfully fuses strong temporal alignment with deep contextual semantics, setting a new state-of-the-art for generating highly natural and responsive full-duplex interactive digital humans. The project page is available at https://warmcongee.github.io/beyond-monologue/ .

虚拟人语音驱动交互生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。