arXiv:2512.14085cs.CLcs.HC2025-12中稿 · presentation at In…被引 1

跨语言预测对话回应时机,揭示日英中文差异

Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study

  • 基于Transformer的多语言帧级模型,联合训练300小时对话数据
  • 多语言模型性能优于单语言基线,捕捉共通与特有节奏模式
  • 适合开发文化敏感的实时对话系统,支持纯CPU运行

我们提出一种针对日语、英语和汉语的多语言连续回应预测模型,并以此研究跨语言回应时机行为。该模型基于Transformer,以帧为单位进行预测,在约300小时双人对话数据上联合训练多个辅助任务。在三种语言中,多语言模型表现不逊于甚至超越单语言基线,表明其同时学习了语言通用线索与语言特定的响应时序特征。两语言训练下的零样本迁移效果有限,凸显显著的跨语言差异。扰动分析显示:日语更依赖短期语言信息,而英语和汉语则更敏感于沉默时长与韵律变化;多语言训练促进了共享但可适应的表征,降低了中文对音高过度依赖。上下文长度研究进一步表明,日语对较短上下文相对鲁棒,而汉语明显受益于更长上下文。最后,我们将训练好的模型集成至实时处理软件中,实现纯CPU推理。这些发现提供了统一模型与实证证据,揭示不同语言回应时机的差异,有助于设计更自然、更具文化意识的语音对话系统。

原文摘要 · Abstract (English)

We present a multilingual, continuous backchannel prediction model for Japanese, English, and Chinese, and use it to investigate cross-linguistic timing behavior. The model is Transformer-based and operates at the frame level, jointly trained with auxiliary tasks on approximately 300 hours of dyadic conversations. Across all three languages, the multilingual model matches or surpasses monolingual baselines, indicating that it learns both language-universal cues and language-specific timing patterns. Zero-shot transfer with two-language training remains limited, underscoring substantive cross-lingual differences. Perturbation analyses reveal distinct cue usage: Japanese relies more on short-term linguistic information, whereas English and Chinese are more sensitive to silence duration and prosodic variation; multilingual training encourages shared yet adaptable representations and reduces overreliance on pitch in Chinese. A context-length study further shows that Japanese is relatively robust to shorter contexts, while Chinese benefits markedly from longer contexts. Finally, we integrate the trained model into a real-time processing software, demonstrating CPU-only inference. Together, these findings provide a unified model and empirical evidence for how backchannel timing differs across languages, informing the design of more natural, culturally-aware spoken dialogue systems.

对话系统多语言语音交互时序预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。