arXiv:2604.23295cs.CLcs.AI2026-04

首个可复现的印地语全双工对话系统,基于真实对话数据训练。

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations

论文配图:Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations
图 1 · 摘自论文原文
  • 用自研印地语分词器改造Moshi架构,保留音频预训练参数。
  • 在2.6万小时真实对话上训练,实现自然抢话与重叠交互。
  • 适合做印度语言实时语音对话系统的研究者与开发者。

全双工语音对话系统能模拟打断、重叠和回应等自然对话行为,但对印度语言的研究仍属空白。本文首次提出开源可复现的印地语全双工语音对话系统,通过改进先进双工语音架构Moshi,采用自研印地语分词器,并在14,695名说话人提供的26,000小时真实自发对话数据上训练,每个说话人有独立声道,直接学习自然互动中的换言与重叠模式。为支持印地语文本生成,替换原英文分词器并重初始化依赖文本词汇的参数,同时保留预训练音频组件。提出两阶段训练策略:大规模预训练后,在1,000小时对话数据上微调。通过提示式对话续写评估,结合自动指标与人工判断,结果表明模型能生成自然且有意义的全双工对话行为。该工作为印地语及其他印度语言的实时双工语音对话系统奠定了基础。

原文摘要 · Abstract (English)

Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open, reproducible full-duplex spoken dialogue system for Hindi by adapting Moshi, a state-of-the-art duplex speech architecture, using a custom Hindi tokeniser and training on 26,000 hours of real spontaneous conversations collected from 14,695 speakers with separate speaker channels, enabling direct learning of turn-taking and overlap patterns from natural interactions. To support Hindi text generation, we replace the original English tokeniser and reinitialise text-vocabulary-dependent parameters while retaining the pre-trained audio components. We propose a two-stage training recipe -- large-scale pre-training followed by fine-tuning on 1,000 hours of conversational data. Evaluation through the prompted dialogue continuation paradigm with both automatic metrics and human judgments demonstrates that the resulting model generates natural and meaningful full-duplex conversational behaviour in Hindi. This work serves as a first step toward real-time duplex spoken dialogue systems for Hindi and other Indian languages.

语音对话全双工印地语自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。