arXiv:2607.27756cs.SDcs.CL2026-07

让语音助手在嘈杂社交环境中智能判断何时回应、倾听或忽略。

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

论文配图:Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
图 1 · 摘自论文原文
  • 用响应/倾听/忽略三类动作标记控制对话行为,实现精准交互决策。
  • 在真实模拟的多人对话数据上训练,可识别不同说话人意图与干扰声音。
  • 适合开发更自然、主动的智能家居、车载语音系统等复杂场景应用。

语音对话系统通常设计用于干净环境下的两人对话,而现实中的社交对话往往更复杂:多个说话人同时参与,伴随无关话语和背景噪声。每个发言可能针对助手、其他说话人,或完全无关。在此类场景中,助手不仅需决定说什么,还需判断是否开口。本文提出Cocktail-Talker,一种面向嘈杂社交环境的多说话人语音对话建模框架。通过三个动作标记(<|respond|>、<|listen|>、<|ignore|>)控制行为,分别表示回应、倾听或忽略。该框架采用监督微调与强化学习联合训练,仅在<|respond|>模式下生成语音回复。为构建训练数据,我们开发了Cocktail-DialogGen——一个基于大语言模型的数据生成管道,可模拟多样社交场景下的多说话人角色对话。整体方案推动语音对话系统在复杂社会环境中更自然、更主动地交互。

原文摘要 · Abstract (English)

Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant speech and background noise. Each utterance may be directed to the assistant, addressed to another speaker, or completely irrelevant. In such settings, the assistant must decide not only what to say, but also whether to speak at all. In this paper, we introduce Cocktail-Talker, a speech LLM framework for multi-speaker spoken dialog modeling in noisy social environments. We model the assistant's behavior with three action tokens: <|respond|>, <|listen|>, and <|ignore|>, placed before a response or silence. Cocktail-Talker is trained via supervised finetuning and reinforcement learning to generate the appropriate action token and, only in <|respond|> mode, a speech response. To prepare the training data, we develop Cocktail-DialogGen, an LLM-based data pipeline that simulates realistic multi-speaker dialogs with speaker roles across diverse social settings. Together, these components take a step toward spoken dialog systems that interact more naturally and selectively in complex social environments.

语音对话多说话人智能助手噪声环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。