让AI对话更自然:根据场景自动调整说话时机
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues
- 用少量人类偏好数据校准大模型,实现场景自适应发言时机
- 在六种任务中生成的对话更符合人类对发言顺序的偏好
- 适合需要真实交互体验的对话系统研发者
轮换发言是全双工交互的核心。合适的发言时机随场景变化,但现有模型采用统一标准,缺乏情境适配。其根源在于训练数据:人类对话语料虽有自然时间模式,但缺少角色定位和场景特定规范;而启发式或提示生成方法虽可注入发言行为,却未基于人类偏好。我们提出DuplexGen框架,通过少量槽位级人类偏好标注校准大模型预测,实现场景自适应的轮换发言生成。在六项合作与竞争任务中,人类对发言顺序的偏好存在系统差异,而经校准的DuplexGen显著优于未经校准的提示或仅使用通用人类对话数据训练的模型;基于DuplexGen生成数据训练的全双工模型展现出更符合人类偏好的发言行为。结果表明,人类校准而非数据规模或提示设计,才是实现场景化发言合成的关键。
原文摘要 · Abstract (English)
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。