解决单边对话中缺失对话方信息的问题,让AI能从单侧录音中还原对话。
Reading Between the Lines: The One-Sided Conversation Problem
- 通过未来一句信息和语句长度提示,提升对话缺失方的重建精度。
- 小模型需微调,大模型用提示词即可生成高质量重建结果。
- 无需重建对话即可生成高质摘要,适合隐私敏感场景使用。
对话式AI在医疗电话、客服中心和智能眼镜等场景中受限于仅能记录单方对话。本文提出单边对话问题(1SC):从单侧对话中推断并学习完整对话。研究两项任务:实时重建缺失说话人话语,以及基于单侧转录生成摘要。在MultiWOZ、DailyDialog和Candor数据集上,通过人工对比测试与大模型评分,发现引入未来一句信息和语句长度有助于提升重建效果;占位符提示可缓解幻觉;大模型经提示后表现良好,小模型则需微调。此外,无需重建即可生成高质量摘要。本工作提出1SC新挑战,成果为隐私保护型对话AI迈出关键一步。
原文摘要 · Abstract (English)
Conversational AI is constrained in many real-world settings where only one side of a dialogue can be recorded, such as telemedicine, call centers, and smart glasses. We formalize this as the one-sided conversation problem (1SC): inferring and learning from one side of a conversation. We study two tasks: (1) reconstructing the missing speaker's turns for real-time use cases, and (2) generating summaries from one-sided transcripts. Evaluating prompting and finetuned models on MultiWOZ, DailyDialog, and Candor with both human A/B testing and LLM-as-a-judge metrics, we find that access to one future turn and information about utterance length improves reconstruction, placeholder prompting helps to mitigate hallucination, and while large models generate promising reconstructions with prompting, smaller models require finetuning. Further, high-quality summaries can be generated without reconstructing missing turns. We present 1SC as a novel challenge and report promising results that mark a step toward privacy-aware conversational AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。