让对话系统理解空间时间关系,实现更自然的共同认知
Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs
- 用动态环境中的关系指代测试模型建立共同认知能力
- 提出强化学习优化共同认知表示,提升复杂场景理解
- 适合研究对话系统、人机交互与机器人对话的开发者
共同认知在情景化语音对话中至关重要,对话双方需共享对实体、事件和关系的认知,以维持在共享空间与时间上的连贯互动。随着具身对话代理与社交机器人的普及,对话系统正确建立并回溯这种认知内容的能力也日益重要。先前研究显示大模型能执行部分接地行为(如确认)。但较少工作探讨其如何利用已接地信息,尤其在涉及空间与时间的复杂场景中(如“去我们昨天去过的公园旁那家咖啡馆”)。本文评估模型在动态共享环境中使用关系指代建立共同认知的能力,并测试多种共同认知表示方法;进一步通过在合成对话数据上应用强化学习,提出改进性能的新方法。
原文摘要 · Abstract (English)
Common ground plays a critical role in situated spoken dialogs, where interlocutors must establish and maintain shared references to entities, events, and relations to sustain coherent interaction in a shared space and over time. With the increasing presence of embodied conversational agents and social robots, the ability to correctly ground this kind of conversational content in order to refer back later also becomes important for dialog systems. Prior studies have demonstrated that LLMs are capable of performing certain grounding acts like acknowledgments. However, relatively little work has investigated their capacity to leverage the grounded information, like in complex scenarios involving space and time (e.g., "let's go to that café near the park we went to yesterday"). To that end, in this work, we evaluate a model's ability to establish common ground by utilizing these "relational references" in the dynamic and shared environments of situated dialogs. We then test multiple methods for representing common ground and further propose approaches to improve their performance by using reinforcement learning on our synthetically generated dialog data .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。