用自洽模拟生成策略记忆,让对话代理更主动智能。
PRINCIPLES: Synthetic Strategy Memory for Proactive Dialogue Agents
- 通过离线自对弈生成可复用的策略记忆
- 在情感支持与说服场景中均显著优于基线
- 无需额外训练,适合实际部署的主动对话系统
基于大语言模型的对话代理在主动对话任务中表现出色,但现有策略规划方法存在覆盖范围有限、偏好偏差及依赖昂贵训练等问题。为此,我们提出PRINCIPLES:一种用于主动对话代理的合成策略记忆。该方法通过离线自对弈模拟生成,作为推理时的可复用知识,避免了额外训练和数据标注。我们在情感支持与说服两个领域评估了PRINCIPLES,结果表明其性能持续优于强基线模型,并在更长对话和多样化测试环境下保持鲁棒性。
原文摘要 · Abstract (English)
Dialogue agents based on large language models (LLMs) have shown promising performance in proactive dialogue, which requires effective strategy planning. However, existing approaches to strategy planning for proactive dialogue face several limitations: limited strategy coverage, preference bias in planning, and reliance on costly additional training. To address these, we propose PRINCIPLES: a synthetic strategy memory for proactive dialogue agents. PRINCIPLES is derived through offline self-play simulations and serves as reusable knowledge that guides strategy planning during inference, eliminating the need for additional training and data annotation. We evaluate PRINCIPLES in both emotional support and persuasion domains, demonstrating consistent improvements over strong baselines. Furthermore, PRINCIPLES maintains its robustness across extended and more diverse evaluation settings. See our project page at https://huggingface.co/spaces/kimnamssya/Principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。