用AI反馈训练对话系统,让回应更自然、有性格、有共情。
Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression
- 构建12个对话印象指标的奖励模型,通过LLM评估整体对话质量。
- 使用该模型反馈微调对话系统,提升各维度评分与自然度。
- 适合关注对话体验优化的研究者和产品团队。
为提升用户在与对话系统交互时的参与感,需同时优化单轮回复及整体对话的连贯性、人格一致性和共情力。尽管大语言模型(LLM)推动了对话系统快速发展,但基于AI反馈的强化学习(RLAIF)逐渐成为对齐对话印象的关键方法。在RLAIF中,采用另一大模型作为奖励模型,通过零样本/少样本提示生成训练信号。然而,仅通过提示评估完整对话仍具挑战。本研究通过监督微调(SFT)构建了对应12项对话印象指标的奖励模型,用于评估对话响应质量。我们利用该奖励模型的信号作为反馈,对对话模型进行调优。自动评估与人工评测结果表明,基于对话印象奖励模型的微调显著提升了各项指标得分及对话回复的自然度。
原文摘要 · Abstract (English)
To improve user engagement during conversations with dialogue systems, we must improve individual dialogue responses and dialogue impressions such as consistency, personality, and empathy throughout the entire dialogue. While such dialogue systems have been developing rapidly with the help of large language models (LLMs), reinforcement learning from AI feedback (RLAIF) has attracted attention to align LLM-based dialogue models for such dialogue impressions. In RLAIF, a reward model based on another LLM is used to create a training signal for an LLM-based dialogue model using zero-shot/few-shot prompting techniques. However, evaluating an entire dialogue only by prompting LLMs is challenging. In this study, the supervised fine-tuning (SFT) of LLMs prepared reward models corresponding to 12 metrics related to the impression of the entire dialogue for evaluating dialogue responses. We tuned our dialogue models using the reward model signals as feedback to improve the impression of the system. The results of automatic and human evaluations showed that tuning the dialogue model using our reward model corresponding to dialogue impression improved the evaluation of individual metrics and the naturalness of the dialogue response.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。