用多模型评估自动构建偏好数据,提升对话模型的人设一致性与指令遵循能力。
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
- 通过闭源与开源模型自动评估生成回复,构造偏好对。
- 在FoCus数据集上超越多个基线模型,人设一致性和指令遵循性显著提升。
- 无需人工标注,适合需要规模化个性化对话训练的研究与应用。
个性化和上下文连贯性是构建有效人设对话系统的关键。尽管开源大语言模型在流畅性和自然性方面表现良好,但在保持回应与用户身份一致、符合上下文方面仍存在挑战。本文提出PersoDPO,一种可扩展的偏好优化框架,利用闭源与开源模型对生成回复的自动评估结果作为监督信号,微调对话模型。该框架整合了针对连贯性、个性化以及长度格式合规性的评估指标,自动构建高质量偏好对,实现无须人工标注的可复现训练流程。在FoCus数据集上的实验表明,经PersoDPO微调的开源模型,在多个评估维度上持续优于强基线模型及标准DPO方法。
原文摘要 · Abstract (English)
Personalization and contextual coherence are two essential components in building effective persona-grounded dialogue systems. These aspects play a crucial role in enhancing user engagement and ensuring responses are more relevant and consistent with user identity. However, recent studies indicate that open-source large language models (LLMs) continue to struggle to generate responses that are both contextually grounded and aligned with persona cues, despite exhibiting strong general conversational abilities like fluency and naturalness. We present PersoDPO, a scalable preference optimisation framework that uses supervision signals from automatic evaluations of responses generated by both closed-source and open-source LLMs to fine-tune dialogue models. The framework integrates evaluation metrics targeting coherence and personalization, along with a length-format compliance feature to promote instruction adherence. These signals are combined to automatically construct high-quality preference pairs without manual annotation, enabling a scalable and reproducible training pipeline. Experiments on the FoCus dataset show that an open-source language model fine-tuned with the PersoDPO framework consistently outperforms strong open-source baselines and a standard Direct Preference Optimization (DPO) variant across multiple evaluation dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。