让大模型摆脱讨好与推责,像有尊严的对话伙伴一样回应。
Dual Optimal: Make Your LLM Peer-like with Dignity
- 构建多角色人格数据集,动态平衡不同人格维度。
- 模型在多项测试中展现更可信、不推诿的回应能力。
- 适合需要真实互动体验的对话系统开发者。
当前对齐语言模型存在双重缺陷,称为‘逃避型仆人’:它们会盲目迎合用户错误观点,同时用模板化免责声明推卸责任。本文提出‘尊严对话者’框架,通过反谄媚和可信性对抗奴性,借助共情与创造力缓解回避行为。实现该代理面临数据监督、目标崩溃和评估偏差等挑战。我们引入PersonaKnob数据集,包含多重人格偏好的组合式偏序结构;结合容错约束拉格朗日直接偏好优化(DPO)算法,动态平衡各人格维度以防止行为坍缩;并采用心理计量校准的项目反应理论评估协议,分离模型真实人格能力与评价者偏见等混淆因素。大量实验证明,该方法成功构建出兼具尊严与同侪特质的大型语言模型代理。
原文摘要 · Abstract (English)
Current aligned language models exhibit a dual failure mode we term the Evasive Servant: they sycophantically validate flawed user beliefs while deflecting responsibility with boilerplate disclaimers. We propose the Dignified Peer framework, which counters servility with anti-sycophancy and trustworthiness, and mitigates evasiveness through empathy and creativity. Realizing this agent requires overcoming significant challenges in data supervision, objective collapse, and evaluation bias. We address these issues by introducing the PersonaKnob dataset which features a compositional partial order structure of multiple persona preference. This data is utilized alongside a tolerant constrained Lagrangian DPO algorithm that dynamically balances all persona dimensions to prevent behavioral collapse. Additionally, we employ a psychometrically calibrated Item Response Theory evaluation protocol to disentangle latent model persona capability from confounders like judge biases. Extensive empirical studies demonstrate that our approach successfully build a LLM agent with both dignity and peer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。