针对自闭症干预对话中模型谄媚问题,提出精细化优化方法。
TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue

- 通过最小编辑构造对比数据对,精准定位需优化的差异片段。
- 在离线设置下显著降低谄媚行为,同时保持干预能力不下降。
- 适合需要安全可控对话的临床自闭症干预场景使用。
大语言模型在自闭症儿童干预对话中的谄媚行为会增加安全风险。监督微调虽能部分缓解,但仅依赖正例难以识别并修正失败模式。我们观察到谄媚行为通常局限于响应中的有限片段。在此情况下,序列级偏好优化可能过度更新无关令牌,导致干预能力下降。为此,我们提出最小编辑数据增强(MEDA)策略,构建受控、稳定的最小编辑偏好对,并提出基于令牌差异的直接偏好优化(TD-DPO),通过上权重选择与拒绝响应间的差异令牌,下权重共享令牌,抑制背景漂移。多骨干模型与评估器的实验表明,TD-DPO在离线设置下实现了谄媚缓解与干预能力保留之间的更优平衡,凸显其在自闭症干预中的实际对齐潜力。
原文摘要 · Abstract (English)
The sycophancy of large language models can increase the safety risk in intervention dialogue for autistic children. Supervised fine-tuning can somewhat reduce sycophancy, but relying solely on positive examples is often insufficient to identify and correct failure patterns. We observe that sycophancy behaviors can often be localized to a limited span within the model response. In this regime, sequence-level preference optimization can over-update preference-irrelevant tokens and degrade intervention ability. To address this, we propose the \textbf{M}inimal \textbf{E}dit \textbf{D}ata \textbf{A}ugmentation (MEDA) strategy to construct controlled, stable, minimal edit preference pairs and \textbf{T}oken-level \textbf{D}ifference \textbf{D}irect \textbf{P}reference \textbf{O}ptimization (TD-DPO), which upweights difference tokens between chosen and rejected responses while downweighting shared tokens to suppress background drift. Extensive experiments across multiple backbones and evaluators show that TD-DPO achieves a better trade-off between sycophancy mitigation and intervention ability retention in our offline settings, highlighting its potential as a practical alignment approach for autism intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。