提升大模型在公共健康资源分配中的偏好鲁棒性,应对数据少、目标模糊挑战。
Preference Robustness for DPO with Applications to Public Health
- 基于DPO改进,引入轻量分布鲁棒优化应对偏好不确定性
- 在真实母婴健康项目中表现更抗噪声,优于现有DPO方法
- 无需复杂自反思机制,推理成本更低,适合实际部署
我们研究一种大型语言模型微调任务,旨在为公共健康中的序列资源分配问题设计奖励函数,依据自然语言表达的人类偏好。该场景因目标复杂模糊且数据有限,成为对齐技术的严峻考验。我们提出DPO-PRO,一种基于直接偏好优化(DPO)的鲁棒微调算法,采用轻量级分布鲁棒优化(DRO)框架处理偏好分布的不确定性。与以往基于DRO的DPO方法不同,DPO-PRO显著降低保守性。我们在非营利组织ARMMAN运营的真实母婴移动健康项目上评估该方法,并对比标准对齐基准。实验表明,本方法在面对噪声偏好信号时表现出更强的鲁棒性,优于现有DPO变体。同时,其在奖励函数设计上达到与先前自反思基线相当的性能,但推理开销显著更低。
原文摘要 · Abstract (English)
We study an LLM fine-tuning task for designing reward functions for sequential resource allocation problems in public health, guided by human preferences expressed in natural language. This setting presents a challenging testbed for alignment due to complex and ambiguous objectives and limited data availability. We propose DPO-PRO, a robust fine-tuning algorithm based on Direct Preference Optimization (DPO), which accounts for uncertainty in the preference distribution using a lightweight Distributionally Robust Optimization (DRO) formulation. Unlike prior DRO-based DPO methods, DPO-PRO is significantly less conservative. We evaluate DPO-PRO on a real-world maternal mobile health program operated by the non-profit organization ARMMAN, as well as on standard alignment benchmarks. Experimental results demonstrate that our method consistently improves robustness to noisy preference signals compared to existing DPO variants. Moreover, DPO-PRO achieves comparable performance to prior self-reflection-based baseline for reward function design, while requiring significantly lower inference-time cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。