arXiv:2502.14187cs.LGcs.CL2025-02被引 6

KTO在联邦微调中比DPO更优,尤其适合单反馈场景。

Federated Fine-Tuning of Large Language Models: Kahneman-Tversky vs. Direct Preference Optimization

  • 用KTO替代DPO进行联邦微调,支持单响应反馈。
  • KTO在多个基准上均超越DPO, redistributed设置下仍表现稳定。
  • 适合隐私保护、数据异构的分布式场景,应用灵活性强。

本文评估了凯恩曼-特沃斯基优化(KTO)作为大语言模型(LLM)在联邦学习(FL)环境下的微调方法,与直接偏好优化(DPO)进行对比。以Alpaca-7B为基线模型,在真实数据集上进行微调,并通过MT-Bench-1、Vicuna和AdvBench三个基准评估性能。此外,引入一种重分配数据集设置,仅KTO适用,因其可处理单响应反馈,而DPO依赖成对响应。结果表明,无论原始(KTOO)还是重分配(KTOR)配置,KTO在所有基准上均持续优于DPO;在重分配设置下,KTO进一步验证其灵活性与鲁棒性,即使在DPO无法应用时仍保持优异表现。这些发现确立了KTO作为联邦学习中稳健且可扩展的微调方法,推动其在隐私保护、去中心化和异构环境中的应用。

原文摘要 · Abstract (English)

We evaluate Kahneman-Tversky Optimization (KTO) as a fine-tuning method for large language models (LLMs) in federated learning (FL) settings, comparing it against Direct Preference Optimization (DPO). Using Alpaca-7B as the base model, we fine-tune on a realistic dataset under both methods and evaluate performance using MT-Bench-1, Vicuna, and AdvBench benchmarks. Additionally, we introduce a redistributed dataset setup, where only KTO is applicable due to its ability to handle single-response feedback, unlike DPO's reliance on paired responses. Our results demonstrate that KTO, in both its original (KTOO) and redistributed (KTOR) configurations, consistently outperforms DPO across all benchmarks. In the redistributed setup, KTO further validates its flexibility and resilience by maintaining superior performance in scenarios where DPO cannot be applied. These findings establish KTO as a robust and scalable fine-tuning method for FL, motivating its adoption for privacy-preserving, decentralized, and heterogeneous environments.

联邦学习大模型微调偏好优化KTO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。