用用户编辑数据统一优化大模型,提升个性化适应能力
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
- 将用户编辑视为偏好、监督标签和代价三类反馈的统一信号
- 提出集成方法,在两个领域上优于单一反馈学习
- 能稳健适应测试时不同的用户编辑分布,适合个性化应用
我们研究如何利用用户编辑部署数据对大语言模型进行微调,该数据包括上下文、智能体响应及用户修改。这类数据自然生成于基于大模型的写作助手和编程代理等应用场景中。用户编辑的自然来源使其成为适配和个性化大模型的理想数据源。在此设定下,偏好、监督标签和代价这三类通常独立研究的反馈类型得以统一。本文首次对从用户编辑中学习进行理论分析:我们推导了分别基于各类反馈的学习算法的泛化界,并证明其性能权衡取决于用户特征、数据分布和模型类别。随后提出一种简单集成方法,联合利用三类反馈。在基于Gao等人2024年工作的两个领域上,实验表明该集成方法优于仅使用单一反馈的基线方法;进一步验证其在测试时面对不同用户编辑分布仍具鲁棒性。
原文摘要 · Abstract (English)
We study how to fine-tune LLMs using user-edit deployment data consisting of a set of context, an agent's response, and user edits. This deployment data is naturally generated by users in applications such as LLMs-based writing assistants and coding agents. The _natural_ origin of user edits makes it a desired source for adapting and personalizing LLMs. In this setup, there emerges a unification of various feedback types namely preferences, supervised labels, and cost that are typically studied separately in the literature. In this paper, we initiate the theoretical investigation of learning from user edits. We first derive bounds for learning algorithms that learn from each of these feedback types. We prove that these algorithms have different trade-offs depending upon the user, data distribution, and model class. We then propose a simple ensembling procedure to jointly learn from these feedback types. On two domains adapted from Gao et al. 2024, we show our ensembling procedure outperforms these methods that learn from individual feedback. Further, we show that our proposed procedure can robustly adapt to different user-edit distributions at test time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。