用数据增强提升小样本下LLM与人类偏好的对齐效果。
Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation
- 通过数据增强筛选更有效的偏好数据,优化小样本场景下的评估对齐。
- 在数学推理任务中相比基线模型提升30%,与参考评判者相关性提高7%。
- 适合需要个性化评估且数据稀缺的实际应用场景。
当前大语言模型的自动评估虽受关注,但评价任务常具主观性且易受多种因素影响,难以适应长期变化的参考评判者,限制了个性化判断的实现。尽管众多研究展示主流闭源LLM可媲美人类评估,但在持续适应参考评判者方面仍存在挑战。同时,开源LLM作为评估者的尝试也常忽视数据稀缺的现实。个性化评估本质处于数据有限场景,这在实际问题中普遍存在。本文提出一种数据增强技术,旨在从有限数据中选择更有效的样本,以对齐开源LLM与人类偏好。实验表明,该方法在基准上使皮尔逊相关性提升约7%,在数学推理任务中相较基础模型(Llama3.1-8B-Instruct)提升30%,证明有效增强偏好数据选择能显著超越现有基线方法。
原文摘要 · Abstract (English)
Automatic evaluation by large language models (LLMs) is a prominent topic today; however, judgment and evaluation tasks are often subjective and influenced by various factors, making adaptation challenging. While many studies demonstrate the capabilities of state-of-the-art proprietary LLMs in comparison to human evaluators, they often struggle to adapt to reference evaluators over time, a requirement for achieving personalized judgment. Additionally, numerous works have attempted to apply open LLMs as judges or evaluators, but these efforts frequently overlook the limitations of working with scarce data. Personalized judgment is inherently associated with limited data scenarios, which are common in many real-world problems. Our work aims to present a data augmentation technique to select a more effective sample from limited data in order to align an open LLM with human preference. Our work achieves approximately 7% improvements in Pearson correlation with a reference judge over the baseline,and 30% improvement over the base model (Llama3.1-8B-Instruct) in the mathematical reasoning evaluation task. demonstrating that augmenting selecting more effective preference data enables our approach to surpass baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。