用错误答案间的比较来训练大模型,反而让模型更准了。
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
- 用自洽性、概率和大模型评判,生成错误选项间的偏好
- 模型在错误中识别差异,正确率比随机猜高20.9%
- 用错误偏好对齐,能减少错误,甚至产出正确答案
在缺乏可靠标注的复杂任务中,如何利用可能错误的答案拓展大模型能力?我们聚焦两个问题:(1)大模型能否在错误选项间生成可靠偏好?若可,(2)基于此类错误对错误的偏好进行对齐是否有益?我们采用基于自一致性、标记概率和大模型作为裁判的方法,生成错误对错误的偏好,并使用偏好优化方法微调语言模型。在七种大模型与八个数据集上的实验表明:(1)大模型具备初步区分不同错误程度的能力,性能最高比随机猜测提升20.9%;(2)与错误对错误偏好对齐有助于模型输出更少错误,有时甚至产生正确答案,同时整体改善模型校准度。
原文摘要 · Abstract (English)
In the absence of abundant reliable annotations for challenging tasks and contexts, how can we expand the frontier of LLM capabilities with potentially wrong answers? We focus on two research questions: (1) Can LLMs generate reliable preferences among wrong options? And if so, (2) Would alignment with such wrong-over-wrong preferences be helpful? We employ methods based on self-consistency, token probabilities, and LLM-as-a-judge to elicit wrong-over-wrong preferences, and fine-tune language models with preference optimization approaches using these synthesized preferences. Extensive experiments with seven LLMs and eight datasets demonstrate that (1) LLMs do have preliminary capability in distinguishing various shades of wrong, achieving up to 20.9% higher performance than random guess; (2) Alignment with wrong-over-wrong preferences helps LLMs to produce less wrong and sometimes even outright correct answers, while overall improving model calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。