通过人机协作提升大模型偏好学习效率,兼顾数据质量与规模。
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy

- 结合自奖励与主动学习,动态识别需人工验证的样本
- 在GSM8K等三个基准上平均提升13.16%以上
- 适合需要高质量对齐但标注成本高的场景
基于偏好反馈的学习已成为对齐大模型在多样化任务中的有效方法。然而,高质量的人工标注偏好数据仍昂贵且稀缺。现有方法或采用自奖励机制(纯AI生成标签,可扩展但可靠性存疑),或依赖主动学习(通过人工标注保证质量,但无法充分利用未标注数据)。本文提出CoAct框架,通过策略性的人机协同,融合自奖励与主动学习优势。CoAct利用自一致性识别可靠的自标注数据及需人工验证的样本;同时,人工反馈引导模型生成在其能力范围内的新指令。在两个模型家族的三个推理基准上评估,CoAct在GSM8K上平均提升13.25%,在MATH上提升8.19%,在WebInstruct上提升13.16%,持续优于所有基线。
原文摘要 · Abstract (English)
Learning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks. However, high-quality human-annotated preference data remains expensive and scarce. Existing methods address this challenge through either self-rewarding, which scales by using purely AI-generated labels but risks unreliability, or active learning, which ensures quality through oracle annotation but cannot fully leverage unlabeled data. In this paper, we present CoAct, a novel framework that synergistically combines self-rewarding and active learning through strategic human-AI collaboration. CoAct leverages self-consistency to identify both reliable self-labeled data and samples that require oracle verification. Additionally, oracle feedback guides the model to generate new instructions within its solvable capability. Evaluated on three reasoning benchmarks across two model families, CoAct achieves average improvements of +13.25% on GSM8K, +8.19% on MATH, and +13.16% on WebInstruct, consistently outperforming all baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。