推理让大模型更难捕捉人工标注分歧,简单思维链反而更好。
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
- 用不同推理方式测试大模型对标注分歧的建模能力
- 强化学习推理使表现下降,简单思维链提升效果
- 提示研究者警惕用模型替代人工标注的风险
自然语言处理中的标注差异普遍存在,常反映任务主观性与样本模糊性,建模此类差异对敏感应用至关重要。尽管基于可验证奖励的强化学习(RLVR)已提升大模型在多项任务上的表现,但其是否能有效捕捉人类标注中的信息性差异仍不明确。本文系统评估了不同推理设置对大模型分歧建模的影响,涵盖模型规模、分布表达方法与引导策略,共设计60组实验,覆盖3项任务。结果表明,RLVR式推理会降低分歧建模性能,而朴素思维链(CoT)推理则能提升基于人类反馈强化学习(RLHF)模型的表现。该发现警示:在标注分歧重要的场景中,以推理模型取代人工标注存在潜在风险。
原文摘要 · Abstract (English)
Variation in human annotation (i.e., disagreements) is common in NLP, often reflecting important information like task subjectivity and sample ambiguity. Modeling this variation is important for applications that are sensitive to such information. Although RLVR-style reasoning (Reinforcement Learning with Verifiable Rewards) has improved Large Language Model (LLM) performance on many tasks, it remains unclear whether such reasoning enables LLMs to capture informative variation in human annotation. In this work, we evaluate the influence of different reasoning settings on LLM disagreement modeling. We systematically evaluate each reasoning setting across model sizes, distribution expression methods, and steering methods, resulting in 60 experimental setups across 3 tasks. Surprisingly, our results show that RLVR-style reasoning degrades performance in disagreement modeling, while naive Chain-of-Thought (CoT) reasoning improves the performance of RLHF LLMs (RL from human feedback). These findings underscore the potential risk of replacing human annotators with reasoning LLMs, especially when disagreements are important.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。