arXiv:2601.03154cs.CL2026-01ACL被引 7

研究长链思维如何影响模型对人类标注差异的建模能力

Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective

  • 通过交叉链思维实验分离推理文本与模型先验的影响
  • 99%准确率差异由推理内容决定,80%以上分布排序依赖模型内在先验
  • 长链思维能选最优答案却无法精细校准模糊任务的概率分布

利用长链思维(CoT)微调的大语言模型在单答案任务中表现优异,但其对人类标注差异——即捕捉概率模糊性而非消除模糊性——的建模能力尚未充分探索。我们通过分布类任务上的系统性解耦实验,采用跨链思维(Cross-CoT)实验将推理文本的影响与模型固有先验分离。观察到明显“解耦机制”:尽管CoT提升分布对齐效果,但最终准确率主要受CoT内容支配(贡献99%方差),而分布排序则由模型先验主导(贡献超80%)。逐步分析显示,虽然CoT对准确率的影响随推理过程单调上升,但分布结构主要由大模型内在先验决定。结果表明,长链思维可作为顶级选项的决策工具,但在模糊任务中难以充当精细的概率校准器。

原文摘要 · Abstract (English)

Reasoning-tuned LLMs utilizing long Chain-of-Thought (CoT) excel at single-answer tasks, yet their ability to model Human Label Variation--which requires capturing probabilistic ambiguity rather than resolving it--remains underexplored. We investigate this through systematic disentanglement experiments on distribution-based tasks, employing Cross-CoT experiments to isolate the effect of reasoning text from intrinsic model priors. We observe a distinct "decoupled mechanism": while CoT improves distributional alignment, final accuracy is dictated by CoT content (99% variance contribution), whereas distributional ranking is governed by model priors (over 80%). Step-wise analysis further shows that while CoT's influence on accuracy grows monotonically during the reasoning process, distributional structure is largely determined by LLM's intrinsic priors. These findings suggest that long CoT serves as a decisive LLM decision-maker for the top option but fails to function as a granular distribution calibrator for ambiguous tasks.

链式思维模型先验概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。