arXiv:2510.12516cs.CLcs.AI2025-10被引 1

测试时扩展在问答任务中表现不佳,因标注不一致干扰了效果。

BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)

  • 用模型平均、多数投票和最佳N采样三种方法提升大模型推理
  • 数学编码任务有效,但LeWiDi任务中最佳N方法无提升
  • 适合关注模型鲁棒性与标注质量的研究者

测试时扩展是一类通过增加推理阶段计算量来提升大模型输出质量的技术。据我们所知,该技术此前仅应用于具有明确正确答案的领域,如数学与编程。本文将测试时扩展方法应用于LeWiDi-2025任务,以评估标注分歧的影响。实验对比了三种方法:两种基准算法(模型平均与多数投票)以及一种最佳N采样方法。结果显示,模型平均与多数投票在LeWiDi任务上表现稳定提升,但最佳N方法未带来改进。结果表明,最佳N方法尚无法从数学领域有效迁移至存在标注分歧的LeWiDi任务,我们分析了可能原因。

原文摘要 · Abstract (English)

Test-time scaling is a family of techniques to improve LLM outputs at inference time by performing extra computation. To the best of our knowledge, test-time scaling has been limited to domains with verifiably correct answers, like mathematics and coding. We transfer test-time scaling to the LeWiDi-2025 tasks to evaluate annotation disagreements. We experiment with three test-time scaling methods: two benchmark algorithms (Model Averaging and Majority Voting), and a Best-of-N sampling method. The two benchmark methods improve LLM performance consistently on the LeWiDi tasks, but the Best-of-N method does not. Our experiments suggest that the Best-of-N method does not currently transfer from mathematics to LeWiDi tasks, and we analyze potential reasons for this gap.

测试时扩展大模型推理标注分歧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。