用户反馈是大模型可利用的有效信号,但现有评估方法会误判其效果。
User Feedback Provides a Unique Signal that LLMs Can not Detect
- 通过合成数据与真实数据对比,验证反馈能显著提升模型修复问题能力。
- 有反馈时模型修正问题的成功率远高于无反馈的基线版本。
- 评估偏差源于大模型裁判无法识别因反馈而改进的正确回答。
从用户交互中获取的自然反馈为大型语言模型提供了潜在的学习信号。然而,近期研究认为该反馈噪声大且难以有效利用。本文挑战这一观点,表明用户反馈实为极具行动价值的改进信号,其看似无效源于当前评估范式中的系统性偏差。为隔离反馈有效性,我们构建了具有明确真值的合成数据,并结合自然数据验证结果在真实场景中的适用性。通过对比有无反馈时的模型修订,在两种设定下均发现,基于反馈的修订能更显著地解决目标问题。最后揭示评估偏差根源:当模型仅因反馈成功修复问题时,大模型裁判常无法识别真正修正后的输出,反而偏好表现较差的基线结果。
原文摘要 · Abstract (English)
Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large Language Models (LLMs). However, recent studies suggest this feedback is inherently noisy and difficult to leverage effectively. We challenge this conception by demonstrating that user feedback is a highly actionable signal for improvement, and that its perceived ineffectiveness stems from a systematic bias in current evaluation paradigms. To isolate the usefulness of feedback, we construct synthetic data with a definitive ground truth, alongside naturalistic data to validate that our findings hold in real-world scenarios. By comparing model revisions generated with and without access to feedback across both settings, we show that feedback-informed revisions resolve targeted issues at significantly higher rates than baseline revisions. Finally, we expose the root of the evaluation bias: when a model successfully fixes an issue exclusively due to feedback, LLM judges frequently fail to identify the genuinely corrected response, systematically preferring inferior baseline outputs instead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。