arXiv:2507.23158cs.CL2025-07EMNLP被引 22

研究用户与大模型对话中隐式反馈的使用价值与局限。

User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal

  • 从对话日志中提取用户隐式反馈,分析其出现时机与原因。
  • 结合反馈内容和正负极性可提升模型在简单问题上的表现。
  • 对复杂长问题无效,提示隐式反馈存在使用边界。

语言模型部署后可长期与用户交互,理想情况下应根据用户反馈持续进化。直接询问用户反馈可能造成干扰,因此本文研究从用户-大模型对话日志中挖掘隐式反馈。基于WildChat和LMSYS两个数据集,首先分析了用户反馈在对话中的发生规律与动因;其次探究如何从中提取学习信号。具体而言,考察将反馈内容(如用户希望澄清)与反馈极性一同纳入训练是否能提升模型性能。实验发现:在短而设计精良的问题集MTBench上表现提升,但在更长更复杂的WildBench上未见改善。整体揭示了隐式反馈的潜力与局限。

原文摘要 · Abstract (English)

Once language models (LMs) are deployed, they can interact with users long-term, ideally evolving based on their feedback. Asking for direct user feedback can be disruptive; thus, we study harvesting implicit user feedback from user-LM interaction logs. We study two user-LM interaction datasets (WildChat and LMSYS). First, we analyze user feedback in the user-LLM conversation logs, providing insights into when and why such feedback occurs. Second, we study harvesting learning signals from such implicit user feedback. Specifically, we study whether incorporating the contents of user feedback (e.g., user wanted clarification), in addition to the polarity of the feedback, can improve the model performance. We observe mixed results, showing this helps in short human-designed questions (MTBench) but not on longer and more complex questions (WildBench). Together, we provide an in-depth study of implicit user feedback, showing its potential and limitations.

用户反馈大模型隐式信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。