arXiv:2410.13852cs.CLcs.AI2024-10ACL被引 7

通过回顾用户交互中的隐式反馈,让大模型自我改进任务完成率。

Retrospective Learning from Interactions

  • 利用用户重述、抱怨等行为作为隐式反馈信号进行反思学习。
  • 在多轮交互中将任务完成率从31%提升至82%。
  • 无需人工标注,适合需要持续优化的对话系统场景。

大型语言模型与用户的多轮交互自然包含隐式反馈信号。若模型对指令响应异常,用户通常会通过重述请求、表达不满或转向其他任务来传递信号。这些信号具有任务无关性,且集中在语言的有限子空间内,使得模型即使未能完成任务,也能识别它们。我们提出ReSpect方法,通过回顾机制从历史交互中学习这些信号,无需额外标注。我们在一个新提出的多模态交互场景中部署ReSpect,让用户指导多模态LLM解决具有组合解空间的抽象推理任务。通过数千次与人类的交互,我们展示了ReSpect能将任务完成率从31%逐步提升至82%,全程无需外部标注。

原文摘要 · Abstract (English)

Multi-turn interactions between large language models (LLMs) and users naturally include implicit feedback signals. If an LLM responds in an unexpected way to an instruction, the user is likely to signal it by rephrasing the request, expressing frustration, or pivoting to an alternative task. Such signals are task-independent and occupy a relatively constrained subspace of language, allowing the LLM to identify them even if it fails on the actual task. We introduce ReSpect, a method to learn from such signals in past interactions via retrospection without additional annotations. We deploy ReSpect in a new multimodal interaction scenario, where humans instruct a multimodal LLM to solve an abstract reasoning task with a combinatorial solution space. Through thousands of interactions with humans, we show how ReSpect gradually improves task completion rate from 31% to 82%, all without any external annotation.

自适应对话隐式反馈多模态零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。