用ReAct提示让大模型做任务型对话,人评更满意但成功率不高。
Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings
- 用ReAct策略引导大模型在任务对话中思考与行动。
- 模拟中成功率低于现有方法,但真人测试差距缩小。
- 用户更满意其自然自信的回复风格,适合追求体验的场景。
大语言模型(LLMs)因其在非结构化对话中的强大能力而广受欢迎。通过引入推理与行动(ReAct)等先进提示策略(Yao et al., 2022),LLMs 在传统上需要强化学习的复杂任务中展现出潜力。本文将ReAct策略应用于任务型对话(TOD)场景,评估基于ReAct的LLM(ReAct-LLMs)在模拟环境和真实用户中的表现。结果显示,尽管在模拟环境中,ReAct-LLMs的成功率显著低于现有先进方法,但在真人评价中,这一差距明显减小。此外,相比基线模型,用户对ReAct-LLM表现出更高的主观满意度,这很可能归因于其自然且自信的回应方式。
原文摘要 · Abstract (English)
Large language models (LLMs) gained immense popularity due to their impressive capabilities in unstructured conversations. Empowering LLMs with advanced prompting strategies such as reasoning and acting (ReAct) (Yao et al., 2022) has shown promise in solving complex tasks traditionally requiring reinforcement learning. In this work, we apply the ReAct strategy to guide LLMs performing task-oriented dialogue (TOD). We evaluate ReAct-based LLMs (ReAct-LLMs) both in simulation and with real users. While ReAct-LLMs severely underperform state-of-the-art approaches on success rate in simulation, this difference becomes less pronounced in human evaluation. Moreover, compared to the baseline, humans report higher subjective satisfaction with ReAct-LLM despite its lower success rate, most likely thanks to its natural and confidently phrased responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。