让大模型主动引导对话,提升自传访谈质量。
GuideLLM: Exploring LLM-Guided Conversation with Applications in Autobiography Interviewing
- 设计三要素框架:目标导航、上下文管理、共情互动
- 在1.4k轮对话中表现优于GPT-4o等6个主流模型
- 适合研究人机对话引导、情感交互与叙事生成的学者
尽管大语言模型在人类引导对话中表现优异,但由模型主导对话方向的潜力仍待挖掘。本文提出GuideLLM,将其对话能力分解为三大核心:目标导航、上下文管理与共情互动,并构建了用于评估的访谈环境。该环境涵盖多主题,每轮评测包含约1.4千轮对话、18.4万词元及200多个事件。通过用户代理与大模型评分器进行自动评估,并招募45名真人参与者进行对比实验。结果表明,GuideLLM在自动评价和人工评分中均显著领先于包括GPT-4o和Llama-3-70b-Instruct在内的6个基线模型,尤其在访谈质量与自传生成质量方面表现突出。
原文摘要 · Abstract (English)
Although Large Language Models (LLMs) succeed in human-guided conversations such as instruction following and question answering, the potential of LLM-guided conversations-where LLMs direct the discourse and steer the conversation's objectives-remains under-explored. In this study, we first characterize LLM-guided conversation into three fundamental components: (i) Goal Navigation; (ii) Context Management; (iii) Empathetic Engagement, and propose GuideLLM as an installation. We then implement an interviewing environment for the evaluation of LLM-guided conversation. Specifically, various topics are involved in this environment for comprehensive interviewing evaluation, resulting in around 1.4k turns of utterances, 184k tokens, and over 200 events mentioned during the interviewing for each chatbot evaluation. We compare GuideLLM with 6 state-of-the-art LLMs such as GPT-4o and Llama-3-70b-Instruct, from the perspective of interviewing quality, and autobiography generation quality. For automatic evaluation, we derive user proxies from multiple autobiographies and employ LLM-as-a-judge to score LLM behaviors. We further conduct a human-involved experiment by employing 45 human participants to chat with GuideLLM and baselines. We then collect human feedback, preferences, and ratings regarding the qualities of conversation and autobiography. Experimental results indicate that GuideLLM significantly outperforms baseline LLMs in automatic evaluation and achieves consistent leading performances in human ratings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。