用轨迹优化提升问答效率,让AI更懂用户真实需求。
TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human Preference
- 通过轨迹优化生成精准提问序列,避免无效对话。
- 在标准任务上比基线方法提升9.32%准确率。
- 适合需要深度交互与偏好挖掘的应用场景。
大语言模型可通过多轮对话有效获取人类偏好。复杂任务可由模型作为提问者(STaR-GATE;Andukuri et al., 2024)通过迭代澄清问题和最终回应完成。然而,现有基于自教推理的方法难以识别最优对话路径,常产生与任务无关的问题。为此,我们提出TO-GATE,一个新框架,通过轨迹优化增强问题生成,包含两个核心组件:澄清解析器生成最优提问路径,摘要器确保任务对齐的最终回应。轨迹优化使模型能生成针对特定任务的有效提问与总结回应。实验表明,TO-GATE显著优于基线方法,在标准偏好获取任务上提升9.32%。
原文摘要 · Abstract (English)
Large language models (LLMs) can effectively elicit human preferences through multi-turn dialogue. Complex tasks can be accomplished through iterative clarifying questions and final responses generated by an LLM acting as a questioner (STaR-GATE; Andukuri et al., 2024}). However, existing approaches based on self-taught reasoning struggle to identify optimal dialogue trajectories and avoid irrelevant questions to the tasks. To address this limitation, we propose TO-GATE, a novel framework that enhances question generation through trajectory optimization, which consists of two key components: a clarification resolver that generates optimal questioning trajectories, and a summarizer that ensures task-aligned final responses. The trajectory optimization enables the model to produce effective elicitation questions and summary responses tailored to specific tasks. Experimental results demonstrate that TO-GATE significantly outperforms baseline methods, achieving a 9.32% improvement on standard preference elicitation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。