让智能助手主动问问题,精准理解用户偏好完成任务。
ADAPT: Actively Discovering and Adapting to Preferences for any Task
- 用教师-学生框架训练模型主动提问以获取用户偏好
- 在未见用户上比零样本基线提升6.1%的偏好满足率
- 适合需要个性化交互的家居任务类智能体研究
辅助智能体应能在任务描述不明确的情况下完成长时序任务,并尊重用户偏好。我们提出了一个名为ADAPT的基准,用于评估智能体在各类家庭任务中通过主动提问来遵循用户偏好的能力。为此,我们提出Reflection-DPO这一新训练方法,微调‘学生’大语言模型以模仿‘教师’模型的行为,并可选择性地提问以获取必要信息,从而更准确预测教师行为。我们发现,先前使用先进大模型的方法在ADAPT中因提问不足和偏好遵守差而表现不佳。相比之下,Reflection-DPO在未见过的用户上实现了更高的偏好满足率,相比零样本思维链基线提升了6.1%。
原文摘要 · Abstract (English)
Assistive agents should be able to perform under-specified long-horizon tasks while respecting user preferences. We introduce Actively Discovering and Adapting to Preferences for any Task (ADAPT) -- a benchmark designed to evaluate agents' ability to adhere to user preferences across various household tasks through active questioning. Next, we propose Reflection-DPO, a novel training approach for adapting large language models (LLMs) to the task of active questioning. Reflection-DPO finetunes a 'student' LLM to follow the actions of a privileged 'teacher' LLM, and optionally ask a question to gather necessary information to better predict the teacher action. We find that prior approaches that use state-of-the-art LLMs fail to sufficiently follow user preferences in ADAPT due to insufficient questioning and poor adherence to elicited preferences. In contrast, Reflection-DPO achieves a higher rate of satisfying user preferences, outperforming a zero-shot chain-of-thought baseline by 6.1% on unseen users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。