用大模型指导偏好学习,减少人力与计算开销
MAPLE: A Framework for Active Preference Learning Guided by Large Language Models
- 大模型结合自然语言与排序反馈,动态建模偏好分布
- 主动学习机制使样本效率提升,真实场景下查询响应率提高40%
- 适合需要高效人机协作的智能系统设计者
大型语言模型(LLMs)的兴起激发了利用自然语言进行偏好学习的兴趣。然而,现有方法常面临高计算开销、依赖大量人工标注且缺乏可解释性的问题。为此,我们提出MAPLE——一种由大模型引导的贝叶斯主动偏好学习框架。MAPLE利用大模型对偏好函数分布进行建模,同时融合自然语言反馈与传统偏好学习反馈(如轨迹成对排序)。该框架采用主动学习策略系统性降低分布不确定性,并引入语言条件下的主动查询选择机制,以识别信息量大且易回答的查询,从而减轻人类负担。我们在两个基准上评估了MAPLE的样本效率和偏好推断质量,包括使用OpenStreetMap数据的真实世界车辆路径规划任务。结果表明,MAPLE显著加速了学习过程,并有效提升了人类回答查询的能力。
原文摘要 · Abstract (English)
The advent of large language models (LLMs) has sparked significant interest in using natural language for preference learning. However, existing methods often suffer from high computational burdens, taxing human supervision, and lack of interpretability. To address these issues, we introduce MAPLE, a framework for large language model-guided Bayesian active preference learning. MAPLE leverages LLMs to model the distribution over preference functions, conditioning it on both natural language feedback and conventional preference learning feedback, such as pairwise trajectory rankings. MAPLE also employs active learning to systematically reduce uncertainty in this distribution and incorporates a language-conditioned active query selection mechanism to identify informative and easy-to-answer queries, thus reducing human burden. We evaluate MAPLE's sample efficiency and preference inference quality across two benchmarks, including a real-world vehicle route planning benchmark using OpenStreetMap data. Our results demonstrate that MAPLE accelerates the learning process and effectively improves humans' ability to answer queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。