用少量用户反馈让大模型预测人类对机器人行为的评价,提升社交导航适应性。
Few-Shot Inference of Human Perceptions of Robot Performance in Social Navigation Scenarios
- 基于少量示例的上下文学习,让大模型从机器人与人的空间轨迹推断人类感知。
- 仅需传统方法1/10的数据量,就能达到或超过监督学习模型性能。
- 个性化示例能进一步提升预测准确率,适合开发可自适应的智能机器人。
理解人类在人机交互中如何评估机器人行为,对开发符合人类预期的社交机器人至关重要。传统方法依赖用户实验,而现有数据驱动方法需大量标注数据,难以实用。为此,本文利用大语言模型(LLM)的少样本学习能力,提升机器人对用户感知的预测能力,并在社交导航任务中进行实验验证。我们扩展了SEAN TOGETHER数据集,加入真实世界的人机导航视频与用户反馈。基于该数据集,评估多个LLM在仅有少量上下文示例的情况下,依据机器人与周围人运动的时空轨迹,预测人类对机器人表现的感知能力。结果表明,LLM仅需传统方法约1/10的标注样本,即可实现相当或更优的预测性能;且随着上下文示例增加,性能持续提升,验证了方法的可扩展性。通过输入特征消融实验,揭示了模型依赖的关键传感器信息。此外,引入同一用户的个性化示例作为上下文,进一步提高了预测精度。本工作为通过用户反馈实现机器人行为的可扩展优化提供了新路径。
原文摘要 · Abstract (English)
Understanding how humans evaluate robot behavior during human-robot interactions is crucial for developing socially aware robots that behave according to human expectations. While the traditional approach to capturing these evaluations is to conduct a user study, recent work has proposed utilizing machine learning instead. However, existing data-driven methods require large amounts of labeled data, which limits their use in practice. To address this gap, we propose leveraging the few-shot learning capabilities of Large Language Models (LLMs) to improve how well a robot can predict a user's perception of its performance, and study this idea experimentally in social navigation tasks. To this end, we extend the SEAN TOGETHER dataset with additional real-world human-robot navigation episodes and participant feedback. Using this augmented dataset, we evaluate the ability of several LLMs to predict human perceptions of robot performance from a small number of in-context examples, based on observed spatio-temporal cues of the robot and surrounding human motion. Our results demonstrate that LLMs can match or exceed the performance of traditional supervised learning models while requiring an order of magnitude fewer labeled instances. We further show that prediction performance can improve with more in-context examples, confirming the scalability of our approach. Additionally, we investigate what kind of sensor-based information an LLM relies on to make these inferences by conducting an ablation study on the input features considered for performance prediction. Finally, we explore the novel application of personalized examples for in-context learning, i.e., drawn from the same user being evaluated, finding that they further enhance prediction accuracy. This work paves the path to improving robot behavior in a scalable manner through user-centered feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。