把人类行为当作可执行代码来预测,提升AI对人的理解效率。
Modeling Others' Minds as Code
- 将社交行为建模为可执行代码,降低认知负担
- 在网格世界和家庭模拟器中比基线高50%准确率
- 适合需要快速适应人类行为的协作AI系统
精准预测人类行为对实现鲁棒且安全的人机协作至关重要。然而,现有方法往往数据需求大且脆弱,或假设人类完全理性,或计算成本过高难以快速适应。我们提出,许多日常社交互动遵循可预测模式,如“等绿灯再走”,这些可视为降低参与者与观察者认知负荷的高效‘脚本’。本文提出将这些行为模式建模为计算机代码而非基于信念与欲望的策略。引入新型算法ROTE,利用大语言模型(LLMs)生成行为程序的假设空间,并通过概率推理在该空间中处理不确定性。在一系列网格世界任务及大规模具身家庭模拟器中测试,ROTE仅凭稀疏观测即可预测人与AI行为,其样本内准确率和样本外泛化能力相比竞争基线(包括行为克隆与基于LLM的方法)最高提升50%。通过将动作理解视为程序合成问题,ROTE为AI在真实世界中高效、准确预测人类行为开辟了新路径。
原文摘要 · Abstract (English)
Accurate prediction of human behavior is essential for robust and safe human-AI collaboration. However, existing approaches for modeling people are often data-hungry and brittle because they either make unrealistic assumptions about rationality or are too computationally demanding to adapt rapidly. Our key insight is that many everyday social interactions may follow predictable patterns; efficient "scripts" that minimize cognitive load for actors and observers, e.g., "wait for the green light, then go." We propose modeling these routines as behavioral programs instantiated in computer code rather than policies conditioned on beliefs and desires. We introduce ROTE, a novel algorithm that leverages both large language models (LLMs) for synthesizing a hypothesis space of behavioral programs, and probabilistic inference for reasoning about uncertainty over that space. We test ROTE in a suite of gridworld tasks and a large-scale embodied household simulator. ROTE predicts human and AI behaviors from sparse observations, outperforming competitive baselines -- including behavior cloning and LLM-based methods -- by as much as 50% in terms of in-sample accuracy and out-of-sample generalization. By treating action understanding as a program synthesis problem, ROTE opens a path for AI systems to efficiently and effectively predict human behavior in the real-world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。