用世界模型预测个体特质,提升机制设计的效率与效果。
Social World Model-Augmented Mechanism Design Policy Learning
- 构建社会世界模型,从交互轨迹推断个体特质并预测反应
- 在多场景中实现更高累积收益与更少样本需求
- 适合需要高效学习的复杂多智能体系统设计
在人工社会智能中,设计能协调个体与集体利益的自适应机制仍是核心挑战。现有方法难以建模具有持续潜在特质(如能力、偏好)的异质智能体,且难以应对复杂的多智能体系统动态。这一问题因现实交互成本高而需高样本效率更加严峻。世界模型通过学习环境动态,为提升异质复杂系统中的机制设计提供了新路径。本文提出SWM-AP(Social World Model-Augmented Mechanism Design Policy Learning),通过分层学习社会世界模型来增强机制设计。该模型从智能体交互轨迹中推断其特质,并建立基于特质的模型以预测对部署机制的响应。机制设计策略通过与社会世界模型交互获取大量训练轨迹,同时在真实交互中在线推断智能体特质,进一步提升学习效率。在税收政策设计、团队协作和设施选址等多样场景中,SWM-AP在累积收益和样本效率上均优于主流模型基与模型无关强化学习基线。
原文摘要 · Abstract (English)
Designing adaptive mechanisms to align individual and collective interests remains a central challenge in artificial social intelligence. Existing methods often struggle with modeling heterogeneous agents possessing persistent latent traits (e.g., skills, preferences) and dealing with complex multi-agent system dynamics. These challenges are compounded by the critical need for high sample efficiency due to costly real-world interactions. World Models, by learning to predict environmental dynamics, offer a promising pathway to enhance mechanism design in heterogeneous and complex systems. In this paper, we introduce a novel method named SWM-AP (Social World Model-Augmented Mechanism Design Policy Learning), which learns a social world model hierarchically modeling agents' behavior to enhance mechanism design. Specifically, the social world model infers agents' traits from their interaction trajectories and learns a trait-based model to predict agents' responses to the deployed mechanisms. The mechanism design policy collects extensive training trajectories by interacting with the social world model, while concurrently inferring agents' traits online during real-world interactions to further boost policy learning efficiency. Experiments in diverse settings (tax policy design, team coordination, and facility location) demonstrate that SWM-AP outperforms established model-based and model-free RL baselines in cumulative rewards and sample efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。