多智能体通信让模型自发现物理属性的组合表示,无需标签也能精准识别材质、质量等隐含属性。
Emergent Compositional Communication for Latent World Properties

- 通过迭代学习与门控软化瓶颈,多智能体自发形成位置解耦的通信协议。
- 4个智能体在80组种子下100%达成高组合性(位置解耦度0.999),目标属性干扰下降15%。
- 适用于物理推理、动作规划,且在真实视频中仍保持85.6%的质量判断准确率。
多智能体通过带Gumbel-Softmax瓶颈的通信,在无属性标签和消息结构监督下,能否从冻结视频特征中提取不可见物理属性的离散组合表示?我们证明,4个智能体通过迭代学习,可发展出对弹性、摩擦力、质量比等潜伏属性的位置解耦协议,100%的80组随机种子收敛至近完美组合性(PosDis=0.999,保留集准确率98.3%)。对照实验确认,该效应由多智能体结构驱动,而非带宽或时间覆盖。因果干预显示,对特定属性的手术式扰动导致其性能下降约15%,其他属性影响小于3%。在骨干网络对比中,基于空间可见性的斜坡物理任务上,DINOv2表现更优(98.3% vs 95.1%);而仅依赖动态碰撞的任务中,V-JEPA 2领先(87.4% vs 77.7%,d=2.74)。规模匹配(d=3.37)和帧数匹配(d=6.53)控制表明该差距完全源于视频原生预训练。冻结协议支持动作条件规划(91.5%)与反事实速度推理(r=0.780)。在Physics 101真实摄像机画面验证中,对未见物体的质量比较准确率达85.6%,时序动态信息贡献额外11.2%性能提升;4智能体组合性在90%水平可复现;因果干预扩展至真实视频(d=1.87,p=0.022)。
原文摘要 · Abstract (English)
Can multi-agent communication pressure extract discrete, compositional representations of invisible physical properties from frozen video features? We show that agents communicating through a Gumbel-Softmax bottleneck with iterated learning develop positionally disentangled protocols for latent properties (elasticity, friction, mass ratio) without property labels or supervision on message structure. With 4 agents, 100% of 80 seeds converge to near-perfect compositionality (PosDis=0.999, holdout 98.3%). Controls confirm multi-agent structure -- not bandwidth or temporal coverage -- drives this effect. Causal intervention shows surgical property disruption (~15% drop on targeted property, <3% on others). A controlled backbone comparison reveals that the perceptual prior determines what is communicable: DINOv2 dominates on spatially-visible ramp physics (98.3% vs 95.1%), while V-JEPA 2 dominates on dynamics-only collision physics (87.4% vs 77.7%, d=2.74). Scale-matched (d=3.37) and frame-matched (d=6.53) controls attribute this gap entirely to video-native pretraining. The frozen protocol supports action-conditioned planning (91.5%) with counterfactual velocity reasoning (r=0.780). Validation on Physics 101 real camera footage confirms 85.6% mass-comparison accuracy on unseen objects, temporal dynamics contributing +11.2% beyond static appearance, agent-scaling compositionality replicating at 90% for 4 agents, and causal intervention extending to real video (d=1.87, p=0.022).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。