大规模训练让陌生智能体能零样本互懂画图。
Drawing with Strangers: Population Scaling Drives Zero-Shot Mutual Intelligibility in Emergent Sketching

- 用大规模独立训练的智能体画图交流,实现零样本互通。
- 群体规模越大,跨组画风越相似,形成通用表达方式。
- 画得像目标图像才被保留,靠视觉感知建立通用语言。
涌现通信中的泛化通常关注新输入或语言结构,但智能体在无前期接触的情况下与完全独立群体沟通的能力仍缺乏研究。本文提出零样本互懂(ZMI):独立训练群体间无需预训练即可成功交流。通过基于绘制笔触的视觉化通信模态——涌现画图,我们发现大幅增加训练群体规模可显著提升跨群体沟通效果。关键在于,随着群体规模扩大,组内表达差异增加,避免同质化;而组间差异减小,呈现向特定通用结构收敛的趋势。进一步分析表明,这种通用性源于感知基础:大规模群体的画作越来越贴近目标图像的客观视觉特征。结果表明ZMI是涌现通信中独特的泛化维度,为构建社会兼容的人工智能代理提供了路径。
原文摘要 · Abstract (English)
Generalization in emergent communication has largely focused on novel inputs or linguistic structures, yet the capacity for agents to communicate with strangers from strictly disjoint communities remains relatively unexplored. In this work, we formalize this capability as \textit{zero-shot mutual intelligibility (ZMI)}: successful communication between independently trained populations without prior exposure. Leveraging emergent sketching -- in which agents communicate through sets of drawn strokes -- as a visually grounded modality, we find that scaling the training population substantially improves ZMI across independent groups. Crucially, as we scale the population size, in-group communicative variation increases, preventing co-adaptation into homogeneity. Simultaneously, cross-group variation decreases, indicating a structural convergence toward a certain type of universality. Further analysis reveals that this universality is achieved through perceptual grounding: scaled populations increasingly anchor their emergent sketches on the objective visual resemblance of the target images. Together, these results position ZMI as a distinct axis of generalization in emergent communication and suggest a route toward socially interoperable artificial agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。