模型越相似,合作越好但创新越少,早期表征相似性影响最大。
Representational Similarity and Model Behavior in Multi-Agent Interaction

- 通过8种游戏测试276对大模型,分析表征相似性对协作与创新的影响。
- 相似表征的模型合作成功率更高,但创意和新颖性显著降低。
- 早期层的相似性与合作/创新关联最强,提示语义基础共享是关键。
研究人员发现,人类之间的神经表征相似性可预测社交亲密与合作成功,而创新常源于差异个体间的互动。本文探究这一规律是否适用于人工智能,考察大型语言模型间的交互。在8种涵盖合作与创新的游戏场景中,测试了276对模型的交互表现。结果表明,表征空间更相似的模型对在合作任务中表现显著更优,但创新性和创造性更低。该结论在控制性能差距与模型规模后依然成立。此外,早期层的表征相似性对合作与创新的影响最为显著,优于中层与后期层。这暗示,两模型间是否存在共通的词汇与语义基础,可能是决定其交互行为的核心因素。整体而言,表征相似性应被纳入多智能体系统设计的重要考量。
原文摘要 · Abstract (English)
Researchers have shown that neural similarity among humans predicts social closeness and cooperative success, whereas innovation often emerges from interactions among dissimilar individuals. We investigate whether these principles extend to artificial intelligence by examining interactions between large language models. In our experiments, 276 model pairs interact across eight games spanning both cooperation and novelty. We find that pairs with more similar representation spaces achieve significantly higher cooperation but exhibit reduced novelty and creativity. The effects of representational similarity on cooperation and novelty remain robust even after controlling for other factors such as performance disparity and model size. We also find that similarity in the early layers consistently shows the strongest association with cooperation and novelty, compared to the middle and later layers. This suggests that a central factor underlying these patterns could be the extent to which the two models share lexical and semantic grounding. Overall, representational similarity can be an important consideration in multi-agent system design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。