提出多智能体框架,检测大模型在自我身份表达上的不一致问题。
STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs

- 设计角色代理协同探查模型自我身份行为
- 多数模型在身份表述上存在明显不一致性
- 适用于评估模型可解释性与公平性,适合安全研究者
知识蒸馏是训练和微调大语言模型的常用技术,可在显著降低计算成本的同时,实现从大型教师模型向小型学生模型的知识与功能迁移。然而,随着蒸馏规模与复杂性的提升,一个关键问题浮现:教师模型究竟传递了哪些知识?本文认为,除了功能性知识外,学生模型还习得了特定的行为模式,尤其是其对自身身份的表达方式,这引发了输出同质化、模型偏见与责任归属等问题。为此,我们提出STEMMA——一种多模态多智能体框架,通过角色特化的智能体协作探测不同模型在自我身份识别方面的表现。同时,我们手工设计了一组对抗性提示,用于评估大模型的身份一致性。实验结果表明,多数模型在自我表征上存在不同程度的不一致性。
原文摘要 · Abstract (English)
Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while significantly reducing computational costs. However, as the use of distillation increases in both scale and complexity it raises an important question about what kind of knowledge is really transferred from the teacher model. In this work, we argue that apart from the functional knowledge, student models also learn behavioral patterns, specifically how a model represents its own identity raising concerns about output homogeneity, model biases, and accountability. To address this challenge, we introduce STEMMA, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models. We also contribute a set of adversarial prompts designed manually to evaluate identity consistency in LLMs. Our results show that to an extent most models are vulnerable to inconsistencies in self-representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。