arXiv:2509.04504cs.CLcs.AI2025-09被引 6

用诊断提示+自动评估,挖出大模型的内在行为指纹

Behavioral Fingerprinting of Large Language Models

  • 设计诊断提示集和自动评估流程,量化模型认知与交互风格
  • 18个模型中顶尖模型推理趋同,但对齐行为差异显著
  • 发现跨模型默认人格集群,揭示对齐策略决定交互风格

当前大语言模型评测主要关注性能指标,难以捕捉模型间细微的行为差异。本文提出一种新型「行为指纹」框架,通过精心设计的诊断提示集和创新的自动化评估流程(以强大LLM作为公正裁判),分析十八个不同能力层级的模型。结果揭示:尽管顶级模型在抽象与因果推理等核心能力上趋于收敛,但在对齐相关行为(如阿谀奉承、语义鲁棒性)上却存在巨大差异。我们进一步发现跨模型存在默认人格聚类(ISTJ/ESTJ),可能反映共同的对齐激励。这表明模型的交互特征并非规模或推理能力的自然产物,而是特定、高度可变的开发者对齐策略直接导致的结果。该框架为可复现、可扩展地揭示深层行为差异提供了方法。

原文摘要 · Abstract (English)

Current benchmarks for Large Language Models (LLMs) primarily focus on performance metrics, often failing to capture the nuanced behavioral characteristics that differentiate them. This paper introduces a novel ``Behavioral Fingerprinting'' framework designed to move beyond traditional evaluation by creating a multi-faceted profile of a model's intrinsic cognitive and interactive styles. Using a curated \textit{Diagnostic Prompt Suite} and an innovative, automated evaluation pipeline where a powerful LLM acts as an impartial judge, we analyze eighteen models across capability tiers. Our results reveal a critical divergence in the LLM landscape: while core capabilities like abstract and causal reasoning are converging among top models, alignment-related behaviors such as sycophancy and semantic robustness vary dramatically. We further document a cross-model default persona clustering (ISTJ/ESTJ) that likely reflects common alignment incentives. Taken together, this suggests that a model's interactive nature is not an emergent property of its scale or reasoning power, but a direct consequence of specific, and highly variable, developer alignment strategies. Our framework provides a reproducible and scalable methodology for uncovering these deep behavioral differences. Project: https://github.com/JarvisPei/Behavioral-Fingerprinting

行为指纹模型评测对齐研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。