arXiv:2506.02204cs.CL2025-06ACL被引 6

自动发现大模型在细粒度文本中的性能差异

BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models

  • 用感知性能的上下文嵌入,定位模型表现差异的细微文本特征
  • 识别出如'if you were'这类特定语境下模型生成难易度差异
  • 适合想深入理解模型优劣细节的研究者与开发者

语言模型评估极为困难:提示词脆弱,整体困惑度模糊,基准测试选择无穷。发现能体现两模型间有意义、可泛化的差异样本至关重要。能否自动化完成?本文提出行为框(BehaviorBox)方法,利用性能感知的上下文嵌入,自动识别两模型在文本生成难易度上的细粒度差异特征。该方法提取出描述特定词语组在精细语境中的表现差异,例如‘if you were’中的条件性‘were’,或情感陈述后的感叹号使用。我们在不同规模、模型族系及后训练策略的模型间应用该方法,揭示了仅靠整体困惑度无法发现的实质性性能差异,为理解模型优劣提供了新视角。

原文摘要 · Abstract (English)

Language model evaluation is a daunting task: prompts are brittle, corpus-level perplexities are vague, and the choice of benchmarks are endless. Finding examples that show meaningful, generalizable differences between two LMs is crucial to understanding where one model succeeds and another fails. Can this process be done automatically? In this work, we propose methodology for automated comparison of language models that uses performance-aware contextual embeddings to find fine-grained features of text where one LM outperforms another. Our method, which we name BehaviorBox, extracts coherent features that demonstrate differences with respect to the ease of generation between two LMs. Specifically, BehaviorBox finds features that describe groups of words in fine-grained contexts, such as "conditional 'were' in the phrase 'if you were'" and "exclamation marks after emotional statements", where one model outperforms another within a particular datatset. We apply BehaviorBox to compare models that vary in size, model family, and post-training, and enumerate insights into specific contexts that illustrate meaningful differences in performance which cannot be found by measures such as corpus-level perplexity alone.

模型评估细粒度分析语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。