arXiv:2510.08931cs.AIcs.LG2025-10被引 1

用机制可解释性区分大模型是靠记忆还是推理,提升评估可靠性。

RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation

  • 通过提取37个表层与深层特征,分析模型响应机制。
  • 在多样测试集上达到93%准确率,模糊案例仍达76.7%。
  • 适合关注模型真实能力评估的研究者和开发者。

数据污染对大模型评估的可靠性构成重大挑战,模型可能因记忆训练数据而非真正推理而获得高分。我们提出RADAR(通过激活表示检测回忆与推理),一种基于机制可解释性的新框架,通过区分基于回忆与基于推理的模型响应来检测污染。RADAR提取了37个特征,涵盖表面置信度轨迹以及注意力专一性、电路动态和激活流模式等深层机制属性。利用这些特征训练的集成分类器在多样化评估集上实现93%准确率,清晰案例中表现完美,模糊案例准确率为76.7%。本工作展示了机制可解释性在超越传统表面指标方面推动大模型评估的潜力。

原文摘要 · Abstract (English)

Data contamination poses a significant challenge to reliable LLM evaluation, where models may achieve high performance by memorizing training data rather than demonstrating genuine reasoning capabilities. We introduce RADAR (Recall vs. Reasoning Detection through Activation Representation), a novel framework that leverages mechanistic interpretability to detect contamination by distinguishing recall-based from reasoning-based model responses. RADAR extracts 37 features spanning surface-level confidence trajectories and deep mechanistic properties including attention specialization, circuit dynamics, and activation flow patterns. Using an ensemble of classifiers trained on these features, RADAR achieves 93\% accuracy on a diverse evaluation set, with perfect performance on clear cases and 76.7\% accuracy on challenging ambiguous examples. This work demonstrates the potential of mechanistic interpretability for advancing LLM evaluation beyond traditional surface-level metrics.

大模型评估机制可解释数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。