提出可跨领域评估大模型深度知识的新方法,揭示模型在追问下的真实理解力。
DepthCharge: A Domain-Agnostic Framework for Measuring Depth-Dependent Knowledge in Large Language Models
- 基于模型实际回答自动生成追问问题,动态测试知识深度。
- 四领域实测显示模型表现差异大,最高知识深度仅7.55,无模型全面领先。
- 无需预设题库,适合专业场景选型,尤其适合关注领域适配性的研究者。
大型语言模型在回答通用问题时表现良好,但在领域细节追问中常出错。现有方法无法对任意领域的深度知识进行即插即用的评估。本文提出DepthCharge,一种无领域依赖的评估框架,通过三项创新实现:根据模型实际提及的概念生成自适应追问、从权威来源实时验证事实、在每层深度保持固定样本量的存活率统计。该框架可部署于任意具有公开可验证事实的知识领域,无需预先构建测试集或领域专业知识。评估结果依赖于用于答案校验的评测模型,因此适用于模型间的相对比较而非绝对准确度认证。在医学、宪法法、古罗马、量子计算四个不同领域,使用五个前沿模型进行实证验证,结果显示标准基准掩盖了深度依赖的性能差异。预期有效深度(EVD)在3.45至7.55之间,模型排名随领域显著变化,无单一模型全胜。成本-性能分析表明,高成本模型未必拥有更深知识,提示专业应用中应优先采用领域特定评估而非综合基准。
原文摘要 · Abstract (English)
Large Language Models appear competent when answering general questions but often fail when pushed into domain-specific details. No existing methodology provides an out-of-the-box solution for measuring how deeply LLMs can sustain accurate responses under adaptive follow-up questioning across arbitrary domains. We present DepthCharge, a domain-agnostic framework that measures knowledge depth through three innovations: adaptive probing that generates follow-up questions based on concepts the model actually mentions, on-demand fact verification from authoritative sources, and survival statistics with constant sample sizes at every depth level. The framework can be deployed on any knowledge domain with publicly verifiable facts, without requiring pre-constructed test sets or domain-specific expertise. DepthCharge results are relative to the evaluator model used for answer checking, making the framework a tool for comparative evaluation rather than absolute accuracy certification. Empirical validation across four diverse domains (Medicine, Constitutional Law, Ancient Rome, and Quantum Computing) with five frontier models demonstrates that DepthCharge reveals depth-dependent performance variation hidden by standard benchmarks. Expected Valid Depth (EVD) ranges from 3.45 to 7.55 across model-domain combinations, and model rankings vary substantially by domain, with no single model dominating all areas. Cost-performance analysis further reveals that expensive models do not always achieve deeper knowledge, suggesting that domain-specific evaluation is more informative than aggregate benchmarks for model selection in professional applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。