为大模型设计更贴近真实场景的评估框架,让评测结果更有实际参考价值。
DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
- 按使用场景拆解评估维度,区分通用与特定上下文参数
- 提出鲁棒性、连贯性、认知诚实等关键评估指标
- 适合关注落地应用效果的开发者与决策者参考
大语言模型正广泛应用于各类场景,但现有评估方法多依赖脱离实际的基准测试,难以反映真实使用情境。为此,我们提出面向维度与上下文的语言模型评估框架DICE。该框架首先分析现有基准的局限性,指出其在真实应用中的适用性不足;随后提出一系列细粒度评估维度,涵盖不依赖具体场景的通用参数(如鲁棒性、连贯性、认知诚实)和需根据部署环境定制的上下文相关参数。文章进一步探讨了该框架的可操作化路径,并讨论其在语言模型评估领域的机遇与挑战。本工作为面向具体应用场景和利益相关方需求的模型评估提供了可落地的起点。
原文摘要 · Abstract (English)
Language models (LMs) are increasingly being integrated into a wide range of applications, yet the modern evaluation paradigm does not sufficiently reflect how they are actually being used. Current evaluations rely on benchmarks that often lack direct applicability to the real-world contexts in which LMs are being deployed. To address this gap, we propose Dimensional and Contextual Evaluation (DICE), an approach that evaluates LMs on granular, context-dependent dimensions. In this position paper, we begin by examining the insufficiency of existing LM benchmarks, highlighting their limited applicability to real-world use cases. Next, we propose a set of granular evaluation parameters that capture dimensions of LM behavior that are more meaningful to stakeholders across a variety of application domains. Specifically, we introduce the concept of context-agnostic parameters - such as robustness, coherence, and epistemic honesty - and context-specific parameters that must be tailored to the specific contextual constraints and demands of stakeholders choosing to deploy LMs into a particular setting. We then discuss potential approaches to operationalize this evaluation framework, finishing with the opportunities and challenges DICE presents to the LM evaluation landscape. Ultimately, this work serves as a practical and approachable starting point for context-specific and stakeholder-relevant evaluation of LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。