提出一致性-准确率分析框架,量化大模型在输入变化下的稳定表现。
CAT: A Metric-Driven Framework for Analyzing the Consistency-Accuracy Relation of LLMs under Controlled Input Variations
- 构建一致性-准确率曲线,动态展示模型在不同一致要求下的表现
- 引入CORE指标,综合评估模型在准确与稳定间的权衡能力
- 适用于通用及领域模型,支持多类型任务评估
我们提出 extsc{CAT}框架,用于在可控输入变化下评估和可视化大型语言模型(LLMs)的准确率与响应一致性之间的相互关系,以多项选择(MC)基准为案例研究。当前评估主要关注准确率或基准得分,而一致性逐渐被视为高风险应用场景中部署的关键属性。本文认为,尽管需独立评估两者,其相互依赖性也应纳入考量,以实现更细致的评估。 extsc{CAT}的核心是一致性-准确率关系(CAR)曲线,通过最小一致性准确率(MCA)定义的一致性要求,展示准确率随一致性提升的变化趋势。我们进一步提出一致性导向鲁棒性估计(CORE)指数,结合CAR曲线的面积与形状,量化准确率与一致性的权衡。我们在多个通用与领域特定的LLMs上,针对多个MC基准进行了实证演示,并说明 extsc{CAT}可通过可适配评分函数扩展至长文本、开放式评估任务。
原文摘要 · Abstract (English)
We introduce \textsc{CAT}, a framework designed to evaluate and visualize the \emph{interplay} of \emph{accuracy} and \emph{response consistency} of Large Language Models (LLMs) under controllable input variations, using multiple-choice (MC) benchmarks as a case study. Current evaluation practices primarily focus on model capabilities such as accuracy or benchmark scores and, more recently, measuring consistency is being considered an essential property for deploying LLMs in high-stake, real-world applications. We argue in this paper that although both dimensions should still be evaluated independently, their inter-dependency also need to be considered for a more nuanced evaluation of LLMs. At the core of \textsc{CAT} are the \emph{Consistency-Accuracy Relation (CAR)} curves, which visualize how model accuracy varies with increasing consistency requirements, as defined by the \emph{Minimum-Consistency Accuracy (MCA)} metric. We further propose the \emph{Consistency-Oriented Robustness Estimate (CORE)} index, a global metric that combines the area and shape of the CAR curve to quantify the trade-off between accuracy and consistency. We present a practical demonstration of our framework across a diverse set of generalist and domain-specific LLMs, evaluated on multiple MC benchmarks. We also outline how \textsc{CAT} can be extended beyond MC tasks to support long-form, open-ended evaluations through adaptable scoring functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。