arXiv:2411.16391cs.CLcs.AI2024-11被引 5

为高风险场景下的生成式语言模型设计可校准的自动化评估框架

Human-Calibrated Automated Testing and Validation of Generative Language Models

  • 基于分层采样自动生成测试用例,结合嵌入向量度量功能与风险安全
  • 通过两阶段校准使机器评估结果与人工判断高度一致
  • 支持对抗样本、分布外数据等鲁棒性测试,适合金融等严监管领域

本文提出一种针对生成式语言模型(GLMs)的综合评估与验证框架,特别关注在银行等高风险领域部署的检索增强生成(RAG)系统。由于输出开放且质量评估主观,评估极具挑战。利用RAG系统中生成内容基于预定义文档集合的结构特性,本文提出人类校准的自动化测试(HCAT)框架,包含:a)基于分层采样的自动化测试生成;b)基于嵌入的可解释性度量,用于评估功能、风险和安全属性;c)两阶段校准方法,通过概率校准与置信区间预测,使机器评估结果与人工判断对齐。框架还包含鲁棒性测试,以评估模型在对抗样本、分布外数据及多样化输入条件下的表现,并通过边际与双变量分析定位具体薄弱环节。该多层、可校准、透明的评估框架为要求高准确性、透明度与合规性的应用场景提供了可靠、可扩展的解决方案。

原文摘要 · Abstract (English)

This paper introduces a comprehensive framework for the evaluation and validation of generative language models (GLMs), with a focus on Retrieval-Augmented Generation (RAG) systems deployed in high-stakes domains such as banking. GLM evaluation is challenging due to open-ended outputs and subjective quality assessments. Leveraging the structured nature of RAG systems, where generated responses are grounded in a predefined document collection, we propose the Human-Calibrated Automated Testing (HCAT) framework. HCAT integrates a) automated test generation using stratified sampling, b) embedding-based metrics for explainable assessment of functionality, risk and safety attributes, and c) a two-stage calibration approach that aligns machine-generated evaluations with human judgments through probability calibration and conformal prediction. In addition, the framework includes robustness testing to evaluate model performance against adversarial, out-of-distribution, and varied input conditions, as well as targeted weakness identification using marginal and bivariate analysis to pinpoint specific areas for improvement. This human-calibrated, multi-layered evaluation framework offers a scalable, transparent, and interpretable approach to GLM assessment, providing a practical and reliable solution for deploying GLMs in applications where accuracy, transparency, and regulatory compliance are paramount.

语言模型评估RAG系统可信AI金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。