arXiv:2606.02755cs.SEcs.AI2026-06

用验收测试驱动评估,确保LLM系统可靠合规。

Acceptance-Test-Driven Evaluation Protocols for Business-Centric LLM Systems

  • 先定义失败的验收测试,再优化系统直到通过
  • 多维门禁达标后才可发布,提升系统可靠性
  • 适合金融、医疗等对合规性要求高的场景

大型语言模型应用日益需要满足确定性的机构要求,却依赖概率生成组件,传统事后基准测试难以保障系统安全、可靠、可审计且具经济价值。本文提出一种基于验收测试驱动开发、安全工程和业务导向验证的评估协议扩展。该方法在修改提示词、模型、检索或代理前,将利益相关者目标转化为可执行的行为合约、发布门禁、监控信号和证据文档。借鉴测试驱动开发的红-绿-重构模式,采用红-训练-绿生命周期:先设定期望行为的失败验收测试,再通过提示词调整、检索设计、微调、护栏机制或数据增强改进系统,最终仅当多维门禁全部满足时才发布。贡献包括治理导向的指标体系、参考架构及与提示优先、基准后测工作流对比的实证协议。

原文摘要 · Abstract (English)

Large language model (LLM) applications are increasingly expected to satisfy deterministic institutional requirements while relying on probabilistic generative components. This mismatch makes ordinary post-hoc benchmarking insufficient for systems that must be safe, reliable, auditable, and economically useful. This paper contributes an evaluation-protocol extension for operational LLM systems grounded in acceptance-test-driven development, safety engineering, and business-centric validation. The extension translates stakeholder goals into executable behavioral contracts, release gates, monitoring signals, and evidence artifacts before prompt, model, retrieval, or agent changes are accepted. It adapts the red-green-refactor discipline of test-driven development to a red-train-green lifecycle: first define failing acceptance tests for desired behavior, then improve the LLM system through prompt changes, retrieval design, fine-tuning, guardrails, or data augmentation, and finally release only when multidimensional gates are satisfied. The contribution is a governance-oriented metric stack, reference architecture, and empirical protocol for comparing acceptance-test-driven LLM development against prompt-first and benchmark-after workflows.

LLM评估验收测试系统安全业务合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。