arXiv:2509.24212cs.CL2025-09中稿 · presentation at th…

为文本转SQL和RAG设计可审计的合规评估基准

ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG

  • 基于条款证据链评估系统决策与解释
  • 量化判断准确率、检索效果与解释幻觉率
  • 适合合规场景下的AI系统可信度验证

ScenarioBench 是一个政策驱动、轨迹感知的评估基准,用于文本转SQL和检索增强生成在合规场景中的表现。每个 YAML 场景包含无窥视的黄金标准包(预期决策)、最小证据轨迹、管理条款集和标准 SQL,支持端到端评分,涵盖决策结果与理由。系统需使用同一政策文档中的条款编号进行输出解释,确保可验证性与可审计性。评估器报告决策准确率、轨迹质量(完整性、正确性、顺序)、检索有效性、通过结果集等价性判断的 SQL 正确率、政策覆盖率、延迟以及解释幻觉率。引入归一化的场景难度指数(SDI)和预算变体(SDI-R),综合考虑检索难度与时限。相比以往的 Text-to-SQL 或 KILT/RAG 基准,ScenarioBench 在严格无窥视规则下将每个决策与条款级证据绑定,使性能提升更聚焦于解释质量,并在明确时间预算下实现评估。

原文摘要 · Abstract (English)

ScenarioBench is a policy-grounded, trace-aware benchmark for evaluating Text-to-SQL and retrieval-augmented generation in compliance contexts. Each YAML scenario includes a no-peek gold-standard package with the expected decision, a minimal witness trace, the governing clause set, and the canonical SQL, enabling end-to-end scoring of both what a system decides and why. Systems must justify outputs using clause IDs from the same policy canon, making explanations falsifiable and audit-ready. The evaluator reports decision accuracy, trace quality (completeness, correctness, order), retrieval effectiveness, SQL correctness via result-set equivalence, policy coverage, latency, and an explanation-hallucination rate. A normalized Scenario Difficulty Index (SDI) and a budgeted variant (SDI-R) aggregate results while accounting for retrieval difficulty and time. Compared with prior Text-to-SQL or KILT/RAG benchmarks, ScenarioBench ties each decision to clause-level evidence under strict grounding and no-peek rules, shifting gains toward justification quality under explicit time budgets.

文本转SQLRAG评估合规推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。