首个针对国家标准文档的规则审查基准,提升AI对规范性文本的审查能力。
Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

- 构建覆盖5大维度25类错误的分级审查体系
- 14个主流模型平均仅达专家水平的49.4%(CMCS 0.3280)
- 多智能体框架让模型审查准确率提升至0.5094,适合高精度文档领域
大型语言模型在复杂专业任务中日益重要,但在规则密集型文档审查方面仍缺乏有效评估。国家标准文件(如中国GB/T标准)具有篇幅长、结构化强、规则明确等特点,是理想的测试场景。现有基准侧重领域知识和问答,忽视对文档内在质量的系统性审查。此类审查高度依赖人工,成本高且难以扩展。为此,我们提出首个针对国家标准文档的结构化审查基准GB/T-Bench。其GB/T审查分类体系涵盖文档结构、范围一致性、规范语态、术语一致性和规范引用等五个维度,包含25种可诊断错误类型。通过结合确定性规则与约束式LLM重写,将488份文档生成7,306个可追溯的审查错误实例。我们设计了基于精确匹配的诊断评估协议,要求错误位置、审查维度和错误类型完全一致,并引入文档级覆盖率指标。同时提出GB/T-Reviewer多智能体框架,将审查知识转化为专业技能,协调全局检查、精准诊断、规则扫描和结果验证。在14个主流模型上的实验表明,最强模型的CMCS仅为0.3280,远低于专家水平(0.6640)。使用GB/T-Reviewer后,最佳模型提升至0.5094,证明结构化技能协同对规则密集型审查的关键价值。本工作为标准化及其他高风险文档领域中的可信AI铺平道路。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency. Existing benchmarks focus on domain knowledge and question answering, largely overlooking intrinsic quality review for professional documents. Such reviews rely heavily on human experts, making them costly and difficult to scale. To bridge this gap, we introduce GB/T-Bench, the first benchmark for the structured review of national standard documents. Its GB/T Review Taxonomy is a hierarchical schema covering document structure, scope alignment, normative modality, terminology consistency, and normative references, with 25 diagnosable error types. A controllable counterexample generation mechanism combines deterministic rules and constrained LLM rewriting to process 488 documents into 7,306 traceable review error instances for evaluation. We also develop a diagnosis-oriented evaluation protocol requiring exact matches on error location, review dimension, and error type, plus document-level coverage metrics. We further propose GB/T-Reviewer, a multi-agent framework that converts review knowledge into specialized skills and coordinates global inspection, targeted diagnosis, rule scanning, and result verification. Experiments with 14 mainstream LLMs reveal a substantial human-LLM gap: the strongest model achieves only 0.3280 CMCS versus 0.6640 for experts. GB/T-Reviewer raises the best CMCS to 0.5094, showing the value of structured skill coordination for rule-intensive document review. This work paves the way for trustworthy AI in standardization and other high-stakes document domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。