用多大模型挖掘隐性监管知识,自动生成合规测试用例。
Explicating Tacit Regulatory Knowledge from LLMs to Auto-Formalize Requirements for Compliance Test Case Generation
- 通过自适应净化聚合,从多个大模型中提取隐性监管知识。
- 在金融、汽车、电力领域达到专家水平,生成效率提升显著。
- 适合需要自动化合规测试的高监管行业研发团队使用。
高监管领域的合规测试至关重要但高度依赖人工,需领域专家将复杂法规转化为可执行测试用例。尽管大语言模型(LLMs)在自动化方面展现出潜力,但其易产生幻觉的问题限制了可靠性。现有混合方法通过形式化模型约束LLM,但仍需昂贵的手动建模。本文提出RAFT框架,通过挖掘多个LLMs中的隐性监管知识,实现需求的自动形式化与合规测试用例生成。RAFT采用自适应净化-聚合策略,将知识整合为三种产物:领域元模型、形式化需求表示和可测试性约束。这些产物动态注入提示词,引导高精度需求形式化与自动化测试生成。跨金融、汽车、电力领域的实验表明,RAFT性能达专家水平,显著优于当前最先进(SOTA)方法,且整体生成与审查时间大幅减少。
原文摘要 · Abstract (English)
Compliance testing in highly regulated domains is crucial but largely manual, requiring domain experts to translate complex regulations into executable test cases. While large language models (LLMs) show promise for automation, their susceptibility to hallucinations limits reliable application. Existing hybrid approaches mitigate this issue by constraining LLMs with formal models, but still rely on costly manual modeling. To solve this problem, this paper proposes RAFT, a framework for requirements auto-formalization and compliance test generation via explicating tacit regulatory knowledge from multiple LLMs. RAFT employs an Adaptive Purification-Aggregation strategy to explicate tacit regulatory knowledge from multiple LLMs and integrate it into three artifacts: a domain meta-model, a formal requirements representation, and testability constraints. These artifacts are then dynamically injected into prompts to guide high-precision requirement formalization and automated test generation. Experiments across financial, automotive, and power domains show that RAFT achieves expert-level performance, substantially outperforms state-of-the-art (SOTA) methods while reducing overall generation and review time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。