用双路检索增强大模型,提升关键软件合规评估的准确性
Document Retrieval Augmented Fine-Tuning (DRAFT) for safety-critical software assessments
- 通过同时检索文档与标准规范,增强模型对合规问题的理解
- 在GPT-4o-mini上实现正确率提升7%,证据处理更准确
- 适合需要可解释性与监管可信度的工业级软件评估场景
安全关键软件评估需应对复杂的法规框架,传统方法受限于人工评审。本文提出文档检索增强微调(DRAFT),一种基于大语言模型的新型合规评估方法。DRAFT在现有检索增强生成(RAG)基础上,引入双路检索架构,同步访问软件文档与适用参考标准,并设计半自动化数据集生成方法,包含不同数量的相关文档与有意义的干扰项,贴近真实评估场景。实验表明,使用GPT-4o-mini时,DRAFT相比基线模型正确率提升7%,在证据引用、回答结构和领域推理方面均有定性改进。该方法为提升合规评估系统性能提供了可解释、基于证据的实用方案。
原文摘要 · Abstract (English)
Safety critical software assessment requires robust assessment against complex regulatory frameworks, a process traditionally limited by manual evaluation. This paper presents Document Retrieval-Augmented Fine-Tuning (DRAFT), a novel approach that enhances the capabilities of a large language model (LLM) for safety-critical compliance assessment. DRAFT builds upon existing Retrieval-Augmented Generation (RAG) techniques by introducing a novel fine-tuning framework that accommodates our dual-retrieval architecture, which simultaneously accesses both software documentation and applicable reference standards. To fine-tune DRAFT, we develop a semi-automated dataset generation methodology that incorporates variable numbers of relevant documents with meaningful distractors, closely mirroring real-world assessment scenarios. Experiments with GPT-4o-mini demonstrate a 7% improvement in correctness over the baseline model, with qualitative improvements in evidence handling, response structure, and domain-specific reasoning. DRAFT represents a practical approach to improving compliance assessment systems while maintaining the transparency and evidence-based reasoning essential in regulatory domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。