用多模态检索增强生成框架,提升核工程决策的可信度与可追溯性。
RADIANT-LLM: an Agentic Retrieval Augmented Generation Framework for Reliable Decision Support in Safety-Critical Nuclear Engineering
- 构建本地优先的多模态RAG框架,支持文档级和图级精准检索
- 在核设施设计基准下,上下文精度达85–98%,幻觉率显著降低
- 适合需要高可靠性、可审计性的核安全与监管领域应用
核工程中的可靠决策支持需具备可追溯、领域相关的知识检索能力,但现有预训练大模型在专业核领域使用时面临文档碎片化与幻觉问题。本文提出RADIANT-LLM(基于LLM的核技术检索增强智能体),一个面向核安全、安保与保障应用的多模态检索增强生成框架。该框架采用本地优先、模型无关架构,结合多模态文档摄入管道与结构化元数据丰富的知识库,支持从技术文档中实现页面级与图表级检索。其智能体层协调领域专用工具,强制要求引用支撑的回答并追踪出处,支持人工介入验证以降低幻觉风险。为严格评估,我们开发并应用一系列领域感知指标,包括上下文精确度(CoP)、幻觉率(HR)和视觉召回率(ViR),基于已专家标注的使用后核燃料储存设施设计指南基准测试。在不同知识库规模下,CoP与ViR保持在85–98%区间,幻觉率远低于通用大模型部署表现。当相同查询提交至无RAG层的商用大模型平台时,幻觉与引用错误显著上升。结果表明,具备领域特定检索与溯源控制的本地化多模态RAG框架,是满足核工程工作流对事实准确性、透明性与审计性需求的关键。
原文摘要 · Abstract (English)
Reliable decision support in nuclear engineering requires traceable, domain-grounded knowledge retrieval, yet safety and risk analysis workflows remain hampered by fragmented documentation and hallucination when use pre-trained large language model (LLM) in specialized nuclear domains. To address these challenges, this paper presents RADIANT-LLM (Retrival-Augumented, Domain-Intelligent Agent for Nuclear Technologies using LLM), a multi-modal retrieval-augmented generation (RAG) framework designed for nuclear safety, security, and safeguards applications. The framework uses a local-first, model-agnostic architecture that pairs a multi-modal document ingestion pipeline with a structured, metadata-rich knowledge base, supporting page- and figure-level retrieval from technical documents. An agentic layer coordinates domain-specific tools, enforces citation-backed responses with provenance tracking, and supports human-in-the-loop validation to reduce hallucination risks. To rigorously evaluate this framework, we develop and apply a suite of domain-aware metrics, including Context Precision (CoP), Hallucination Rate (HR), and Visual Recall (ViR), to expert-curated benchmarks derived from Used Nuclear Fuel Storage Facility design guidance. Across varying knowledge base sizes, CoP and ViR remain within an 85--98\% band, and hallucination rates are substantially lower than those observed in general-purpose deployments. When the same queries are posed to commercial LLM platforms without the RAG layer, hallucinations and citation errors increase markedly. These results indicate that a locally controlled, multi-modal RAG framework with domain-specific retrieval and provenance enforcement is necessary to achieve the factual accuracy, transparency, and auditability that nuclear engineering workflows demand.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。