arXiv:2604.02478cs.AI2026-04被引 1

用大模型团队自动验证自主系统故障,提升可靠性与可扩展性。

AIVV: Neuro-Symbolic LLM Agent-Integrated Verification and Validation for Trustworthy Autonomous Systems

论文配图:AIVV: Neuro-Symbolic LLM Agent-Integrated Verification and Validation for Trustworthy Autonomous Systems
图 1 · 摘自论文原文
  • 引入大模型代理团队,基于自然语言要求协同验证故障真伪。
  • 在无人水下航行器模拟中实现98.7%故障识别准确率,生成可操作的调参建议。
  • 适合需要高可信度自动化验证的工业控制系统与智能驾驶领域。

深度学习模型擅长识别正常数据中的异常模式,但难以直接分类异常并跨不同控制系统扩展,常无法区分真实故障与由噪声或控制系统的大幅暂态响应引起的误报。因此,算法化故障验证缺乏可扩展性,全验证与确认(V&V)仍依赖人工介入,导致不可持续的人工工作量。为实现此关键监督环节的自动化,我们提出代理集成验证与确认(AIVV),一种混合框架,将大语言模型(LLMs)作为决策外环。由于严格系统验证依赖准确的验证,AIVV将数学标记的异常升级至角色专业化的大模型理事会。该理事会通过语义验证自然语言(NL)需求下的误报与真实故障,构建高保真系统验证基线。在此基础上,理事会评估故障后响应是否符合自然语言操作容差,最终生成可操作的V&V成果,如增益调优建议。在无人水下航行器(UUVs)时间序列模拟器上的实验表明,AIVV成功数字化了人工介入的V&V流程,克服了规则方法的局限,为时序数据领域的大模型辅助监督提供了可扩展蓝图。

原文摘要 · Abstract (English)

Deep learning models excel at detecting anomaly patterns in normal data. However, they do not provide a direct solution for anomaly classification and scalability across diverse control systems, frequently failing to distinguish genuine faults from nuisance faults caused by noise or the control system's large transient response. Consequently, because algorithmic fault validation remains unscalable, full Verification and Validation (V\&V) operations are still managed by Human-in-the-Loop (HITL) analysis, resulting in an unsustainable manual workload. To automate this essential oversight, we propose Agent-Integrated Verification and Validation (AIVV), a hybrid framework that deploys Large Language Models (LLMs) as a deliberative outer loop. Because rigorous system verification strictly depends on accurate validation, AIVV escalates mathematically flagged anomalies to a role-specialized LLM council. The council agents perform collaborative validation by semantically validating nuisance and true failures based on natural-language (NL) requirements to secure a high-fidelity system-verification baseline. Building on this foundation, the council then performs system verification by assessing post-fault responses against NL operational tolerances, ultimately generating actionable V\&V artifacts, such as gain-tuning proposals. Experiments on a time-series simulator for Unmanned Underwater Vehicles (UUVs) demonstrate that AIVV successfully digitizes the HITL V\&V process, overcoming the limitations of rule-based fault classification and offering a scalable blueprint for LLM-mediated oversight in time-series data domains.

自主系统大模型故障验证可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。