arXiv:2607.24539cs.AIcs.SY2026-07

检测并修复电力系统诊断中多模态大模型的证据依赖偏差

Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis

  • 通过自报告、干预行为与预注册重要性对比,评估模型证据使用是否合理
  • 在IEEE 39和118节点系统上发现并修正了不同规模模型的可靠性问题
  • 适合关注电力系统AI可解释性与可信推理的研究者与工程师

多模态大语言模型可融合拓扑结构、测量数据与故障文本进行电网故障诊断,但回答准确并不等于使用了恰当证据。本文提出一种任务条件下的可信度审计框架,通过比较自报告依赖、受控模态删减下的行为变化及预注册工程重要性来识别问题。框架首先登记任务特定的证据要求,并与模型自述依赖及行为变化进行对比。针对发现的偏差,设计了证据约束下的修正与重审计机制:在证据限制下重生成失败响应,并独立重新删减模态,验证改进的合理性且不损失性能。案例研究在IEEE 39-和118-节点系统上评估三种不同规模的LLM,结果验证了该框架在检测、诊断与纠正任务条件下的可信度失效方面的能力。

原文摘要 · Abstract (English)

Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with self-reported reliance and behavioral changes under controlled modality ablations. To resolve detected discrepancies, we design an evidence-gated correction and re-audit mechanism that regenerates failed responses under evidence constraints and independently re-ablates them to verify improved grounding without performance loss. Case studies evaluate three differently scaled LLMs on IEEE 39- and 118-bus scenarios. These results validate the framework ability to detect, diagnose, and correct task-conditional faithfulness failures.

多模态模型电网诊断可信度审计大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。