arXiv:2604.05371cs.AI2026-04

用大模型当裁判,自动检查无人机巡线分割结果是否靠谱。

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection

论文配图:LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection
图 1 · 摘自论文原文
  • 用大模型评估无人机巡线分割结果的语义可靠性
  • 相同输入下判断一致,环境变差时信心下降明显
  • 适合对安全要求高的无人机巡检系统使用

将轻量级分割模型部署于无人机进行自主输电线路巡检面临严峻挑战:在实际环境中,模型性能可能因与训练数据差异而不可预测地下降,带来安全隐患。尽管如U-Net等紧凑架构可实现机载实时推理,其分割输出在恶劣环境下仍可能失控。本文研究利用大语言模型(LLM)作为语义裁判,评估无人机搭载模型产生的分割结果可靠性。不构建新系统,而是提出一种离线监控机制,由外部LLM对分割叠加图进行评估,并检验该裁判是否具有一致性与感知一致性。为此设计两种评估协议:一是通过重复查询相同输入和固定提示,测量质量评分与置信度的稳定性;二是引入可控视觉退化(雾、雨、雪、阴影、强光),分析裁判输出随分割质量逐步退化的响应。结果显示,相同条件下LLM给出高度一致的分类判断,且随视觉可靠性下降,其置信度相应降低。此外,即便在复杂条件下,裁判仍能识别出缺失或误标线路。这些发现表明,在严格约束下,LLM可作为可靠语义裁判,用于监控高安全要求的空中巡检任务中的分割质量。

原文摘要 · Abstract (English)

The deployment of lightweight segmentation models on drones for autonomous power line inspection presents a critical challenge: maintaining reliable performance under real-world conditions that differ from training data. Although compact architectures such as U-Net enable real-time onboard inference, their segmentation outputs can degrade unpredictably in adverse environments, raising safety concerns. In this work, we study the feasibility of using a large language model (LLM) as a semantic judge to assess the reliability of power line segmentation results produced by drone-mounted models. Rather than introducing a new inspection system, we formalize a watchdog scenario in which an offboard LLM evaluates segmentation overlays and examine whether such a judge can be trusted to behave consistently and perceptually coherently. To this end, we design two evaluation protocols that analyze the judge's repeatability and sensitivity. First, we assess repeatability by repeatedly querying the LLM with identical inputs and fixed prompts, measuring the stability of its quality scores and confidence estimates. Second, we evaluate perceptual sensitivity by introducing controlled visual corruptions (fog, rain, snow, shadow, and sunflare) and analyzing how the judge's outputs respond to progressive degradation in segmentation quality. Our results show that the LLM produces highly consistent categorical judgments under identical conditions while exhibiting appropriate declines in confidence as visual reliability deteriorates. Moreover, the judge remains responsive to perceptual cues such as missing or misidentified power lines, even under challenging conditions. These findings suggest that, when carefully constrained, an LLM can serve as a reliable semantic judge for monitoring segmentation quality in safety-critical aerial inspection tasks.

大模型应用无人机巡检语义评估可靠性监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。