用多模态大模型实现全球光伏分布精准识别,跨区域适应性强。
Cross-Domain Generalization of Multimodal LLMs for Global Photovoltaic Assessment
- 用结构化提示+微调,统一完成光伏检测、定位与量化。
- 跨区域测试中性能下降最小,优于传统视觉模型和变压器基线。
- 适合需要全球可迁移、可解释光伏地图的能源规划者。
分布式光伏系统快速扩张给电网管理带来挑战,许多安装未被记录。卫星影像虽具全球覆盖能力,但传统计算机视觉模型(如CNN、U-Net)需大量标注数据,且跨区域泛化能力差。本研究探索多模态大语言模型(MLLM)在全局光伏评估中的跨域泛化能力。通过结构化提示与微调,模型在统一框架内实现检测、定位与量化。基于ΔF1指标的跨区域评估显示,该模型在未见区域表现退化最小,优于传统CV与Transformer基线。结果表明,多模态大模型在领域偏移下具有强鲁棒性,具备可扩展、可迁移、可解释的全球光伏制图潜力。
原文摘要 · Abstract (English)
The rapid expansion of distributed photovoltaic (PV) systems poses challenges for power grid management, as many installations remain undocumented. While satellite imagery provides global coverage, traditional computer vision (CV) models such as CNNs and U-Nets require extensive labeled data and fail to generalize across regions. This study investigates the cross-domain generalization of a multimodal large language model (LLM) for global PV assessment. By leveraging structured prompts and fine-tuning, the model integrates detection, localization, and quantification within a unified schema. Cross-regional evaluation using the $Δ$F1 metric demonstrates that the proposed model achieves the smallest performance degradation across unseen regions, outperforming conventional CV and transformer baselines. These results highlight the robustness of multimodal LLMs under domain shift and their potential for scalable, transferable, and interpretable global PV mapping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。