arXiv:2604.21102cs.CVcs.AI2026-04

用AI从街景图自动评估美国建筑状况,速度快精度高。

Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery

论文配图:Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
图 1 · 摘自论文原文
  • 用微调的Gemma 3 27B模型分析街景图,匹配人工评分。
  • 小模型速度提升30倍,准确率仍接近大模型。
  • 适合城市规划、房主自查等大规模评估场景。

我们提出一种新框架,通过大型语言模型(LLMs)和谷歌街景(GSV)图像,实现对美国全国建筑状况的自动化评估。通过对一个规模较小的人工标注数据集微调Gemma 3 27B,该方法在与人类平均评分(MOS)的对齐上表现优异,在SRCC和PLCC指标上甚至超过个别评估者。为提升效率,采用知识蒸馏技术,将Gemma 3 27B的能力迁移至更小的Gemma 3 4B模型,实现3倍加速且性能相近。进一步蒸馏到基于CNN的EfficientNetV2-M和基于Transformer的SwinV2-B,分别实现30倍提速,性能仍接近原模型。此外,通过人-智能体对齐研究,探索了LLMs对大量建成环境与住房属性的评估能力,并开发了可视化仪表板,供房主进行后续分析。该框架提供了一种灵活高效的规模化建筑状况评估方案,以极低的人工标注成本实现高精度。

原文摘要 · Abstract (English)

We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV) imagery. By fine-tuning Gemma 3 27B on a modest human-labeled dataset, our approach achieves strong alignment with human mean opinion scores (MOS), outperforming even individual raters on SRCC and PLCC relative to the MOS benchmark. To enhance efficiency, we apply knowledge distillation, transferring the capabilities of Gemma 3 27B to a smaller Gemma 3 4B model that achieves comparable performance with a 3x speedup. Further, we distill the knowledge into a CNN-based model (EfficientNetV2-M) and a transformer (SwinV2-B), delivering close performance while achieving a 30x speed gain. Furthermore, we investigate LLMs' capabilities for assessing an extensive list of built environment and housing attributes through a human-AI alignment study and develop a visualization dashboard that integrates LLM assessment outcomes for downstream analysis by homeowners. Our framework offers a flexible and efficient solution for large-scale building condition assessment, enabling high accuracy with minimal human labeling effort.

建筑评估多模态模型街景分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。