对比不同模型在街景图像上的表现,给出高效构建社区环境指标的实用方案。
How to Build Robust, Scalable Models for GSV-Based Indicators in Neighborhood Research
- 用无监督训练提升小规模标注数据下的模型性能
- 实证发现迁移学习结合无监督微调效果最佳
- 适合社会健康研究者快速部署视觉分析工具
大量健康研究证明社区环境与健康结果密切相关。近年来,计算机视觉技术被用于大规模系统化刻画社区建成环境。然而,现有视觉模型在跨域场景(如ImageNet到Google街景图像)中的泛化能力仍不明确。在社会健康研究中,关键问题包括:应选用何种模型、是否采用无监督训练策略、在计算资源受限下可接受的训练规模,以及此类策略对下游任务的增益程度。这些决策成本高且需专业知识。本文通过实证分析回答上述问题,提供针对小样本、弱标注数据集的实用建议:利用大規模未标注数据进行无监督训练,适配基础模型。研究包含定量与可视化对比分析,比较模型在无监督适应前后的性能表现。
原文摘要 · Abstract (English)
A substantial body of health research demonstrates a strong link between neighborhood environments and health outcomes. Recently, there has been increasing interest in leveraging advances in computer vision to enable large-scale, systematic characterization of neighborhood built environments. However, the generalizability of vision models across fundamentally different domains remains uncertain, for example, transferring knowledge from ImageNet to the distinct visual characteristics of Google Street View (GSV) imagery. In applied fields such as social health research, several critical questions arise: which models are most appropriate, whether to adopt unsupervised training strategies, what training scale is feasible under computational constraints, and how much such strategies benefit downstream performance. These decisions are often costly and require specialized expertise. In this paper, we answer these questions through empirical analysis and provide practical insights into how to select and adapt foundation models for datasets with limited size and labels, while leveraging larger, unlabeled datasets through unsupervised training. Our study includes comprehensive quantitative and visual analyses comparing model performance before and after unsupervised adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。