用大模型零训练解码社区环境,准确率超88%。
Decoding Neighborhood Environments with Large Language Models
- 用YOLOv11模型精准检测6类环境要素,平均准确率达99.13%
- 四款大模型经提示工程后识别准确率超88%,无需训练
- 适合城市规划、公共健康研究者快速评估社区环境
邻里环境包含住房质量、道路和人行道等物理与环境因素,显著影响人类健康与福祉。传统评估方法如实地调研和地理信息系统(GIS)成本高,难以大规模应用。尽管机器学习有潜力实现自动化分析,但标注数据耗时且缺乏可访问模型限制了其扩展性。本研究探索大语言模型(LLM)如ChatGPT、Gemini在规模化解码邻里环境(如人行道、电力线)方面的可行性。我们训练了一个稳健的YOLOv11模型,在六类环境指标(路灯、人行道、电力线、公寓、单车道、多车道)上达到99.13%的平均准确率。随后评估了四款LLM(ChatGPT、Gemini、Claude、Grok),考察其在提示策略与微调下的可行性、鲁棒性及局限性。通过前三名模型的多数投票,实现超过88%的准确率,表明无需训练即可利用大模型有效解码邻里环境。
原文摘要 · Abstract (English)
Neighborhood environments include physical and environmental conditions such as housing quality, roads, and sidewalks, which significantly influence human health and well-being. Traditional methods for assessing these environments, including field surveys and geographic information systems (GIS), are resource-intensive and challenging to evaluate neighborhood environments at scale. Although machine learning offers potential for automated analysis, the laborious process of labeling training data and the lack of accessible models hinder scalability. This study explores the feasibility of large language models (LLMs) such as ChatGPT and Gemini as tools for decoding neighborhood environments (e.g., sidewalk and powerline) at scale. We train a robust YOLOv11-based model, which achieves an average accuracy of 99.13% in detecting six environmental indicators, including streetlight, sidewalk, powerline, apartment, single-lane road, and multilane road. We then evaluate four LLMs, including ChatGPT, Gemini, Claude, and Grok, to assess their feasibility, robustness, and limitations in identifying these indicators, with a focus on the impact of prompting strategies and fine-tuning. We apply majority voting with the top three LLMs to achieve over 88% accuracy, which demonstrates LLMs could be a useful tool to decode the neighborhood environment without any training effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。