用强化学习让大模型更准地从图像推断城市经济状况。
CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning
- 通过强化学习引导大模型关注有意义的视觉线索。
- 在未见城市和指标上显著提升预测准确率与泛化能力。
- 适合城市规划、政策研究者及多模态模型开发者。
城市社会经济感知对推动全球可持续发展目标至关重要。随着大型视觉语言模型(LVLMs)的发展,这一挑战可被视作多模态感知与推理任务。然而,现有研究显示,LVLMs在从视觉数据中做出准确且可解释的社会经济预测方面仍存在不足。为克服这些局限并充分挖掘LVLM潜力,我们提出CityRiSE——一种基于强化学习(RL)的都市社会经济状态推理框架。通过精心构建的多模态数据集与可验证的奖励设计,该方法引导LVLM聚焦于语义上有意义的视觉特征,实现结构化、目标导向的推理,以支持通用型社会经济状态预测。实验表明,具备涌现推理能力的CityRiSE显著优于现有基线,在多样城市环境中表现更优,尤其在未见城市与未见指标上提升明显。本工作凸显了结合强化学习与LVLM在可解释、通用性城市社会经济感知中的前景。
原文摘要 · Abstract (English)
Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language Models (LVLMs), new opportunities have emerged to address this challenge by framing it as a multi-modal perception and reasoning task. However, recent studies show that LVLMs still struggle to make accurate and interpretable socio-economic predictions from visual data. To overcome these limitations and fully exploit the potential of LVLMs, we propose CityRiSE, a novel framework for Reasoning urban Socio-Economic status in LVLMs via reinforcement learning (RL). With carefully curated multi-modal dataset and verifiable reward design, our approach guides the LVLM to focus on semantically meaningful visual cues, enabling structured and goal-oriented reasoning for generalist socio-economic status prediction. Experiments demonstrate that CityRiSE, equipped with emergent reasoning, significantly outperforms existing baselines, improving both prediction accuracy and generalization across diverse urban contexts, especially on unseen cities and unseen indicators. This work highlights the promise of combining RL and LVLMs for interpretable and generalist urban socio-economic sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。