用强化学习训练城市AI,减少地域偏见,提升跨区域通用能力。
Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence
- 基于强化学习优化多模态模型的城市推理能力
- 在多个地区测试中显著降低地理偏见,提升跨区域泛化表现
- 适合关注城市智能公平性与可信赖AI的研究者
快速城市化推动了城市通用智能(UGI)的需求,即能够理解并推理复杂城市环境的AI系统。现有研究通过监督微调(SFT)构建城市基础模型,但这些模型存在持续的地理偏见,导致预测结果区域性失衡且泛化能力有限。为此,我们提出Urban-R1,一种基于强化学习的后训练框架,用于对齐多模态大模型(MLLMs)与UGI目标。Urban-R1采用组相对策略优化(GRPO),在不同地理群体间优化推理过程,并利用城市区域画像作为代理任务,从多模态城市数据中获取可度量奖励信号。在多种区域和任务上的广泛实验表明,Urban-R1有效缓解地理偏见,显著提升跨区域泛化性能,优于传统SFT训练及闭源模型。结果表明,强化学习对齐是实现公平可信城市智能的可行路径。
原文摘要 · Abstract (English)
Rapid urbanization intensifies the demand for Urban General Intelligence (UGI), referring to AI systems that can understand and reason about complex urban environments. Recent studies have built urban foundation models using supervised fine-tuning (SFT) of LLMs and MLLMs, yet these models exhibit persistent geospatial bias, producing regionally skewed predictions and limited generalization. To this end, we propose Urban-R1, a reinforcement learning-based post-training framework that aligns MLLMs with the objectives of UGI. Urban-R1 adopts Group Relative Policy Optimization (GRPO) to optimize reasoning across geographic groups and employs urban region profiling as a proxy task to provide measurable rewards from multimodal urban data. Extensive experiments across diverse regions and tasks show that Urban-R1 effectively mitigates geo-bias and improves cross-region generalization, outperforming both SFT-trained and closed-source models. Our results highlight reinforcement learning alignment as a promising pathway toward equitable and trustworthy urban intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。