arXiv:2508.13564cs.CVcs.AI2025-08ICCV被引 26

AI城市挑战赛9年聚焦真实场景,四赛道推动视觉算法落地。

The 9th AI City Challenge

  • 四赛道并行,覆盖交通、仓储、安全等多场景视觉任务。
  • 245支队伍来自15国,数据集下载超3万次,参与度提升17%。
  • 强调轻量部署与可复现性,推动真实世界应用落地。

第九届AI城市挑战赛持续推动计算机视觉与人工智能在交通、工业自动化及公共安全领域的实际应用。2025年赛事设四个赛道,参与团队达245支,来自15个国家,较往年增长17%。挑战赛公开数据集至今下载量超3万次。第一赛道聚焦多类3D多摄像头跟踪,涵盖人、类人机器人、自主移动机器人和叉车,使用精细标定与3D边界框标注。第二赛道解决交通安全隐患的视频问答,结合3D注视标签增强多摄像头事件理解。第三赛道针对动态仓库环境中的细粒度空间推理,要求系统基于RGB-D输入回答融合感知、几何与语言的空间问题,数据集由NVIDIA Omniverse生成。第四赛道关注鱼眼相机下的高效道路目标检测,支持边缘设备的轻量化实时部署。评估框架限制提交次数,并采用部分保留测试集以确保公平性。最终排名在比赛结束后公布,促进结果可复现并减少过拟合。多个团队取得顶尖表现,在多项任务中设立新基准。

原文摘要 · Abstract (English)

The ninth AI City Challenge continues to advance real-world applications of computer vision and AI in transportation, industrial automation, and public safety. The 2025 edition featured four tracks and saw a 17% increase in participation, with 245 teams from 15 countries registered on the evaluation server. Public release of challenge datasets led to over 30,000 downloads to date. Track 1 focused on multi-class 3D multi-camera tracking, involving people, humanoids, autonomous mobile robots, and forklifts, using detailed calibration and 3D bounding box annotations. Track 2 tackled video question answering in traffic safety, with multi-camera incident understanding enriched by 3D gaze labels. Track 3 addressed fine-grained spatial reasoning in dynamic warehouse environments, requiring AI systems to interpret RGB-D inputs and answer spatial questions that combine perception, geometry, and language. Both Track 1 and Track 3 datasets were generated in NVIDIA Omniverse. Track 4 emphasized efficient road object detection from fisheye cameras, supporting lightweight, real-time deployment on edge devices. The evaluation framework enforced submission limits and used a partially held-out test set to ensure fair benchmarking. Final rankings were revealed after the competition concluded, fostering reproducibility and mitigating overfitting. Several teams achieved top-tier results, setting new benchmarks in multiple tasks.

AI城市多摄像头边缘计算视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。