无需激光雷达,用视觉实现道路设施的3D跟踪与语言推理。
RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning

- 纯图像方法结合掩码一致性和标定信息,实现跨视角身份关联
- 在6个路口20段视频上达成66.9的MOTA指标
- 构建了含3.3万条问答的数据集,支持未来物体与交互推理评估
我们提出RISE框架,融合度量3D跟踪与结构化视觉-语言推理,用于路边序列理解。在3D跟踪方面,该图像驱动方法结合SAM3视频身份与校准引导的掩码一致性,实现无激光雷达、无需特定3D训练的持续3D轨迹重建;其校准依赖几何特性使系统可在不同标定多摄像头路口部署而无需布局重训。在六处路口共20段人工审核视频中,生成轨迹在多视角评估范围内达到66.9 MOTA。在结构化视觉-语言推理方面,通过人工审核的多模态大模型流水线提取高价值片段,并使用受限全上下文Oracle构建基于边界框的预测问答任务,不暴露未来信息。由此产生的RISE-VQA数据集包含557段视频、16个路口、61个视角下的33,910对问答。其交叉路口保留的RISE-Bench通过确定性任务指标评估语义选择、坐标、未来边界框及交互集合。实验表明领域自适应和时序上下文普遍带来提升,但空间定位、未来定位与交互推理仍存在显著挑战。
原文摘要 · Abstract (English)
We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method combines SAM3 video identities with calibration-guided mask agreement for multi-view identity association, recovering persistent 3D tracks without LiDAR or task-specific 3D training. Its calibration-conditioned geometry allows the procedure to be instantiated at different calibrated multi-camera intersections without layout-specific retraining. On 20 human-reviewed clips from six intersections, the generated tracks achieve 66.9 MOTA within the defined multi-view evaluation scope. For structured vision-language reasoning, a human-reviewed MLLM pipeline mines high-value clips and uses a constrained full-context Oracle to construct bbox-grounded predictive QA without exposing future evidence to evaluated models. The resulting RISE-VQA dataset contains 33,910 QA pairs from 557 clips across 16 intersections and 61 roadside views. Its intersection-held-out RISE-Bench evaluates semantic choices, coordinates, future boxes, and interaction sets with deterministic task-specific metrics. Experiments show consistent benefits from domain adaptation and generally from temporal context, while revealing persistent challenges in spatial grounding, future localization, and interaction reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。