用场景图+多模态推理理解交通视频,回答复杂问题更透明。
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
- 构建交通场景图,融合符号查询与视觉输入进行推理
- 在TUMTraffic VideoQA上多类问题准确率表现优秀
- 适合需要可解释性交通视频理解的智能系统研究
我们提出基于场景图的多模态交通代理(SGTA),一个模块化框架,用于交通视频理解。该框架通过检测、跟踪和车道提取从路边视频构建交通场景图,并基于工具执行对符号图查询和视觉输入的推理。SGTA采用ReAct机制,将大语言模型的推理轨迹与工具调用交替处理,实现复杂视频问题的可解释决策。在选定的TUMTraffic VideoQA数据集样本上的实验表明,SGTA在多种问题类型上达到具有竞争力的准确率,同时提供清晰的推理步骤。结果表明,结构化场景表示与多模态代理结合在交通视频理解中具有巨大潜力。
原文摘要 · Abstract (English)
We present Scene-Graph Based Multi-Modal Traffic Agent (SGTA), a modular framework for traffic video understanding that combines structured scene graphs with multi-modal reasoning. It constructs a traffic scene graph from roadside videos using detection, tracking, and lane extraction, followed by tool-based reasoning over both symbolic graph queries and visual inputs. SGTA adopts ReAct to process interleaved reasoning traces from large language models with tool invocations, enabling interpretable decision-making for complex video questions. Experiments on selected TUMTraffic VideoQA dataset sample demonstrate that SGTA achieves competitive accuracy across multiple question types while providing transparent reasoning steps. These results highlight the potential of integrating structured scene representations with multi-modal agents for traffic video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。