arXiv:2602.19768cs.CV2026-02被引 2

让AI像人一样看图说话,能追踪视线路径并解释描述与区域的关系。

TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

  • 用轨迹信息融合视觉特征,实现双向感知
  • 在多个任务上达到当前最优效果,提升空间理解能力
  • 适合需要可解释视觉分析的场景,如医疗影像、自动驾驶

近期大型视觉语言模型在图像理解和自然语言生成方面表现优异,但现有方法主要关注全局图像理解,难以模拟人类视觉注意力轨迹,也难以解释描述与特定区域之间的关联。我们提出TraceVision,一个统一的端到端视觉语言模型,集成轨迹感知的空间理解能力。TraceVision采用轨迹感知视觉感知(TVP)模块,实现视觉特征与轨迹信息的双向融合。通过几何简化从原始轨迹中提取语义关键点,并设计三阶段训练流程,使轨迹引导描述生成与区域定位。我们将TraceVision扩展至轨迹引导分割和视频场景理解,支持跨帧跟踪与时间注意力分析。我们构建了基于推理的交互式局部叙事(RILN)数据集,以增强逻辑推理与可解释性。大量实验表明,在轨迹引导描述、文本引导轨迹预测、理解与分割等任务上,TraceVision均达到当前最优性能,为直观空间交互与可解释视觉理解奠定了基础。

原文摘要 · Abstract (English)

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, struggling to simulate human visual attention trajectories and explain associations between descriptions and specific regions. We propose TraceVision, a unified vision-language model integrating trajectory-aware spatial understanding in an end-to-end framework. TraceVision employs a Trajectory-aware Visual Perception (TVP) module for bidirectional fusion of visual features and trajectory information. We design geometric simplification to extract semantic keypoints from raw trajectories and propose a three-stage training pipeline where trajectories guide description generation and region localization. We extend TraceVision to trajectory-guided segmentation and video scene understanding, enabling cross-frame tracking and temporal attention analysis. We construct the Reasoning-based Interactive Localized Narratives (RILN) dataset to enhance logical reasoning and interpretability. Extensive experiments on trajectory-guided captioning, text-guided trajectory prediction, understanding, and segmentation demonstrate that TraceVision achieves state-of-the-art performance, establishing a foundation for intuitive spatial interaction and interpretable visual understanding.

视觉语言模型空间理解轨迹感知可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。