arXiv:2609.08938cs.CV2026-09

让无人机视频理解更准:从光流中提取稳定空间证据,提升推理能力。

EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

论文配图:EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning
图 1 · 摘自论文原文
  • 将双向光流转化为运动归一化视觉证据,无需姿态信息
  • 在SIS-Bench上达89.9%感知、76.2%综合准确率,自知感知提升显著
  • 适配Qwen模型,不增视觉令牌数,可解释性强

无人机视频问答需区分相机运动与场景变化,但仅用RGB的多模态模型缺乏明确稳定的参考。本文提出EgoSIS,一种无姿态依赖的适配器,通过三阶段将RGB导出的双向光流转换为运动归一化视觉证据。因子化视觉自运动(FVET)拟合鲁棒图像平面过渡,揭示运动、残差支持与可靠性因子;可靠性门控自运动记忆(ReTEM)采用权重更新维护有界历史,在切变或持续不确定性处重锚定;自对齐空间证据(EASE)将支持的视觉特征映射至各片段局部锚点,并通过零初始化残差注入每视觉切片四个空间证据令牌,不改变Qwen的视觉令牌数。在SIS-Bench上,EgoSIS-8B实现89.9%感知、82.5%感知+记忆、76.2%整体准确率,最大提升集中于自知感知与记忆任务。该适配器为光流与空间推理间提供了可解释接口。

原文摘要 · Abstract (English)

UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-derived bidirectional flow into motion-canonical visual evidence in three stages. Factorized Visual Ego-Transitions (FVET) fits a robust image-plane transition and exposes motion, residual-support, and reliability factors. Reliability-Gated Ego-Transition Memory (ReTEM) uses reliability-weighted updates for a bounded history and re-anchors it at cuts or sustained uncertainty. Ego-Aligned Spatial Evidence (EASE) warps supported visual features into each segment's local anchor and injects four spatial evidence tokens per visual slice through zero-initialized residuals, without changing Qwen's visual-token count. On SIS-Bench, EgoSIS-8B obtains 89.9\% perception, 82.5\% perception-plus-memory, and 76.2\% overall accuracy, with the largest gains concentrated in self-awareness perception and memory. The adapter thus provides an interpretable interface between optical flow and spatial reasoning.

无人机视觉空间推理光流多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。