arXiv:2608.09656cs.CV2026-08

模仿人脑分层视觉机制,提升第一视角视频中目标定位精度

EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization

论文配图:EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization
图 1 · 摘自论文原文
  • 分层设计:先提取前景候选,再通过判别性滤波精确定位
  • 在多个基准上达到当前最佳性能,3D定位依赖可信视点融合
  • 适合第一视角视频理解、智能助手等需要精准定位的场景

视觉查询定位(VQL)旨在从第一视角视频中检索并重新定位目标物体,但当物体边界模糊且全局上下文无法有效指导细粒度定位时仍具挑战。人类视觉通过分层处理应对此类模糊性:快速筛选前景候选,选择性关注目标,利用全局与局部反馈精细感知,并在单视角不可靠时依据可信度融合多视角证据。受此启发,我们提出EgoHieraLoc,一个统一的VQL-2D与VQL-3D框架。判别性分割模块利用分割先验提取前景感知的查询表征;查询感知模块通过可变形建模的判别相关滤波实现鲁棒定位;区域自适应模块将多尺度上下文反馈至局部区域以恢复精确边界。为扩展至3D定位,引入几何-语义联合置信度(GSJC),通过乘法耦合分割置信度、局部深度一致性、多视角反投影一致性及三角化基线质量,仅当视角在语义和几何上均可信时才参与3D估计。大量实验表明,该方法在VQL-2D与VQL-3D基准上均达到领先性能。

原文摘要 · Abstract (English)

Visual query localization (VQL) aims to retrieve and re-localize a queried object in egocentric videos, yet remains challenging when object boundaries are ambiguous and global context cannot effectively guide fine-grained localization. Human vision handles such ambiguity through a hierarchical process: it rapidly screens foreground candidates, selectively attends to the target despite distractors, refines perception via feedback between global context and local detail, and, when a single view is unreliable, integrates evidence across viewpoints according to its credibility. Inspired by these competencies, we propose \textbf{EgoHieraLoc}, a unified framework for VQL-2D and VQL-3D. A Discriminative Parsing Module first extracts foreground-aware query representations using segmentation priors; a Query-Aware Module then performs robust target localization through discriminative correlation filtering with deformable modeling; and a Regional Adaptation Module feeds multi-scale context back into local regions to recover precise object boundaries. To extend this perceptual hierarchy to 3D localization, we introduce Geometric-Semantic Joint Confidence (GSJC), which multiplicatively couples segmentation confidence with local depth consistency, multi-view back-projection consistency, and triangulation-baseline quality, so that a viewpoint contributes to the 3D estimate only when it is credible both semantically and geometrically. Extensive experiments demonstrate state-of-the-art performance on both VQL-2D and -3D benchmarks.

视觉定位第一视角分层建模3D定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。