利用手部轨迹提升第一人称视频中自然语言查询的定位精度
Hand Trajectory Fusion for Egocentric Natural Language Query Grounding

- 通过手部骨骼序列生成高语义手部运动特征
- 在手物交互和数量状态类查询上准确率提升4.32点
- 特别适合需要理解动作意图的视频理解任务
第一人称自然语言查询(NLQ)定位任务要求模型在长时序第一人称视频中定位回答自由文本查询的时间区间。现有方法融合视频视觉外观与查询文本,但忽略了手部运动信息。事实上,约41%的Ego4D NLQ查询的答案出现在手-物体操作或其直接结果时刻。本文提出一种手部轨迹编码器,将手部骨骼序列转化为高语义的手部运动特征,并通过带有自适应门控的交叉注意力融合策略,与预训练的视频-文本特征对齐结合。在Ego4D NLQ v2验证集上,手物交互类查询的R1@IoU=0.3提升2.54点,数量/状态类查询提升4.32点,表明手部轨迹提供了超越视觉外观的定位线索。
原文摘要 · Abstract (English)
Egocentric Natural Language Query (NLQ) grounding asks a model to localize, in a long first-person video, the temporal interval that answers a free-form text query. Existing methods fuse video appearance with the query but ignore hand motion, despite the fact that roughly 41% of Ego4D NLQ queries are answered at a moment of hand--object manipulation or their immediate outcomes.We propose a hand-trajectory encoder for converting a sequence of hand skeletons into highly-semantic hand kinematic features, which are then aligned and combined with pretrained video--text features through a cross-attention fusion strategy with adaptive gating. On the Ego4D NLQ v2 validation split, the clearest gains appear for Hand-Object Interaction queries (+2.54 R1@IoU=0.3) and Quantity/State queries (+4.32 R1@IoU=0.3), indicating that hand trajectory provides grounding cues beyond appearance alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。