arXiv:2607.26215cs.CV2026-07

考虑双手动作延迟,提升手势分割精度。

Lag-aware cross-hand alignment for dual-hand action segmentation

论文配图:Lag-aware cross-hand alignment for dual-hand action segmentation
图 1 · 摘自论文原文
  • 显式建模双手特征流的时间偏移分布
  • 在两个数据集上分别提升均值F1达2.1和1.9
  • 轻量设计适合实时系统部署

双手动作分割通常在相同时间点融合左右手特征,但协调动作可能存在非零且动态变化的延迟。本文提出轻量级模块LACA,显式估计双手特征流间的方向性时间偏移分布。LACA从估计偏移中检索跨手信息,并引入可学习的空状态,在无匹配过渡时抑制信息传递。对HA-ViD和ATTACH训练标注的分析显示,44.7%与48.9%的过渡锚点存在显著非零跨手匹配,而时移对照组仅为18.6%与21.3%。集成至Polyphony后,LACA使HA-ViD的两双手均值F1@50从40.4升至42.5,边界F1从56.5升至59.6;在ATTACH上分别提升至21.8与47.9,相对复现基线。仅增加约0.0086百万可训练参数。进一步提出无需未来信息的LACA-C,其在ATTACH上实现83.6%过渡提示召回率,平均可用延迟233毫秒,每分钟0.72个误报提示,分割阶段吞吐率达每秒224.9个当前位置预测。结果表明,显式跨手时间对齐能同时提升动作分割与边界定位,并支持及时的未来无关感知。

原文摘要 · Abstract (English)

Dual-hand action segmentation commonly fuses left- and right-hand representations at identical temporal indices, although coordinated hand transitions may occur with nonzero and time-varying delays. We introduce Lag-Aware Cross-Hand Alignment (LACA), a lightweight module that explicitly estimates directional temporal-offset distributions between hand-specific feature streams. LACA retrieves cross-hand information from the estimated offsets and incorporates a learned null state to suppress transfer when no compatible cross-hand transition is supported. Alignment is supervised using compatibility-aware targets derived automatically from frame-level training annotations, without requiring additional labels. Analysis of the HA-ViD and ATTACH training annotations reveals robust nonzero cross-hand matches for 44.7% and 48.9% of transition anchors, respectively, compared with 18.6% and 21.3% under temporally shifted controls. When integrated into Polyphony, LACA improves the two-hand mean F1@50 from 40.4 to 42.5 and boundary F1 from 56.5 to 59.6 on HA-ViD, and from 19.9 to 21.8 and 44.7 to 47.9, respectively, on ATTACH, relative to our reproduced Polyphony baseline. These gains require only approximately 0.0086 million additional trainable parameters. We further introduce LACA-C, a future-free variant that restricts alignment and the complete inference pipeline to current and past observations. On ATTACH, LACA-C achieves 83.6% transition-cue recall, a seed-averaged median availability delay of 233~ms, 0.72 false cues per minute, and segmentation-stage throughput of 224.9 current-position predictions per second. These results demonstrate that explicit cross-hand temporal alignment improves both action segmentation and boundary localization while supporting timely future-free perception.

动作分割双手交互时间对齐轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。