用图约束推理实现无训练的视角动作流在线解析
VidParse: Online Parsing of Egocentric Procedures Like a Pro

- 基于冻结模型提取手物交互特征,动态构建时序相似矩阵识别语义转换
- 通过束搜索解码器在程序图上约束有效动作跳转,避免错误分割和结构坍塌
- 无需训练即可实现复杂多步任务10倍精度提升,适合真实场景动作理解
将连续、嘈杂的视角视频流转化为离散、时序有序的动作步骤面临巨大视觉挑战。强烈的自我运动、瞬时遮挡以及非脚本化人-物交互的高类内差异,导致标准帧级在线时序模型表现不佳,常出现严重过分割和结构坍塌。为弥合不稳定底层感知与高层程序逻辑之间的差距,我们提出VidParse,一种无需训练的在线框架,将活动理解建模为图约束推理问题。不依赖学习的时序滤波器,而是通过操纵锚定特征上的时序相似矩阵动态识别语义转换,这些特征来自冻结的基础模型,优先关注前景手-物交互。束搜索解码器利用生成的程序任务图显式强制合法动作转移并剔除不可能轨迹。通过将鲁棒视觉片段锚定于严格的程序约束,该方法保持了长程状态转移,相比强在线基线,在复杂多步解析中准确率最高提升10倍,且无需任何梯度更新。
原文摘要 · Abstract (English)
Translating continuous, noisy egocentric video streams into discrete, temporally ordered action steps is fraught with visual challenges. Heavy ego-motion, transient occlusions, and the high intra-class variability of unscripted human-object interactions cause standard frame-level online temporal models to struggle, often resulting in severe over-segmentation and structural collapse. To bridge the gap between unstable low-level perception and high-level procedural logic, we present VidParse, an online, training-free framework that treats activity understanding as a graph-constrained inference problem. Rather than relying on learned temporal filters, we dynamically identify semantic transitions using a temporal similarity matrix over manipulation-anchored features, which are extracted from frozen foundation models to prioritize foreground hand-object interactions. A beam search decoder then leverages an induced procedural task graph to explicitly enforce valid action transitions and prune impossible trajectories. By anchoring robust visual segments to hard procedural constraints, our approach preserves long-range state transitions and achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines, all without requiring a single gradient update.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。