用动态图变压器捕捉复杂行为,让视频描述更自然准确。
Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning
- 构建多尺度时序模块与语义感知模块,融合时空特征
- 在MSVD和MSR-VTT上显著提升描述质量
- 适合需要精细行为理解的视频生成任务
现有视频描述方法对物体行为的表征较为浅显,导致描述模糊。为此,提出一种动态动作语义感知图变压器。首先设计多尺度时序建模模块,灵活学习长短期潜在动作特征,兼顾局部细节,增强表示的一致性与敏感性;其次引入视觉-动作语义感知模块,自适应捕获与行为相关的语义信息,提升动作表征的丰富性与准确性。结合二者获得丰富的行为表征,并构建时序对象-动作图,输入图变压器以建模对象与动作间的复杂时序依赖。为避免推理阶段复杂度增加,通过知识蒸馏将行为知识压缩至轻量网络。在MSVD和MSR-VTT数据集上的实验表明,该方法在多个指标上均取得显著性能提升。
原文摘要 · Abstract (English)
Existing video captioning methods merely provide shallow or simplistic representations of object behaviors, resulting in superficial and ambiguous descriptions. However, object behavior is dynamic and complex. To comprehensively capture the essence of object behavior, we propose a dynamic action semantic-aware graph transformer. Firstly, a multi-scale temporal modeling module is designed to flexibly learn long and short-term latent action features. It not only acquires latent action features across time scales, but also considers local latent action details, enhancing the coherence and sensitiveness of latent action representations. Secondly, a visual-action semantic aware module is proposed to adaptively capture semantic representations related to object behavior, enhancing the richness and accurateness of action representations. By harnessing the collaborative efforts of these two modules,we can acquire rich behavior representations to generate human-like natural descriptions. Finally, this rich behavior representations and object representations are used to construct a temporal objects-action graph, which is fed into the graph transformer to model the complex temporal dependencies between objects and actions. To avoid adding complexity in the inference phase, the behavioral knowledge of the objects will be distilled into a simple network through knowledge distillation. The experimental results on MSVD and MSR-VTT datasets demonstrate that the proposed method achieves significant performance improvements across multiple metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。