arXiv:2412.04189cs.CV2024-12被引 7

专注手部动作的图文视频生成,提升复杂手势的清晰度。

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation

  • 根据文本和视觉上下文自动定位手部运动区域,无需人工标注。
  • 引入手部精细化损失,显著改善手部姿态的流畅性和一致性。
  • 在EpicKitchens和Ego4D数据集上效果领先,适合真实场景动作生成。

尽管视频生成技术取得进展,当前主流方法仍难以处理视觉细节。尤其在手部动作复杂、背景稳定且易分散注意力的视频中,准确呈现复杂动作及其效果尤为困难。为此,我们提出一种以手部为中心的视频生成新方法。基于扩散模型,提出两项创新:第一,设计自动方法,结合视觉上下文与动作文本提示,生成手部活动区域,避免依赖人工标注;第二,引入关键的手部精细化损失,引导模型关注手部姿态的平滑与一致。我们在基于EpicKitchens和Ego4D的增强数据集上进行评估,结果表明,在多种环境和动作下,本方法在动作清晰度,尤其是目标区域内手部运动表现方面,显著优于现有最先进方法。视频演示见https://excitedbutter.github.io/project_page。

原文摘要 · Abstract (English)

Despite the recent strides in video generation, state-of-the-art methods still struggle with elements of visual detail. One particularly challenging case is the class of videos in which the intricate motion of the hand coupled with a mostly stable and otherwise distracting environment is necessary to convey the execution of some complex action and its effects. To address these challenges, we introduce a new method for video generation that focuses on hand-centric actions. Our diffusion-based method incorporates two distinct innovations. First, we propose an automatic method to generate the motion area -- the region in the video in which the detailed activities occur -- guided by both the visual context and the action text prompt, rather than assuming this region can be provided manually as is now commonplace. Second, we introduce a critical Hand Refinement Loss to guide the diffusion model to focus on smooth and consistent hand poses. We evaluate our method on challenging augmented datasets based on EpicKitchens and Ego4D, demonstrating significant improvements over state-of-the-art methods in terms of action clarity, especially of the hand motion in the target region, across diverse environments and actions. Video results can be found in https://excitedbutter.github.io/project_page

视频生成扩散模型手部追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。