基于指令生成第一人称视角下物体状态变化的中间帧。
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos

- 用多步推理模型理解初始与目标状态间的变换过程。
- 生成符合指令且物体外观一致的连贯过渡帧序列。
- 适合研究动作建模与人机交互的学者与开发者。
理解物理变换过程对人类认知和人工智能系统至关重要,尤其从第一人称视角出发,是连接人类与机器在动作建模中的关键桥梁。我们将其定义为第一人称指令视觉状态转换(Egocentric Instructed Visual State Transition, EIVST),即在简短动作指令下生成描述物体从初始状态到目标状态之间变换的中间帧。当前生成模型面临两大挑战:(1) 从第一人称视角理解初始与目标状态的视觉场景,并推理出变换步骤;(2) 生成遵循指令且保持物体外观一致的连贯过渡过程。为此,我们提出EgoIn框架。首先利用在自建数据集上微调的TransitionVLM模型,推断两状态间的多步转换过程,以减少幻觉信息。随后,通过提出的过渡条件模块生成一系列满足转换条件的帧。此外,引入物体感知辅助监督机制,以保持整个转换过程中物体外观的一致性。在人-物及机器人-物交互数据集上的大量实验表明,EgoIn在生成语义合理、视觉连贯的变换序列方面表现卓越。
原文摘要 · Abstract (English)
Understanding physical transformation processes is crucial for both human cognition and artificial intelligence systems, particularly from an egocentric perspective, which serves as a key bridge between humans and machines in action modeling. We define this modeling process as Egocentric Instructed Visual State Transition (EIVST), which involves generating intermediate frames that depict object transformations between initial and target states under a brief action instruction. EIVST poses two challenges for current generative models: (1) understanding the visual scenes of the initial and target states and reasoning about transformation steps from an egocentric view, and (2) generating a consistent intermediate transition that follows the given instruction while preserving object appearance across the two visual states. To address these challenges, we propose the EgoIn framework. It first infers the multi-step transition process between two given states using TransitionVLM, fine-tuned on our curated dataset to better adapt to this task and reduce hallucinated information. It then generates a sequence of frames based on transition conditions produced by the proposed Transition Conditioning module. Additionally, we introduce Object-aware Auxiliary Supervision to preserve consistent object appearance throughout the transition. Extensive experiments on human-object and robot-object interaction datasets demonstrate EgoIn's superior performance in generating semantically meaningful and visually coherent transformation sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。