首个百万级人物交互视频数据集,解决生成难题
HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation
- 用多模态大模型自动筛选+人工精修,构建高质量视频集
- 提出混合专家描述策略,生成精准无幻觉的交互描述
- 设计细粒度评估指标,适合研究视频生成与交互建模者
文本到视频(T2V)生成在复杂场景生成方面取得显著进展,但当前模型难以精确生成人物与物体交互(HOI),主要因缺乏大规模、标注准确的HOI视频数据。为此,我们提出HOIGen-1M,首个面向HOI生成的大规模数据集,包含超过一百万条高质量视频,来源多样。为确保视频质量,我们首先设计高效框架,利用强大的多模态大语言模型(MLLM)自动筛选并预处理视频,再经人工标注进一步清洗。为获取精准的文本描述,我们提出基于多模态专家混合(MoME)策略的新视频描述方法,不仅生成丰富描述,还能有效抑制单个MLLM的幻觉问题。此外,针对生成视频缺乏评估体系的问题,我们提出两种从粗到细的评估指标。大量实验表明,现有T2V模型在生成高质量HOI视频方面表现不佳,验证了HOIGen-1M对提升该任务的关键作用。
原文摘要 · Abstract (English)
Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale videos with accurate captions for HOI. To address this issue, we introduce HOIGen-1M, the first largescale dataset for HOI Generation, consisting of over one million high-quality videos collected from diverse sources. In particular, to guarantee the high quality of videos, we first design an efficient framework to automatically curate HOI videos using the powerful multimodal large language models (MLLMs), and then the videos are further cleaned by human annotators. Moreover, to obtain accurate textual captions for HOI videos, we design a novel video description method based on a Mixture-of-Multimodal-Experts (MoME) strategy that not only generates expressive captions but also eliminates the hallucination by individual MLLM. Furthermore, due to the lack of an evaluation framework for generated HOI videos, we propose two new metrics to assess the quality of generated videos in a coarse-to-fine manner. Extensive experiments reveal that current T2V models struggle to generate high-quality HOI videos and confirm that our HOIGen-1M dataset is instrumental for improving HOI video generation. Project webpage is available at https://liuqi-creat.github.io/HOIGen.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。