arXiv:2512.05557cs.CVcs.AI2025-12被引 1

构建首个可精准控制角色与情节一致性的大规模叙事数据集。

2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency

  • 通过人机协同流程生成2000个独特角色、1万条连贯故事。
  • 分离角色身份与姿态表情等瞬时属性,确保跨画面一致性。
  • 适合研究可控图像生成、故事化视觉合成的团队使用。

在可控视觉叙事中,如何在精确控制瞬时属性的前提下保持角色身份的一致性,仍是长期难题。现有数据集缺乏足够保真度,未能有效解耦稳定身份与瞬时属性,限制了对姿态、表情和场景构图的结构化控制,进而制约了可靠序列生成。为此,我们提出 extbf{2K-Characters-10K-Stories},一个包含 extbf{2{,}000} 个独特风格化角色、分布在 extbf{10{,}000} 幅插画故事中的多模态叙事数据集。该数据集是首个将大规模唯一身份与显式解耦控制信号相结合的资源。我们设计了 extbf{人机协同流水线(HiL)},结合专家验证的角色模板与大语言模型引导的故事规划,生成高度对齐的结构化数据。采用 extbf{解耦控制机制},将持久身份与瞬时属性(姿态、表情)分离;通过融合 MMLM 评估、自动提示调优与局部图像编辑的 extbf{质量门控循环},实现像素级一致性保障。大量实验表明,基于本数据集微调的模型,在生成视觉叙事方面表现接近闭源模型。

原文摘要 · Abstract (English)

Sequential identity consistency under precise transient attribute control remains a long-standing challenge in controllable visual storytelling. Existing datasets lack sufficient fidelity and fail to disentangle stable identities from transient attributes, limiting structured control over pose, expression, and scene composition and thus constraining reliable sequential synthesis. To address this gap, we introduce \textbf{2K-Characters-10K-Stories}, a multi-modal stylized narrative dataset of \textbf{2{,}000} uniquely stylized characters appearing across \textbf{10{,}000} illustration stories. It is the first dataset that pairs large-scale unique identities with explicit, decoupled control signals for sequential identity consistency. We introduce a \textbf{Human-in-the-Loop pipeline (HiL)} that leverages expert-verified character templates and LLM-guided narrative planning to generate highly-aligned structured data. A \textbf{decoupled control} scheme separates persistent identity from transient attributes -- pose and expression -- while a \textbf{Quality-Gated loop} integrating MMLM evaluation, Auto-Prompt Tuning, and Local Image Editing enforces pixel-level consistency. Extensive experiments demonstrate that models fine-tuned on our dataset achieves performance comparable to closed-source models in generating visual narratives.

叙事生成角色一致性数据集可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。