用少量参考图生成风格一致的无限角色,支持创意设计扩展。
Few-shot multi-token DreamBooth with LoRa for style-consistent character generation
- 用多标记+LoRA实现小样本角色风格迁移
- 在5个数据集上生成高质量多样角色,保留原风格特征
- 适合动画、游戏领域快速角色创作,无需大量训练数据
音视频行业正经历深刻变革,AI不仅用于自动化任务,更推动新艺术形式诞生。本文解决如何基于少量人工设计的角色参考,生成数量近乎无限、且保持统一艺术风格与视觉共性的新角色,从而拓展动画、游戏等领域的创作可能。方法基于成熟的文本到图像扩散模型微调技术DreamBooth,针对两个核心挑战:捕捉超出文本提示的复杂角色细节,以及训练数据的少样本特性。提出多标记策略,通过聚类为每个角色及其整体风格分配独立标记,并结合LoRA实现参数高效微调。通过移除类别特定正则化项,并在生成时引入随机标记与嵌入,实现无限角色生成的同时保持学习到的风格一致性。我们在五个小型专用数据集上评估该方法,对比多个基线,采用定量指标与人类评估相结合。结果表明,本方法能生成高质量、多样化角色,同时忠实保留参考角色的独特美学特征,人类评估进一步验证其有效性,凸显该方法潜力。
原文摘要 · Abstract (English)
The audiovisual industry is undergoing a profound transformation as it is integrating AI developments not only to automate routine tasks but also to inspire new forms of art. This paper addresses the problem of producing a virtually unlimited number of novel characters that preserve the artistic style and shared visual traits of a small set of human-designed reference characters, thus broadening creative possibilities in animation, gaming, and related domains. Our solution builds upon DreamBooth, a well-established fine-tuning technique for text-to-image diffusion models, and adapts it to tackle two core challenges: capturing intricate character details beyond textual prompts and the few-shot nature of the training data. To achieve this, we propose a multi-token strategy, using clustering to assign separate tokens to individual characters and their collective style, combined with LoRA-based parameter-efficient fine-tuning. By removing the class-specific regularization set and introducing random tokens and embeddings during generation, our approach allows for unlimited character creation while preserving the learned style. We evaluate our method on five small specialized datasets, comparing it to relevant baselines using both quantitative metrics and a human evaluation study. Our results demonstrate that our approach produces high-quality, diverse characters while preserving the distinctive aesthetic features of the reference characters, with human evaluation further reinforcing its effectiveness and highlighting the potential of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。