用检索代替训练,让图像生成自动匹配个性化风格需求。
Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs

- 从6500个模型和7.5万个适配器中检索并组合最相关组件。
- 零训练实现可控、对齐的个性化图像生成,支持百万级风格需求。
- 适合需要快速响应多样化风格指令的开发者与创作者。
用户日益期望图像生成模型能快速适应高度多样且个性化的生成需求,如特定风格或特征的图像。传统方法依赖微调,成本高且难以扩展。社区已积累大量针对特定需求的微调模块与适配器,形成可复用的资源池。本文提出Polaris,一个智能检索框架,可根据用户指令自动从模型库中选取并集成合适的组件,无需额外训练。其核心在于:从超过6500个检查点和75000个适配器中高效检索最相关模块,并有效对齐以实现指令驱动的生成与编辑。该方法实现了可扩展、可控制、精准对齐的个性化图像生成,全面释放现有模型生态潜力。
原文摘要 · Abstract (English)
Users increasingly expect image generation models to quickly adapt to highly diverse and personalized requirements, such as producing images with distinctive styles or characteristics. Traditional approaches rely on fine-tuning, which is costly and difficult to scale. To cope with these limitations, the community has accumulated a growing library of fine-tuned modules and adapters, where each component targets specific generation needs and collectively serves as a foundation for handling new demands. This naturally raises a question: instead of repeatedly training new models, can we systematically exploit this expanding ecosystem to better fulfill user instructions? To this end, we present Polaris, an intelligent retrieval framework that automatically selects and integrates suitable models from the model library based on a user's instructions. The key insight is that harnessing such a massive and heterogeneous pool requires not only finding the most relevant modules among thousands of candidates, but also aligning them effectively for instruction-driven generation and editing. Polaris addresses this challenge by indexing over 6,500 checkpoints and 75,000 adapters, and retrieving the most relevant components given a user's input and instruction. In doing so, it delivers scalable, controllable, and well-aligned generation -- without any additional training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。