arXiv:2512.19300cs.CV2025-12AAAI被引 1

用强化学习让不同类别的文字概念融合生成新图像。

RMLer: Synthesizing Novel Objects across Diverse Categories via Reinforcement Mixing Learning

  • 将跨类别概念融合建模为强化学习问题,动态调整文本特征混合系数。
  • 在多个数据集上生成的图像更连贯、视觉质量更高,优于现有方法。
  • 适合影视、游戏和设计领域需要创新视觉概念的场景。

通过整合来自不同类别的文本概念进行新型物体合成,仍是文本到图像(T2I)生成中的重大挑战。现有方法常存在概念混合不足、评估不严谨及输出质量不佳的问题,表现为概念失衡、表面组合或简单拼接。为此,我们提出强化混合学习(RMLer),将跨类别概念融合建模为强化学习问题:混合特征作为状态,混合策略作为动作,视觉结果作为奖励。具体地,设计一个MLP策略网络,预测跨类别文本嵌入的动态混合系数;进一步引入基于语义相似性和融合对象与其组成概念之间构图平衡的视觉奖励,通过近端策略优化(PPO)训练策略。推理时,利用这些奖励选择最优融合结果。大量实验表明,RMLer在从多样类别中合成连贯、高保真图像方面优于现有方法。本工作为生成新颖视觉概念提供了稳健框架,具有在影视、游戏和设计领域的广阔应用前景。

原文摘要 · Abstract (English)

Novel object synthesis by integrating distinct textual concepts from diverse categories remains a significant challenge in Text-to-Image (T2I) generation. Existing methods often suffer from insufficient concept mixing, lack of rigorous evaluation, and suboptimal outputs-manifesting as conceptual imbalance, superficial combinations, or mere juxtapositions. To address these limitations, we propose Reinforcement Mixing Learning (RMLer), a framework that formulates cross-category concept fusion as a reinforcement learning problem: mixed features serve as states, mixing strategies as actions, and visual outcomes as rewards. Specifically, we design an MLP-policy network to predict dynamic coefficients for blending cross-category text embeddings. We further introduce visual rewards based on (1) semantic similarity and (2) compositional balance between the fused object and its constituent concepts, optimizing the policy via proximal policy optimization. At inference, a selection strategy leverages these rewards to curate the highest-quality fused objects. Extensive experiments demonstrate RMLer's superiority in synthesizing coherent, high-fidelity objects from diverse categories, outperforming existing methods. Our work provides a robust framework for generating novel visual concepts, with promising applications in film, gaming, and design.

图像生成强化学习文本融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。