提升文生图模型多样性,避免生成结果偏倚。
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

- 设计多轴max@K强化学习机制,按类别分配奖励
- 在三个评估器上公平性得分提升0.23-0.36
- 适合关注生成多样性与公平性的研究者
文生图(T2I)模型虽能生成符合提示的逼真图像,但相同提示下生成的样本常局限于少数视觉模式,限制了多样性,尤其在人物相关提示中可能放大人口分布偏倚。本文将该问题形式化为预定义语义模式的覆盖问题,称为目标模式覆盖。提出多轴max@K方法,一种基于组的强化学习目标,用于改进扩散模型的覆盖率。给定一组样本和每个目标类别的评分,多轴max@K首先对每个类别取样本间最大评分,再求和这些类别最大值。该信用分配机制仅当某样本使类别组最大值提升时才给予正权重,使不同样本可贡献于不同类别。我们在合成混合数据及SD3.5-M上验证了该机制,使用确定性像素级颜色奖励。进一步在感知外观公平性上评估,跨三个自动评估器,在保留图像质量和文本一致性的前提下,多轴max@K相较基线模型公平性得分提升0.23–0.36。
原文摘要 · Abstract (English)
Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits the diversity of images, and for person-centric prompts, can reflect or amplify demographic skew. We formalize this problem as coverage of a predefined set of semantically specified modes, which we call target-mode coverage. We then propose multi-axis max@K, a group-based reinforcement learning objective for improving such coverage in diffusion-based T2I models. Given a group of samples and one score per target category, multi-axis max@K first takes the maximum score across samples for each category and then sums these category-wise maxima. The resulting credit assignment gives a sample positive weight on a category only when it increases that category's group-wise maximum, allowing different samples to contribute to different categories. We first validate the credit-assignment mechanism on a synthetic mixture and on SD3.5-M using deterministic pixel-based color rewards. We then evaluate the same objective on perceived-appearance fairness. Across three automatic evaluators on held-out prompts, multi-axis max@K improves the Fairness Score by 0.23-0.36 relative to the base model, while maintaining image quality and text alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。