arXiv:2604.01864cs.CV2026-04中稿 · AMME 2025

让AI画画更符合人类审美且能应对模糊指令

MAR-MAER: Metric-Aware and Ambiguity-Adaptive Autoregressive Image Generation

  • 用质量指标引导模型学习,让生成图像更贴近人眼偏好
  • 在模糊提示下生成更多样且连贯的图像,提升创意多样性
  • 适合需要高质量、多风格图像生成的研究者与设计师

自回归模型在文生图任务中表现优异,但常面临图像质量不达预期、难以处理模糊提示的问题。为此,我们提出MAR-MAER,一种分层自回归框架,包含两个核心组件:度量感知嵌入正则化方法和概率潜在建模机制。前者通过轻量级投影头与自适应核回归损失函数,使模型内部表征对齐人类偏好的质量指标(如CLIPScore和HPSv2);后者引入条件变分模块,在层级令牌生成中加入可控随机性,增强对模糊或开放性提示的语义灵活性。在COCO和新构建的模糊提示基准测试中,MAR-MAER在指标一致性与语义灵活性上均优于基线Hi-MAR模型,分别提升+1.6(CLIPScore)和+5.3(HPSv2)。对于模糊输入,生成输出范围显著更广。人工评估与自动指标结果一致验证了其有效性。

原文摘要 · Abstract (English)

Autoregressive (AR) models have demonstrated significant success in the realm of text-to-image generation. However, they usually face two major challenges. Firstly, the generated images may not always meet the quality standards expected by humans. Furthermore, these models face difficulty when dealing with ambiguous prompts that could be interpreted in several valid ways. To address these issues, we introduce MAR-MAER, an innovative hierarchical autoregressive framework. It combines two main components. It is a metric-aware embedding regularization method. The other one is a probabilistic latent model used for handling ambiguous semantics. Our method utilizes a lightweight projection head, which is trained with an adaptive kernel regression loss function. This aligns the model's internal representations with human-preferred quality metrics, such as CLIPScore and HPSv2. As a result, the embedding space that is learned more accurately reflects human judgment. We are also introducing a conditional variational module. This approach incorporates an aspect of controlled randomness within the hierarchical token generation process. This capability allows the model to produce a diverse array of coherent images based on ambiguous or open-ended prompts. We conducted extensive experiments using COCO and a newly developed Ambiguous-Prompt Benchmark. The results show that MAR-MAER achieves excellent performance in both metric consistency and semantic flexibility. It exceeds the baseline Hi-MAR model's performance, showing an improvement of +1.6 in CLIPScore and +5.3 in HPSv2. For unclear inputs, it produces a notably wider range of outputs. These findings have been confirmed through both human evaluation and automated metrics.

图像生成自回归模糊提示质量优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。