arXiv:2605.28733cs.AI2026-05

让商品图生成更懂销量,直接优化用户需求。

Utility-Aware Multimodal Contrastive Learning for Product Image Generation

论文配图:Utility-Aware Multimodal Contrastive Learning for Product Image Generation
图 1 · 摘自论文原文
  • 引入需求感知的对比学习损失,引导生成高需求图像。
  • 在亚马逊和爱彼迎上提升点击率,同时保持图文一致。
  • 适合电商、营销场景,可嵌入主流生成模型。

商品图像对在线市场中的消费者决策有显著影响。尽管多模态对比学习使生成式AI能生成与文本提示高度匹配的图像,但现有模型未直接优化市场表现。本文提出一种‘需求感知的多模态对比学习’框架,通过创新的‘需求感知InfoNCE损失’将消费者需求纳入训练目标。该目标促使图像-文本表征空间向需求驱动的视觉特征迁移,并通过理论边界验证其有效性。在亚马逊和爱彼迎的下游应用中,本方法生成或编辑的商品图像在提升需求的同时,保持了高保真度与图文一致性。尤其值得注意的是,该框架能保留美学与独特性等属性的倒U型需求曲线,在提升商业效果的同时维持内容质量。人类实验进一步验证其商业价值。随着生成式AI发展,该组件可灵活集成至新兴模型,增强其直接商业应用能力。

原文摘要 · Abstract (English)

Product images strongly influence consumer decision-making in online marketplaces. Empowered by multimodal contrastive learning, generative AI can output images that closely align with text prompts. Yet existing generative AI models do not directly optimize marketplace performance. This is a critical gap, since semantic alignment alone does not guarantee that an image will sell. To address this limitation, we propose a \textit{utility-aware multimodal contrastive learning} framework that incorporates consumer demand into a novel Utility-Aware InfoNCE loss. Optimizing this utility-aware objective guides generation toward images that are both semantically coherent and demand-enhancing. This effect arises directly from a shift in the learned image-text representation space toward demand-driven visual cues, which we also validate through the theoretical bound of the proposed objective. In downstream applications on Amazon and Airbnb, product images generated and edited by our method outperform state-of-the-art models in increasing demand and preserving fidelity, while maintaining text-image consistency. Notably, our utility-aware framework preserves inverse U-shaped demand patterns for attributes such as aesthetics and uniqueness, improving demand-based performance while preserving fidelity and semantic consistency. Human-subject experiments further validate its commercial effectiveness. As generative AI technology continues to evolve, our utility-aware component can be flexibly embedded into emerging generative models to improve direct commercial use.

图像生成多模态电商应用需求优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。