arXiv:2511.20957cs.CV2025-11

让图像合成更艺术:从真实感转向用户表达意图

Beyond Realism: Learning the Art of Expressive Composition with StickerNet

  • 分两阶段预测贴纸位置、透明度等参数,基于真实平台编辑行为训练
  • 在180万条真实用户编辑数据上训练,表现接近人类放置水平
  • 适合关注创意表达而非真实感的视觉生成与内容创作场景

在现代内容创作中,图像合成常用于追求艺术性、趣味性或社交传播效果,而非单纯追求视觉真实。为此,本文提出表达性合成任务,强调风格多样性和灵活布局。我们构建了包含180万条真实用户编辑动作的数据集,来自匿名在线视觉创作平台,每条记录均反映社区认可的放置决策。基于此,提出StickerNet框架,先判断合成类型,再预测透明度、掩码、位置和缩放等参数。用户测试与量化评估表明,该模型性能优于主流基线,且接近人类放置行为,验证了从真实编辑模式学习的有效性。本工作开辟了以表达力和用户意图为核心的视觉理解新方向。

原文摘要 · Abstract (English)

As a widely used operation in image editing workflows, image composition has traditionally been studied with a focus on achieving visual realism and semantic plausibility. However, in practical editing scenarios of the modern content creation landscape, many compositions are not intended to preserve realism. Instead, users of online platforms motivated by gaining community recognition often aim to create content that is more artistic, playful, or socially engaging. Taking inspiration from this observation, we define the expressive composition task, a new formulation of image composition that embraces stylistic diversity and looser placement logic, reflecting how users edit images on real-world creative platforms. To address this underexplored problem, we present StickerNet, a two-stage framework that first determines the composition type, then predicts placement parameters such as opacity, mask, location, and scale accordingly. Unlike prior work that constructs datasets by simulating object placements on real images, we directly build our dataset from 1.8 million editing actions collected on an anonymous online visual creation and editing platform, each reflecting user-community validated placement decisions. This grounding in authentic editing behavior ensures strong alignment between task definition and training supervision. User studies and quantitative evaluations show that StickerNet outperforms common baselines and closely matches human placement behavior, demonstrating the effectiveness of learning from real-world editing patterns despite the inherent ambiguity of the task. This work introduces a new direction in visual understanding that emphasizes expressiveness and user intent over realism.

图像合成表达性生成用户行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。