arXiv:2604.10454cs.CV2026-04被引 1

首个针对情感图像编辑的细粒度评测基准,解决现有方法情绪偏差问题。

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control

  • 构建双路径情感建模框架,融合情绪分类与连续维度控制。
  • 收集800张高质量样本,覆盖8类情绪和5种编辑类型。
  • 提出反向重绘数据引擎,生成4万条平衡指令数据提升模型表现。

情感图像编辑(AIM)旨在通过精准修改激发特定情绪。现有图像编辑基准多聚焦于通用场景下的对象级修改,缺乏捕捉情感维度的细粒度能力。为此,我们提出首个专为AIM设计的基准——AIM-Bench。该基准基于双路径情感建模方案,结合Mikels情绪分类体系与效价-唤醒-支配(Valence-Arousal-Dominance)框架,实现高层语义与细粒度连续控制。通过分层人机协同流程,最终构建包含800个高质量样本的数据集,覆盖8类情绪和5种编辑类型。为有效评估性能,我们设计了融合规则与模型的综合评价体系,全面衡量指令一致性、美学质量与情感表达力。大量实验表明,当前编辑模型存在显著正向偏差,源于训练数据分布固有的不平衡。为此,我们提出可扩展的数据引擎,采用逆向重绘策略构建平衡的指令微调数据集AIM-40k(共4万样本)。具体而言,通过生成式重绘建立高保真真实标签,并合成具有差异情绪及精确指令的输入图像。在基准模型上微调后,整体性能相对提升9.15%,验证了AIM-40k的有效性。相关数据与代码将尽快开源。

原文摘要 · Abstract (English)

Affective Image Manipulation (AIM) aims to evoke specific emotions through targeted editing. Current image editing benchmarks primarily focus on object-level modifications in general scenarios, lacking the fine-grained granularity to capture affective dimensions. To bridge this gap, we introduce the first benchmark designed for AIM termed AIM-Bench. This benchmark is built upon a dual-path affective modeling scheme that integrates the Mikels emotion taxonomy with the Valence-Arousal-Dominance framework, enabling high-level semantic and fine-grained continuous manipulation. Through a hierarchical human-in-the-loop workflow, we finally curate 800 high-quality samples covering 8 emotional categories and 5 editing types. To effectively assess performance, we also design a composite evaluation suite combining rule-based and model-based metrics to holistically assess instruction consistency, aesthetics, and emotional expressiveness. Extensive evaluations reveal that current editing models face significant challenges, most notably a prevalent positivity bias, which stemming from inherent imbalances in training data distribution. To tackle this, we propose a scalable data engine utilizing an inverse repainting strategy to construct AIM-40k, a balanced instruction-tuning dataset comprising 40k samples. Concretely, we enhance raw affective images via generative redrawing to establish high-fidelity ground truths, and synthesize input images with divergent emotions and paired precise instructions. Fine-tuning a baseline model on AIM-40k yields a 9.15% relative improvement in overall performance, demonstrating the effectiveness of our AIM-40k. Our data and related code will be made open soon.

情感编辑图像生成数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。