arXiv:2607.23920cs.CV2026-07

让图像自己告诉你能怎么改情绪,更自然真实。

What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape

论文配图:What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape
图 1 · 摘自论文原文
  • 先分析图像能支持什么情绪表达,再制定专属编辑策略
  • 用户偏好度达88.1%,显著优于现有方法
  • 支持交互式调整策略,适合需要精细控制的创作场景

情感图像编辑不仅需应用情感滤镜或修改预设视觉因子,更需识别特定图像所能承载的目标情绪。现有方法多依赖预定义因素分类、知识库或通用编辑模板,在限定策略空间内操作,常忽略图像特异性与上下文关联策略。本文提出EmoScope,一种多智能体框架,将任务从“如何编辑”重构为“能编辑什么”。EmoScope通过情感条件下的可操作性推理,发现图像特有的可编辑空间,并利用语义层级的锚点、变量与上下文,在内容一致性与情绪表现力间取得平衡,完成并验证编辑。其计划以图像特异的可操作性表达,而非检索模板,因此可在计划层提供用户交互式优化界面。在覆盖8个Mikels情感类别的大规模人类评估中(共4,693个有效响应,1,824对比较),参与者平均88.1%偏好EmoScope胜过两个基线。属性分析显示,EmoScope选择适应目标情绪的策略,而非统一模板。同一可操作性计划亦支持轻量级交互式用户修正。最后,我们揭示分类器指标在非刻板、情境化编辑上存在情感条件盲区,并构建相对内容-情绪偏好亲和力图谱,表明EmoScope优势在不同图像-情绪组合中系统性变化。

原文摘要 · Abstract (English)

Emotional image editing requires more than applying affective filters or modifying predefined visual factors: an effective edit must identify what a particular image can afford for a target emotion. Existing affective image manipulation methods, including recent agentic variants, largely operate within bounded strategy spaces based on predefined factor taxonomies, knowledge libraries, or conventional editing templates, and therefore often miss image-specific, context-grounded strategies. We introduce EmoScope, a multi-agent framework that reframes the task from "how should I edit?" to "what can I edit?" EmoScope first discovers an image-specific editable space through emotion-conditioned affordance reasoning, then uses a semantic hierarchy of anchors, variables, and context to balance content consistency and emotional expressiveness before executing and verifying the edit. Because its plans are expressed as image-specific affordances rather than retrieved templates, EmoScope also exposes the editing strategy as an interactive surface for user refinement at the plan level. In a large-scale human evaluation covering all eight Mikels emotion categories, with 4,693 valid responses across 1,824 pairwise questions, participants preferred EmoScope over two competitive baselines by 88.1% on average. Attribution analysis further shows that EmoScope selects target-emotion-adaptive strategies rather than applying a uniform template. The same affordance-level plan also supports lightweight user refinement in an interactive pilot. Finally, we show that classifier-based metrics exhibit emotion-conditional blind spots toward non-stereotypical, context-grounded edits, and present a relative content-emotion preference-affinity landscape showing that EmoScope's advantage varies systematically across image-emotion combinations.

情感编辑多智能体交互设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。