arXiv:2603.09286cs.CV2026-03

让AI图像生成更符合用户心理意图,可精准调控情绪与记忆效果。

CogBlender: Towards Continuous Cognitive Intervention in Text-to-Image Generation

  • 分两阶段控制图像认知属性:先生成极端状态提示,再通过速度场插值实现连续调节。
  • 在4种心理属性上均实现有效干预,可精确调整图像的情绪价值与记忆度。
  • 适合需要情感化或高记忆性图像的创意设计、心理研究等场景。

除了传递语义信息外,图像还具备引发特定心理反应的认知属性,如记忆编码或情绪反应。尽管现代文本到图像(T2I)模型能有效生成语义一致的内容,但在控制认知属性(如正负性、唤醒度、支配感和记忆性)方面仍表现不佳,难以匹配用户的心理意图。为此,我们提出CogBlender,一种通过新颖两阶段方法实现认知属性连续多维干预的算法。首先,构建离散的认知感知重写提示——即代表不同极端认知状态的输入提示变体;其次,通过在流匹配模型的速度场空间内插值,将这些离散提示转化为连续控制信号。通过根据目标认知评分动态混合来自这些提示的速度场,CogBlender能够平滑引导生成轨迹,实现最终图像所需的认知属性。在四种认知属性(正负性、唤醒度、支配感和记忆性)上的大量实验表明,CogBlender实现了有效的认知干预。

原文摘要 · Abstract (English)

Beyond conveying semantic information, images also possess cognitive properties that elicit specific psychological responses from viewers, such as memory encoding or emotional reactions. Although modern text-to-image (T2I) models generate semantically coherent content effectively, they struggle to control cognitive properties (e.g., valence, memorability) and often fail to align with the user's psychological intent. To bridge the gap, we introduce CogBlender, an algorithm that enables continuous and multi-dimensional intervention on cognitive properties through a novel two-stage approach. First, we construct discrete cognition-aware rewritten prompts-variants of the input prompt that represent distinct extreme cognitive states. Second, we translate these discrete prompts into continuous control signals by interpolating within the velocity-field domain of flow-matching models. By dynamically blending the velocity fields predicted from these prompts according to the target cognitive scores, CogBlender smoothly steers the generative trajectory to realize the desired cognitive properties in the final image. Extensive experiments across four cognitive properties (i.e., valence, arousal, dominance, and memorability) demonstrate that CogBlender achieves effective cognitive intervention.

图像生成认知控制心理意图流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。