让视觉模型学会生成有笑点的网络梗图,靠的是分步推理和群体偏好对齐。
From Perception to Punchline: Empowering VLM with the Art of In-the-wild Meme
- 采用分层多路径思维链,逐步推导出符合语境的幽默梗图结构。
- 在同模板梗图中进行成对奖励建模,提升幽默感判断的可靠性。
- 适用于需要主观审美判断的开放式多模态生成任务,如内容创作。
生成幽默梗图是一项挑战性的多模态任务,超越了简单的图像到标题监督。它需要对视觉内容、上下文线索和主观幽默感进行精细推理。为弥合视觉感知与幽默结尾生成之间的差距,我们提出 HUMOR 框架,通过分层推理和群体层面的人类偏好对齐来引导视觉语言模型(VLM)。首先,HUMOR 采用分层多路径思维链(CoT):模型先识别模板级意图,再在不同语境下探索多种推理路径,最后锚定高质量、上下文相关的路径。这种从真实标题回溯的 CoT 监督提升了推理多样性。我们进一步分析表明,在高质量路径保持显著概率质量的前提下,多路径探索加锚定能维持较高的预期幽默质量。其次,为捕捉主观幽默感,我们训练了一个基于组内成对比较的奖励模型,该模型在共享同一模板的梗图组上运行。依据既有理论,此方法即使在主观且含噪标签下也能提供稳定可靠的偏好代理。奖励模型随后支持组内强化学习优化,确保在信任区域内单调提升。大量实验表明,HUMOR 赋能多种 VLM,实现了更优的推理多样性、更可靠的偏好对齐以及更高的整体梗图质量。本研究还提出了一个通用训练范式,适用于以人类偏好为导向的开放式多模态生成任务,其成功依赖于同质输出组内的比较判断。
原文摘要 · Abstract (English)
Generating humorous memes is a challenging multimodal task that moves beyond direct image-to-caption supervision. It requires a nuanced reasoning over visual content, contextual cues, and subjective humor. To bridge this gap between visual perception and humorous punchline creation, we propose HUMOR}, a novel framework that guides VLMs through hierarchical reasoning and aligns them with group-wise human preferences. First, HUMOR employs a hierarchical, multi-path Chain-of-Thought (CoT): the model begins by identifying a template-level intent, then explores diverse reasoning paths under different contexts, and finally anchors onto a high-quality, context-specific path. This CoT supervision, which traces back from ground-truth captions, enhances reasoning diversity. We further analyze that this multi-path exploration with anchoring maintains a high expected humor quality, under the practical condition that high-quality paths retain significant probability mass. Second, to capture subjective humor, we train a pairwise reward model that operates within groups of memes sharing the same template. Following established theory, this approach ensures a consistent and robust proxy for human preference, even with subjective and noisy labels. The reward model then enables a group-wise reinforcement learning optimization, guaranteeing providing a theoretical guarantee for monotonic improvement within the trust region. Extensive experiments show that HUMOR empowers various VLMs with superior reasoning diversity, more reliable preference alignment, and higher overall meme quality. Beyond memes, our work presents a general training paradigm for open-ended, human-aligned multimodal generation, where success is guided by comparative judgment within coherent output group.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。