arXiv:2604.18091cs.CLcs.CV2026-04

让AI根据文化背景生成贴切又搞笑的图文说明。

Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts

论文配图:Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
图 1 · 摘自论文原文
  • 分阶段训练模型,先学西方文化再适配东方文化。
  • 在文化适配上提升明显,幽默与图像相关性更平衡。
  • 新评测框架涵盖六维度,更全面评估生成质量。

近期多模态大模型在生成图像幽默标题方面展现潜力,但仍难以稳定控制显性文化语境,导致在特定文化背景下难以同时保证图像相关性、语境恰当性和幽默质量。为此,我们提出一种新任务——文化感知幽默生成,要求模型基于输入图像和目标文化背景生成幽默标题。不同文化下的标题不应形式相同,但需基于相似视觉情境或幽默逻辑。为支持该任务,我们建立六维评估框架,涵盖图像相关性、语境契合度、语义丰富度、合理性、幽默感和创意性。我们进一步提出分阶段对齐框架:首先在西方文化下用高资源监督初始化模型,接着通过基于判官的GRPO进行多维度偏好对齐,并引入降级感知原型排斥约束以缓解开放生成中的奖励滥用问题;最后仅用少量标注数据将模型适配至东方文化。实验表明,该方法在新评估框架下整体表现更强,尤其在语境契合度上提升显著,且在文化约束下实现了图像相关性与幽默性的更好平衡。

原文摘要 · Abstract (English)

Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over explicit cultural context, making it difficult to jointly maintain image relevance, contextual appropriateness, and humor quality under a specified cultural background. To address this limitation, we introduce a new multimodal generation task, culture-aware humorous captioning, which requires a model to generate a humorous caption conditioned on both an input image and a target cultural context. Captions generated under different cultural contexts are not expected to share the same surface form, but should remain grounded in similar visual situations or humorous rationales.To support this task, we establish a six-dimensional evaluation framework covering image relevance, contextual fit, semantic richness, reasonableness, humor, and creativity. We further propose a staged alignment framework that first initializes the model with high-resource supervision under the Western cultural context, then performs multi-dimensional preference alignment via judge-based GRPO with a Degradation-aware Prototype Repulsion Constraint to mitigate reward hacking in open-ended generation, and finally adapts the model to the Eastern cultural context with a small amount of supervision. Experimental results show that our method achieves stronger overall performance under the proposed evaluation framework, with particularly large gains in contextual fit and a better balance between image relevance and humor under cultural constraints.

多模态幽默生成文化适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。