用信息瓶颈原理自动挖掘跨模态幽默背后的有用知识
BottleHumor: Self-Informed Humor Explanation using the Information Bottleneck Principle
- 基于信息瓶颈原理,从多模态模型中迭代提取相关世界知识
- 在三个数据集上优于多个基线方法,无需标注数据
- 适合研究幽默理解、多模态推理及知识引导生成的学者
幽默广泛存在于在线交流中,常依赖多种模态(如漫画和表情包)。在多模态情境下解释幽默需要调用多元知识,包括隐喻、社会文化及常识。但如何识别最有用的知识仍是未解问题。本文提出\method{},一种受信息瓶颈原理启发的方法,可从视觉与语言模型中无监督地提取并迭代优化相关世界知识,用于生成幽默解释。在三个数据集上的实验验证了该方法优于多种基线。未来可拓展至其他需利用相关知识的任务,开辟新研究方向。
原文摘要 · Abstract (English)
Humor is prevalent in online communications and it often relies on more than one modality (e.g., cartoons and memes). Interpreting humor in multimodal settings requires drawing on diverse types of knowledge, including metaphorical, sociocultural, and commonsense knowledge. However, identifying the most useful knowledge remains an open question. We introduce \method{}, a method inspired by the information bottleneck principle that elicits relevant world knowledge from vision and language models which is iteratively refined for generating an explanation of the humor in an unsupervised manner. Our experiments on three datasets confirm the advantage of our method over a range of baselines. Our method can further be adapted in the future for additional tasks that can benefit from eliciting and conditioning on relevant world knowledge and open new research avenues in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。