用语音+文字联合提示,提升大模型理解笑话能力
Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor
- 给大模型同时提供笑话文本和语音,利用语音节奏辅助理解
- 多模态提示下,幽默解释准确率显著优于纯文本提示
- 适合研究多模态理解、幽默生成或人机交互的学者
尽管大语言模型在各类文本任务中表现出色,但在理解幽默方面仍面临挑战。幽默常具多模态特性,依赖语音中的语调歧义、节奏与时机传递含义。本文提出一种简单有效的多模态提示方法:将笑话的文本与通过现成文本转语音(TTS)系统生成的语音一并输入大模型。实验表明,在所有测试数据集上,结合文本与语音的多模态提示,相比仅使用文本提示,能显著提升模型对幽默的解释能力。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have demonstrated impressive natural language understanding capabilities across various text-based tasks, understanding humor has remained a persistent challenge. Humor is frequently multimodal, relying on phonetic ambiguity, rhythm and timing to convey meaning. In this study, we explore a simple multimodal prompting approach to humor understanding and explanation. We present an LLM with both the text and the spoken form of a joke, generated using an off-the-shelf text-to-speech (TTS) system. Using multimodal cues improves the explanations of humor compared to textual prompts across all tested datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。