构建首个中文细粒度多模态反讽数据集,提升AI理解反讽的精准度。
CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark

- 设计三层次标注:识别反讽、定位对象、生成解释。
- 含2796组图文对,发现解释标注能有效指导AI生成反讽图像。
- 提出强化学习优化例证选择方法,适合反讽理解研究者使用。
多模态反讽检测近年受到关注,但现有基准存在标注粗粒度和文化覆盖有限的问题,制约了细粒度语义理解研究。为此,我们构建了面向中文社交媒体的首个细粒度多模态反讽数据集CFMS,包含2,796个高质量图文对,并采用三层次标注框架:反讽识别、目标识别与解释生成。实验表明,细粒度解释标注能有效引导AI生成具有明确反讽意图的图像。此外,我们构建了高一致性中英隐喻平行子集(各200条),揭示当前模型在隐喻推理方面存在显著局限。为突破传统检索方法限制,我们提出基于强化学习的上下文学习策略(PGDS),实现例证动态优化。大量实验表明,CFMS为构建可靠的多模态反讽理解系统提供了坚实基础,且PGDS方法在关键任务上显著优于现有基线。数据与代码已公开于https://anonymous.4open.science/r/CFMS-E8F9。
原文摘要 · Abstract (English)
Multimodal sarcasm detection has recently garnered significant attention. However, existing benchmarks suffer from coarse-grained annotations and limited cultural coverage, which hinder research into fine-grained semantic understanding. To address this, we construct CFMS, the first fine-grained multimodal sarcasm dataset tailored for Chinese social media. It comprises 2,796 high-quality image-text pairs and provides a triple-level annotation framework: sarcasm identification, target recognition, and explanation generation. We find that the fine-grained explanation annotations effectively guide AI in generating images with explicit sarcastic intent. Furthermore, we curate a high-consistency parallel Chinese-English metaphor subset (200 entries each), revealing significant limitations of current models in metaphoric reasoning. To overcome the constraints of traditional retrieval methods, we propose a Reinforcement Learning-augmented In-Context Learning strategy (PGDS) to dynamically optimize exemplar selection. Extensive experiments demonstrate that CFMS provides a solid foundation for building reliable multimodal sarcasm understanding systems, and the PGDS method significantly outperforms existing baselines on key tasks. Our data and code are available at https://anonymous.4open.science/r/CFMS-E8F9.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。