自动生成含表情包的中文多轮对话数据集,提升对话表达力。
MemeCMD: An Automatically Generated Chinese Multi-turn Dialogue Dataset with Contextually Retrieved Memes
- 用双智能体生成对话,结合大模型标注的表情包库自动检索匹配
- 通过自适应阈值确保表情包与语境贴合,使用自然且多样
- 适合研究多模态对话、社交机器人及中文语境下的情感表达
表情包广泛用于线上社交互动,能生动直观地传递意图与情绪。现有对话数据集多为人工标注或纯文本对话,缺乏多模态交互的表达力与语境细节。为此,我们提出 MemeCMD,一个自动构建的中文多轮对话数据集,支持上下文相关表情包插入。该数据集融合大规模多模态大模型标注的表情包库与双智能体在多样化场景下自动生成的对话。我们设计了检索框架与自适应阈值机制,确保表情包在语境上相关且分布自然。实验表明,该方法能有效生成贴合语境且多样化的带表情包对话,为多模态对话AI提供可扩展、隐私友好的资源。
原文摘要 · Abstract (English)
Memes are widely used in online social interactions, providing vivid, intuitive, and often humorous means to express intentions and emotions. Existing dialogue datasets are predominantly limited to either manually annotated or pure-text conversations, lacking the expressiveness and contextual nuance that multimodal interactions provide.To address these challenges, we introduce MemeCMD, an automatically generated Chinese Multi-turn Dialogue dataset with contextually retrieved memes. Our dataset combines a large-scale, MLLM-annotated meme library with dialogues auto-generated by dual agents across diverse scenarios. We introduce a retrieval framework and adaptive threshold to ensure contextually relevant, naturally spaced meme usage. Experiments demonstrate the effectiveness of our approach in generating contextually appropriate and diverse meme-incorporated dialogues, offering a scalable and privacy-preserving resource for advancing multimodal conversational AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。