arXiv:2505.18903cs.CL2025-05EMNLP被引 10

构建首个多语言单口喜剧幽默检测数据集,支持跨语言幽默理解研究。

StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos

  • 提出词级序列标注新范式,捕捉自然对话中的连续幽默标记。
  • 包含超330小时多语言视频数据,覆盖7种语言,是当前最大最多样数据集。
  • 开源自动笑声检测优化方法,适合多模态幽默研究与跨语言应用者。

为提升幽默检测的计算模型性能,我们构建了一个涵盖七种语言(英语、法语、西班牙语、意大利语、葡萄牙语、匈牙利语、捷克语)的多模态单口喜剧数据集,总时长超过330小时,目前是该任务下最大且最多样化的数据集。整个数据集通过自动方式标注观众笑声,用于模型验证的子集则由人工标注。不同于现有二元序列分类范式,本研究将幽默检测任务定义为词级序列标注,以更好捕捉自然对话中连续的幽默结构。同时,提出一种基于语音识别错误改进自动笑声检测的方法。代码与数据已公开:https://tinyurl.com/EMNLPHumourStandUpPublic。

原文摘要 · Abstract (English)

Aiming towards improving current computational models of humor detection, we propose a new multimodal dataset of stand-up comedies, in seven languages: English, French, Spanish, Italian, Portuguese, Hungarian and Czech. Our dataset of more than 330 hours, is at the time of writing the biggest available for this type of task, and the most diverse. The whole dataset is automatically annotated in laughter (from the audience), and the subpart left for model validation is manually annotated. Contrary to contemporary approaches, we do not frame the task of humor detection as a binary sequence classification, but as word-level sequence labeling, in order to take into account all the context of the sequence and to capture the continuous joke tagging mechanism typically occurring in natural conversations. As par with unimodal baselines results, we propose a method for e propose a method to enhance the automatic laughter detection based on Audio Speech Recognition errors. Our code and data are available online: https://tinyurl.com/EMNLPHumourStandUpPublic

幽默检测多语言视频分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。