arXiv:2506.00955cs.CLcs.SD2025-06被引 6

用大模型自动生成讽刺语音数据集,提升检测效果。

Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection

  • 用GPT-4o和LLaMA 3生成初始标注,再经人工校验。
  • 构建的PodSarc数据集使检测模型达到73.63% F1分数。
  • 适合研究语音讽刺检测但缺乏标注数据的研究者。

讽刺通过语调和语境根本性地改变语义,但因数据稀缺,语音中的讽刺检测仍具挑战性。现有系统常依赖多模态数据,限制了仅语音场景的应用。为此,我们提出一种利用大语言模型(LLMs)生成讽刺语音数据集的标注流程。基于公开的讽刺主题播客,采用GPT-4o和LLaMA 3进行初步标注,再通过人工验证解决分歧。通过协作门控架构在公开数据集上验证该方法的标注质量与检测性能。最终构建了名为PodSarc的大规模讽刺语音数据集。检测模型在该数据集上达到73.63% F1分数,证明其作为讽刺检测研究基准的潜力。

原文摘要 · Abstract (English)

Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, existing detection systems often rely on multimodal data, limiting their applicability in contexts where only speech is available. To address this, we propose an annotation pipeline that leverages large language models (LLMs) to generate a sarcasm dataset. Using a publicly available sarcasm-focused podcast, we employ GPT-4o and LLaMA 3 for initial sarcasm annotations, followed by human verification to resolve disagreements. We validate this approach by comparing annotation quality and detection performance on a publicly available sarcasm dataset using a collaborative gating architecture. Finally, we introduce PodSarc, a large-scale sarcastic speech dataset created through this pipeline. The detection model achieves a 73.63% F1 score, demonstrating the dataset's potential as a benchmark for sarcasm detection research.

讽刺检测语音分析大模型标注数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。