首个孟加拉-英语混用语料库,助力识别隐含情感与讽刺。
MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification
- 构建9087条人工标注的混用语句数据集,覆盖幽默、讽刺等多标签。
- 模型在幽默识别上表现良好,但讽刺等复杂语义仍面临挑战。
- 适合研究跨语言文化语义、多语言社交文本分析的学者使用。
孟加拉-英语混用在南亚社交媒体中广泛存在,但针对此类语境下隐含意义识别的资源仍十分稀缺。现有情感与讽刺识别模型多集中于纯英语或高资源语言,难以应对拼写变异、文化引用及句子内部语言切换问题。为此,我们提出MixSarc,首个公开可用的孟加拉-英语代码混用语料库,用于隐含意义识别。该数据集包含9,087条经人工标注的句子,涵盖幽默、讽刺、冒犯性与粗俗性四类标签。通过定向社交媒体采集、系统化筛选与多标注员验证构建。我们对基于Transformer的模型进行基准测试,并在结构化提示下评估零样本大模型表现。结果显示,幽默检测性能优异,但讽刺、冒犯与粗俗性识别因类别不平衡与语用复杂性出现显著下降。零样本模型取得有竞争力的微平均F1分数,但精确匹配准确率较低。进一步分析表明,在外部数据集中超过42%的负面情感实例具有讽刺特征。MixSarc为文化敏感型自然语言处理提供基础资源,支持代码混用环境下的更可靠多标签建模。
原文摘要 · Abstract (English)
Bangla-English code-mixing is widespread across South Asian social media, yet resources for implicit meaning identification in this setting remain scarce. Existing sentiment and sarcasm models largely focus on monolingual English or high-resource languages and struggle with transliteration variation, cultural references, and intra-sentential language switching. To address this gap, we introduce MixSarc, the first publicly available Bangla-English code-mixed corpus for implicit meaning identification. The dataset contains 9,087 manually annotated sentences labeled for humor, sarcasm, offensiveness, and vulgarity. We construct the corpus through targeted social media collection, systematic filtering, and multi-annotator validation. We benchmark transformer-based models and evaluate zero-shot large language models under structured prompting. Results show strong performance on humor detection but substantial degradation on sarcasm, offense, and vulgarity due to class imbalance and pragmatic complexity. Zero-shot models achieve competitive micro-F1 scores but low exact match accuracy. Further analysis reveals that over 42\% of negative sentiment instances in an external dataset exhibit sarcastic characteristics. MixSarc provides a foundational resource for culturally aware NLP and supports more reliable multi-label modeling in code-mixed environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。