arXiv:2503.22585cs.CLcs.AI2025-03被引 5

用大模型提升19世纪西语报纸反讽识别,构建新数据集并优化标注流程

Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish

  • 结合历史语境设计半自动标注法,融合人类专家判断
  • 构建首个19世纪拉美报刊情感与反讽标注数据集
  • 适合数字人文、历史文本分析研究者参考

本研究探索大型语言模型(LLMs)在19世纪拉丁美洲报纸中反讽检测的应用。采用两种策略评估BERT和GPT-4o在多分类与二分类任务中的表现:一是通过丰富情感与上下文线索增强数据集,但对历史语言分析效果有限;二是实施半自动标注流程,有效缓解类别不平衡问题,并生成高质量标注数据。尽管反讽识别面临语言复杂性挑战,本工作仍作出两项关键贡献:构建首个针对19世纪西班牙语的带情感与反讽标签的数据集,并提出以历史与文化背景为核心特征的半自动标注方法,强调人类专家在优化模型结果中的关键作用。

原文摘要 · Abstract (English)

This study explores the use of large language models (LLMs) to enhance datasets and improve irony detection in 19th-century Latin American newspapers. Two strategies were employed to evaluate the efficacy of BERT and GPT-4o models in capturing the subtle nuances nature of irony, through both multi-class and binary classification tasks. First, we implemented dataset enhancements focused on enriching emotional and contextual cues; however, these showed limited impact on historical language analysis. The second strategy, a semi-automated annotation process, effectively addressed class imbalance and augmented the dataset with high-quality annotations. Despite the challenges posed by the complexity of irony, this work contributes to the advancement of sentiment analysis through two key contributions: introducing a new historical Spanish dataset tagged for sentiment analysis and irony detection, and proposing a semi-automated annotation methodology where human expertise is crucial for refining LLMs results, enriched by incorporating historical and cultural contexts as core features.

反讽检测历史文本大模型应用数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。