arXiv:2503.20588cs.CL2025-03被引 1

用大模型生成伪数据提升跨领域隐含语义关系识别效果,但实验显示效果不显著。

Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation Recognition

  • 用大模型基于目标域无标注数据生成语义连贯的文本延续
  • 在大规模测试集上未观察到性能显著提升,优于基线不足1%
  • 提醒评估时需兼顾统计显著性与可比性,适合关注模型可靠性研究者

隐含语义关系识别(IDRR)需要深层语义理解。近期研究表明,零样本或少样本方法远落后于有监督模型,而大语言模型(LLMs)可用于合成数据增强,即生成符合特定语义关系的第二个文本片段。本文在跨领域设置下应用该方法:使用未标注的目标域数据生成语义延续,以适应在源域标注数据上训练的基线模型。在大规模测试集上的评估表明,不同变体均未带来显著性能提升。结论是,大模型常无法生成对IDRR任务有用的样本,并强调在评估IDRR模型时需同时考虑统计显著性和可比性。

原文摘要 · Abstract (English)

Implicit discourse relation recognition (IDRR) -- the task of identifying the implicit coherence relation between two text spans -- requires deep semantic understanding. Recent studies have shown that zero- or few-shot approaches significantly lag behind supervised models, but LLMs may be useful for synthetic data augmentation, where LLMs generate a second argument following a specified coherence relation. We applied this approach in a cross-domain setting, generating discourse continuations using unlabelled target-domain data to adapt a base model which was trained on source-domain labelled data. Evaluations conducted on a large-scale test set revealed that different variations of the approach did not result in any significant improvements. We conclude that LLMs often fail to generate useful samples for IDRR, and emphasize the importance of considering both statistical significance and comparability when evaluating IDRR models.

隐含关系合成数据跨域学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。