研究科学写作中上下文如何提升论断关系分类效果
On the Role of Context for Discourse Relation Classification in Scientific Writing
- 用预训练模型分析科学文本的论断结构,考察上下文作用
- 实验表明,基于语篇结构的上下文普遍有助于提升分类性能
- 发现特定论断类型更受益于上下文信息,适合研究AI辅助科研者
随着生成式人工智能在科研工作流中的广泛应用,我们关注如何利用语篇层面的信息来验证AI生成的科学主张。实现这一目标的第一步是探究科学写作中语篇结构的推断任务。本文初步考察了预训练语言模型(PLM)和大语言模型(LLM)在科学出版物中的论断关系分类(DRC)表现,该领域此前研究较少。通过实验分析上下文(以语篇结构定义)对DRC任务的帮助,结果表明上下文通常具有积极作用。同时,我们还分析了哪些科学论断关系类型最能从上下文中获益。
原文摘要 · Abstract (English)
With the increasing use of generative Artificial Intelligence (AI) methods to support science workflows, we are interested in the use of discourse-level information to find supporting evidence for AI generated scientific claims. A first step towards this objective is to examine the task of inferring discourse structure in scientific writing. In this work, we present a preliminary investigation of pretrained language model (PLM) and Large Language Model (LLM) approaches for Discourse Relation Classification (DRC), focusing on scientific publications, an under-studied genre for this task. We examine how context can help with the DRC task, with our experiments showing that context, as defined by discourse structure, is generally helpful. We also present an analysis of which scientific discourse relation types might benefit most from context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。