分层教学法提升法律文书修辞角色标注准确率
HiCuLR: Hierarchical Curriculum Learning for Rhetorical Role Labeling of Legal Documents
- 外层按文档难易度排序,内层逐步细化修辞角色区分
- 在4个数据集上效果优于传统方法,提升显著
- 适合法律文本分析与自动摘要研究者使用
法律文书的修辞角色标注(RRL)对摘要生成、语义案例检索和论点挖掘等下游任务至关重要。现有方法常忽视法律文本话语风格和修辞角色内在难度差异。本文提出HiCuLR,一种分层课程学习框架,包含两层课程:外层为修辞角色层级课程(RC),内层为文档层级课程(DC)。DC根据文档偏离标准话语结构的程度等指标划分难度,模型按由易到难顺序学习;RC则逐步引导模型识别粗粒度到细粒度的修辞角色差异。在4个RRL数据集上的实验表明,该方法有效,且DC与RC具有互补性。
原文摘要 · Abstract (English)
Rhetorical Role Labeling (RRL) of legal documents is pivotal for various downstream tasks such as summarization, semantic case search and argument mining. Existing approaches often overlook the varying difficulty levels inherent in legal document discourse styles and rhetorical roles. In this work, we propose HiCuLR, a hierarchical curriculum learning framework for RRL. It nests two curricula: Rhetorical Role-level Curriculum (RC) on the outer layer and Document-level Curriculum (DC) on the inner layer. DC categorizes documents based on their difficulty, utilizing metrics like deviation from a standard discourse structure and exposes the model to them in an easy-to-difficult fashion. RC progressively strengthens the model to discern coarse-to-fine-grained distinctions between rhetorical roles. Our experiments on four RRL datasets demonstrate the efficacy of HiCuLR, highlighting the complementary nature of DC and RC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。