arXiv:2603.03856cs.CL2026-03Conference of the …

通过层次结构融合局部与全局语义原型,提升法律文本修辞角色标注准确率。

Coupling Local Context and Global Semantic Prototypes via a Hierarchical Architecture for Rhetorical Roles Labeling

  • 设计双原型机制:软原型正则化与条件调制,连接局部上下文与全局语义
  • 在低频角色上实现4点宏观F1提升,跨法律、医学、科学三领域均有效
  • 首个三级标注的美国最高法院判决书数据集,适合法律NLP与修辞分析研究

修辞角色标注(RRL)旨在识别文档中每句话的功能角色,是法律、医学等领域篇章理解的关键任务。尽管层次模型能有效捕捉局部依赖,却难以建模全局、语料库级别的特征。为此,我们提出两种基于原型的方法,将局部上下文与全局表示相结合。原型正则化(PBR)通过基于距离的辅助损失学习软原型,以结构化潜在空间;原型条件调制(PCM)构建语料级原型,并在训练和推理阶段注入。鉴于RRL资源稀缺,我们引入SCOTUS-Law,首个标注了美国最高法院判例修辞角色的数据集,涵盖三个粒度层级:类别、修辞功能与步骤。在法律、医学及科学基准上的实验表明,相比强基线模型表现持续提升,尤其在低频角色上取得4点宏观F1增益。我们进一步探讨大模型时代下的影响,并通过专家评估验证结果可靠性。

原文摘要 · Abstract (English)

Rhetorical Role Labeling (RRL) identifies the functional role of each sentence in a document, a key task for discourse understanding in domains such as law and medicine. While hierarchical models capture local dependencies effectively, they are limited in modeling global, corpus-level features. To address this limitation, we propose two prototype-based methods that integrate local context with global representations. Prototype-Based Regularization (PBR) learns soft prototypes through a distance-based auxiliary loss to structure the latent space, while Prototype-Conditioned Modulation (PCM) constructs corpus-level prototypes and injects them during training and inference. Given the scarcity of RRL resources, we introduce SCOTUS-Law, the first dataset of U.S. Supreme Court opinions annotated with rhetorical roles at three levels of granularity: category, rhetorical function, and step. Experiments on legal, medical, and scientific benchmarks show consistent improvements over strong baselines, with 4 Macro-F1 gains on low-frequency roles. We further analyze the implications in the era of Large Language Models and complement our findings with expert evaluation.

修辞标注法律NLP原型学习层次模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。