用标签语义重排硬样本预测,显著提升文本功能角色标注准确率。
Semantic Reranking at Inference Time for Hard Examples in Rhetorical Role Labeling

- 基于标签语义构建对比学习表示,推理时重排低置信度结果。
- 在8个领域数据集上,硬样本宏F1平均提升9.15点。
- 无需重训练模型,适合医疗法律等高可靠性场景使用。
修辞角色标注(RRL)为文档中每句话赋予功能角色,广泛应用于法律、医学和科学领域。尽管语言模型(LMs)在整体表现上较强,但在低置信度的硬样本上仍不可靠。现有方法通常隐式处理不确定性,将标签视为离散标识符,忽略了标签名称中的语义信息。我们提出RISE,一种推理时的语义重排框架,利用标签语义对硬实例的预测进行优化。RISE自动识别低置信度预测,并通过对比学习获得的标签表征对模型输出进行重排,无需重新训练或修改底层模型。在包含7种语言模型(涵盖编码器与因果架构)的8个特定领域RRL数据集上实验表明,硬样本的宏F1平均提升9.15点。为进一步增强可解释性,我们还引入人工难度标注,从模型与人类双重视角分析困难性,结果显示二者一致性中等,科恩卡帕系数为0.40。
原文摘要 · Abstract (English)
Rhetorical Role Labeling (RRL) assigns a functional role to each sentence in a document and is widely used in legal, medical, and scientific domains. While language models (LMs) achieve strong average performance, they remain unreliable on hard examples, where prediction confidence is low. Existing approaches typically handle uncertainty implicitly and treat labels as discrete identifiers, overlooking the semantic information encoded in label names. We introduce RISE, an inference-time semantic reranking framework that leverages label semantics to refine predictions on hard instances. RISE automatically identifies low-confidence predictions and reranks model outputs using contrastively learned label representations, without retraining or modifying the underlying model. Experiments on eight domain-specific RRL datasets with seven LMs, including encoder-based and causal architectures, show an average gain of +9.15 macro-F1 points on hard examples. For explainability, we further propose manual hardness annotations to study difficulty from both model and human perspectives, revealing a moderate agreement with Cohen's kappa = 0.40.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。