通过结构化约束提升文档级关系抽取的准确性与泛化能力
Ontology-Driven Structural Regularization for Document-Level Relation Extraction
- 基于本体构建结构一致性约束,识别并修正关系三元组中的逻辑矛盾
- 在DocRED distant数据上发现显著结构噪声,强制正则化后模型性能显著提升
- 适合需要利用远监督数据提升效果的研究者,尤其关注数据质量与可解释性
文档级关系抽取(DocRE)严重依赖昂贵的人工标注数据,而像DocRED distant这样的大规模远监督资源因存在噪声未被充分使用。我们发现,关系三元组中的结构不一致——包括违反本体约束和逻辑矛盾——是关键但被忽视的噪声来源。为此,提出一种基于本体的框架,量化并强制执行文档级数据的结构一致性。分析显示,DocRED distant中存在大量结构噪声,且这些不一致会传播至模型预测。在训练中引入结构正则化,显著减少逻辑矛盾,并持续提升模型泛化性能。结果表明,结构一致性是文档级关系抽取中缺失的监督维度,结构正则化是规模化利用远监督数据的有效策略。
原文摘要 · Abstract (English)
Document-Level Relation Extraction (DocRE) relies heavily on costly manually annotated datasets, while large distant supervision resources such as DocRED distant remain underexploited due to noise. We show that a critical yet overlooked source of noise lies in structural inconsistencies within relational triples, including violations of ontology constraints and logical contradictions. We introduce an ontology-driven framework to quantify and enforce structural consistency in DocRE datasets. Our analysis reveals substantial structural noise in DocRED distant and demonstrates that such inconsistencies propagate to model predictions. Enforcing structural well-formedness during training significantly reduces logical contradictions and consistently improves generalization performance. These findings establish structural consistency as a missing axis of supervision in DocRE and highlight structural regularization as an effective strategy for leveraging distant data at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。