arXiv:2604.21469cs.CLcs.LG2026-04

通过精准选数据,让合规检测模型跨法规更可靠。

Cross-Domain Data Selection and Augmentation for Automatic Compliance Detection

论文配图:Cross-Domain Data Selection and Augmentation for Automatic Compliance Detection
图 1 · 摘自论文原文
  • 从大语料中选关键样本做数据增强
  • 选对数据可显著减少跨领域误差
  • 适合做法律AI、合规自动化的人看

自动化监管合规检测仍具挑战,因法律文本复杂多变,基于某一法规训练的模型难以泛化到其他法规。本文将合规检测建模为自然语言推理(NLI)任务,研究数据选择策略以缓解负迁移问题。评估四种从大规模源域中选取增强数据的方法:随机采样、Moore-Lewis交叉熵差异、重要性加权和基于嵌入的检索。系统性地改变所选数据比例,分析其对跨域适应的影响。结果表明,有针对性的数据选择能显著降低负迁移,为在异构法规间实现可扩展、可靠的合规自动化提供了实用路径。

原文摘要 · Abstract (English)

Automating the detection of regulatory compliance remains a challenging task due to the complexity and variability of legal texts. Models trained on one regulation often fail to generalise to others. This limitation underscores the need for principled methods to improve cross-domain transfer. We study data selection as a strategy to mitigate negative transfer in compliance detection framed as a natural language inference (NLI) task. Specifically, we evaluate four approaches for selecting augmentation data from a larger source domain: random sampling, Moore-Lewis's cross-entropy difference, importance weighting, and embedding-based retrieval. We systematically vary the proportion of selected data to analyse its effect on cross-domain adaptation. Our findings demonstrate that targeted data selection substantially reduces negative transfer, offering a practical path toward scalable and reliable compliance automation across heterogeneous regulations.

合规检测数据选择NLI跨域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。