arXiv:2609.05296cs.CLcs.LG2026-09

提出LexFlip诊断法,检验法律条款简化是否保留原意。

LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics

  • 设计保持字面形式但反转法律效力的微小修改样本
  • 现有指标仅0.022~0.039范围响应法律效力变化,表现极弱
  • 适合评估法律文本简化工具的语义保真度

当前法律条款简化评估方法无法验证语义是否保留:要求相同对得分最高、无关对最低,会将词汇重叠与法律效力捆绑,导致任何单调函数都满足。本文提出解耦诊断法,通过保持表面形式不变而改变法律效力的最小扰动来检验。我们发布LexFlip数据集,包含373个魁北克法规法语的微小修改样本,法律效力反转的同时保留0.93的词元(tokens)。该数据集用于评估多种指标、回归器和人工裁判。测试发现,七种嵌入模型与BERTScore指标在该编辑上的响应范围仅为0.022至0.039,远低于双向NLI的0.670;在FrJudge上,人类评估上限为相关系数r=0.597,而仅长度特征就超越所有语义指标,且误差最小。

原文摘要 · Abstract (English)

Does a simplified legal clause still say what the original said? The checks in current use cannot establish that it does: requiring an identical pair to score highest and an unrelated pair lowest moves lexical overlap and legal force together, so any monotone function of token overlap satisfies both. Our remedy is a dissociation, an item holding surface form fixed while legal force moves. We release LexFlip, 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving 0.93 of the tokens, with a harness scoring metrics, regressors and prompted judges alike. The seven embedding and BERTScore metrics we test spend only 0.022 to 0.039 of their identical-to-unrelated range on such an edit, against 0.670 for bidirectional NLI, the one family the identical-pair check would disqualify. On FrJudge, against a measured human ceiling of r=0.597, a bare length feature outscores every semantic metric and has the lowest margin we measure.

法律文本语义保真评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。