构建跨文字历史人物识别基准,验证来源证据对身份判断的关键作用。
When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World
- 基于来源证据构建人物名称配对基准,区分同名不同人。
- 加入来源信息后准确率提升12.96至94.44个百分点。
- 适合历史语言处理、数字人文与可信推理研究者参考。
历史人物在不同语言、文字和转写传统中可能呈现各异形式,而不同个体也可能拥有高度相似甚至相同的姓名。这使得历史身份辨识远超字符串匹配或转写问题。我们提出MHER,一个受来源控制的蒙古世界人物名称配对辨识基准。MHER包含396对仅基于名称的核心数据集(覆盖84位主要历史人物)和更严格的160对基于来源证据的子集,且开发与测试集实体互不重叠。在五种生成系统中,使用来源证据使测试准确率提升12.96至94.44个百分点。在5个表面相同但实际不同的人物案例中,仅靠名称时所有模型均失败(0/25正确),而引入来源证据后24/25可正确分辨,剩余结果为放弃判断。上下文消融实验显示历史描述常含关键身份信息,显式错误来源控制导致性能显著下降。此外发现:名称并非始终有益——对Qwen3-8B而言,恢复表面形式会将10个原本正确的上下文区分误合并为同一身份。结果表明,历史实体辨识不仅依赖表面一致性,更取决于对来源控制证据的恰当响应。MHER为此类研究提供可控框架,用于分析证据使用、放弃判断与失败模式。
原文摘要 · Abstract (English)
Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or even identical names. This makes historical identity reconciliation more than a problem of string matching or transliteration. We introduce MHER, a provenance-controlled benchmark for pairwise reconciliation of person-name attestations from the Mongol world. MHER contains a balanced 396-pair Name-only core over 84 primary historical persons and a stricter 160-pair Source-grounded subset constructed from mention-by-source evidence, with entity-disjoint development and test splits. Across five generative systems, correctly Source-grounded evidence improves paired TEST accuracy by 12.96 to 94.44 percentage points relative to Name-only input. On five identical-surface different-person cases, all models fail under names alone (0/25 model-item decisions), whereas Source-grounded evidence yields 24/25 correct resolutions, with the remaining output an abstention. Context-only ablations show that historical descriptions often carry substantial identity information, while explicitly signaled misgrounding controls produce substantially lower performance. We also find that names are not uniformly beneficial: for Qwen3-8B, restoring surface forms converts ten otherwise correct Context-only distinctions into false identity merges. These results show that historical entity reconciliation depends not only on surface correspondence, but on whether identity judgments respond appropriately to provenance-controlled historical evidence. MHER therefore provides a controlled framework for studying evidence use, abstention, and failure modes in historical NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。