arXiv:2609.03597cs.CL2026-09

首个针对孟加拉法律地契的多模态大模型评测基准,揭示现有模型根本无法读懂当地手写地契。

KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records

论文配图:KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records
图 1 · 摘自论文原文
  • 构建四任务基准,涵盖符号识别、十六进制转十进制、字段提取与法律问答
  • 5个问答类别中39.3%问题所有模型全错,算术任务均低于基础均值
  • 真实人工标注+律师验证,专为多跳推理设计,适合法律与低资源语言研究者

孟加拉国土地所有权记录采用一种名为Ana-Ganda-Kora-Kranti-Til的十六进制位置分数系统,该系统使用专用Unicode字符,无主流字体支持,且未被任何OCR或分词器覆盖。手写地契(RS Khatian)是数百万地块的权威权属证明,常引发民事诉讼,但尚未有评测能检验机器是否可读取。我们提出KhatianDoc,基于孟加拉国门希甘杰地区瓦米地籍办公室的107份真实地契,构建四任务基准:符号识别、十六进制转十进制、结构化字段提取及法律文档问答(共1,634对问答)。真值由人工逐字转录,并经土地法执业者验证达成完全一致,通过位置标记匿名化以保留多跳问题依赖的指代关系。在固定零样本协议下评估六种多模态大模型(8B至72B+,开源与闭源)。5个问答类别中39.3%的问题所有模型回答错误;算术任务中,所有输出数值的模型表现均劣于常数均值基线,精确与近似匹配得分一致:非近似,而是完全脱节。审计自身指标发现两个反向偏差:修正拒绝评分漏洞并报告修正前后分数,标记膨胀元数据指标为上限。KhatianDoc揭示的并非性能差距,而是能力缺失,提供经验证的真值供未来系统使用。代码与数据已公开,图像已脱敏发布。

原文摘要 · Abstract (English)

Land ownership in Bangladesh is recorded in Ana-Ganda-Kora-Kranti-Til, a base-16 positional fraction system with dedicated Unicode glyphs, no mainstream font, and no coverage in any OCR pipeline or tokenizer. The handwritten records that carry these fractions, RS Khatians, are the authoritative title record for millions of parcels and a frequent subject of civil litigation, yet no benchmark has asked whether a machine can read one. We introduce KhatianDoc, a four-task benchmark built from 107 real RS Khatian records from the Vumi (land) Office of Munshiganj, Bangladesh: symbol recognition, base-16-to-decimal conversion, structured field extraction, and legal document question answering over 1,634 QA pairs. Ground truth was transcribed by hand, verified by a land-law practitioner to full agreement, and anonymized through positional tokens that keep the referential distinctions multi-hop questions depend on. We evaluate six multimodal LLMs (8B to 72B+, open and closed) under a fixed zero-shot protocol. Five QA categories, 39.3% of our stratified set, return zero correct answers from every model; on the arithmetic task, every model that emits a number does worse than a constant-mean baseline, with exact- and near-match scores coinciding: decorrelation, not approximation. Auditing our own metrics surfaced two artifacts in opposite directions: we correct a refusal-scoring bug and report the fixed scores beside the originals, and flag an inflated metadata metric as an upper bound. KhatianDoc documents not a performance gap but the absence of a capability, with verified ground truth for future systems. Code and data, with a redacted image release, are publicly available.

多模态法律文本低资源语言地契识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。