arXiv:2608.20391cs.CL2026-08

构建美国移民上诉案结构化数据集,助力法律AI研究

ImmigrationReason: A Structured Dataset of U.S. Immigration Appeals for Legal Reasoning Research

  • 从1.2万份移民裁决中提取法律框架与证据评估结果
  • 发现近9000条裁决错误实例,涵盖21年政策变迁
  • 适合法律AI、监管决策系统研究者使用

现有法律自然语言处理资源多基于联邦判例,聚焦粗粒度分类,忽视了绝大多数政府决策发生于行政裁决领域。本文提出ImmigrationReason,一个大规模结构化数据集,源自美国公民及移民服务局行政上诉办公室(USCIS AAO)2005至2026年间12,375份非先例裁决。每条记录包含适用法律框架、五类标准下的证据充分性判断、裁决官批评原文引述、全部引用文献及最终裁决结果,并附高精度Claude转录文本。通过三阶段验证流程(双模态对比+Opus 4.7提示比对+领域专家抽样核查),确保数据质量。数据集记录近9,000条AAO识别的法律错误,覆盖2016年Dhanasar规则变更这一自然法律制度转型,涵盖21年行政裁决实践。本文详析数据特征,提出可开展的科研方向,如裁决结果预测、裁决官错误分析、高风险监管领域的智能代理设计。

原文摘要 · Abstract (English)

Most legal NLP resources draw from federal case law and focus on coarse classification, leaving administrative adjudication, where the vast majority of government decisions occur, essentially unaddressed. We introduce ImmigrationReason, a large-scale structured dataset derived from 12,375 non-precedent decisions of the U.S. Citizenship and Immigration Services (USCIS) Administrative Appeals Office (AAO) spanning 2005 to 2026. Each record captures the applicable legal framework, per-criterion evidence-sufficiency findings under a five-category label, verbatim adjudicator-criticism quotes, all citations, and final dispositions, alongside high-quality Claude-transcribed source text. Extraction quality is validated through a three-pass pipeline combining two independent modalities with comparison-prompt adjudication by Opus 4.7, and verified by domain experts on a 500-record sample. The dataset documents nearly 9,000 verbatim instances of AAO-identified legal errors, spans a natural legal-regime transition (the 2016 Dhanasar rule change), and covers 21 years of adjudication. We analyze the dataset in detail and outline research directions it enables, from outcome prediction and adjudicator-error analysis to agent design for high-stakes regulatory domains.

法律AI数据集移民政策结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。