用AI统一解析各类走失人员文档,提升侦查效率与数据质量。
LLM-based Schema-Guided Extraction and Validation of Missing-Person Intelligence from Heterogeneous Data Sources
- 分步处理多源文档,先规则后大模型,确保结构一致。
- 大模型路径提取准确率高达0.866,关键字段完整率达96.97%。
- 适合公安、救援机构在高要求场景中使用,兼顾准确与可审计。
走失人员与儿童安全调查依赖结构化表格、公告海报及网络叙述性资料等异构文档。布局、术语与数据质量差异阻碍快速筛查、大规模分析与搜寻规划。本文提出Guardian Parser Pack,一种基于AI的解析与标准化流水线,将多源调查文档转化为统一、符合架构的表示形式,适用于操作审查与下游空间建模。系统整合:(i) 多引擎PDF文本提取与OCR备用;(ii) 基于规则的来源识别与专用解析器;(iii) 以架构为先的统一与验证;(iv) 可选的大语言模型(LLM)辅助提取路径,含验证引导修复与共享地理编码服务。我们展示系统架构、关键实现决策与输出设计,并通过黄金对齐提取指标与语料级运营指标评估性能。在75例人工对齐样本上,LLM路径提取质量显著优于确定性对比方案(F1=0.8664 vs. 0.2578),在517条记录中,关键字段完整性也更高(96.97% vs. 93.23%)。确定性路径更快(平均0.03秒/条记录,对比LLM路径3.95秒/条记录)。所有LLM输出均通过初始架构验证,因此验证引导修复未贡献增益,仅作为内置安全机制。结果支持在高风险调查场景中,采用有架构约束的可审计概率型AI。
原文摘要 · Abstract (English)
Missing-person and child-safety investigations rely on heterogeneous case documents, including structured forms, bulletin-style posters, and narrative web profiles. Variations in layout, terminology, and data quality impede rapid triage, large-scale analysis, and search-planning workflows. This paper introduces the Guardian Parser Pack, an AI-driven parsing and normalization pipeline that transforms multi-source investigative documents into a unified, schema-compliant representation suitable for operational review and downstream spatial modeling. The proposed system integrates (i) multi-engine PDF text extraction with Optical Character Recognition (OCR) fallback, (ii) rule-based source identification with source-specific parsers, (iii) schema-first harmonization and validation, and (iv) an optional Large Language Model (LLM)-assisted extraction pathway incorporating validator-guided repair and shared geocoding services. We present the system architecture, key implementation decisions, and output design, and evaluate performance using both gold-aligned extraction metrics and corpus-level operational indicators. On a manually aligned subset of 75 cases, the LLM-assisted pathway achieved substantially higher extraction quality than the deterministic comparator (F1 = 0.8664 vs. 0.2578), while across 517 parsed records per pathway it also improved aggregate key-field completeness (96.97\% vs. 93.23\%). The deterministic pathway remained much faster (mean runtime 0.03 s/record vs. 3.95 s/record for the LLM pathway). In the evaluated run, all LLM outputs passed initial schema validation, so validator-guided repair functioned as a built-in safeguard rather than a contributor to the observed gains. These results support controlled use of probabilistic AI within a schema-first, auditable pipeline for high-stakes investigative settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。