用大模型识别家庭迁徙,提升混乱数据中人员关联的检测效果。
Household Movement Detection in Mixed-Format Occupancy Data Using LLM-Based Entity Resolution
- 用提示词驱动的大模型提取姓名地址,免去繁琐预处理。
- 结合语义嵌入与图推理,发现多人共同迁移的间接关联。
- 在真实模拟数据上提升召回率8-15%,F1得分提高6-8%。
实体消歧通常依赖记录间的成对相似性比较,难以捕捉人口居住数据中的间接关系。一种重要间接模式是家庭迁徙:多个个体共同跨地址迁移,但因记录格式混杂、噪声、重复及缺乏稳定标识,检测困难。本文提出一种基于大语言模型的框架,用于在非标准化姓名-地址数据中识别与家庭迁徙相关的间接实体关联。该方法整合了提示词驱动的LLM命名实体识别以提取姓名与地址,无需复杂预处理;利用语义文本嵌入实现鲁棒的相似性计算;并通过图推理推断群体级迁移动态。在使用合成居住生成器构建的SPX基准数据集(S8-S12)上的实验表明,引入家庭迁徙间接证据可使召回率提升8-15%,同时保持高精度,相较强基线的F1分数提升6-8%。
原文摘要 · Abstract (English)
Entity resolution (ER) typically relies on pairwise similarity comparisons between records, which limits its ability to capture indirect relationships present in demographic occupancy data. An important indirect pattern arises from household movement, where multiple individuals relocate together across addresses, but detecting such patterns is difficult due to mixed-format records, noise, duplication, and the absence of stable identifiers. This paper proposes an AI-enhanced framework for detecting indirect entity links associated with household movement in unstandardized name-address data. The approach integrates prompt-based large language model (LLM) named entity recognition for extracting personal names and addresses without extensive preprocessing, semantic text embeddings for robust similarity computation, and graph-based reasoning to infer group-level movement patterns. Experimental evaluation on SPX benchmark datasets (S8-S12) generated using the Synthetic Occupancy Generator demonstrates that incorporating indirect household movement evidence improves recall by 8-15% while maintaining high precision, yielding F1-score gains of 6-8% over a strong pairwise baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。