arXiv:2512.07865cs.LG2025-12KDD被引 1

用瑞典居民轨迹文本预测搬家,让社会科学研究更精准。

Using Text-Based Life Trajectories from Swedish Register Data to Predict Residential Mobility with Pretrained Transformers

  • 将690万个人的注册数据转为带语义的文本轨迹
  • Transformer模型预测搬家准确率显著高于传统方法
  • 适合做社会学、人口学与序列建模的研究者参考

我们将大规模瑞典注册数据转化为文本化的生活轨迹,以解决类别变量高基数和编码方案随时间不一致的长期挑战。基于涵盖690万个体(2001–2013年)的完整人口注册数据,将其转换为包含每年居住、工作、教育、收入和家庭状况变化的语义丰富文本,并用于预测后续年份(2013–2017年)的居住迁移。通过对比LSTM、DistilBERT、BERT和Qwen等多种NLP架构,发现序列与基于Transformer的模型能更有效地捕捉时间与语义结构。结果表明,文本化的注册数据保留了个体发展路径的重要信息,支持复杂且可扩展的建模。由于极少国家拥有覆盖广、精度高的纵向微观数据,该数据集为开发与评估新型序列建模方法提供了严格测试平台。总体而言,结合语义丰富的注册数据与现代语言模型,可显著推进社会科学中的纵向分析。

原文摘要 · Abstract (English)

We transform large-scale Swedish register data into textual life trajectories to address two long-standing challenges in data analysis: high cardinality of categorical variables and inconsistencies in coding schemes over time. Leveraging this uniquely comprehensive population register, we convert register data from 6.9 million individuals (2001-2013) into semantically rich texts and predict individuals' residential mobility in later years (2013-2017). These life trajectories combine demographic information with annual changes in residence, work, education, income, and family circumstances, allowing us to assess how effectively such sequences support longitudinal prediction. We compare multiple NLP architectures (including LSTM, DistilBERT, BERT, and Qwen) and find that sequential and transformer-based models capture temporal and semantic structure more effectively than baseline models. The results show that textualized register data preserves meaningful information about individual pathways and supports complex, scalable modeling. Because few countries maintain longitudinal microdata with comparable coverage and precision, this dataset enables analyses and methodological tests that would be difficult or impossible elsewhere, offering a rigorous testbed for developing and evaluating new sequence-modeling approaches. Overall, our findings demonstrate that combining semantically rich register data with modern language models can substantially advance longitudinal analysis in social sciences.

社会计算文本建模轨迹预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。