用大模型分析人生轨迹数据,预测更准还发现社会不平等模式。
LifeSentence: Language models can encode human life course trajectories from longitudinal panel data

- 将人生事件转为自然语言,用240亿参数大模型进行指令微调
- 仅用6.5万德国样本即实现事件与时间预测三倍提升,排序准确率达91.2%
- 无需标注就能发现教育溢价、性别工资差等社会规律,适合社科研究
预测人类生命历程对理解健康长寿至关重要。传统统计方法因忽略生命历程的时序结构而精度有限。现代Transformer需海量数据,但多数纵向面板研究无法满足。本文提出LifeSentence,将大语言模型与纵向面板数据结合:将每个生命事件表示为结构化自然语言记录,并在18个任务评估体系(涵盖预测、鲁棒性与推理)上指令微调一个240亿参数预训练模型。该模型利用预训练中已编码的分布知识,仅需约6.5万名德国社会经济面板(SOEP)个体数据——远少于以往基于Transformer的方法(约45倍),仍全面超越经典与深度学习基线。在联合事件与时间预测上较最优基线提升三倍,从无时间戳事件集中重构时序顺序的Kendall's tau达91.2%。无需显式监督,模型即可从离散事件序列中恢复出已知的社会分层模式,如教育溢价、性别工资差距和母亲惩罚。自然语言接口支持全新研究问题,例如将早年经历与特定晚年结果关联,使LifeSentence兼具预测工具与反事实探索探针功能。
原文摘要 · Abstract (English)
Forecasting human life outcomes is important to gain insights into how individuals attain long and healthy lives. Conventional statistical approaches yield limited accuracy, potentially due to discarding the sequential structure of the life course. Modern methods such as transformer architectures require large scale training data that most longitudinal panel studies lack. Here we introduce LifeSentence, a model for life-course reasoning that bridges large language models with longitudinal panel data. By representing each life event as a structured natural-language record and instruction-tuning a pretrained 24-billion-parameter language model across an 18-task evaluation taxonomy spanning prediction, robustness and reasoning, LifeSentence supplements panel data with distributional knowledge already encoded during pretraining. Trained on approximately 65,000 individuals from the German Socio-Economic Panel - roughly 45 times fewer than prior transformer-based approaches - LifeSentence outperforms classical and deep learning baselines across all task families, achieving a threefold improvement in joint event-and-timing prediction from best baselines and 91.2% Kendall's tau when reconstructing chronological order from timestamp-stripped event sets. Without explicit supervision, the model recovers documented patterns of social stratification, including the education premium, the gender wage gap and the motherhood penalty, from discrete event sequences alone. A natural-language interface further enables qualitatively new research queries, such as connecting an early-life history to a specified late-life endpoint, establishing LifeSentence as both a predictive tool and a probe for counterfactual exploration of human biographies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。