arXiv:2609.05189cs.CL2026-09

用大模型预测灵活就业者养老金参保行为,效果优于主流模型。

Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers

  • 注入政策规则和边际效应的提示词增强模型推理
  • 在盲测数据上达0.9316复合F1,超越多数基线模型
  • 可解释决策路径,适合政策制定者验证方案效果

评估社会政策变化影响是政策制定者的普遍难题。传统计量方法在假设情景下不可靠,实地试点成本高昂。本文提出将大语言模型(LLM)作为政策评估工具,基于通用模型定制化开发了面向中国灵活就业人员养老金参保预测的领域专用模型FlexPension-LLM。引入DKI-RDistill方法,在提示中注入基于Probit的边际效应与户籍-省份养老金规则等政策线索,并通过LoRA/SFT将带有理由增强的监督信号蒸馏至开源权重的MoE学生模型,对教师错误案例以真实标签重生成修正。在CHFS 2019盲测集上,FlexPension-LLM取得0.9316复合F1,超过其Claude Sonnet 4.5教师及17个基线中的15个,与Claude Opus 4.6无统计差异。在四个外部调查中平均复合F1为0.7549,性能波动最小。组件分析表明,提升主要来自政策线索注入与错误过滤的监督,且理由提供可验证的决策轨迹。

原文摘要 · Abstract (English)

Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometric methods can be unreliable when extrapolating to hypothetical scenarios, while field pilot programs are highly costly. In this paper, we propose using large language models (LLMs) as policy-assessment tools adapted from general-purpose models. We present FlexPension-LLM, the first domain-specialized large language model for a hierarchical pension-enrollment prediction task among flexible workers in China, and introduce DKI-RDistill, which injects policy-grounded cues into the prompt, including Probit-derived marginal effects and hukou-province pension rules. The method then uses LoRA/SFT to distill rationale-augmented supervision into an open-weight MoE student, with teacher errors corrected by regenerating those cases under ground-truth labels. On a CHFS 2019 blind split, FlexPension-LLM achieves 0.9316 Composite F1, surpassing its Claude Sonnet 4.5 teacher and 15 of 17 baselines, and is statistically indistinguishable from Claude Opus 4.6. Across four external surveys, it averages 0.7549 Composite F1 and shows the narrowest performance range among the strongest systems. Component analysis shows that gains come mainly from policy-grounded cue injection and error-filtered supervision, while rationales provide decision traces that can be checked against policy rules.

大模型政策评估养老金可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。