arXiv:2409.19308cs.CLcs.AI2024-09被引 6

用人口数据微调大模型,让模拟公众意见更精准。

Designing Domain-Specific Large Language Models: The Critical Role of Fine-Tuning in Public Opinion Simulation

  • 结合年龄、收入等人口特征微调模型,生成更具代表性的观点。
  • 在卡方检验等指标上显著优于预训练模型,提升观点多样性捕捉能力。
  • 适合政策研究、社会学分析者,推动公平决策的可信模拟。

大型语言模型虽已革新自然语言处理,但在环境政策公众意见模拟等专业任务中仍面临挑战。本文提出一种新颖的微调方法,融合英国住户纵向研究的人口统计数据,利用年龄、性别、收入、教育水平和区域等特征进行建模。通过生成多样化的合成用户画像,微调后的模型在捕捉人口统计细节方面显著优于原始预训练模型。评估采用卡方检验、余弦相似度、杰卡德指数和KL散度等指标,结果显示合成观点与真实世界意见高度一致。该研究展示了针对社会背景定制微调的大模型在政策模拟中的潜力,未来可拓展至医疗、教育等领域,促进研究与实践中的包容性与数据驱动决策。

原文摘要 · Abstract (English)

Large language models (LLMs) have transformed natural language processing, yet face challenges in specialized tasks such as simulating opinions on environmental policies. This paper introduces a novel fine-tuning approach that integrates socio-demographic data from the UK Household Longitudinal Study, uniquely using profiling factors, such as age, gender, income, education, and region. This method enhances the accuracy and representation of generated views. By emulating diverse synthetic profiles, the fine-tuned models significantly outperform pre-trained counterparts, achieving measurable improvements in capturing demographic nuances. Evaluation metrics, including Chi-Squared, Cosine Similarity, Jaccard Index, and KL-divergence, reveal a strong alignment between synthetic and real-world opinions. This work demonstrates the potential of fine-tuned LLMs tailored to societal contexts to enable more ethical and precise policy simulations. Its broader implications include deploying LLMs in domains like healthcare and education, fostering inclusive and data-driven decision-making in both research and practice.

大模型微调公众意见模拟社会仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。