arXiv:2510.16547cs.LGcs.AI2025-10被引 26

用机器学习和可解释AI预测幸福感,准确率达93.8%。

Predicting life satisfaction using machine learning and explainable AI

  • 基于丹麦1.9万人数据,用特征学习提取27个关键问题。
  • 模型准确率93.80%,宏F1达73.00%,健康是跨龄最重要因素。
  • 首次验证生物医学大模型比临床模型更适配幸福感预测。

幸福感是人类福祉的核心维度。传统测量方法复杂且易出错,影响结果可信度。本研究利用丹麦19000名16-64岁居民的政府调查数据,通过机器学习算法实现幸福感预测,准确率达93.80%,宏F1分数为73.00%。采用特征学习技术提取27个关键评估问题,提升模型可复现性与可解释性。进一步探索将表格数据转换为自然语言输入临床与生物医学大语言模型(LLMs),分别获得93.74%准确率与73.21%宏F1。结果显示,幸福感预测与生物医学领域关联更紧密。消融实验表明数据重采样与特征选择对性能有显著影响。年龄分组分析揭示健康状况是各年龄段最重要的决定因素。该研究展示了机器学习、大模型与可解释AI协同提升对主观福祉理解的潜力,对相关领域学者与从业者具有重要参考价值。

原文摘要 · Abstract (English)

Life satisfaction is a crucial facet of human well-being. Hence, research on life satisfaction is incumbent for understanding how individuals experience their lives and influencing interventions targeted at enhancing mental health and well-being. Life satisfaction has traditionally been measured using analog, complicated, and frequently error-prone methods. These methods raise questions concerning validation and propagation. However, this study demonstrates the potential for machine learning algorithms to predict life satisfaction with a high accuracy of 93.80% and a 73.00% macro F1-score. The dataset comes from a government survey of 19000 people aged 16-64 years in Denmark. Using feature learning techniques, 27 significant questions for assessing contentment were extracted, making the study highly reproducible, simple, and easily interpretable. Furthermore, clinical and biomedical large language models (LLMs) were explored for predicting life satisfaction by converting tabular data into natural language sentences through mapping and adding meaningful counterparts, achieving an accuracy of 93.74% and macro F1-score of 73.21%. It was found that life satisfaction prediction is more closely related to the biomedical domain than the clinical domain. Ablation studies were also conducted to understand the impact of data resampling and feature selection techniques on model performance. Moreover, the correlation between primary determinants with different age brackets was analyzed, and it was found that health condition is the most important determinant across all ages. This study demonstrates how machine learning, large language models and XAI can jointly contribute to building trust and understanding in using AI to investigate human behavior, with significant ramifications for academics and professionals working to quantify and comprehend subjective well-being.

幸福感机器学习可解释AI大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。