arXiv:2506.19258cs.CLcs.LG2025-06被引 1

用大模型分析人生故事,精准预测五大性格特质。

Personality Prediction from Life Stories using Language Models

  • 分段提取文本嵌入,再用带注意力的RNN整合长文本依赖
  • 在超2000词的叙事文本上,准确率优于现有长文本模型
  • 兼顾效果与可解释性,适合心理评估与人机交互研究

自然语言处理为性格评估提供了新路径,通过分析丰富的开放文本,超越传统问卷。本研究聚焦于处理超过2000个标记(tokens)的长篇叙述访谈,以预测五因素模型(Five-Factor Model, FFM)人格特质。提出两步法:首先利用滑动窗口微调预训练语言模型提取上下文嵌入;随后采用带注意力机制的循环神经网络(RNN)建模长程依赖并提升可解释性。该混合方法有效结合了预训练变换器与序列建模的优势,成功处理长文本。通过消融实验及与LLaMA、Longformer等前沿长文本模型对比,验证了在预测精度、效率和可解释性上的提升。结果表明,融合语言特征与长文本建模,能显著推进基于人生叙事的性格评估。

原文摘要 · Abstract (English)

Natural Language Processing (NLP) offers new avenues for personality assessment by leveraging rich, open-ended text, moving beyond traditional questionnaires. In this study, we address the challenge of modeling long narrative interview where each exceeds 2000 tokens so as to predict Five-Factor Model (FFM) personality traits. We propose a two-step approach: first, we extract contextual embeddings using sliding-window fine-tuning of pretrained language models; then, we apply Recurrent Neural Networks (RNNs) with attention mechanisms to integrate long-range dependencies and enhance interpretability. This hybrid method effectively bridges the strengths of pretrained transformers and sequence modeling to handle long-context data. Through ablation studies and comparisons with state-of-the-art long-context models such as LLaMA and Longformer, we demonstrate improvements in prediction accuracy, efficiency, and interpretability. Our results highlight the potential of combining language-based features with long-context modeling to advance personality assessment from life narratives.

性格预测长文本建模语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。