通过筛选相关文本提升人格预测精度
BIG5-TPoT: Predicting BIG Five Personality Traits, Facets, and Items Through Targeted Preselection of Texts
- 基于语义相关性预选文本输入模型
- 在意识流作文数据集上误差与准确率均提升
- 适合需要精准人格分析的研究者使用
从生成文本中预测个体人格是一项挑战,尤其当文本量较大时。本文提出一种简单而有效的策略——目标文本预筛选(TPoT),该方法对输入文本进行语义过滤,配合专用于预测大五人格特质、分面或条目的深度学习模型(BIG5-TPoT)。通过选择与特定特质、分面或条目语义相关的文本,该策略不仅缓解了大语言模型的输入文本容量限制,还在意识流作文数据集上提升了预测的平均绝对误差和准确率。
原文摘要 · Abstract (English)
Predicting an individual's personalities from their generated texts is a challenging task, especially when the text volume is large. In this paper, we introduce a straightforward yet effective novel strategy called targeted preselection of texts (TPoT). This method semantically filters the texts as input to a deep learning model, specifically designed to predict a Big Five personality trait, facet, or item, referred to as the BIG5-TPoT model. By selecting texts that are semantically relevant to a particular trait, facet, or item, this strategy not only addresses the issue of input text limits in large language models but also improves the Mean Absolute Error and accuracy metrics in predictions for the Stream of Consciousness Essays dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。