用生成数据+跨语言注意力蒸馏,提升多语言人格识别效果
Cross-Lingual Attention Distillation with Personality-Informed Generative Augmentation for Multilingual Personality Recognition

- 用大模型生成多语言人格文本数据,结合人格特征引导增强质量
- 跨语言注意力蒸馏使模型在各语言上平均准确率提升5.7%~9.7%
- 适合做多语言人格分析或数据稀缺场景的研究者参考
尽管人格识别研究进展显著,但多语言数据集匮乏仍是未解难题。为此,我们提出ADAM(跨语言注意力蒸馏与人格引导生成增强的多语言人格识别方法),以英文人格数据集为基础,利用大语言模型进行翻译增强,并引入人格感知生成增强(PIGA)技术,在日语、中文、马来语和法语中生成高质量训练数据。通过详尽分析验证了该增强方法的有效性。在此基础上,ADAM采用跨语言注意力蒸馏(CLAD)训练模型,实现跨语言人格理解,弥合语言与文化差异。实验表明,结合PIGA增强后,CLAD在所有语言和人格维度上均显著优于标准BCE损失,平均准确率分别提升0.0573(论文数据集)和0.0968(Kaggle数据集)。该模型具备强泛化能力,性能达到当前领先编码器模型水平。模型权重、数据集及代码已公开。
原文摘要 · Abstract (English)
While significant work has been done on personality recognition, the lack of multilingual datasets remains an unresolved challenge. To address this, we propose ADAM (Cross-Lingual (A)ttention (D)istillation with Personality-Guided Generative (A)ugmentation for (M)ultilingual Personality Recognition), a state-of-the-art approach designed to advance multilingual personality recognition. Our approach leverages an existing English-language personality dataset as the primary source and employs a large language model (LLM) for translationbased augmentation, enhanced by Personality-Informed Generative Augmentation (PIGA), to generate high-quality training data in multiple languages, including Japanese, Chinese, Malay, and French. We provide a thorough analysis to justify the effectiveness of these augmentation techniques. Building on these advancements, ADAM integrates Cross-Lingual Attention Distillation (CLAD) to train a model capable of understanding and recognizing personality traits across languages, bridging linguistic and cultural gaps in personality analysis. This research presents a thorough evaluation of the proposed augmentation method, incorporating an ablation study on recognition performance to ensure fair comparisons and robust validation. Overall, with PIGA augmentation, the findings demonstrate that CLAD significantly outperforms the standard BCE across all languages and personality traits, achieving notable improvements in average BA scores - 0.6332 (+0.0573) on the Essays dataset and 0.7448 (+0.0968) on the Kaggle dataset. The CLAD-trained model also demonstrated strong generalizability and achieved benchmark performance comparable to current leading encoder models. The model weight, dataset, and algorithm repository are available at https://research.jingjietan.com/?q=ADAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。