arXiv:2510.09158cs.CL2025-10中稿 · the First Workshop…

用说话前的内心独白增强对话数据,让大模型更像真人

Augmenting Dialog with Think-Aloud Utterances for Modeling Individual Personality Traits by LLM

  • 在对话中加入说话前的内心独白,提升模型对人格的模仿能力
  • 加入内心独白后,模型在宜人性和神经质维度上更贴近真实人类
  • 内心独白质量直接影响模型表现,适合人格建模研究者

本研究提出通过在文本对话中加入说话前的内心独白(Think-Aloud Utterances, TAU)来增强对话数据,以更好建模个体人格特征。TAU是说话人表达话语前的思维外化。我们假设使用TAU增强数据训练的“人格大模型”(persona LLM)能更准确地模拟说话者的个性特质。实验检验了这些模型在五大性格特质(Big Five)框架下是否与人类性格一致。结果表明,使用TAU增强数据训练的模型,在宜人性(Agreeableness)和神经质(Neuroticism)维度上比仅用原始对话数据训练的模型更接近真实说话者。同时发现,TAU增广的质量显著影响人格模型的表现。

原文摘要 · Abstract (English)

This study proposes augmenting dialog data with think-aloud utterances (TAUs) for modeling individual personalities in text chat by LLM. TAU is a verbalization of a speaker's thought before articulating the utterance. We expect "persona LLMs" trained with TAU-augmented data can mimic the speaker's personality trait better. We tested whether the trained persona LLMs obtain the human personality with respect to Big Five, a framework characterizing human personality traits from five aspects. The results showed that LLMs trained with TAU-augmented data more closely align to the speakers' Agreeableness and Neuroticism of Big Five than those trained with original dialog data. We also found that the quality of TAU-augmentation impacts persona LLM's performance.

人格建模大模型对话增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。