四款主流大模型在危机求助对话中过度发言,缺乏倾听。
All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers
- 用心理特质明确的虚拟求助者测试模型持续对话能力
- 模型平均说话比例超1,且均未能有效稳定情绪
- 可识别求助者性格特质,适合心理辅助场景研究
大型语言模型越来越多地被用于心理危机时刻,但单轮评估无法检验持续交流或区分用户差异。我们构建了一个基于人格特质的评估框架,让四款广泛使用的模型为多名具有心理测量学定义人格特征的合成求助者提供建议,情境为照护者得知亲属患痴呆症的急性危机。盲评审计员仅凭对话内容即可高一致性地恢复出指定人格维度(ICC(2,4) = 0.91;各量表0.79-0.96;量表得分相关系数r = 0.78),不仅适用于大五人格,也涵盖应对风格、应对自我效能、韧性与逆反性等此前词典方法未覆盖的维度。该评估超越了五因素模型,触及动机、调节与自我评价倾向。四款模型在情绪稳定化方面表现无异,均呈现三种共性:冗长表达、说多听少(说话/倾听比>1)、在未充分理解情境前即启动问题解决。
原文摘要 · Abstract (English)
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。