对比4种大模型在日语心理咨询评估中的可靠性与专家意见一致性。
Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling
- 用4个大模型对18场模拟咨询进行三轮评分,评估其稳定性与专家判断匹配度。
- 模型平均得分普遍高于专家,但不同模型间差异显著,最高一致性达0.96。
- 模型可靠性高不等于专家意见一致,适合用于评估自动化评价系统的构建质量。
大型语言模型(LLMs)越来越多地用于生成对话的评估,但重复评分并不一定与专业判断一致。本观察性固定基准研究比较了四种配置的LLM评估系统(GPT-5.5、Gemini 3.5 Flash、Claude Opus 4.8和Fable 5)与15位心理咨询专家对18场模拟日语人机咨询会话的综合评分,这些会话涵盖三种咨询师条件下的六种预设来访者画像。每个系统对每份转录文本独立评分三次,涵盖四种以动机访谈为指导的维度及整体质量。所有四个系统在软化持续谈话(sustain talk)和整体质量维度上的得分均高于专家小组,但差异程度因系统和维度而异。单次运行的组内相关系数(ICC)范围为.33至.96,表明高重测信度并不保证更接近专家判断。Claude Opus 4.8的平均绝对差最小,而Fable 5处于中等水平。在二次基准分析中,使用结构化多步对话提示生成的GPT-4-turbo会话获得更高专家评分,相比仅含简化指令(促进改变谈话、伙伴关系、共情)的情况;但在软化持续谈话方面差异仍不确定。固定基准设计每类组合仅包含一个会话,因此结论仅适用于这些具体会话,而非所有可能的随机重生成。运行间可靠性、专家一致性与条件区分能力是自动咨询评估中相互独立的属性。该基准支持基于构念层面的LLM评估器与专业判断的对照评估。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly evaluate generated dialogue, but repeatable scores do not necessarily align with professional judgment. This observational fixed-benchmark study compared four configured LLM evaluator systems (GPT-5.5, Gemini 3.5 Flash, Claude Opus 4.8, and Fable 5) with aggregated ratings from 15 counseling experts on 18 complete simulated AI-to-AI counseling sessions conducted in Japanese. The sessions represented three counselor conditions across six prespecified client profiles. Each system scored every transcript three times on four motivational interviewing-informed dimensions and overall quality. All four systems assigned higher scores than the expert panel for softening sustain talk and overall quality, although differences varied across systems and constructs. Single-run intraclass correlation coefficients ranged from .33 to .96, showing that high run-to-run reliability did not ensure closer expert-panel alignment. Claude Opus 4.8 had the smallest mean absolute difference, whereas Fable 5 had an intermediate difference. In a secondary benchmark analysis, GPT-4-turbo sessions generated with the Structured Multi-step Dialogue Prompt received higher expert ratings than sessions generated by the same model with a minimal instruction for cultivating change talk, partnership, empathy, and overall quality; the softening sustain talk contrast remained uncertain. The fixed benchmark contained one session per counselor-condition-by-profile cell, so inference concerns these sessions rather than all possible stochastic regenerations. Run-to-run reliability, expert-panel alignment, and condition discrimination are separate properties of automated counseling evaluation. This benchmark supports construct-level assessment of LLM evaluators against professional judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。