LLM评分能大致判断模型优劣,但不能替代人工评估。
How Trustworthy Are LLM-as-Judge Ratings for Interpretive Responses? Implications for Qualitative Research Workflows
- 用LLM当裁判,评估五个模型的解读质量。
- 模型间评分趋势相似,但具体分数差异大,尤其在细微理解上。
- 适合用于筛选差模型,不适合完全取代人类判断。
随着定性研究者越来越多使用自动化工具辅助解释性分析,大型语言模型(LLM)常未经系统评估便直接引入分析流程。本研究考察了LLM作为评判者对解释性质量的评估是否与人类判断一致,并能否支持模型选择。基于712段来自中小学数学教师半结构化访谈的对话片段,采用五种主流推理模型(Command R+、Gemini 2.5 Pro、GPT-5.1、Llama 4 Scout-17B Instruct、Qwen 3-32B Dense)生成一句式解释性回应。通过AWS Bedrock的LLM-as-judge框架在五个指标上进行自动评估,并由受训人类评审员对其中分层抽样的响应独立评价解释准确性、细微性保留和解释连贯性。结果显示,尽管模型层面的评分趋势与人类判断方向一致,但在分数幅度上存在显著偏差;其中连贯性指标与人类综合评分相关性最强,而忠实性和正确性在非字面、复杂解释中出现系统性偏差。安全类指标与解释质量无关。研究建议:LLM-as-judge更适合用于淘汰表现差的模型,而非替代人工判断,为定性研究中的模型选型提供实证指导。
原文摘要 · Abstract (English)
As qualitative researchers show growing interest in using automated tools to support interpretive analysis, a large language model (LLM) is often introduced into an analytic workflow as is, without systematic evaluation of interpretive quality or comparison across models. This practice leaves model selection largely unexamined despite its potential influence on interpretive outcomes. To address this gap, this study examines whether LLM-as-judge evaluations meaningfully align with human judgments of interpretive quality and can inform model-level decision making. Using 712 conversational excerpts from semi-structured interviews with K-12 mathematics teachers, we generated one-sentence interpretive responses using five widely adopted inference models: Command R+ (Cohere), Gemini 2.5 Pro (Google), GPT-5.1 (OpenAI), Llama 4 Scout-17B Instruct (Meta), and Qwen 3-32B Dense (Alibaba). Automated evaluations were conducted using AWS Bedrock's LLM-as-judge framework across five metrics, and a stratified subset of responses was independently rated by trained human evaluators on interpretive accuracy, nuance preservation, and interpretive coherence. Results show that LLM-as-judge scores capture broad directional trends in human evaluations at the model level but diverge substantially in score magnitude. Among automated metrics, Coherence showed the strongest alignment with aggregated human ratings, whereas Faithfulness and Correctness revealed systematic misalignment at the excerpt level, particularly for non-literal and nuanced interpretations. Safety-related metrics were largely irrelevant to interpretive quality. These findings suggest that LLM-as-judge methods are better suited for screening or eliminating underperforming models than for replacing human judgment, offering practical guidance for systematic comparison and selection of LLMs in qualitative research workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。