arXiv:2602.02360cs.CL2026-02

用多智能体提示框架自动评分面试,效果媲美专家。

Automated Multiple Mini Interview (MMI) Scoring

  • 拆解评估流程:先精炼访谈文本,再分维度打分
  • 平均加权肯德尔系数达0.62,远超微调模型的0.32
  • 无需额外训练,通用性强,适合教育医疗选拔场景

评估共情、伦理判断和沟通等软技能在竞争性选拔中至关重要,但人工评分常存在不一致和偏见。尽管大语言模型(LLMs)提升了自动作文评分(AES)性能,但现有基于理由的微调方法难以处理多重迷你面试(MMI)中抽象且依赖上下文的特质,忽略候选人叙事中的隐含信号。本文提出一种多智能体提示框架,将评估过程分解为转录稿优化与特定标准打分两步。采用三样本上下文学习与大型指令微调模型,该方法在多个指标上超越专用微调基线(平均加权肯德尔系数0.62对比0.32),达到与人类专家相当的可靠性。进一步在ASAP基准测试中验证了其泛化能力,无需额外训练即可媲美领域特定的最先进模型。结果表明,对于复杂主观推理任务,结构化提示工程或可成为数据密集型微调的可扩展替代方案,重新定义了大模型在自动化评估中的应用方式。

原文摘要 · Abstract (English)

Assessing soft skills such as empathy, ethical judgment, and communication is essential in competitive selection processes, yet human scoring is often inconsistent and biased. While Large Language Models (LLMs) have improved Automated Essay Scoring (AES), we show that state-of-the-art rationale-based fine-tuning methods struggle with the abstract, context-dependent nature of Multiple Mini-Interviews (MMIs), missing the implicit signals embedded in candidate narratives. We introduce a multi-agent prompting framework that breaks down the evaluation process into transcript refinement and criterion-specific scoring. Using 3-shot in-context learning with a large instruct-tuned model, our approach outperforms specialised fine-tuned baselines (Avg QWK 0.62 vs 0.32) and achieves reliability comparable to human experts. We further demonstrate the generalisability of our framework on the ASAP benchmark, where it rivals domain-specific state-of-the-art models without additional training. These findings suggest that for complex, subjective reasoning tasks, structured prompt engineering may offer a scalable alternative to data-intensive fine-tuning, altering how LLMs can be applied to automated assessment.

自动评分大模型应用面试评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。