arXiv:2603.09995cs.CLcs.AI2026-03

人机协作比自动思维链更高效提升面试回答质量

Context Over Compute Human-in-the-Loop Outperforms Iterative Chain-of-Thought Prompting in Interview Answer Quality

  • 引入人类参与迭代优化,显著提升回答真实性和信心
  • 人类介入仅需1次迭代,准确率100%,自动方法需5次且成功率84%
  • 适合用于面试培训,尤其关注真实反馈与个性化改进

使用大语言模型进行行为面试评估面临结构化评判、模拟真实面试官行为及对候选人训练具有教学价值等挑战。本文通过两组控制实验,基于50个行为面试问答对,研究了思维链提示在面试回答评估与优化中的应用。结果表明:人机协同模式在评分提升上显著优于自动化链式思维,信心评分从3.16升至4.16(p<0.001),真实性从2.94升至4.53(p<0.001,Cohen's d=3.21)。人机模式仅需平均1.0次迭代(自动为5.0次,p<0.001),并实现完整个人细节融合。两者收敛迅速,均值迭代低于1次,但人机模式对初始表现弱的回答成功率达100%,自动化为84%(Cohen's h=0.82,大效应)。额外迭代收益递减,说明瓶颈在于上下文信息而非算力。本文还提出基于负面偏见模型的对抗性挑战机制‘bar raiser’,以模拟真实面试官行为,量化验证待后续工作。

原文摘要 · Abstract (English)

Behavioral interview evaluation using large language models presents unique challenges that require structured assessment, realistic interviewer behavior simulation, and pedagogical value for candidate training. We investigate chain of thought prompting for interview answer evaluation and improvement through two controlled experiments with 50 behavioral interview question and answer pairs. Our contributions are threefold. First, we provide a quantitative comparison between human in the loop and automated chain of thought improvement. Using a within subject paired design with n equals 50, both approaches show positive rating improvements. The human in the loop approach provides significant training benefits. Confidence improves from 3.16 to 4.16 (p less than 0.001) and authenticity improves from 2.94 to 4.53 (p less than 0.001, Cohen's d is 3.21). The human in the loop method also requires five times fewer iterations (1.0 versus 5.0, p less than 0.001) and achieves full personal detail integration. Second, we analyze convergence behavior. Both methods converge rapidly with mean iterations below one, with the human in the loop approach achieving a 100 percent success rate compared to 84 percent for automated approaches among initially weak answers (Cohen's h is 0.82, large effect). Additional iterations provide diminishing returns, indicating that the primary limitation is context availability rather than computational resources. Third, we propose an adversarial challenging mechanism based on a negativity bias model, named bar raiser, to simulate realistic interviewer behavior, although quantitative validation remains future work. Our findings demonstrate that while chain of thought prompting provides a useful foundation for interview evaluation, domain specific enhancements and context aware approach selection are essential for realistic and pedagogically valuable results.

面试生成人机协作思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。