arXiv:2502.18650cs.CL2025-02中稿 · the Fourth Worksho…被引 2

用双提示生成面试对话,质量显著优于单提示。

Single- vs. Dual-Prompt Dialogue Generation with LLMs for Job Interviews in Human Resources

  • 双代理对话生成法让对话更像真人
  • 生成成本增加6倍,但辨识率提升2至10倍
  • 适合作为人力资源智能面试系统开发参考

优化语言模型用于对话代理需要大量对话样本。在难以获取真实人类数据的人力资源领域,越来越多采用大语言模型(LLMs)合成对话。本文比较了两种基于LLM的招聘面试对话生成方法:一种是单提示生成完整对话,另一种是使用两个代理相互对话。通过让判断模型对成对对话进行对比,评估其真实性。实验发现,双提示方法虽使词元数量增加六倍,但生成对话的胜率比单提示方法高出2至10倍,该结果在使用GPT-4o或Llama 3.3 70B进行生成或判断时均保持一致。

原文摘要 · Abstract (English)

Optimizing language models for use in conversational agents requires large quantities of example dialogues. Increasingly, these dialogues are synthetically generated by using powerful large language models (LLMs), especially in domains where obtaining authentic human data is challenging. One such domain is human resources (HR). In this context, we compare two LLM-based dialogue generation methods for producing HR job interviews, and assess which method generates higher-quality dialogues, i.e., those more difficult to distinguish from genuine human discourse. The first method uses a single prompt to generate the complete interview dialogue. The second method uses two agents that converse with each other. To evaluate dialogue quality under each method, we ask a judge LLM to determine whether AI was used for interview generation, using pairwise interview comparisons. We empirically find that, at the expense of a sixfold increase in token count, interviews generated with the dual-prompt method achieve a win rate 2 to 10 times higher than those generated with the single-prompt method. This difference remains consistent regardless of whether GPT-4o or Llama 3.3 70B is used for either interview generation or quality judging.

对话生成面试模拟LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。