arXiv:2608.26171cs.CYcs.CL2026-08

用提示词防护和人工审核降低大模型招聘流水线中的虚构信息风险

Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints

论文配图:Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints
图 1 · 摘自论文原文
  • 在简历优化后加入人工审核节点,可彻底消除身份虚构
  • 提示词防护使虚假内容密度下降86%,但仍有一半输出含造假
  • 适合关注招聘安全、合规的HR与AI系统设计者参考

多阶段大模型招聘流程(简历优化、面试题生成、回答反馈)会编造资质、夸大条件、虚构经历。我们评估了提示词防护和人工介入检查点两种缓解策略,对比全自动基线。在受控实验中(10份合成简历×2个职位描述×3次重复×3种条件;共180次运行),基线(C1)在96.7%输出中至少存在一条无依据主张(平均6.80处/输出)。提示词防护(C2)使主张密度下降86%(从6.80降至0.92/输出),但仍有50.0%输出含伪造内容,表明仅靠提示词不足以解决问题。在简历优化后加入人工检查点(C3)彻底消除身份虚构,主张密度下降59%(6.88降至2.82/输出),项目级伪造率从96.7%降至75.0%(p=.022),捕获职位描述嵌入陷阱要求的比例从47%降至2%(对比防护组为5%)。探索性分析显示,跨专业简历的污染程度随领域距离增加而单调上升,转行者尤为易受影响。人工审查发现所有明显造假,但对细微降级和合理新主张仅能剔除54.5%。两种干预均未损害产出质量:两者下主张保留率均超99%。二者互补:提示词防护低成本消除无提示添加和资格虚高,人工检查点提供对最严重错误(如虚构身份、诱导性回应)的近保证。结果支持结合提示词防护与人工检查点的分层架构。补充实验使用新一代模型(基线造假率90.0%)表明,模型进步无法单独解决该问题。

原文摘要 · Abstract (English)

Multi-stage LLM hiring pipelines (resume improvement, interview question generation, answer feedback) can fabricate credentials, inflate qualifiers, and invent experience. We evaluate two mitigations, prompt guardrails and human-in-the-loop (HITL) checkpoints, against a fully automated baseline. In a controlled experiment (10 synthetic resumes x 2 job descriptions x 3 repetitions x 3 conditions; 180 runs), the baseline (C1) produced at least one unsupported claim in 96.7% of outputs (mean 6.80 findings/output). Prompt guardrails (C2) reduced finding density by 86% (6.80 to 0.92/output), but 50.0% of outputs still contained a fabrication, showing prompt-level mitigation alone is insufficient. A human checkpoint after resume improvement (C3) eliminated all identity fabrications, reduced finding density by 59% (6.88 to 2.82/output), reduced item-level fabrication from 96.7% to 75.0% (p=.022), and cut capture of JD-embedded trap requirements from 47% to 2% (vs. 5% under the guardrail). An exploratory analysis of multi-specialty resumes shows contamination rising monotonically with domain distance between specialties, suggesting career changers are especially exposed. The reviewer in this study caught all flagrant fabrications, but subtle qualifier drops and plausible new claims survived review roughly half the time (54.5% removal). Neither mitigation degraded the deliverable: claim retention exceeded 99% under both. The interventions are complementary: the guardrail eliminates unprompted additions and qualifier inflation cheaply, while the checkpoint gives near-categorical guarantees against the most severe failures, invented identities and JD-baited claims. These results support a layered architecture combining guardrails with a human checkpoint. A supplementary run with a newer-generation model (90.0% baseline fabrication rate) suggests the problem is not resolved by model progress alone.

大模型安全招聘系统幻觉检测人工审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。