arXiv:2609.02231cs.AI2026-09

用可追溯证据自动评估视频面试,得分更准且透明。

PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment

论文配图:PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment
图 1 · 摘自论文原文
  • 构建多模态语义图作记忆,跨模态验证行为证据
  • 在VInterview-2025上达91.50%等级准确率,超越大模型
  • 适合需透明评分的招聘场景,支持人工复核证据

视频面试评估需基于行为证据进行逐项判断,但申请人量激增使人工评估成本高且不一致,现有AI方法评分模糊无依据。本文提出PhoenixNest-Video,一种基于证据的多模态智能体框架,通过构建语义视频图作为结构化工作记忆,实现跨视觉、音频和文本流的鲁棒检索与验证,并生成锚定在候选人材料上的分项评分。评分器采用基于量表的强化学习训练,双奖励机制同时优化量表对齐与评分差异性。在VInterview-2025数据集上达到91.50%的等级准确率,显著优于更大规模的商用模型。该紧凑型量表驱动智能体评分结果与专家小组高度一致,且可暴露每项得分的原始证据供人工审查。

原文摘要 · Abstract (English)

Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agent framework for automated video interview assessment. It builds a semantic video graph as structured working memory, performs rubric-conditioned retrieval with cross-modal verification across visual, audio, and textual streams, and produces per-criterion scores anchored to the candidate's materials. A Scorer trained via Rubrics-based Reinforcement Learning with dual rewards for rubric alignment and score-level differentiation internalizes the discriminative structure of multi-level rubrics. PhoenixNest-Video attains 91.50\% grade-level accuracy on VInterview-2025, outperforming substantially larger proprietary models. A compact, rubric-grounded agent therefore scores candidates in closer agreement with an expert panel than direct prompting of much larger models, and exposes the evidence behind each score for human review.

视频面试多模态可解释性智能评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。