融合音视频文本三模态,365维度评估面试表现
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
- 三模态数据+六次回答+五维评价,构建全面评估框架
- 多层感知机压缩特征,均方误差达0.1824,夺冠挑战赛
- 适合自动化招聘、人才评估系统研发者参考
面试表现评估对判断候选人适配度至关重要。为实现全面且公平的评估,我们提出一个全新框架,通过整合视频、音频、文本三模态数据,每名候选人提供六次回应,从五个关键维度进行评估,共探索365个方面。框架采用模态专用特征提取器编码异构数据流,并通过共享压缩多层感知机融合,将多模态嵌入压缩至统一潜在空间,促进高效特征交互。为提升预测鲁棒性,引入两级集成学习策略:(1) 各回应独立使用回归头预测分数,(2) 通过均值池化聚合所有回应预测结果,生成最终五维评分。通过捕捉显性与隐性线索,该方法实现全面、无偏评估。在AVI Challenge 2025中,多维平均均方误差达0.1824,获得第一名,验证了其有效性与鲁棒性。完整代码已公开于https://github.com/MSA-LMC/365Aspects。
原文摘要 · Abstract (English)
Interview performance assessment is essential for determining candidates' suitability for professional positions. To ensure holistic and fair evaluations, we propose a novel and comprehensive framework that explores ``365'' aspects of interview performance by integrating \textit{three} modalities (video, audio, and text), \textit{six} responses per candidate, and \textit{five} key evaluation dimensions. The framework employs modality-specific feature extractors to encode heterogeneous data streams and subsequently fused via a Shared Compression Multilayer Perceptron. This module compresses multimodal embeddings into a unified latent space, facilitating efficient feature interaction. To enhance prediction robustness, we incorporate a two-level ensemble learning strategy: (1) independent regression heads predict scores for each response, and (2) predictions are aggregated across responses using a mean-pooling mechanism to produce final scores for the five target dimensions. By listening to the unspoken, our approach captures both explicit and implicit cues from multimodal data, enabling comprehensive and unbiased assessments. Achieving a multi-dimensional average MSE of 0.1824, our framework secured first place in the AVI Challenge 2025, demonstrating its effectiveness and robustness in advancing automated and multimodal interview performance assessment. The full implementation is available at https://github.com/MSA-LMC/365Aspects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。