融合语音、文本与心理语言学特征,提升犹豫与矛盾情绪识别准确率。
Audio-Text Cross-Attention with Psycholinguistic Support Features for Ambivalence/Hesitancy Recognition
- 用跨模态注意力融合音频与文本,结合74个心理语言学特征
- 在525视频测试集上达到0.875平均精度和0.722宏F1
- 适合情绪分析、人机交互领域研究者参考
我们提出一种帧无关的音视频联合系统,用于第11届Affective & Behavior Analysis in-the-Wild(ABAW)研讨会第三届犹豫与矛盾情绪识别挑战赛。视频被划分为与转录时间对齐的重叠5秒窗口,每个窗口整合了韵律性音频描述符、情感导向的RoBERTa嵌入以及74个代表不确定、缓和与态度冲突的心理语言学特征。通过时序交叉注意力融合音频与文本信息,并利用支持特征调控门控多实例学习(MIL)池化。五组种子集成在525个视频的公开测试集上取得0.875的平均精度和0.722的宏F1。值得注意的是,我们的提交在官方排行榜中位列第三,宏F1达0.7455。源代码已开源:https://github.com/Liga-de-IA-PUCPR/abaw-11-ah-challenge/
原文摘要 · Abstract (English)
We present a frame-independent audio-text system for the 3rd Ambivalence/Hesitancy Video Recognition Challenge at the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop. Videos are divided into overlapping 5-s windows aligned with transcript timestamps. Each window combines prosodic audio descriptors, emotion-oriented RoBERTa embeddings, and 74 psycholinguistic features representing uncertainty, hedging, and attitudinal conflict. Temporal cross-attention fuses audio and text, while the support features condition gated Multiple Instance Learning (MIL) pooling. A five-seed ensemble achieves an average precision of 0.875 and a macro-F1 of 0.722 on the 525-video labeled public-test split. Notably, our submission ranked third overall on the official challenge leaderboard, with a macro-F1 of 0.7455. Source code is available at https://github.com/Liga-de-IA-PUCPR/abaw-11-ah-challenge/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。