通过分块处理面部区域,提升疼痛识别准确率。
ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment

- 将面部划分为四个区域分别处理,优化空间表示
- 视频仅凭此方法达56.00%准确率,创基准新高
- 单区处理仅用1/4计算量,仍保持良好性能
自动从面部视频评估疼痛仍具挑战,因疼痛相关面部线索存在空间异质性。本文提出ReFace,一种空间重组织流程:在分 token 前将面部输入划分为四个空间象限,而非整体处理。在AI4Pain数据集上,该方法仅使用视频即达56.00%测试准确率,在固定基准协议下优于所有对比方法。值得注意的是,四象限配置与全脸输入共享相同像素预算,却获得更高准确率,表明空间重组织可提升该分 token 设计下的性能。单个象限仅处理四分之一像素,计算成本大幅降低,但仍具备竞争力。
原文摘要 · Abstract (English)
Automatic pain assessment from facial video remains challenging due to the spatial heterogeneity of pain-related facial cues. This study proposes ReFace, a spatial reorganization pipeline that divides facial input into four spatial quadrants before tokenization, rather than processing the entire face as a single region. Evaluated on the AI4Pain dataset, the proposed approach achieves $56.00\%$ accuracy on the test set using video only, achieving the highest reported accuracy under the fixed AI4Pain benchmark protocol among the compared methods. Notably, the four-quadrant configuration processes the same total pixel budget as the full-face input, yet achieves higher accuracy, suggesting that spatial reorganization can improve performance under the proposed tokenization design. A single quadrant region, processing just one quarter of those pixels, remains competitive at a fraction of the computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。