提出新方法提升跨域面部动作单元检测准确率
Decoupled Doubly Contrastive Learning for Cross Domain Facial Action Unit Detection
- 分离表情与领域特征,专注表情信息适应
- 跨域合成图像质量高,平均F1提升6%-14%
- 适合缺乏多样数据的跨域表情识别任务
尽管现有基于视觉的面部动作单元(AU)检测方法表现优异,但对不同域间差异仍高度敏感,跨域检测方法研究尚不充分。为此,我们提出解耦双重对比学习(D²CA)方法,旨在学习语义对齐的纯净AU表示。具体而言,将潜在表示分解为与AU相关和无关的成分,仅在AU相关子空间内促进适配。通过在跨域场景中修改AU或领域属性时评估合成人脸质量,实现特征解耦。为强化解耦效果,尤其在AU数据多样性有限时,D²CA采用图像与特征级双重对比学习机制,确保合成质量并减少特征歧义。该框架实现自动、专一的AU与领域因素分离,支持直观、尺度可控的跨域人脸图像生成。大量实验表明,D²CA成功解耦了AU与领域因素,生成视觉上令人满意的跨域合成图像;同时在多种跨域场景中持续优于当前最优方法,平均F1分数提升6%–14%。
原文摘要 · Abstract (English)
Despite the impressive performance of current vision-based facial action unit (AU) detection approaches, they are heavily susceptible to the variations across different domains and the cross-domain AU detection methods are under-explored. In response to this challenge, we propose a decoupled doubly contrastive adaptation (D$^2$CA) approach to learn a purified AU representation that is semantically aligned for the source and target domains. Specifically, we decompose latent representations into AU-relevant and AU-irrelevant components, with the objective of exclusively facilitating adaptation within the AU-relevant subspace. To achieve the feature decoupling, D$^2$CA is trained to disentangle AU and domain factors by assessing the quality of synthesized faces in cross-domain scenarios when either AU or domain attributes are modified. To further strengthen feature decoupling, particularly in scenarios with limited AU data diversity, D$^2$CA employs a doubly contrastive learning mechanism comprising image and feature-level contrastive learning to ensure the quality of synthesized faces and mitigate feature ambiguities. This new framework leads to an automatically learned, dedicated separation of AU-relevant and domain-relevant factors, and it enables intuitive, scale-specific control of the cross-domain facial image synthesis. Extensive experiments demonstrate the efficacy of D$^2$CA in successfully decoupling AU and domain factors, yielding visually pleasing cross-domain synthesized facial images. Meanwhile, D$^2$CA consistently outperforms state-of-the-art cross-domain AU detection approaches, achieving an average F1 score improvement of 6\%-14\% across various cross-domain scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。