构建首个大规模直播人脸吸引力数据集,提升模型预测准确性。
Facial Attractiveness Prediction in Live Streaming: A New Benchmark and Multi-modal Method
- 提出多模态融合方法,结合面部先验与美学语义特征。
- 在10,000张直播人脸图像上获得20万条主观评分标注。
- 适用于直播美颜、内容推荐等实际场景,适合视觉算法研究者。
人脸吸引力预测(FAP)是计算机视觉的重要任务,可广泛应用于直播中的美颜处理、内容推荐等场景。然而,以往的FAP数据集普遍存在规模小、闭源或多样性不足的问题,且对应模型泛化能力有限。为此,本文提出了首个面向直播场景的大规模FAP数据集LiveBeauty,直接从直播平台采集10,000张人脸图像,并通过精心设计的主观实验获取200,000条吸引力标注,成为当前规模最大、公开可用的直播场景FAP数据集。此外,提出一种多模态FAP方法,通过个性化吸引力先验模块(PAPM)提取整体面部先验知识,利用多模态吸引力编码模块(MAEM)捕获多模态美学语义特征,再经跨模态融合模块(CMFM)整合,实现精准预测。在LiveBeauty及其他开源数据集上的大量实验表明,该方法达到当前最优性能。数据集将尽快开放。
原文摘要 · Abstract (English)
Facial attractiveness prediction (FAP) has long been an important computer vision task, which could be widely applied in live streaming for facial retouching, content recommendation, etc. However, previous FAP datasets are either small, closed-source, or lack diversity. Moreover, the corresponding FAP models exhibit limited generalization and adaptation ability. To overcome these limitations, in this paper we present LiveBeauty, the first large-scale live-specific FAP dataset, in a more challenging application scenario, i.e., live streaming. 10,000 face images are collected from a live streaming platform directly, with 200,000 corresponding attractiveness annotations obtained from a well-devised subjective experiment, making LiveBeauty the largest open-access FAP dataset in the challenging live scenario. Furthermore, a multi-modal FAP method is proposed to measure the facial attractiveness in live streaming. Specifically, we first extract holistic facial prior knowledge and multi-modal aesthetic semantic features via a Personalized Attractiveness Prior Module (PAPM) and a Multi-modal Attractiveness Encoder Module (MAEM), respectively, then integrate the extracted features through a Cross-Modal Fusion Module (CMFM). Extensive experiments conducted on both LiveBeauty and other open-source FAP datasets demonstrate that our proposed method achieves state-of-the-art performance. Dataset will be available soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。