arXiv:2502.10642cs.AIcs.CV2025-02被引 3

用多模态模型分析人脸和语言数据,提升社交机器人对用户特征的识别能力。

Demographic User Modeling for Social Robotics with Multimodal Pre-trained Models

  • 基于人脸与语言数据构建两个专用用户画像数据集。
  • 微调后CLIP模型显著提升预测能力,但仍难捕捉细微差异。
  • 提出掩码图像建模策略,增强对微妙用户特征的敏感度。

本文研究多模态预训练模型在基于视觉-语言人口统计学数据的用户画像任务中的表现。这类模型对社交机器人适应用户需求与偏好至关重要,可提供个性化回应并提升交互质量。首先,我们构建了两个专门针对用户面部图像衍生的人口统计特征数据集。随后,在这些数据集上评估了代表性对比型多模态预训练模型CLIP的性能,包括其原始状态及微调后的表现。初步结果表明,未微调的CLIP在匹配图像与人口统计描述时表现不佳。尽管微调显著提升了其预测能力,模型仍存在有效泛化细微人口统计差异的局限。为此,我们提出采用掩码图像建模策略,以改善泛化能力并更精准捕捉细微的人口统计属性。该方法为提升多模态用户建模任务中的人口敏感性提供了新路径。

原文摘要 · Abstract (English)

This paper investigates the performance of multimodal pre-trained models in user profiling tasks based on visual-linguistic demographic data. These models are critical for adapting to the needs and preferences of human users in social robotics, thereby providing personalized responses and enhancing interaction quality. First, we introduce two datasets specifically curated to represent demographic characteristics derived from user facial images. Next, we evaluate the performance of a prominent contrastive multimodal pre-trained model, CLIP, on these datasets, both in its out-of-the-box state and after fine-tuning. Initial results indicate that CLIP performs suboptimal in matching images to demographic descriptions without fine-tuning. Although fine-tuning significantly enhances its predictive capacity, the model continues to exhibit limitations in effectively generalizing subtle demographic nuances. To address this, we propose adopting a masked image modeling strategy to improve generalization and better capture subtle demographic attributes. This approach offers a pathway for enhancing demographic sensitivity in multimodal user modeling tasks.

社交机器人多模态用户建模人脸识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。