让机器人读懂用户个性,实时调整互动方式。
USER-VLM 360: Personalized Vision Language Models with User-aware Tuning for Social Human-Robot Interactions
- 用视觉语言信号动态调整交互策略,实现个性化响应。
- 在个性化问答和人脸特征理解上分别提升35.3%和47.5%的准确率。
- 兼顾公平性,降低15%偏见,适合社交机器人研发者使用。
将视觉语言模型融入机器人系统,显著提升了机器与环境的自然交互能力。现有方法缺乏对用户个体差异的适应性,常采用通用交互模式,忽视行为、情境及情感层面的个性化需求。当尝试定制时,用户数据中的未缓解偏见又引发伦理风险,可能导致排斥或不公。为此,我们提出User-VLM 360°,一个整合多模态用户建模与偏见感知优化的综合性框架。其核心包括:(1)基于视觉-语言信号实现用户感知的实时交互调优;(2)通过偏好优化实现偏见缓解;(3)构建含人口统计、情绪与关系元数据标注的360°社会情感交互数据集。在八个基准测试中验证,该框架达到顶尖性能:个性化VQA的F1提升35.3%,面部特征理解的F1提升47.5%,偏见降低15%,推理速度比基线快30倍。消融实验确认各组件有效性,且在Pepper机器人上的部署验证了跨用户的实时适应能力。我们开源了参数高效训练的3B/10B模型及伦理验证框架,支持负责任的个性化适配。
原文摘要 · Abstract (English)
The integration of vision-language models into robotic systems constitutes a significant advancement in enabling machines to interact with their surroundings in a more intuitive manner. While VLMs offer rich multimodal reasoning, existing approaches lack user-specific adaptability, often relying on generic interaction paradigms that fail to account for individual behavioral, contextual, or socio-emotional nuances. When customization is attempted, ethical concerns arise from unmitigated biases in user data, risking exclusion or unfair treatment. To address these dual challenges, we propose User-VLM 360°, a holistic framework integrating multimodal user modeling with bias-aware optimization. Our approach features: (1) user-aware tuning that adapts interactions in real time using visual-linguistic signals; (2) bias mitigation via preference optimization; and (3) curated 360° socio-emotive interaction datasets annotated with demographic, emotion, and relational metadata. Evaluations across eight benchmarks demonstrate state-of-the-art results: +35.3% F1 in personalized VQA, +47.5% F1 in facial features understanding, 15% bias reduction, and 30X speedup over baselines. Ablation studies confirm component efficacy, and deployment on the Pepper robot validates real-time adaptability across diverse users. We open-source parameter-efficient 3B/10B models and an ethical verification framework for responsible adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。