arXiv:2505.16815cs.CVcs.RO2025-05被引 3

为机器人视觉设计新图像质量评估方法,提升真实场景应用能力。

Image Quality Assessment for Embodied AI

  • 构建感知-认知-决策-执行流程,定义机器人视角的主观评分体系。
  • 建立含36,000对图像的Embodied-IQA数据集,提供超500万细粒度标注。
  • 揭示现有方法不适用于机器人任务,推动更精准质量评估研究。

近年来,具身智能发展迅速,但主要仍局限于实验室环境,真实世界中的各类失真限制了其应用。传统图像质量评估(IQA)方法用于预测人类对失真图像的偏好,但缺乏针对机器人任务可用性的评估标准。为此,我们首次提出具身AI图像质量评估(Embodied-IQA)问题:(1)基于默顿体系与元认知理论,构建感知-认知-决策-执行流程,设计全面的主观评分收集机制;(2)建立Embodied-IQA数据库,包含超过36,000对参考/失真图像,由视觉语言模型、视觉语言动作模型及真实机器人提供超过500万条细粒度标注;(3)在Embodied-IQA上训练并验证主流IQA方法性能,证明需开发更精确的质量指标以支持具身智能在复杂失真环境中的部署。我们希望通过评估推动具身智能在真实世界的广泛应用。

原文摘要 · Abstract (English)

Embodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application. Traditionally, Image Quality Assessment (IQA) methods are applied to predict human preferences for distorted images; however, there is no IQA method to assess the usability of an image in embodied tasks, namely, the perceptual quality for robots. To provide accurate and reliable quality indicators for future embodied scenarios, we first propose the topic: IQA for Embodied AI. Specifically, we (1) based on the Mertonian system and meta-cognitive theory, constructed a perception-cognition-decision-execution pipeline and defined a comprehensive subjective score collection process; (2) established the Embodied-IQA database, containing over 36k reference/distorted image pairs, with more than 5m fine-grained annotations provided by Vision Language Models/Vision Language Action-models/Real-world robots; (3) trained and validated the performance of mainstream IQA methods on Embodied-IQA, demonstrating the need to develop more accurate quality indicators for Embodied AI. We sincerely hope that through evaluation, we can promote the application of Embodied AI under complex distortions in the Real-world. Project page: https://github.com/lcysyzxdxc/EmbodiedIQA

具身智能图像评估机器人视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。