arXiv:2605.06311cs.RO2026-05被引 1

构建高保真视觉仿真基准,提升机器人抓取评估与现实的匹配度。

Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation

论文配图:Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation
图 1 · 摘自论文原文
  • 用物理渲染材料+多模态大模型自动生成逼真3D资产
  • 仿真与现实性能相关性达0.92,显著缩小虚实差距
  • 适合评估视觉-语言-动作模型的仿真测试场景

可靠的机器人操作仿真评估可作为真实世界表现的高保真代理。尽管现有基准涵盖多种任务类型,但缺乏视觉真实性,导致仿真与现实之间存在巨大域差距,削弱了仿真评估预测真实性能的可靠性。为缓解这一问题,我们系统分析光照与材质的影响,发现它们在几何推理和空间定位中起关键作用,却在现有基准中被忽略。基于此,我们提出VISER——一个面向机器人操作的视觉真实仿真基准。该基准包含超过1000个3D资产的高保真数据集,采用基于物理的渲染(PBR)材质,并通过精心设计的布局或生成方式构建3D场景。为此,我们提出一种自动化管道,利用多模态大语言模型(MLLMs)实现材质感知的部件分割与材质检索,支持可扩展的物理合理资产生成。基于该高保真3D资产数据集,我们构建了抓取、放置及长时序任务等多样化评估任务,实现对视觉-语言-动作(VLA)模型的可扩展、可复现评估。实验表明,该基准在不同策略下与真实世界性能的平均皮尔逊相关系数达0.92。

原文摘要 · Abstract (English)

Reliable simulation evaluation of robot manipulation policies serves as a high-fidelity proxy for real-world performance. Although existing benchmarks cover a wide range of task categories, they lack visual realism, creating a large domain gap between simulation and reality. This undermines the reliability of simulation-based evaluation in predicting real-world performance. To mitigate the sim-to-real visual gap, we conduct a systematic analysis to isolate the effects of lighting and material. Our results show that these factors play a critical role in geometric reasoning and spatial grounding, yet are largely overlooked in existing benchmarks. Motivated by the analysis, we propose VISER, a visually realistic benchmark for evaluating robot manipulation in simulation. VISER features a high-fidelity dataset of over 1,000 3D assets with physically-based rendering (PBR) materials, along with 3D scenes created from these assets through curated layouts or generation. To this end, we propose an automated pipeline leveraging Multi-modal Large Language Models (MLLMs) for material-aware part segmentation and material retrieval, enabling scalable generation of physically plausible assets. Building on the high-fidelity 3D asset dataset, we construct diverse evaluation tasks, such as grasping, placing, and long-horizon tasks, enabling scalable and reproducible assessment of Vision-Language-Action (VLA) models. Our benchmark shows a strong correlation between simulation and real-world performance, achieving an average Pearson correlation coefficient of 0.92 across different policies.

机器人仿真视觉真实评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。