揭露3D大模型评测中的2D作弊问题,提出真实评估3D能力的新标准。
Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities?
- 发现3D评测可被2D视觉模型通过渲染图像轻松骗过
- 多基准测试显示视觉模型在3D任务上表现接近甚至超越3D模型
- 建议明确区分3D与1D/2D能力,避免评估混淆
本文揭示了3D大模型评测中的“2D作弊”问题:许多任务可通过将点云渲染为图像后由视觉语言模型(VLM)解决,导致对3D模型独特3D能力的评估失效。我们在多个3D LLM基准上测试了VLM的表现,并以此为参考,提出了更有效的评估原则。研究强调应明确区分3D能力与1D或2D特性,在评估3D大模型时避免混淆。代码与数据已公开于https://github.com/LLM-class-group/Revisiting-3D-LLM-Benchmarks。
原文摘要 · Abstract (English)
In this work, we identify the "2D-Cheating" problem in 3D LLM evaluation, where these tasks might be easily solved by VLMs with rendered images of point clouds, exposing ineffective evaluation of 3D LLMs' unique 3D capabilities. We test VLM performance across multiple 3D LLM benchmarks and, using this as a reference, propose principles for better assessing genuine 3D understanding. We also advocate explicitly separating 3D abilities from 1D or 2D aspects when evaluating 3D LLMs. Code and data are available at https://github.com/LLM-class-group/Revisiting-3D-LLM-Benchmarks
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。