arXiv:2601.23281cs.CV2026-01中稿 · IEEE VR 2026: GenA…被引 1

研究用户提示对XR中开放集目标检测的影响并提升鲁棒性

User Prompting Strategies and Prompt Enhancement Methods for Open-Set Object Detection in XR Environments

  • 模拟四种提示类型,测试模型在真实交互中的表现
  • 模糊提示下性能下降,增强策略使mIoU提升超55%
  • 提出适用于XR场景的提示优化方法,适合交互系统设计者

开放集目标检测(OSOD)在定位物体的同时识别并拒绝未知类别。尽管现有OSOD模型在基准测试中表现良好,但在真实用户提示下的行为仍缺乏研究。在交互式扩展现实(XR)环境中,用户生成的提示常存在模糊、信息不足或过度详细等问题。为研究提示条件下的鲁棒性,我们在真实XR图像上评估了GroundingDINO和YOLO-E两款OSOD模型,并利用视觉语言模型模拟多样化的用户提示行为。考虑四种提示类型:标准、信息不足、过度详细和语用模糊,并分析两种增强策略的效果。结果表明,两种模型在信息不足和标准提示下表现稳定,但在模糊提示下性能显著下降;过度详细提示主要影响GroundingDINO。提示增强在模糊情境下大幅提升了鲁棒性,使mIoU提升超过55%,平均置信度提升41%。基于此,我们提出了适用于XR环境的提示策略与增强方法。

原文摘要 · Abstract (English)

Open-set object detection (OSOD) localizes objects while identifying and rejecting unknown classes at inference. While recent OSOD models perform well on benchmarks, their behavior under realistic user prompting remains underexplored. In interactive XR settings, user-generated prompts are often ambiguous, underspecified, or overly detailed. To study prompt-conditioned robustness, we evaluate two OSOD models, GroundingDINO and YOLO-E, on real-world XR images and simulate diverse user prompting behaviors using vision-language models. We consider four prompt types: standard, underdetailed, overdetailed, and pragmatically ambiguous, and examine the impact of two enhancement strategies on these prompts. Results show that both models exhibit stable performance under underdetailed and standard prompts, while they suffer degradation under ambiguous prompts. Overdetailed prompts primarily affect GroundingDINO. Prompt enhancement substantially improves robustness under ambiguity, yielding gains exceeding 55% mIoU and 41% average confidence. Based on the findings, we propose several prompting strategies and prompt enhancement methods for OSOD models in XR environments.

开放集检测提示工程XR交互视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。