arXiv:2510.06218cs.CVcs.AI2025-10中稿 · ICLR被引 18

首个夜间第一人称视觉理解基准,揭示光照对模型性能的显著影响。

EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark

  • 构建昼夜对齐视频数据集,提升夜间标注质量
  • 包含3658个问答对,覆盖12类问题,耗时超300小时人工标注
  • 适合研究夜间视觉、跨光照泛化模型的学者使用

现有第一人称视觉理解基准多聚焦日间场景,忽视了真实应用中不可避免的低光条件。为填补这一空白,我们提出EgoNight,首个面向夜间第一人称视觉的综合性基准,以视觉问答(VQA)为核心任务。其关键特征是引入昼夜对齐视频,利用日间数据提升夜间标注质量,并揭示不同光照下的明显性能差距。我们通过Blender渲染合成视频与真实世界拍摄相结合,确保场景与动作在视觉和时间上对齐。基于这些配对视频,构建EgoNight-VQA,采用新颖的昼增强夜自动标注引擎,并经大量人工验证进行优化,每组问答均由标注员双重检查以保证可靠性。总计包含3658个问答对,覆盖90段视频、12种多样化问题类型,人工标注工作量超过300小时。对先进多模态大模型(MLLMs)的评估显示,从白天迁移到夜间时性能显著下降,凸显了低光条件下推理的挑战。除VQA外,EgoNight还引入两个辅助任务:昼夜对应检索与夜间第一人称深度估计,进一步探索现有模型的边界。我们认为EgoNight-VQA为推动面向应用的第一人称视觉研究及开发跨光照泛化模型提供了坚实基础。代码与数据见https://github.com/dehezhang2/EgoNight。

原文摘要 · Abstract (English)

Most existing benchmarks for understanding egocentric vision focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first comprehensive benchmark for nighttime egocentric vision, with visual question answering (VQA) as the core task. A key feature of EgoNight is the introduction of day-night aligned videos, which enhance night annotation quality using the daytime data and reveal clear performance gaps between lighting conditions. To achieve this, we collect both synthetic videos rendered by Blender and real-world recordings, ensuring that scenes and actions are visually and temporally aligned. Leveraging these paired videos, we construct EgoNight-VQA, supported by a novel day-augmented night auto-labeling engine and refinement through extensive human verification. Each QA pair is double-checked by annotators for reliability. In total, EgoNight-VQA contains 3658 QA pairs across 90 videos, spanning 12 diverse QA types, with more than 300 hours of human work. Evaluations of state-of-the-art multimodal large language models (MLLMs) reveal substantial performance drops when transferring from day to night, underscoring the challenges of reasoning under low-light conditions. Beyond VQA, EgoNight also introduces two auxiliary tasks, day-night correspondence retrieval and egocentric depth estimation at night, that further explore the boundaries of existing models. We believe EgoNight-VQA provides a strong foundation for advancing application-driven egocentric vision research and for developing models that generalize across illumination domains. The code and data can be found at https://github.com/dehezhang2/EgoNight.

第一人称视觉夜间理解视觉问答跨光照

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。