arXiv:2606.23118cs.CV2026-06

提出新数据集与自适应网络,提升夜间动作识别准确率

LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions

论文配图:LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions
图 1 · 摘自论文原文
  • 基于光照自适应的专家混合网络,动态增强低光视频
  • 在ELLAR上达55.13%准确率,在新数据集上达75.95%
  • 适合做夜间视觉、智能监控方向研究者参考

由于光照不足、噪声放大、动作模糊及场景多样性,低光下的人体动作识别仍具挑战。现有低光数据集普遍存在动作类别少、真实感不足或类别分布不均的问题。为此,本文提出LUMINA-26:一个包含26个动作类别的数据集,涵盖6,784段视频,由22名受试者在20个室内外场景中于自然低光条件下采集。同时提出Illumi-Net:一种光照自适应的专家混合网络,利用视频级光照线索引导自适应增强,并结合Transformer进行时空特征提取,通过专家条件化决策融合。该方法在ELLAR上达到Top-1: 55.13%、Top-5: 78.87%性能,为LUMINA-26建立75.95%(Top-1)和93.58%(Top-5)强基线,为未来低光动作识别研究提供实用基准。

原文摘要 · Abstract (English)

Low-light human action recognition remains a challenging problem due to poor illumination, amplified noise, motion ambiguity, and diverse real-world scenes. Existing low-light datasets often lack sufficient action diversity, capture realism, or balanced class distribution, limiting the development of robust models. To address this, we introduce LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions, comprising 6,784 clips across 26 action classes, recorded from 22 subjects across 20 indoor and outdoor locations under naturally occurring low-light conditions. We also propose Illumi-Net: An Illumination-Adaptive Mixture-of-Experts Network, which leverages video-level illumination cues to guide adaptive enhancement and transformer-based spatio-temporal feature extraction, with expert-conditioned decision fusion. Our method surpasses previous state-of-the-art performance on ELLAR (Top-1: 55.13%, Top-5: 78.87%) and establishes a strong baseline on LUMINA-26 (Top-1: 75.95%, Top-5: 93.58%), offering a practical benchmark for future low-light action recognition research.

动作识别低光图像视频理解自适应网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。