用事件相机捕捉唇动细节,跨场景识别更准更稳。
NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition

- 基于事件流设计时序-结构双增强模块,精准建模唇部运动动态。
- 在未知视角下识别准确率超71%,低光条件下近76%,显著优于旧方法。
- 专为跨场景泛化设计,适合需要鲁棒视觉说话人识别的系统应用。
基于唇动的视觉说话人识别提供了一种无声、无需双手、行为驱动的生物特征方案,在声学线索缺失时依然有效。与依赖外观的旧方法不同,唇动蕴含由一致发音模式和肌肉协同带来的个体化行为动态,具有环境变化下的天然稳定性。但传统帧相机因运动模糊和动态范围低,难以捕捉此类精细动态。为此,我们提出NeuroLip,一种事件驱动的时空学习框架,采用严格且实用的跨场景协议:仅在单一受控条件下训练,却需在未见视点与光照条件下实现泛化。NeuroLip包含三个核心组件:1)具自适应事件加权的时序感知体素编码模块;2)结构感知空间增强器,通过抑制噪声并保留垂直运动结构信息来放大判别性行为模式;3)极性一致性正则化机制,保留事件极性所编码的运动方向信息。为系统评估,我们构建了DVSpeaker数据集,涵盖50名受试者在四种不同视点与光照条件下的事件数据。大量实验表明,NeuroLip在匹配场景下达到近乎完美的准确率,并在未见视点下保持超71%准确率,低光条件下接近76%,相比现有方法至少提升8.54%。代码与数据集已公开于https://github.com/JiuZeongit/NeuroLip。
原文摘要 · Abstract (English)
Visual speaker recognition based on lip motion offers a silent, hands-free, and behavior-driven biometric solution that remains effective even when acoustic cues are unavailable. Compared to traditional methods that rely heavily on appearance-dependent representations, lip motion encodes subject-specific behavioral dynamics driven by consistent articulation patterns and muscle coordination, offering inherent stability across environmental changes. However, capturing these robust, fine-grained dynamics is challenging for conventional frame-based cameras due to motion blur and low dynamic range. To exploit the intrinsic stability of lip motion and address these sensing limitations, we propose NeuroLip, an event-based framework that captures fine-grained lip dynamics under a strict yet practical cross-scene protocol: training is performed under a single controlled condition, while recognition must generalize to unseen viewing and lighting conditions. NeuroLip features a 1) Temporal-aware Voxel Encoding module with adaptive event weighting, 2) Structure-aware Spatial Enhancer that amplifies discriminative behavioral patterns by suppressing noise while preserving vertically structured motion information, and 3) Polarity Consistency Regularization mechanism to retain motion-direction cues encoded in event polarities. To facilitate systematic evaluation, we introduce DVSpeaker, a comprehensive event-based lip-motion dataset comprising 50 subjects recorded under four distinct viewpoint and illumination scenarios. Extensive experiments demonstrate that NeuroLip achieves near-perfect matched-scene accuracy and robust cross-scene generalization, attaining over 71% accuracy on unseen viewpoints and nearly 76% under low-light conditions, outperforming representative existing methods by at least 8.54%. The dataset and code are publicly available at https://github.com/JiuZeongit/NeuroLip.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。