arXiv:2605.26672cs.MMcs.SD2026-05

用事件相机生成带情绪的语音,解决传统摄像头模糊问题

Can We Hear from Events? Generating Speech from Event Camera

论文配图:Can We Hear from Events? Generating Speech from Event Camera
图 1 · 摘自论文原文
  • 用微秒级精确的事件数据替代传统摄像头,对齐语音动态
  • 在真实事件数据上表现优于现有模型,保留细微情感特征
  • 适合研究神经形态感知与多模态语音生成的学者

传统基于RGB的语音生成因固定曝光时间导致时间粒度不匹配,不可避免地模糊了高频率发音瞬变,影响情感表达。为突破此瓶颈,我们提出EventSpeech,首个利用类脑事件数据进行文本条件语音生成的框架。事件数据具有微秒级精度,天然契合声波动态。模型采用专用事件编码器处理稀疏事件信号,结合多尺度音频编码器及分层小波上下文模块(HWC),通过双向对齐机制实现语言内容、视觉动态与密集声学特征的无缝同步。我们构建了首个基准数据集EVT-SPK,包含大规模合成数据与真实事件硬件采集的录音。大量实验表明,EventSpeech显著优于现有基线,在保留精细情感特征和抗运动模糊方面表现优异,建立了多模态语音生成新范式。代码与演示见https://xrfang-0102.github.io/EventSpeechWeb/。

原文摘要 · Abstract (English)

Traditional RGB-based speech generation faces Temporal Granularity Mismatch since fixed camera exposure times inevitably blur the high-frequency articulatory transients essential for rendering emotional speech. To break this ceiling, we propose EventSpeech as a novel text-conditioned framework pioneering the use of neuromorphic events for expressive speech generation, since these microsecond-precise events naturally align with acoustic waveform dynamics. Our architecture integrates a dedicated Event Encoder to model sparse neuromorphic events alongside a multi-scale Audio Encoder featuring a Hierarchical Wavelet Contextualizer (HWC). A bidirectional alignment mechanism seamlessly synchronizes linguistic content and visual dynamics with dense acoustic features. Furthermore, we construct EVT-SPK as the first benchmark comprising large-scale synthetic data and real-world recordings from specialized neuromorphic hardware. Extensive evaluations demonstrate that EventSpeech significantly outperforms current baselines by preserving fine-grained emotions and resisting motion blur to establish a new paradigm for multimodal speech generation. Code and demo are available at https://xrfang-0102.github.io/EventSpeechWeb/.

语音生成事件相机类脑计算多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。