解决音视频融合中早期决策过早的问题,提升后期表示可靠性
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning

- 引入可感知的准备度缺陷指标,定位融合瓶颈
- 通过跨层跨模态证据修正不充分的融合状态
- 适用于语音分离、事件定位等多类音视频任务
分阶段音视频编码器在层间传递融合中间状态,使后续表示依赖于早期融合状态的就绪程度。局部音视频一致性提供有效对应证据,但融合状态还需足够的跨层和跨模态支持才能可靠引导后续融合。本文通过传播感知的表示就绪度研究此问题,将过早感知承诺视为就绪度不足导致的现象,即在中间阶段同时出现局部合理性、传播影响力与支持不足。提出延迟感知承诺网络(DPC-Net),一种编码器级框架,可估计可观测的就绪度缺陷代理,定位干预敏感瓶颈,并利用跨层跨模态证据进行支持感知修正。DPC-Net保持任务特定头、损失函数、解码模块和评估协议不变,通过编码器侧干预适用于不同音视频任务。在音视频语音分离、音视频事件定位和音视频语音识别任务上的实验显示,在重建、定位和识别场景中均获得一致性能提升。对组件贡献、选择标准、反事实干预及就绪度轨迹的进一步分析验证了就绪度引导瓶颈修正的有效性。
原文摘要 · Abstract (English)
Stage-wise audio-visual encoders propagate fused intermediate states across layers, making the formation of later representations depend on the readiness of earlier fusion states. Strong local audio-visual agreement provides useful correspondence evidence, yet a fused state also needs sufficient cross-layer and cross-modal support before it can reliably guide later fusion. This paper studies this issue through propagation-aware representation readiness and formulates premature perceptual commitment as a readiness-deficiency problem, where local plausibility, propagation influence, and support insufficiency jointly appear at an intermediate stage. We propose the Delayed Perceptual Commitment Network (DPC-Net), an encoder-level framework that estimates an observable readiness-deficiency surrogate, localizes the intervention-sensitive bottleneck, and applies support-aware correction with cross-layer and cross-modal evidence. DPC-Net preserves task-specific heads, losses, decoding modules, and evaluation protocols, making it applicable to different audio-visual tasks through encoder-side intervention. Experiments on audio-visual speech separation, audio-visual event localization, and audio-visual speech recognition show consistent improvements across reconstruction, localization, and recognition regimes. Further analyses on component contribution, selection criteria, counterfactual intervention, and readiness trajectories support the effectiveness of readiness-guided bottleneck correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。