用分层语音标签器实现家庭环境下的婴儿音频精准理解
Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

- 基于微调的Whisper与轻量级目标说话人感知模型,支持长时序推理
- 引入分层语音标记机制,实现跨层级帧级预测与时间一致性增强
- 通过共享层级令牌与家庭特异性偏移设计,降低家庭偏差,提升泛化能力
近期在模型设计和自监督音频表征方面的进展提升了语音与音频理解能力,但针对婴儿为中心的自然录音仍面临标注数据有限、信噪比低以及跨家庭域偏移等挑战。本文提出一种家庭条件化的多层级音频标签器,结合LoRA微调的Whisper编码器与轻量级目标说话人感知Transformer,实现长上下文推理与跨层级帧级预测。为提升时间连贯性,引入简单的序列级平滑损失;为增强跨家庭鲁棒性,提出因子化说话人令牌设计,包含共享层级令牌与学习得到的家庭特异性偏移,有效减少家庭偏差,促进可泛化表示。该方法实现了对家庭环境中全天候音频记录的高效且有效的婴儿中心音频标记。
原文摘要 · Abstract (English)
Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。