跨数据集验证发现,婴儿哭声模型迁移效果差,但统一标签体系可显著提升性能。
How Far Do Foundation Models Transfer to Infant Signals? A Cross-Dataset Transfer Audit with a Unified Need Ontology

- 构建统一需求语义体系,对四个哭声数据集进行多层级去重审计。
- 跨数据集迁移平均为负,但噪声最大数据集内迁移效果持续正向。
- 统一标签后联合训练比直接拼接数据提升最高37个F1点,适合小样本研究者。
公开的婴儿哭声数据集规模小、标注不兼容,且通常单个数据集独立评估。本文通过多层级去重审计(字节级、嵌入级去重及同数据集近似重复检测),在统一五类需求语义体系下,测试四个冻结编码器与手工基线模型。审计揭示:单数据集评估掩盖了真实性能差异——同一编码器域内宏F1波动达0.57-0.80;跨数据集迁移平均为负(负迁移率0.19-0.35,30个方向中有18个显著,BH-FDR校正);存在349组内容相同的音频片段在不同数据集中标签冲突。然而,审计也指出可行路径:在匹配训练量和移除近似重复后,向最嘈杂数据集迁移始终正向有效。冻结探针在少量标签下饱和,而稳定微调在全标签时更优;领域自适应预训练在5-10样本下显著优于稳定微调(1样本优势对优化种子敏感),但在50样本及以上无显著优势。在二分类共享标签设置中,语义映射联合训练在所有编码器-目标组合上表现最佳,而直接合并未映射标签导致最高37点F1下降。论文发布语义本体、映射代码与审计流水线,使不兼容哭声数据集可整合为联合训练资源。
原文摘要 · Abstract (English)
Public infant cry corpora are small, label-incompatible, and almost always evaluated one corpus at a time. We ask what this practice hides and what fixes it. Across four cry corpora screened by a multi-level leakage audit (byte-level and embedding-level deduplication plus a within-corpus train-test near-duplicate audit), we probe four frozen encoders and a handcrafted baseline under a unified five-class need ontology and shared task formulations. The audit exposes what single-corpus evaluation conceals: within-domain macro-F1 swings by 0.57-0.80 for the same encoder, cross-corpus transfer is negative on average (negative-transfer ratio 0.19-0.35, significant in 18 of 30 directed cells, BH-FDR), and 349 content-identical clip groups carry conflicting metadata labels across corpus distributions. The same audit, however, reveals a consistent way forward. Transfer into the noisiest corpus is consistently positive in effect size at matched training size and after near-duplicate removal, offering a practical recipe for small, noisy corpora. Frozen probes saturate at modest label budgets, while stabilized fine-tuning wins with full labels; domain-adaptive pretraining significantly beats stabilized fine-tuning at 5-10-shot (the 1-shot advantage is not robust to optimization-seed variance) but shows no significant advantage at 50-shot or beyond. In the tested binary, shared-label settings, ontology-mapped joint training wins in all four encoder-by-target combinations, whereas naively merging unmapped labels costs up to 37 F1 points. We release the ontology, mapping code, and audit pipeline, turning incompatible cry corpora into a usable joint-training resource.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。