模型预测声学参数的高精度,往往来自评估方式而非模型本身。
What does the model actually see? Evaluation protocols and input availability in data-driven prediction of room acoustic parameters

- 用不同划分数据和输入条件测试模型,发现评估协议影响结果
- 真实部署场景下模型准确率降至0.09-0.57,远低于报告值
- 模型实际依赖位置指纹,非通用声学信息,适合已测位置插值
机器学习模型被广泛用于从稀疏测量中预测ISO 3382-1房间声学参数,报告的决定系数常高于0.85。本文表明,这些高分多由评估协议造成而非模型性能。在264座会议厅与180座音乐厅的多条件测量实验中,对三类模型在因子设计的评估协议下进行测试:验证集按行划分或按接收器位置分组,输入特征包含实测值或仅限源-接收器几何与环境状态。按行划分并使用实测输入时,核心参数平均R²达0.81;而按位置分组且限制输入时,准确率降至0.09–0.57,参数类别排序改变。以目标脉冲响应为输入的混合CNN实际上利用其作为位置指纹,而非可迁移声学信息;训练阶段仅访问信号对任何参数均无增益,包括混响时间。在符合部署条件的协议下,随机森林、混合CNN与反距离加权间的差距仅为原协议下同一模型差异的十分之一;学习模型在声强与混响时间上仍具真实优势,原始高精度再次出现在已测位置的条件插值任务中(带内均值0.80–0.88),此任务具有实际应用价值。
原文摘要 · Abstract (English)
Machine-learnt models are increasingly used to predict ISO 3382-1 room acoustic parameters from sparse measurements, with reported coefficients of determination frequently above 0.85. This paper shows that such figures are often determined by the evaluation protocol rather than by the model. Using a multi-condition measurement campaign in a 264-seat conference hall and a 180-seat concert hall, three model families were evaluated under a factorial protocol ablation: validation splits either row-based or grouped by receiver position, and input features either including measured-at-test quantities or restricted to source-receiver geometry and environmental state. Row-based splits with measured-at-test inputs reproduce the high reported accuracies (mean $R^2$ 0.81 for the core parameters); grouping the splits by position and restricting inputs to information available at an unmeasured position reduces these to 0.09-0.57, reordering the apparent difficulty of parameter classes. A hybrid CNN evaluated with the target's own impulse response as input is shown to exploit it as a position fingerprint rather than as transferable acoustic information; training-only signal access yields no gain for any parameter tested, including reverberation time. Under the deployment-consistent protocol, the spread between Random Forest, the hybrid CNN, and inverse-distance weighting is an order of magnitude smaller than the spread between protocols for a fixed model; the learnt models retain a genuine advantage for sound strength and reverberation time, and the high accuracy of the original pipelines re-emerges as condition interpolation at measured positions (band means 0.80-0.88), a distinct and operationally useful task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。