评估深度模型训练推理的能耗与碳排放,发现中等规模模型更环保
Assessing the Energy and Carbon Emissions of Neural Speaker Verification Model in Training and Inference

- 测试不同深度宽度的ResNet在VoxCeleb2上的能效表现
- 更深更宽模型增益小但耗能飙升,存在收益递减拐点
- 推荐使用ResNet-50等中等规模模型以平衡性能与碳排放
深度学习说话人验证(SV)越来越多依赖深层神经网络主干,其环境影响尚未被充分记录。本文对在VoxCeleb2数据集上训练的ResNet架构进行评估,调整深度、通道宽度和阶段分布,利用节点级传感器测量能耗与碳足迹。结果表明存在明显的收益递减点:更深层或更宽的模型仅带来微小准确率提升,但能耗急剧上升。相比之下,中等规模网络如ResNet-50及阶段集中型变体在性能与环境影响之间取得了良好权衡。这些发现为设计节能型语音验证系统提供了可操作指导。
原文摘要 · Abstract (English)
Deep-learning speaker verification (SV) increasingly relies on deep neural network backbones, whose environmental impact remains largely undocumented. In this paper, we conduct an evaluation of ResNet architectures trained on VoxCeleb2, varying depth, channel width, and stage distribution, and measure energy consumption and carbon footprint using node-level sensors. Results show a clear point of diminishing returns: deeper or wider models bring only marginal accuracy gains while energy consumption grows steeply. In contrast, mid-sized networks such as ResNet-50 and stage-concentrated variants achieve favorable trade-offs between performance and environmental impact. These findings provide actionable guidelines for designing energy-efficient SV systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。