对比知识蒸馏在机器翻译中的质量与能耗,发现部署量影响碳足迹
Is Knowledge Distillation Actually Greener? A Case Study in Machine Translation
- 用生命周期评估工具衡量蒸馏全过程的计算开销
- 部署量需达数万次才能抵消蒸馏的环境成本
- 适合关注绿色AI和模型效率的研究者参考
知识蒸馏(KD)常用于将大型教师模型压缩为小型学生模型。在机器翻译中,现有评估主要关注翻译质量和推理效率,却未综合考虑蒸馏系统生产与部署的环境成本。本文基于定制化MT模型与大语言模型,结合机器学习生命周期评估工具,全面评估典型KD方法的翻译质量与计算开销。关键发现:实现碳中和所需的部署量依赖于服务场景,且在批处理条件下可相差数个数量级。研究提供可在质量与算力约束下选择、开发与评估KD方法的具体建议。
原文摘要 · Abstract (English)
Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student. In machine translation, KD is commonly evaluated through translation quality and inference efficiency, without jointly accounting for the environmental costs of producing and deploying the distilled system. We evaluate representative KD methods both on bespoke MT models and LLMs, by considering both translation quality and computational cost, using the Machine Learning Life Cycle Assessment tool, which accounts for costs throughout the KD model life cycle. Our key finding is that the deployment volume required to amortize KD is serving-dependent and can shift by several orders of magnitude under batching. We include actionable guidance for selecting, developing, and evaluating KD methods under quality and compute-induced constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。