大模型越安全,耗能可能越高,60倍差距令人警醒
Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs
- 用临床安全分与碳足迹数据对比47种模型配置
- 安全提升2.61分,能耗飙升60倍,非线性关系显著
- 额外计算未必更安全,动态选型或更可持续
在心理健康领域部署大语言模型时,临床安全与环境成本之间的关系备受关注。本文通过结合K-Bench临床安全评分与EcoLogits生命周期评估,分析了47种模型配置在能源消耗、碳排放、水耗和矿物耗竭四个维度的表现。结果显示,在安全分数较高的区间存在非线性权衡:临床安全得分每提高2.61个百分点,每百万输出词元的估算能耗约增加60倍。逐行分析表明,额外的推理阶段计算并未持续提升安全性,某些配置甚至导致安全分下降。这表明仅依赖更大模型或更多推理算力并非提升治疗类AI系统安全性的高效策略。我们讨论了可持续部署的启示,并提出动态模型选择(如模型级联)作为在高风险场景中兼顾性能与环保的潜在路径。
原文摘要 · Abstract (English)
The deployment of large language models (LLMs) in mental health contexts raises questions about the relationship between clinical safety and environmental cost. In this paper, we examine this relationship by combining K-Bench clinical safety scores with EcoLogits life-cycle assessment estimates across 47 supported model configurations. We evaluate model performance and environmental impact across four dimensions: energy use, carbon emissions, water consumption, and abiotic depletion. The results indicate a non-linear trade-off at the upper end of the safety distribution: a 2.61 percentage-point increase in clinical safety score corresponded to an approximately 60-fold increase in estimated energy use per million output tokens. Row-level analyses further suggest that additional test-time compute did not consistently improve clinical safety and, in some configurations, was associated with lower clinical safety scores. These findings suggest that relying solely on larger models or additional inference-time computation may be an inefficient strategy for improving safety in therapeutic AI systems. We discuss the implications for sustainable deployment and highlight dynamic model selection, including model cascading, as a potential approach for reducing environmental impact while preserving clinical performance in higher-risk cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。