从生产日志中学习大模型服务的运行指纹,提升运维预测能力。
Beyond Capability Benchmarks: Learning Operational Fingerprints of LLM Cloud Services from Production Incident Metadata

- 基于结构化支持案例数据,无文本构建服务运行特征向量。
- 在3.3万+案例上验证,能准确捕捉版本与家族级运行规律。
- 适合模型上线评估、故障预警和跨模型运维分析场景。
托管式大模型服务已广泛应用于真实生产系统,但模型选型与服务规划仍严重依赖能力基准测试,这些测试难以揭示部署后的运行行为。本文提出运营嵌入(OpEmbed)框架,通过结构化且隐私保护的支持案例元数据,无需使用案例文本即可学习大模型云服务的紧凑运行指纹。OpEmbed将模型-时间窗口聚合为八通道运行签名,并通过时序对比学习、跨视图重建和生成序数正则化学习低维表示。在谷歌云平台覆盖七类大模型、26个月、超3.3万条生产支持案例上评估,OpEmbed成功恢复可解释的家族与版本级结构,相比非学习基线显著提升留一模型外的运行预测性能,在早期数据有限条件下依然有效,并支持跨模型故障类型迁移。本文还总结了工具在模型接入、支持就绪评估和运行监控中的实践经验。
原文摘要 · Abstract (English)
Managed LLM services are now part of real production systems, but model selection and service planning still rely heavily on capability benchmarks that reveal little about operational behavior after deployment. We present Operational Embedding (OpEmbed), a framework for learning compact operational fingerprints of LLM cloud services from structured, privacy-preserving support-case metadata, without using case text. OpEmbed aggregates model--time windows into an eight-channel operational signature and learns a low-dimensional representation via temporal contrastive learning, cross-view reconstruction, and generational-ordinality regularization. Evaluated on more than 33,000 production support cases spanning seven LLM families over 26 months at Google Cloud, OpEmbed recovers interpretable family- and version-level structure, improves leave-one-model-out operational forecasting over non-learned baselines, remains useful under limited early-window data, and supports cross-model fault-type transfer. We report the practical lessons learned from building and evaluating this tool for model onboarding, support readiness assessment, and operational monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。