通过压缩与并行验证,让大模型广告生成更快更实用。
Efficient LLM-based Advertising via Model Compression and Parallel Verification

- 用动态分组量化+分层稀疏化降低计算开销
- 在真实广告场景中实现显著加速,质量损失可接受
- 适合追求低延迟广告生成的工业部署团队
大语言模型在广告创意生成和精准投放等场景中展现出巨大潜力,但其高推理延迟和计算成本制约了实时系统中的部署。本文提出高效生成式定向框架(Efficient Generative Targeting),融合自适应分组量化、层级自适应稀疏化和前缀树并行验证技术,在保持生成质量的前提下加速LLM推理。在两个真实广告场景的大量实验表明,该框架实现了显著提速,同时质量下降在可接受范围内,具备实际部署可行性。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown remarkable potential in advertising scenarios such as ad creative generation and targeted advertising. However, deploying LLMs in real-time advertising systems poses significant challenges due to their high inference latency and computational cost. In this paper, we propose an Efficient Generative Targeting framework that integrates adaptive group quantization, layer-adaptive hierarchical sparsification, and prefix-tree parallel verification to accelerate LLM inference while preserving generation quality. Extensive experiments on two real-world advertising scenarios demonstrate that our framework achieves significant speedup with acceptable quality degradation, making it operationally viable for practical deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。