通过回放驱动验证,加速芯片粒架构下CPU-GPU系统集成。
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
- 用统一数据库实现仿真与仿真中的确定性波形回放。
- 支持复杂GPU工作负载和协议序列的可靠复现。
- 适合芯片粒系统预硅验证,提升集成效率与调试速度。
CPU与GPU技术的融合是现代AI与图形工作负载的关键,结合控制处理与大规模并行计算能力。随着系统向芯片粒架构演进,紧密耦合的CPU-GPU子系统在硅前验证面临巨大挑战:验证框架搭建复杂、设计规模大、并发度高、执行非确定性强,且芯片粒边界协议交互复杂,常导致集成周期长。本文针对基于ODIN集成芯片粒架构的SoC基础模块,提出一种回放驱动的验证方法,整合了CPU子系统、多个Xe GPU核心及可配置网络-on-_chip(NoC)。通过单一设计数据库,在仿真与仿真中实现确定性波形捕获与回放,可靠重现复杂GPU工作负载与协议序列。该方法显著加速调试,提升集成信心,并在单个季度内完成端到端系统启动与工作负载执行,验证了回放式验证在芯片粒系统中的可扩展有效性。
原文摘要 · Abstract (English)
Integration of CPU and GPU technologies is a key enabler for modern AI and graphics workloads, combining control-oriented processing with massive parallel compute capability. As systems evolve toward chiplet-based architectures, pre-silicon validation of tightly coupled CPU-GPU subsystems becomes increasingly challenging due to complex validation framework setup, large design scale, high concurrency, non-deterministic execution, and intricate protocol interactions at chiplet boundaries, often resulting in long integration cycles. This paper presents a replay-driven validation methodology developed during the integration of a CPU subsystem, multiple Xe GPU cores, and a configurable Network-on-Chip (NoC) within a foundational SoC building block targeting the ODIN integrated chiplet architecture. By leveraging deterministic waveform capture and replay across both simulation and emulation using a single design database, complex GPU workloads and protocol sequences can be reproduced reliably at the system level. This approach significantly accelerates debug, improves integration confidence, and enables end-to-end system boot and workload execution within a single quarter, demonstrating the effectiveness of replay-based validation as a scalable methodology for chiplet-based systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。