用AI多智能体自动生成数据中心硬件测试方案,效率提升显著。
Automated Hardware Validation Test Plan Generation for Large Scale AI Datacenter Platforms Using a Generative AI Multi-Agents Architecture

- 构建多智能体系统,从文档和物料清单自动生成结构化测试用例。
- 在两个真实平台测试中覆盖率分别提升74.2%和51.4%,编写时间从天级降至小时级。
- 输出可追溯、可复用,适合大规模硬件验证团队使用。
大规模AI数据中心平台包含数千个异构硬件组件,其验证需全面的故障注入测试计划。当前依赖人工撰写:工程师查阅自愈验证文档与物料清单,逐项列举故障模式,生成单层测试用例列表。该过程耗时、易错,且依赖经验;覆盖盲区常后期显现,溯源不明确,重复工作量大。本文提出一种生成式AI多智能体架构,可从两类标准输入自动生成结构化硬件验证测试计划:自愈验证文档(列出各可更换单元的已知故障模式及其检测与修复行为)及组件物料清单(BOM)。一个摄入智能体将异构输入归一化为统一表示;分类智能体通过上下文推理对部件进行功能域划分;生成智能体结合归一化故障模式与分类数据,补全遗漏并生成边界情况。输出符合标准格式,可直接导入内部验证软件。在两个生产平台上的评估显示,覆盖率分别提升74.2%和51.4%,编写时间由数日缩短至数小时。结果具备完全可追溯性,支持跨平台复用。人工与专家评估确认100%提取准确率,并高度接受新场景,验证该框架作为人机协同增强工具的鲁棒性。
原文摘要 · Abstract (English)
Large-scale AI datacenter platforms comprise thousands of heterogeneous hardware components whose validation requires comprehensive fault injection test plans. Today these plans are authored manually: engineers review hardware self-healing validation documents and bills of materials, enumerate failure modes per field-replaceable unit, and produce flat lists of single-layer test cases. This process is labor-intensive, error-prone, and dependent on institutional knowledge; coverage gaps surface late, traceability to source specifications is implicit, and the effort is largely repeated per platform. This paper presents a generative AI multi-agent architecture that automates the generation of structured hardware validation test plans from two canonical inputs: self-healing validation documents, which enumerate known failure modes and their detection and remediation behaviors per field-replaceable unit, and component Bills of Material. An ingestion agent normalizes heterogeneous inputs into a canonical representation; a classification agent maps components to functional domains via contextual reasoning over part descriptions and sub-category hierarchies; and a generation agent synthesizes test cases by combining normalized failure modes with domain-classified data, filling gaps and producing edge cases. The output conforms to a standardized schema for direct import into internal validation software. Evaluated on two production platforms against manual baselines, the framework achieves coverage expansions of 74.2% and 51.4%, cutting authoring from days to hours. It yields fully traceable mappings from each test case to its source specification, and its multi-agent decomposition is portable across platform generations. Automated and expert evaluations confirm 100% extraction fidelity and high acceptance of new scenarios, validating the framework as a robust human-in-the-loop force multiplier.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。