arXiv:2607.06839cs.LGcs.CV2026-07被引 19

构建超大规模神经网络多样性数据集,支持跨模态与真实设备验证。

LEMUR 2: Unlocking Neural Network Diversity for AI

论文配图:LEMUR 2: Unlocking Neural Network Diversity for AI
图 1 · 摘自论文原文
  • 通过代码变异、遗传算法与LLM生成1.4万+独特模型架构。
  • 覆盖75万条训练记录,涵盖图像、文本、多模态任务性能数据。
  • 支持移动端与虚拟现实设备部署,实现真实硬件性能评估。

现有神经网络架构搜索基准(如NAS-Bench、NATS-Bench)仅覆盖狭窄、任务特定的架构空间,缺乏跨领域和部署感知的评估。LEMUR 2提出一个大规模、可扩展的框架,统一生成、评估与部署流程,以解锁神经网络多样性。该框架包含超过14,000个独立架构及超过750,000条结构化训练记录,涵盖模型性能、超参数与任务结果。这些模型通过基于抽象语法树(AST)的代码变异、遗传与强化学习演化、分形架构生成以及大语言模型(LLM)引导合成生成,其中使用检索增强系统NN-RAG,从900多个公开PyTorch模块中提取并应用架构模式。LEMUR 2还集成NN-VR与NN-Lite流水线,实现异构移动与Unity VR平台上的自动化部署与延迟基准测试,提供真实设备性能元数据。其涵盖多模态任务,包括图像描述、文生图合成与语言建模,支持架构迁移性的跨领域分析。通过关联多样化架构、任务与部署数据,为大语言模型微调及大规模、跨平台实证验证提供数据基础。该数据集定义了可复现、数据驱动的AI设计新范式,推动大语言模型驱动的AutoML与跨模态、跨硬件的架构泛化发展。

原文摘要 · Abstract (English)

Existing NAS benchmarks (e.g., NAS-Bench, NATS-Bench) cover only narrow, task-specific regions of the architectural design space and lack cross-domain or deployment-aware evaluation. LEMUR 2 introduces a large-scale, extensible framework unifying generative, evaluative, and deployment pipelines to unlock neural-network diversity. It comprises over 14,000 distinct architectures and more than 750,000 structured training records documenting model performance, hyperparameters, and task outcomes. These models were produced through AST-based code mutation, genetic and reinforcement-learning evolution, generation of fractal architectures, and synthesis guided by a Large Language Model (LLM). This includes deep models generated with the retrieval-augmented system NN-RAG, which derived and used architectural motifs from over 900 PyTorch modules extracted from public repositories. LEMUR 2 further employs NN-VR and NN-Lite pipelines for automated deployment and latency benchmarking on heterogeneous mobile and Unity-based VR platforms, providing real-device performance metadata. It spans multimodal tasks, image captioning, text-to-image synthesis, and language modeling, supporting cross-domain analysis of architectural transferability. By linking diverse architectures, tasks, and deployment data, LEMUR 2 provides the data foundation for LLM fine-tuning and coupling diverse architectural origins with large-scale, cross-platform empirical validation. This dataset defines a new basis for reproducible and data-driven AI design, advancing the emerging paradigm of LLM-driven AutoML and architectural generalization across modalities and hardware.

神经网络架构AutoML大模型跨平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。