arXiv:2508.10553cs.CL2025-08被引 2

欧洲首个远程可解释性研究平台,支持大模型分析

eDIF: A European Deep Inference Fabric for Remote Interpretability of LLM

  • 构建基于GPU的集群,通过NNsight API实现远程模型检查
  • 16位欧洲研究人员参与测试,验证平台稳定性和可用性
  • 适合从事大模型可解释性研究的学者,推动社区协作

本文开展了一项可行性研究,部署欧洲深度推理架构(eDIF),这是一种符合NDIF标准的基础设施,旨在支持大型语言模型(LLM)的机制可解释性研究。该倡议由欧洲对广泛可及的LLM可解释性基础设施的需求驱动,致力于为研究社区民主化先进模型分析能力。项目在安斯巴赫应用技术大学部署了一个基于GPU的计算集群,并与合作机构互联,通过NNsight API实现远程模型检测。一项包含来自欧洲16名研究人员的结构化试点研究评估了平台的技术性能、易用性与科学实用性。用户对GPT-2和DeepSeek-R1-70B等模型进行了激活修补、因果追踪和表示分析等干预操作。研究发现用户参与度逐步提升,平台运行稳定,远程实验能力获得积极评价。同时,该平台也开启了用户社区的建设进程。识别出的局限包括激活数据下载时间过长及执行偶发中断,已在后续发展路线图中规划改进。此项工作标志着欧洲在广泛获取LLM可解释性基础设施方面迈出重要一步,为更大范围部署、工具扩展和持续社区合作奠定了基础。

原文摘要 · Abstract (English)

This paper presents a feasibility study on the deployment of a European Deep Inference Fabric (eDIF), an NDIF-compatible infrastructure designed to support mechanistic interpretability research on large language models. The need for widespread accessibility of LLM interpretability infrastructure in Europe drives this initiative to democratize advanced model analysis capabilities for the research community. The project introduces a GPU-based cluster hosted at Ansbach University of Applied Sciences and interconnected with partner institutions, enabling remote model inspection via the NNsight API. A structured pilot study involving 16 researchers from across Europe evaluated the platform's technical performance, usability, and scientific utility. Users conducted interventions such as activation patching, causal tracing, and representation analysis on models including GPT-2 and DeepSeek-R1-70B. The study revealed a gradual increase in user engagement, stable platform performance throughout, and a positive reception of the remote experimentation capabilities. It also marked the starting point for building a user community around the platform. Identified limitations such as prolonged download durations for activation data as well as intermittent execution interruptions are addressed in the roadmap for future development. This initiative marks a significant step towards widespread accessibility of LLM interpretability infrastructure in Europe and lays the groundwork for broader deployment, expanded tooling, and sustained community collaboration in mechanistic interpretability research.

模型可解释性远程实验欧洲科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。