arXiv:2507.19419cs.CL2025-07EMNLP

让大模型训练数据的编辑与分析变得简单高效。

TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability

  • 提供可视化界面和模块化后端,支持数据搜索、查看、导出等操作。
  • 无需修改训练代码即可结构化编辑预训练数据,提升调试效率。
  • 适配GPT-NeoX等主流框架,助力研究人员快速验证数据影响。

理解预训练过程中训练数据与模型行为的关系至关重要,但现有工作流程繁琐、分散,难以被研究者使用。我们提出TokenSmith,一个开源库,用于Megatron风格预训练框架(如GPT-NeoX、Megatron、NVIDIA NeMo)中数据集的交互式编辑、检查与分析。TokenSmith支持搜索、查看、导入、导出、检查、采样等多种操作,通过简洁用户界面和模块化后端实现。它可在不修改训练代码的前提下实现预训练数据的结构化编辑,简化数据调试、验证与实验。TokenSmith设计为即插即用组件,可无缝集成至现有大模型预训练流程,推动生产级数据工具普及。项目托管于GitHub,附带文档、教程及演示视频(YouTube可访问)。

原文摘要 · Abstract (English)

Understanding the relationship between training data and model behavior during pretraining is crucial, but existing workflows make this process cumbersome, fragmented, and often inaccessible to researchers. We present TokenSmith, an open-source library for interactive editing, inspection, and analysis of datasets used in Megatron-style pretraining frameworks such as GPT-NeoX, Megatron, and NVIDIA NeMo. TokenSmith supports a wide range of operations including searching, viewing, ingesting, exporting, inspecting, and sampling data, all accessible through a simple user interface and a modular backend. It also enables structured editing of pretraining data without requiring changes to training code, simplifying dataset debugging, validation, and experimentation. TokenSmith is designed as a plug-and-play addition to existing large language model pretraining workflows, thereby democratizing access to production-grade dataset tooling. TokenSmith is hosted on GitHub, with accompanying documentation, tutorials, and a demonstration video (available on YouTube).

数据编辑大模型训练工具库可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。