arXiv:2608.24065cs.LGcs.AI2026-08

用可解释的神经电路控制数据生成,让好数据更可控、更高效。

Mechanistic Circuit Identification for Controllable Data Generation

论文配图:Mechanistic Circuit Identification for Controllable Data Generation
图 1 · 摘自论文原文
  • 基于模型内部电路识别数据质量三维度:可学性、挑战性、对齐度。
  • 通过电路操控生成目标数据,多样性比提示词方法提升23%以上。
  • 适合需要精准调控数据质量的研究者,尤其在问答任务中表现突出。

尽管近期数据合成技术致力于构建高质量数据集,但多数生成流程仍依赖启发式提示控制。这种黑箱范式难以揭示样本与模型学习动态的交互机制。为此,我们提出一种基于电路的框架,将训练动态的数据估值与机械可解释性(MI)相结合。具体地,我们从可学性、挑战性和对齐度三个互补维度定义数据质量。首先,我们发现模型内部专门电路会因果影响这些质量信号;随后,突破启发式提示,利用这些电路作为可调控接口,主动引导生成符合特定质量目标的数据。基于此能力,我们提出阶段感知的机械调度方法(SAMS),根据模型优化进程动态调度电路驱动的数据。在多项选择题问答任务上的实验表明,该方法生成的数据具有更高多样性,且相比提示基线显著提升下游性能与校准效果。本工作建立了可解释数据生成的白盒范式,首次将机械可解释性作为实际可控接口,而非仅分析工具。

原文摘要 · Abstract (English)

While recent advances in data synthesis aim to curate high-quality datasets, most generation pipelines still rely on heuristic prompt-based control. This black-box paradigm provides limited insight into how individual samples interact with a model's underlying learning dynamics. To bridge this gap, we propose a circuit-grounded framework that connects training-dynamics-based data valuation with mechanistic interpretability (MI). Specifically, we conceptualize data quality along three complementary utility axes, learnability, challenge, and alignment. First, we uncover specialized model-internal circuits that causally govern these utility signals. Then, moving beyond heuristic prompting toward mechanistic control, we leverage these circuits as controllable interfaces, actively steering generation to produce utility-targeted data. Building on this capability, we introduce SAMS (Stage-Aware Mechanistic Scheduling), which schedules circuit-steered data according to the model's evolving optimization needs. Experiments on multiple-choice QA tasks demonstrate that our approach yields precisely controlled data with greater diversity than prompt-based baselines, consistently improving downstream performance and calibration. Ultimately, this work establishes a principled white-box paradigm for interpretable data generation, pioneering the use of MI not just as an analytical tool, but as a practical, controllable interface.

数据生成可解释性神经电路可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。