arXiv:2605.10831cs.LGcs.AI2026-05被引 1

让大模型分子编辑更可控且可解释,通过稀疏特征精准调节分子属性。

SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing

论文配图:SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing
图 1 · 摘自论文原文
  • 用可学习门控的稀疏自编码器分解隐藏状态,提取与属性对齐的稀疏特征。
  • 在四个模型架构上对八种属性测试,编辑成功率提升最高达42.4点。
  • 特征稀疏性支持属性调控过程的可解释分析,适合需要可控生成的研究者。

大语言模型具备强大的化学推理能力,使其成为有效的分子编辑工具。然而,属性相关的信息隐含于密集的隐藏状态中,缺乏显式控制手段:大量编辑未能改善甚至损害目标属性。为此,我们提出SLIM(Sparse Latent Interpretable Molecular editing)——一种即插即用框架,通过带可学习重要性门控的稀疏自编码器,将编辑器的隐藏状态分解为稀疏且与属性对齐的特征。在该稀疏特征空间中精确激活属性相关维度,实现属性控制,无需修改模型参数。相同的稀疏基底还支持对编辑行为的可解释分析。在MolEditRL基准上对四种模型架构和八种分子属性的实验表明,相比基线方法持续取得提升,最高改进达42.4点。

原文摘要 · Abstract (English)

Large language models possess strong chemical reasoning capabilities, making them effective molecular editors. However, property-relevant information is implicitly entangled across their dense hidden states, providing no explicit handle for property control: a substantial fraction of edits fail to improve or even degrade target properties. To address these issues, we propose SLIM (Sparse Latent Interpretable Molecular editing), a plug-and-play framework that decomposes the editor's hidden states into sparse, property-aligned features via a Sparse Autoencoder with learnable importance gates. Steering in this sparse feature space precisely activates property-relevant dimensions, improving editing success rate without modifying model parameters. The same sparse basis further supports interpretable analysis of editing behavior. Experiments on the MolEditRL benchmark across four model architectures and eight molecular properties show consistent gains over baselines, with improvements of up to 42.4 points.

分子生成大模型属性控制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。