让蛋白质编辑可解释:分开控制结构与功能属性。
DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing
- 用双语监督分离蛋白的结构和功能表征。
- 多属性编辑任务中同时成功率达61.7%。
- 适合需要精准调控蛋白特性的研究者。
我们提出DisProtEdit,一种可控的蛋白质编辑框架,通过双通道自然语言监督学习结构与功能属性的解耦表征。不同于依赖整体嵌入的已有方法,DisProtEdit显式分离语义因子,实现模块化与可解释性控制。为此,我们构建了SwissProtDis数据集,其中每个蛋白序列配有结构与功能两段文本描述,由大语言模型自动分解。DisProtEdit通过对齐与均匀性目标对齐蛋白与文本嵌入,并使用解耦损失促进结构与功能语义的独立性。推理时,通过修改一个或两个文本输入并从更新后的隐空间解码完成蛋白质编辑。在蛋白编辑与表征学习基准测试中,DisProtEdit表现优于现有方法,且具备更强可解释性与可控性。在新构建的多属性编辑基准上,模型达到最高61.7%的双属性成功率,验证其协同编辑能力。
原文摘要 · Abstract (English)
We introduce DisProtEdit, a controllable protein editing framework that leverages dual-channel natural language supervision to learn disentangled representations of structural and functional properties. Unlike prior approaches that rely on joint holistic embeddings, DisProtEdit explicitly separates semantic factors, enabling modular and interpretable control. To support this, we construct SwissProtDis, a large-scale multimodal dataset where each protein sequence is paired with two textual descriptions, one for structure and one for function, automatically decomposed using a large language model. DisProtEdit aligns protein and text embeddings using alignment and uniformity objectives, while a disentanglement loss promotes independence between structural and functional semantics. At inference time, protein editing is performed by modifying one or both text inputs and decoding from the updated latent representation. Experiments on protein editing and representation learning benchmarks demonstrate that DisProtEdit performs competitively with existing methods while providing improved interpretability and controllability. On a newly constructed multi-attribute editing benchmark, the model achieves a both-hit success rate of up to 61.7%, highlighting its effectiveness in coordinating simultaneous structural and functional edits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。