arXiv:2603.01396cs.AIcs.CE2026-03被引 2

自动应对单细胞扰动数据的语义与分布差异,实现无需人工干预的建模。

HarmonyCell: Automating Single-Cell Perturbation Modeling under Semantic and Distribution Shifts

  • 用大模型自动统一不同数据集的元数据语义,避免手动映射。
  • 通过自适应搜索生成最优模型结构,在分布偏移下仍保持高精度。
  • 适合需要跨数据集自动化建模的研究者,尤其在数据异构场景中。

单细胞扰动研究面临双重异质性瓶颈:(i) 语义异质性——相同生物学概念在不同数据集中以不兼容的元数据模式编码;(ii) 统计异质性——生物变异导致的数据分布偏移,需依赖数据特定归纳偏置。我们提出 HarmonyCell,一个端到端智能体框架,通过专用机制分别解决两类问题:由大语言模型驱动的语义统一体可自主将异构元数据映射为标准接口,无需人工干预;基于分层动作空间的自适应蒙特卡洛树搜索引擎,可合成具备最优统计归纳偏置的模型架构以应对分布偏移。在多种扰动任务中评估,面对语义与分布双重偏移,HarmonyCell 在异构输入数据集上实现了95%的有效执行率(普通智能体为0%),且在严格的分布外测试中达到或超过专家设计基线性能。该双轨协同机制实现了无需数据集定制的可扩展自动虚拟细胞建模。

原文摘要 · Abstract (English)

Single-cell perturbation studies face dual heterogeneity bottlenecks: (i) semantic heterogeneity--identical biological concepts encoded under incompatible metadata schemas across datasets; and (ii) statistical heterogeneity--distribution shifts from biological variation demanding dataset-specific inductive biases. We propose HarmonyCell, an end-to-end agent framework resolving each challenge through a dedicated mechanism: an LLM-driven Semantic Unifier autonomously maps disparate metadata into a canonical interface without manual intervention; and an adaptive Monte Carlo Tree Search engine operates over a hierarchical action space to synthesize architectures with optimal statistical inductive biases for distribution shifts. Evaluated across diverse perturbation tasks under both semantic and distribution shifts, HarmonyCell achieves a 95% valid execution rate on heterogeneous input datasets (versus 0% for general agents) while matching or even exceeding expert-designed baselines in rigorous out-of-distribution evaluations. This dual-track orchestration enables scalable automatic virtual cell modeling without dataset-specific engineering.

单细胞自动化建模分布偏移语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。