arXiv:2608.05900cs.LG2026-08

研究单细胞注释中邻近细胞移除对结果的干扰,发现可不改变目标细胞而操纵注释结果。

CohortHijack: Robustness of Single Cell Annotation to Companion Cell Removal

  • 通过移除非目标细胞测试注释鲁棒性,保持目标表达谱不变
  • 多起点搜索使19.67%~24.33%的注释结果被改变,且影响小
  • 适用于评估注释工具在数据扰动下的可靠性,尤其关注细胞群体组成

许多单细胞注释工具通过邻近细胞或聚类级投票来优化初始标签。我们研究这种优化是否可在不改变目标细胞的情况下被操纵。提出CohortHijack鲁棒性审计方法:移除特定非目标细胞,同时保留目标表达谱、初始预测和训练模型。在PBMC3K和Paul15数据集上,使用逻辑回归与校准线性SVM分类器,评估随机与结构化移除策略,以及贪婪、多起点和束搜索方法。在Paul15上,结构化移除优于随机移除。多起点搜索仅移除少量细胞,却使线性SVM的24.33%和逻辑回归的19.67%的目标注释发生改变,平均附带影响低于0.4%。消融实验显示,当邻域精修机制关闭后,该效应消失。此外,在CellTypist多数投票中,独立预测未变,但融合标签在少量同伴细胞移除后发生变化。结果表明,查询群体组成是目标保持型攻击面。

原文摘要 · Abstract (English)

Many single-cell annotation tools refine an initial cell label using nearby cells or cluster-level voting. We study whether this refinement can be manipulated without changing the target cell. We introduce CohortHijack, a robustness audit that removes selected non-target cells from the query cohort while preserving the target expression profile, base prediction, and trained model. We evaluate random and structured removal methods, together with greedy, multi-start, and beam search, on PBMC3K and Paul15 using logistic regression and calibrated linear SVM classifiers. Structured removal was consistently stronger than random removal on Paul15. Multi-start search changed 24.33% of linear-SVM targets and 19.67% of logistic-regression targets while removing a small fraction of the cohort and keeping mean collateral changes below 0.4%. Ablations confirmed that the effect disappeared when neighborhood refinement was disabled. We also evaluated CellTypist majority voting, where independent predictions remained unchanged across all evaluations, but refined labels changed after small companion-cell removals. These findings identify query cohort composition as a target-preserving attack surface in single-cell annotation.

单细胞注释鲁棒性数据安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。