用张量补全技术修复毒理基因组数据库,提升预测精度。
Completion of the DrugMatrix Toxicogenomics Database using 3-Dimensional Tensors
- 保留组织、处理、转录组的三维结构,用张量方法补全数据。
- 相比传统方法,均方误差和绝对误差更低,更贴近原始分布。
- 可发现组织间关联,适用于跨物种药物研究。
我们探索使用张量补全方法完善DrugMatrix毒理基因组学数据集。假设通过保持数据的三维结构(组织、处理、转录组测量),并结合机器学习框架,能优于现有最先进方法。结果表明,新张量方法更准确地反映原始数据分布,有效捕捉器官特异性差异。相比传统的CP分解和二维矩阵分解方法,该方法在均方误差和平均绝对误差上均更低。此外,非负张量补全实现了组织间关系的揭示。研究不仅以更高精度填补了全球最大的活体毒理基因组数据库,也为未来跨物种药物研究(如从大鼠到人类)提供了有前景的方法支持。
原文摘要 · Abstract (English)
We explore applying a tensor completion approach to complete the DrugMatrix toxicogenomics dataset. Our hypothesis is that by preserving the 3-dimensional structure of the data, which comprises tissue, treatment, and transcriptomic measurements, and by leveraging a machine learning formulation, our approach will improve upon prior state-of-the-art results. Our results demonstrate that the new tensor-based method more accurately reflects the original data distribution and effectively captures organ-specific variability. The proposed tensor-based methodology achieved lower mean squared errors and mean absolute errors compared to both conventional Canonical Polyadic decomposition and 2-dimensional matrix factorization methods. In addition, our non-negative tensor completion implementation reveals relationships among tissues. Our findings not only complete the world's largest in-vivo toxicogenomics database with improved accuracy but also offer a promising methodology for future studies of drugs that may cross species barriers, for example, from rats to humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。