用小波表示法直接从实验轨迹中高效学习异常扩散机制
Data-Efficient Learning of Anomalous Diffusion with Wavelet Representations: Enabling Direct Learning from Experimental Trajectories
- 用六种小波族处理轨迹生成小波模量谱,构建统一表示
- 仅需1200条实验轨迹即超越用百万条模拟数据训练的模型
- 适合缺乏大量模拟数据的生物物理实验场景
机器学习已成为分析异常扩散轨迹的强大工具,但现有方法多基于大规模模拟数据训练。而真实实验轨迹(如单粒子追踪SPT)数量稀少且与理想模型差异显著,导致性能下降甚至失效。为此,我们提出一种基于小波的异常扩散表示方法,可直接从实验记录中实现数据高效学习。该表示通过将六种互补小波族应用于每条轨迹,并融合所得小波模量标量图构建。在andi-datasets基准上,该方法在仅1000条训练轨迹时即优于传统特征与轨迹基方法,且在大样本下仍具优势。进一步应用于荧光微粒在F-肌动蛋白网络中的SPT实验轨迹,其在扩散指数回归与网格尺寸分类任务中均优于现有方法。尤其在预测实验轨迹扩散指数时,使用1200条实验数据训练的模型误差显著低于纯用10⁶条模拟数据训练的顶尖深度学习模型。这种数据效率归因于小波谱中出现的区分性尺度指纹,能解耦底层扩散机制。
原文摘要 · Abstract (English)
Machine learning (ML) has become a versatile tool for analyzing anomalous diffusion trajectories, yet most existing pipelines are trained on large collections of simulated data. In contrast, experimental trajectories, such as those from single-particle tracking (SPT), are typically scarce and may differ substantially from the idealized models used for simulation, leading to degradation or even breakdown of performance when ML methods are applied to real data. To address this mismatch, we introduce a wavelet-based representation of anomalous diffusion that enables data-efficient learning directly from experimental recordings. This representation is constructed by applying six complementary wavelet families to each trajectory and combining the resulting wavelet modulus scalograms. We first evaluate the wavelet representation on simulated trajectories from the andi-datasets benchmark, where it clearly outperforms both feature-based and trajectory-based methods with as few as 1000 training trajectories and still retains an advantage on large training sets. We then use this representation to learn directly from experimental SPT trajectories of fluorescent beads diffusing in F-actin networks, where the wavelet representation remains superior to existing alternatives for both diffusion-exponent regression and mesh-size classification. In particular, when predicting the diffusion exponents of experimental trajectories, a model trained on 1200 experimental tracks using the wavelet representation achieves significantly lower errors than state-of-the-art deep learning models trained purely on $10^6$ simulated trajectories. We associate this data efficiency with the emergence of distinct scale fingerprints disentangling underlying diffusion mechanisms in the wavelet spectra.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。