用张量分解生成多维数据,大幅降低模拟成本。
Efficiently Generating Multidimensional Calorimeter Data with Tensor Decomposition Parameterization
- 用小规模张量因子代替完整张量,减少模型输出和参数
- 实验表明生成数据仍具实用性,显著降低资源消耗
- 适合需要高效生成复杂多维模拟数据的科研人员
生成大规模复杂仿真数据往往耗时且耗费资源。尤其在实验成本高昂的情况下,使用合成数据进行下游任务愈发合理。近年来,生成对抗网络或扩散模型等生成式机器学习方法被广泛采用。本文在这些生成模型中引入内部张量分解,进一步降低计算成本。针对多维数据(即张量),我们不直接生成完整张量,而是生成其更小的张量因子,从而显著减少模型输出量与总参数量。实验表明,该方法在大幅降低生成成本的同时,仍能保持生成数据的有效性。张量分解具有提升生成模型效率的潜力,尤其适用于多维数据生成任务。
原文摘要 · Abstract (English)
Producing large complex simulation datasets can often be a time and resource consuming task. Especially when these experiments are very expensive, it is becoming more reasonable to generate synthetic data for downstream tasks. Recently, these methods may include using generative machine learning models such as Generative Adversarial Networks or diffusion models. As these generative models improve efficiency in producing useful data, we introduce an internal tensor decomposition to these generative models to even further reduce costs. More specifically, for multidimensional data, or tensors, we generate the smaller tensor factors instead of the full tensor, in order to significantly reduce the model's output and overall parameters. This reduces the costs of generating complex simulation data, and our experiments show the generated data remains useful. As a result, tensor decomposition has the potential to improve efficiency in generative models, especially when generating multidimensional data, or tensors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。