arXiv:2603.05220eess.IVcs.IT2026-03

用自适应采样让DNA存图可渐进读取,降低检索成本。

Adaptive Sampling for Storage of Progressive Images on DNA

  • 按分辨率分层编码图像,仅读取所需部分DNA序列。
  • 实验表明,只需10%序列即可还原低分辨率图像,节省90%读取成本。
  • 适合需要高效检索的长期图像存档场景。

传统数据存储介质寿命短,而存储需求呈指数增长,长期归档已成为核心挑战。DNA分子凭借高密度、长寿命和低能耗,成为理想的长期存储方案。然而当前技术在成本与可靠性上仍面临瓶颈,编码率与抗错能力是规模化关键。此外,不同文件的DNA片段常混合在同一寡核苷酸池中,缺乏池内随机访问能力,导致解码特定文件需全量测序,显著增加读取成本。本文提出一种高效图像DNA存储方案,基于JPEG2000的渐进解码特性,将每层分辨率编码为一组寡核苷酸,使用JPEG DNA VM编码器确保高可靠性。根据目标分辨率动态调整需测序的寡核苷酸集合与数量,结合纳米孔测序仪的自适应采样技术,在无需PCR条件下实现随机访问,大幅降低读取开销。

原文摘要 · Abstract (English)

The short lifespan of traditional data storage media, coupled with an exponential increase in storage demand, has made long-term archival a fundamental problem in the data storage industry and beyond. Consequently, researchers are looking for innovative media solutions that can store data over long time periods at a very low cost. DNA molecules, with their high density, long lifespan, and low energy needs, have emerged as a viable alternative to digital data archival. However, current DNA data storage technologies are facing challenges with respect to cost and reliability. Thus, coding rate and error robustness are critical to scale DNA storage and make it technologically and economically achievable. Moreover, the molecules of DNA that encode different files are often located in the same oligo pool. Without random access solutions at the oligo level, it is very impractical to decode a specific file from these mixed pools, as all oligos need to first be sequenced and decoded before a target file can be retrieved, which greatly deteriorates the read cost. This paper introduces a solution to efficiently encode and store images into DNA molecules, that aims at reducing the read cost necessary to retrieve a resolution-reduced version of an image. This image storage system is based on the Progressive Decoding Functionality of the JPEG2000 codec but can be adapted to any conventional progressive codec. Each resolution layer is encoded into a set of oligos using the JPEG DNA VM codec, a DNA-based coder that aims at retrieving a file with a high reliability. Depending on the desired resolution to be read, the set of oligos as well as the portion of the oligos to be sequenced and decoded are adjusted accordingly. These oligos will be selected at sequencing time, with the help of the adaptive sampling method provided by the Nanopore sequencers, making it a PCR-free random access solution.

DNA存储图像编码渐进读取自适应采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。