首个端到端视频DNA压缩模型,用令牌直接映射碱基,提升存储效率。
From Pixels to Nucleotides: End-to-End Token-Based Video Compression for DNA Storage

- 用令牌表示视频,直接对应ATCG碱基,实现压缩与编码统一优化。
- 达到1.91比特/核苷酸的压缩率,优于传统两阶段方法。
- 适合生物存储、神经编码交叉研究者,推动视频存入DNA技术落地。
基于DNA的存储已成为应对全球数据危机的有前景方案,具备分子级密度、千年级稳定性和低维护成本。过去十年中,文本、图像和文件已成功存入DNA,但视频仍是一个开放挑战。难点不仅在于技术:有效视频DNA存储需要从头开始协同设计压缩与分子编码,这涉及两个长期独立发展的领域。本文提出HELIX,首个端到端神经网络,联合优化视频压缩与DNA编码——以往方法分两阶段处理,导致生化约束与压缩目标严重错位。关键洞察:令牌化表示天然契合DNA四碱基字母表,离散语义单元可直接映射为ATCG。我们引入TK-SCONE(Token-Kronecker Structured Constraint-Optimized Neural Encoding),通过克罗内克结构混合打破空间相关性,结合基于有限状态机(FSM)的映射确保生化约束。与两阶段方法不同,HELIX同时学习令牌分布,优化视觉质量、掩码预测能力及DNA合成效率。首次证明,学习型压缩与分子存储在令牌层面自然收敛,提示一种新范式:神经视频编解码器应从底层为生物介质设计。
原文摘要 · Abstract (English)
DNA-based storage has emerged as a promising approach to the global data crisis, offering molecular-scale density and millennial-scale stability at low maintenance cost. Over the past decade, substantial progress has been made in storing text, images, and files in DNA -- yet video remains an open challenge. The difficulty is not merely technical: effective video DNA storage requires co-designing compression and molecular encoding from the ground up, a challenge that sits at the intersection of two fields that have largely evolved independently. In this work, we present HELIX, the first end-to-end neural network jointly optimizing video compression and DNA encoding -- prior approaches treat the two stages independently, leaving biochemical constraints and compression objectives fundamentally misaligned. Our key insight: token-based representations naturally align with DNA's quaternary alphabet -- discrete semantic units map directly to ATCG bases. We introduce TK-SCONE (Token-Kronecker Structured Constraint-Optimized Neural Encoding), which achieves 1.91 bits per nucleotide through Kronecker-structured mixing that breaks spatial correlations and FSM-based mapping that guarantees biochemical constraints. Unlike two-stage approaches, HELIX learns token distributions simultaneously optimized for visual quality, prediction under masking, and DNA synthesis efficiency. This work demonstrates for the first time that learned compression and molecular storage converge naturally at token representations -- suggesting a new paradigm where neural video codecs are designed for biological substrates from the ground up.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。