arXiv:2604.09657cs.CVcs.HC2026-04

用校验和向量检测老式磁带图像中的重复与变体,助力数字遗产自动修复。

Prints in the Magnetic Dust: Robust Similarity Search in Legacy Media Images Using Checksum Count Vectors

论文配图:Prints in the Magnetic Dust: Robust Similarity Search in Legacy Media Images Using Checksum Count Vectors
图 1 · 摘自论文原文
  • 基于校验和计数向量构建特征表示,适合老旧媒体数据
  • 损坏数据中识别变体准确率达58%,发现副本准确率97%
  • 适用于历史数字文物的自动化修复与知识挖掘

数字化包含计算机数据的磁性介质只是保存早期家用计算时代文物的第一步。音频磁带图像必须解码、验证、必要时修复、测试并记录。若该过程部分可自动化,志愿者便可专注于贡献背景与历史知识,而非困于技术工具。为此,我们提出一种基于校验和计数向量的特征表示方法,并评估其在大型数据存储中检测录音重复与变体的适用性。该方法在4902个已解码磁带图像上测试,对缺失高达75%记录的受损数据,变体检测准确率为58%,替代副本识别准确率为97%。这些结果标志着通过序列匹配、自动修复与知识发现实现历史数字文物全自动恢复、去重与语义整合的重要进展。

原文摘要 · Abstract (English)

Digitizing magnetic media containing computer data is only the first step towards the preservation of early home computing era artifacts. The audio tape images must be decoded, verified, repaired if necessary, tested, and documented. If parts of this process could be effectively automated, volunteers could focus on contributing contextual and historical knowledge rather than struggling with technical tools. We therefore propose a feature representation based on Checksum Count Vectors and evaluate its applicability to detecting duplicates and variants of recordings within a large data store. The approach was tested on a collection of decoded tape images (n=4902), achieving 58\% accuracy in detecting variants and 97% accuracy in identifying alternative copies, for damaged recordings with up to 75% of records missing. These results represent an important step towards fully automated pipelines for restoration, de-duplication, and semantic integration of historical digital artifacts through sequence matching, automatic repair and knowledge discovery.

数字考古特征表示老式媒体自动修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。