首个古代植物种子图像数据集及分类模型,助力考古研究智能化
Towards Ancient Plant Seed Classification: A Benchmark Dataset and Baseline Model
- 构建首个古代植物种子图像数据集,含8340张图片
- 提出APSNet模型,准确率达90.5%,超越现有方法
- 适合考古、计算机交叉研究者参考
理解古代社会的饮食偏好及其在时空上的演变,对揭示人地关系至关重要。种子作为重要考古遗存,是植物考古研究的核心对象。然而传统研究高度依赖专家经验,效率低下。尽管智能分析在考古学其他领域取得进展,但在植物考古,尤其是古代植物种子分类方面仍存在数据与方法的空白。为此,我们构建了首个古代植物种子图像分类(APS)数据集,包含来自中国18个考古遗址的17个属或种级别的种子图像共8,340张。同时,我们设计了专为该任务定制的APSNet框架,通过引入种子尺度特征(大小)来辅助网络学习细粒度信息,从而发现关键分类依据。具体地,在编码器部分设计了尺寸感知与嵌入(SPE)模块,显式提取尺寸信息以补充细粒度特征。此外,提出基于渐进学习的异步解耦解码(ADD)架构,从通道和空间两个维度解码特征,实现判别性特征的高效学习。定量与定性分析均表明,本方法优于现有先进图像分类方法,准确率达到90.5%。这证明了该工作为大规模、系统性的考古研究提供了有效工具。
原文摘要 · Abstract (English)
Understanding the dietary preferences of ancient societies and their evolution across periods and regions is crucial for revealing human-environment interactions. Seeds, as important archaeological artifacts, represent a fundamental subject of archaeobotanical research. However, traditional studies rely heavily on expert knowledge, which is often time-consuming and inefficient. Intelligent analysis methods have made progress in various fields of archaeology, but there remains a research gap in data and methods in archaeobotany, especially in the classification task of ancient plant seeds. To address this, we construct the first Ancient Plant Seed Image Classification (APS) dataset. It contains 8,340 images from 17 genus- or species-level seed categories excavated from 18 archaeological sites across China. In addition, we design a framework specifically for the ancient plant seed classification task (APSNet), which introduces the scale feature (size) of seeds based on learning fine-grained information to guide the network in discovering key "evidence" for sufficient classification. Specifically, we design a Size Perception and Embedding (SPE) module in the encoder part to explicitly extract size information for the purpose of complementing fine-grained information. We propose an Asynchronous Decoupled Decoding (ADD) architecture based on traditional progressive learning to decode features from both channel and spatial perspectives, enabling efficient learning of discriminative features. In both quantitative and qualitative analyses, our approach surpasses existing state-of-the-art image classification methods, achieving an accuracy of 90.5%. This demonstrates that our work provides an effective tool for large-scale, systematic archaeological research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。