用大模型自动识别古尸X光片中的骨骼信息,提升考古影像分析效率。
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
- 通过精心设计提示词,让大视觉语言模型从原始DICOM文件中识别骨骼、视角和左右侧。
- 在100张样本上准确率达92%(主骨)、80%(视角)、100%(侧别),模糊案例可标记。
- 适合需要快速处理海量古尸影像的考古与人类学研究者使用。
古尸放射学利用现代成像技术研究考古与人类学遗骸,揭示人类健康演变规律。然而,野外采集的放射图像存在高度异质性:骨骼分离、摆放随意、缺乏侧别标记,且死亡年龄、骨骼年龄、性别及设备差异带来显著变异。导致内容检索如定位特定投影视角图像耗时费力,成为专家分析瓶颈。本文提出一种零样本提示策略,利用先进大视觉语言模型(LVLM)自动识别图像中的主骨、投影视角和侧别。流程将原始DICOM文件转为骨窗PNG图,输入经优化提示词的LVLM,获得结构化JSON输出,解析后生成可用于验证的表格。在由专家认证的古尸放射学家随机抽取的100张图像测试中,系统主骨识别准确率达92%,投影视角准确率80%,侧别识别准确率100%,对模糊案例可标注低或中置信度。结果表明,LVLM能显著加速大规模古尸数据集的关键词标注,为未来人类学工作流提供高效内容导航支持。
原文摘要 · Abstract (English)
Paleoradiology, the use of modern imaging technologies to study archaeological and anthropological remains, offers new windows on millennial scale patterns of human health. Unfortunately, the radiographs collected during field campaigns are heterogeneous: bones are disarticulated, positioning is ad hoc, and laterality markers are often absent. Additionally, factors such as age at death, age of bone, sex, and imaging equipment introduce high variability. Thus, content navigation, such as identifying a subset of images with a specific projection view, can be time consuming and difficult, making efficient triaging a bottleneck for expert analysis. We report a zero shot prompting strategy that leverages a state of the art Large Vision Language Model (LVLM) to automatically identify the main bone, projection view, and laterality in such images. Our pipeline converts raw DICOM files to bone windowed PNGs, submits them to the LVLM with a carefully engineered prompt, and receives structured JSON outputs, which are extracted and formatted onto a spreadsheet in preparation for validation. On a random sample of 100 images reviewed by an expert board certified paleoradiologist, the system achieved 92% main bone accuracy, 80% projection view accuracy, and 100% laterality accuracy, with low or medium confidence flags for ambiguous cases. These results suggest that LVLMs can substantially accelerate code word development for large paleoradiology datasets, allowing for efficient content navigation in future anthropology workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。