首个用于甲骨文研究的多模态智能代理,提升信息检索与分析效率。
OracleAgent: A Multimodal Reasoning Agent for Oracle Bone Script Research
- 构建多模态工具链,结合大模型实现甲骨文多任务协同处理
- 拥有140万字符拓片图像和8万条释读文本的知识库,支持跨模态检索
- 实测性能超越GPT-4o,显著缩短专家研究时间,适合古文字学者使用
作为最早的文字系统之一,甲骨文承载着古代文明的文化与智慧遗产。然而,当前甲骨文研究面临两大挑战:(1) 解读过程涉及多个串行与并行子任务的复杂流程;(2) 信息组织与检索效率低下,学者需耗费大量时间查找、整理与管理资源。为此,我们提出OracleAgent,首个专为甲骨文研究设计的结构化信息管理与检索智能体。该系统融合多种甲骨文分析工具,依托大语言模型(LLMs)实现灵活编排。同时,我们构建了一个涵盖140万张单字拓片图像和8万条释读文本的领域专用多模态知识库,经过多年数据收集、清洗与专家标注。OracleAgent通过多模态工具,辅助专家完成字符、文献、释读文本及拓片图像的检索任务。大量实验表明,OracleAgent在多项多模态推理与生成任务中表现优异,优于主流多模态大模型(如GPT-4o)。案例研究进一步证明其可显著降低甲骨文研究的时间成本。该成果标志着甲骨文辅助研究与自动化解读系统的实用化迈出关键一步。
原文摘要 · Abstract (English)
As one of the earliest writing systems, Oracle Bone Script (OBS) preserves the cultural and intellectual heritage of ancient civilizations. However, current OBS research faces two major challenges: (1) the interpretation of OBS involves a complex workflow comprising multiple serial and parallel sub-tasks, and (2) the efficiency of OBS information organization and retrieval remains a critical bottleneck, as scholars often spend substantial effort searching for, compiling, and managing relevant resources. To address these challenges, we present OracleAgent, the first agent system designed for the structured management and retrieval of OBS-related information. OracleAgent seamlessly integrates multiple OBS analysis tools, empowered by large language models (LLMs), and can flexibly orchestrate these components. Additionally, we construct a comprehensive domain-specific multimodal knowledge base for OBS, which is built through a rigorous multi-year process of data collection, cleaning, and expert annotation. The knowledge base comprises over 1.4M single-character rubbing images and 80K interpretation texts. OracleAgent leverages this resource through its multimodal tools to assist experts in retrieval tasks of character, document, interpretation text, and rubbing image. Extensive experiments demonstrate that OracleAgent achieves superior performance across a range of multimodal reasoning and generation tasks, surpassing leading mainstream multimodal large language models (MLLMs) (e.g., GPT-4o). Furthermore, our case study illustrates that OracleAgent can effectively assist domain experts, significantly reducing the time cost of OBS research. These results highlight OracleAgent as a significant step toward the practical deployment of OBS-assisted research and automated interpretation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。