arXiv:2410.18641cs.IRcs.AI2024-10被引 1

用AI智能提取和分类旅游工具,提升数据更新效率

Smart ETL and LLM-based contents classification: the European Smart Tourism Tools Observatory experience

  • 结合PDF抓取与大模型实现自动化内容分类
  • 初步验证大模型在文本分类上的有效性
  • 适合需批量处理多源数据的智能系统开发者

本研究旨在通过引入人工智能驱动的智能ETL流程,改进欧洲智能旅游工具(STTs)观测站的内容更新与分类。从PDF目录中提取二维码、图像、链接和文本信息,利用智能提取转换加载技术去除重复项,并基于文本内容使用大语言模型(LLMs)进行分类。最终数据按广泛接受且灵活的都柏林核心元数据结构进行转换。初步结果表明,大语言模型在基于文本的内容分类方面具有显著潜力。该方法不仅适用于智能旅游领域,也具备向其他领域扩展的可行性。未来工作将聚焦于优化分类流程。

原文摘要 · Abstract (English)

Purpose: Our research project focuses on improving the content update of the online European Smart Tourism Tools (STTs) Observatory by incorporating and categorizing STTs. The categorization is based on their taxonomy, and it facilitates the end user's search process. The use of a Smart ETL (Extract, Transform, and Load) process, where \emph{Smart} indicates the use of Artificial Intelligence (AI), is central to this endeavor. Methods: The contents describing STTs are derived from PDF catalogs, where PDF-scraping techniques extract QR codes, images, links, and text information. Duplicate STTs between the catalogs are removed, and the remaining ones are classified based on their text information using Large Language Models (LLMs). Finally, the data is transformed to comply with the Dublin Core metadata structure (the observatory's metadata structure), chosen for its wide acceptance and flexibility. Results: The Smart ETL process to import STTs to the observatory combines PDF-scraping techniques with LLMs for text content-based classification. Our preliminary results have demonstrated the potential of LLMs for text content-based classification. Conclusion: The proposed approach's feasibility is a step towards efficient content-based classification, not only in Smart Tourism but also adaptable to other fields. Future work will mainly focus on refining this classification process.

智能旅游大模型数据抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。