用边缘仓库构建分布式文档系统,提升司法文本的语义探索能力
Elevating Semantic Exploration: A Novel Approach Utilizing Distributed Repositories
- 通过边缘仓库分布存储与处理文本数据,实现高效协同
- 在意大利司法部实际场景中验证,显著增强语义分析能力
- 适合对数据敏感性高、需高可用性的政府机构使用
集中式与分布式系统是信息技术基础设施组织的两种主要方式,各有优劣。集中式系统将资源集中于单一位置,便于管理但存在单点故障风险;分布式系统将资源分布在多个节点,具备更好的可扩展性和容错性,但管理更复杂。选择取决于应用需求、可扩展性及数据敏感性。集中式系统适用于对可扩展性要求低且需集中控制的应用,而分布式系统在需要高可用性和高性能的大规模环境中表现更优。本文探讨了一个为意大利司法部开发的分布式文档仓库系统,利用边缘仓库对文本数据和元数据进行分析,显著提升了语义探索能力。
原文摘要 · Abstract (English)
Centralized and distributed systems are two main approaches to organizing ICT infrastructure, each with its pros and cons. Centralized systems concentrate resources in one location, making management easier but creating single points of failure. Distributed systems, on the other hand, spread resources across multiple nodes, offering better scalability and fault tolerance, but requiring more complex management. The choice between them depends on factors like application needs, scalability, and data sensitivity. Centralized systems suit applications with limited scalability and centralized control, while distributed systems excel in large-scale environments requiring high availability and performance. This paper explores a distributed document repository system developed for the Italian Ministry of Justice, using edge repositories to analyze textual data and metadata, enhancing semantic exploration capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。