分析互联网多语言信息可用性及技术瓶颈
Information availability in different languages and various technological constraints related to multilinguism on the Internet
- 考察不同语言在互联网上的信息可得性差异
- 指出仅20%-25%人口以英语为母语,存在显著语言壁垒
- 适合关注数字公平与多语言AI的读者
过去二十年间,互联网使用量呈指数级增长,用户数从1995年的1600万增至2010年的16.5亿。互联网已成为覆盖几乎所有领域的信息库。由于互联网起源于以英语为主的美国,英语在万维网上占据主导地位。尽管英语是全球通用语言,但全球仅有20%-25%的人口以英语为母语,许多人因语言障碍无法访问互联网。近年来,随着非英语使用者数量快速增长,语言多样性问题愈发突出。虽然已有多种解决方案尝试消除语言壁垒,但仍存在巨大差距。本文旨在分析不同语言在互联网上的信息可用性需求,以及与多语言相关的各种技术限制。
原文摘要 · Abstract (English)
The usage of Internet has grown exponentially over the last two decades. The number of Internet users has grown from 16 Million to 1650 Million from 1995 to 2010. It has become a major repository of information catering almost every area. Since the Internet has its origin in USA which is English speaking country there is huge dominance of English on the World Wide Web. Although English is a globally acceptable language, still there is a huge population in the world which is not able to access the Internet due to language constraints. It has been estimated that only 20-25% of the world population speaks English as a native language. More and more people are accessing the Internet nowadays removing the cultural and linguistic barriers and hence there is a high growth in the number of non-English speaking users over the last few years on the Internet. Although many solutions have been provided to remove the linguistic barriers, still there is a huge gap to be filled. This paper attempts to analyze the need of information availability in different languages and the various technological constraints related to multi-linguism on the Internet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。