免费开源的越英双语文本分析工具,支持分词、情感分析与摘要生成。
FreeTxt-Vi: A Benchmarked Vietnamese-English Toolkit for Segmentation, Sentiment, and Summarisation
- 融合VnCoreNLP与BPE分词策略,构建统一双语处理流程
- 在越语和英语上均达到或超过主流基线性能
- 无需编程即可分析多领域质性文本数据
FreeTxt-Vi 是一个免费开源的基于网页的越英双语文本分析工具包,结合语料语言学与自然语言处理技术,使用户无需编程即可构建、探索和解读自由文本数据。系统集成共现分析、关键词提取、词关系探索与交互可视化功能,并搭载基于Transformer的文本情感分析与摘要生成组件。核心贡献在于设计了统一的双语处理流水线:采用混合VnCoreNLP与字节对编码(BPE)分词策略,微调TabularisAI情感分类器,以及微调Qwen2.5模型实现抽象式摘要。不同于现有平台,FreeTxt-Vi通过三部分评估(分词、情感分析、摘要)验证其性能,在越南语和英语任务中均达到或超越广泛使用的基线模型。该工具降低了多语言文本分析的技术门槛,支持可复现研究,推动越南语这一广泛使用但资源匮乏的语言在自然语言处理中的发展。适用于教育、数字人文、文化遗产及社会科学等以质性文本为主的领域。
原文摘要 · Abstract (English)
FreeTxt-Vi is a free and open source web based toolkit for creating and analysing bilingual Vietnamese English text collections. Positioned at the intersection of corpus linguistics and natural language processing NLP it enables users to build explore and interpret free text data without requiring programming expertise. The system combines corpus analysis features such as concordancing keyword analysis word relation exploration and interactive visualisation with transformer based NLP components for sentiment analysis and summarisation. A key contribution of this work is the design of a unified bilingual NLP pipeline that integrates a hybrid VnCoreNLP and Byte Pair Encoding BPE segmentation strategy a fine tuned TabularisAI sentiment classifier and a fine tuned Qwen2.5 model for abstractive summarisation. Unlike existing text analysis platforms FreeTxt Vi is evaluated as a set of language processing components. We conduct a three part evaluation covering segmentation sentiment analysis and summarisation and show that our approach achieves competitive or superior performance compared to widely used baselines in both Vietnamese and English. By reducing technical barriers to multilingual text analysis FreeTxt Vi supports reproducible research and promotes the development of language resources for Vietnamese a widely spoken but underrepresented language in NLP. The toolkit is applicable to domains including education digital humanities cultural heritage and the social sciences where qualitative text data are common but often difficult to process at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。