融合抽取与生成的多语言情感分析方法,提升低资源语言效果。
Hybrid Extractive Abstractive Summarization for Multilingual Sentiment Analysis
- 用TF-IDF提取关键句,再由XLM-R模型生成摘要
- 英文准确率达0.90,低资源语言达0.84
- 效率比传统方法高22%,适合实时监控应用
我们提出一种结合抽取式与生成式摘要的混合多语言情感分析方法,以克服单一方法的局限。模型整合基于TF-IDF的抽取策略与微调后的XLM-R生成模块,并引入动态阈值和文化适应机制。在10种语言上的实验显示,该方法显著优于基线模型,在英文上达到0.90的准确率,在低资源语言中达0.84。相比传统方法,计算效率提升22%。实际应用包括实时品牌监测与跨文化话语分析。未来工作将通过8位量化优化低资源语言性能。
原文摘要 · Abstract (English)
We propose a hybrid approach for multilingual sentiment analysis that combines extractive and abstractive summarization to address the limitations of standalone methods. The model integrates TF-IDF-based extraction with a fine-tuned XLM-R abstractive module, enhanced by dynamic thresholding and cultural adaptation. Experiments across 10 languages show significant improvements over baselines, achieving 0.90 accuracy for English and 0.84 for low-resource languages. The approach also demonstrates 22% greater computational efficiency than traditional methods. Practical applications include real-time brand monitoring and cross-cultural discourse analysis. Future work will focus on optimization for low-resource languages via 8-bit quantization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。