用多语言大模型精准提取评论中的观点关键词,提升跨语种社交分析效果。
Improving Multilingual Social Media Insights: Aspect-based Comment Analysis
- 通过微调多语言大模型生成评论中的观点词,引导模型关注关键信息。
- 在两个下游任务中显著提升社交媒体话语理解效果,跨语言性能均衡。
- 首次发布英、中、马来、印尼语评论观点词标注数据集,适合多语言研究者。
社交媒体评论语言自由、观点分散,给评论聚类、摘要和情感分析等任务带来挑战。为此,我们提出在细粒度层面从单条评论中识别并生成观点词,以引导模型注意力。具体而言,利用经过监督微调的多语言大语言模型进行评论观点词生成(CAT-G),并通过直接偏好优化(DPO)使模型预测更贴近人类判断。实验表明该方法在两项下游NLP任务中有效提升了对社交媒体话语的理解能力。此外,本文构建了首个涵盖英语、中文、马来语和印度尼西亚语的多语言CAT-G测试集,支持不同语言间基于大模型能力差异的性能对比分析。
原文摘要 · Abstract (English)
The inherent nature of social media posts, characterized by the freedom of language use with a disjointed array of diverse opinions and topics, poses significant challenges to downstream NLP tasks such as comment clustering, comment summarization, and social media opinion analysis. To address this, we propose a granular level of identifying and generating aspect terms from individual comments to guide model attention. Specifically, we leverage multilingual large language models with supervised fine-tuning for comment aspect term generation (CAT-G), further aligning the model's predictions with human expectations through DPO. We demonstrate the effectiveness of our method in enhancing the comprehension of social media discourse on two NLP tasks. Moreover, this paper contributes the first multilingual CAT-G test set on English, Chinese, Malay, and Bahasa Indonesian. As LLM capabilities vary among languages, this test set allows for a comparative analysis of performance across languages with varying levels of LLM proficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。