arXiv:2409.05530cs.CL2024-09被引 1

用BERT分析葡萄牙语校园对话,判断发言是否相关,准确率超95%。

QiBERT -- Classifying Online Conversations Messages with BERT as a Feature

  • 用SBERT嵌入作为特征,结合监督学习分类对话是否偏离主题。
  • 在真实校园对话数据上达到平均准确率95%以上。
  • 适合研究在线讨论行为、教育社交互动的社会科学家。

在线交流的快速发展催生了大量短文本数据,对这类内容进行分类具有重要意义。在线辩论尤其值得关注,因其能反映用户的意见、立场与偏好。本文利用来自葡萄牙学校在线社交对话的短文本数据,研究学生在被激发时是否持续参与讨论。通过使用基于BERT的先进机器学习方法,将SBERT嵌入作为特征,采用监督学习对发言是否偏离议题进行分类。实验结果显示,该模型在分类任务上平均准确率超过0.95。这一成果有助于社会科学家更深入理解人类沟通、行为、讨论与说服机制。

原文摘要 · Abstract (English)

Recent developments in online communication and their usage in everyday life have caused an explosion in the amount of a new genre of text data, short text. Thus, the need to classify this type of text based on its content has a significant implication in many areas. Online debates are no exception, once these provide access to information about opinions, positions and preferences of its users. This paper aims to use data obtained from online social conversations in Portuguese schools (short text) to observe behavioural trends and to see if students remain engaged in the discussion when stimulated. This project used the state of the art (SoA) Machine Learning (ML) algorithms and methods, through BERT based models to classify if utterances are in or out of the debate subject. Using SBERT embeddings as a feature, with supervised learning, the proposed model achieved results above 0.95 average accuracy for classifying online messages. Such improvements can help social scientists better understand human communication, behaviour, discussion and persuasion.

文本分类BERT在线对话教育数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。