针对中文网络毒评的复杂表达,提出细粒度分类框架。
NLP-Based Review for Toxic Comment Detection Tailored to the Chinese Cyberspace
- 构建可扩展的中文毒评定义与分类框架
- 梳理现有数据集局限并提出标注策略
- 适合中文NLP与内容安全研究者参考
随着移动互联网深入融合与社交平台广泛普及,中文网络空间用户生成内容呈爆炸式增长。其中,毒评泛滥对个体心理健康、社区氛围与社会信任构成严重挑战。由于中文网络语言具有强上下文依赖性、文化特异性及快速演变特性,毒评常以谐音、隐喻等复杂形式呈现,传统检测方法面临显著局限。本文聚焦基于自然语言处理的中文毒评检测,系统梳理该领域研究进展与核心挑战。首先界定中文毒评内涵与特征,分析其依托的平台生态与传播机制;其次综述现有公开数据集的构建方法与局限,提出新型细粒度、可扩展的毒评定义与分类框架,配套数据标注与质量评估策略;系统总结检测模型从传统方法到深度学习的演进路径,强调模型可解释性的重要性;最后深入探讨当前研究面临的关键开放问题,并提出未来研究方向建议。
原文摘要 · Abstract (English)
With the in-depth integration of mobile Internet and widespread adoption of social platforms, user-generated content in the Chinese cyberspace has witnessed explosive growth. Among this content, the proliferation of toxic comments poses severe challenges to individual mental health, community atmosphere and social trust. Owing to the strong context dependence, cultural specificity and rapid evolution of Chinese cyber language, toxic expressions are often conveyed through complex forms such as homophones and metaphors, imposing notable limitations on traditional detection methods. To address this issue, this review focuses on the core topic of natural language processing based toxic comment detection in the Chinese cyberspace, systematically collating and critically analyzing the research progress and key challenges in this field. This review first defines the connotation and characteristics of Chinese toxic comments, and analyzes the platform ecology and transmission mechanisms they rely on. It then comprehensively reviews the construction methods and limitations of existing public datasets, and proposes a novel fine-grained and scalable framework for toxic comment definition and classification, along with corresponding data annotation and quality assessment strategies. We systematically summarize the evolutionary path of detection models from traditional methods to deep learning, with special emphasis on the importance of interpretability in model design. Finally, we thoroughly discuss the open challenges faced by current research and provide forward-looking suggestions for future research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。