构建波兰语情色文本数据集,助力本地化内容安全检测
Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse
- 构建24000+条波兰语情色文本标注数据,涵盖模糊性、暴力等多维度标签
- 专用波兰语模型表现优于通用多语言模型,变压器架构在类别不平衡下更优
- 为复杂屈折语言的内容检测提供可复用框架,适合研究者与平台方参考
在线内容激增催生了对鲁棒检测系统的需求,尤其在非英语语境中现有工具存在显著局限。我们提出forePLay,一个全新的波兰语情色内容检测数据集,包含超过24,000条带标注的句子,采用多维分类体系,涵盖模糊性、暴力与社会不可接受性等维度。全面评估表明,专用波兰语模型在性能上优于多语言模型,基于Transformer的架构在处理类别不平衡时表现尤为突出。该数据集与配套分析建立了发展语言感知型内容审核系统的必要框架,同时揭示了将此类能力扩展至形态复杂的语言时的关键考量。
原文摘要 · Abstract (English)
The surge in online content has created an urgent demand for robust detection systems, especially in non-English contexts where current tools demonstrate significant limitations. We present forePLay, a novel Polish language dataset for erotic content detection, featuring over 24k annotated sentences with a multidimensional taxonomy encompassing ambiguity, violence, and social unacceptability dimensions. Our comprehensive evaluation demonstrates that specialized Polish language models achieve superior performance compared to multilingual alternatives, with transformer-based architectures showing particular strength in handling imbalanced categories. The dataset and accompanying analysis establish essential frameworks for developing linguistically-aware content moderation systems, while highlighting critical considerations for extending such capabilities to morphologically complex languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。