识别直播中故意变音规避监管的行为,提升电商直播合规性
Chinese Morph Resolution in E-commerce Live Streaming Scenarios
- 将变音检测转为文本生成任务,利用大模型生成训练数据
- 构建首个包含8.6万条样本的直播变音数据集
- 适用于监管机构和平台内容安全团队
中国电商平台(如抖音)的直播带货已成为主要销售渠道,但主播常通过变音手段规避监管、进行虚假宣传。本研究提出直播语音变音识别(LiveAMR)任务,针对医疗健康类直播中的发音规避行为。与以往聚焦社交媒体文本逃避的研究不同,该工作首次面向语音层面的变音问题。我们构建了首个包含86,790个样本的LiveAMR数据集,并将任务转化为文本到文本生成问题。通过大语言模型(LLMs)生成额外训练数据,显著提升检测性能,证明变音识别可有效增强直播内容监管能力。
原文摘要 · Abstract (English)
E-commerce live streaming in China, particularly on platforms like Douyin, has become a major sales channel, but hosts often use morphs to evade scrutiny and engage in false advertising. This study introduces the Live Auditory Morph Resolution (LiveAMR) task to detect such violations. Unlike previous morph research focused on text-based evasion in social media and underground industries, LiveAMR targets pronunciation-based evasion in health and medical live streams. We constructed the first LiveAMR dataset with 86,790 samples and developed a method to transform the task into a text-to-text generation problem. By leveraging large language models (LLMs) to generate additional training data, we improved performance and demonstrated that morph resolution significantly enhances live streaming regulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。