arXiv:2507.11966cs.CLcs.AI2025-07被引 1

保留新加坡式英语中的毒害性表达,提升低资源语种翻译质量

Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation

  • 人工验证少样本提示工程,捕捉俚语与语气细节
  • 多模型对比优化,提升翻译的语义相似度与效率
  • 适合跨文化AI安全与区域内容治理研究者

随着在线交流越来越多地涉及未充分代表的语言和方言,标准翻译系统常无法保留本地俚语、混用语言及有害言论的文化特征。低资源语言对间翻译有毒内容面临平行数据稀缺与安全过滤器净化攻击性表达的双重挑战。本文提出可复现的两阶段框架,在代码混杂的新加坡式英语安全语料上进行验证。首先,通过人工标注迭代筛选并排序使用者选中的新马语-目标语例句,以捕捉细微俚语、语气与毒性特征;其次,基于直接翻译与反向翻译的语义相似性,对多个大模型进行提示对优化。定量人评证实了该流程在有效性和效率上的优势。本框架不仅提升翻译质量,还通过支持文化敏感的监督机制与低资源场景基准测试,促进多元文化大模型的安全性。以新加坡式英语为试验场,强调在内容审核与区域平台治理等实际应用中保留社会语言细微差别的必要性。

原文摘要 · Abstract (English)

As online communication increasingly incorporates under-represented languages and colloquial dialects, standard translation systems often fail to preserve local slang, code-mixing, and culturally embedded markers of harmful speech. Translating toxic content between low-resource language pairs poses additional challenges due to scarce parallel data and safety filters that sanitize offensive expressions. In this work, we propose a reproducible, two-stage framework for toxicity-preserving translation, demonstrated on a code-mixed Singlish safety corpus. First, we perform human-verified few-shot prompt engineering: we iteratively curate and rank annotator-selected Singlish-target examples to capture nuanced slang, tone, and toxicity. Second, we optimize model-prompt pairs by benchmarking several large language models using semantic similarity via direct and back-translation. Quantitative human evaluation confirms the effectiveness and efficiency of our pipeline. Beyond improving translation quality, our framework contributes to the safety of multicultural LLMs by supporting culturally sensitive moderation and benchmarking in low-resource contexts. By positioning Singlish as a testbed for inclusive NLP, we underscore the importance of preserving sociolinguistic nuance in real-world applications such as content moderation and regional platform governance.

低资源翻译毒性检测多语种提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。