arXiv:2606.05857cs.CL2026-06

提出新方法在音频检索中抑制仇恨内容,既保效果又防有害输出。

Forgive or forget: Understanding the context of hate in audio retrieval systems

论文配图:Forgive or forget: Understanding the context of hate in audio retrieval systems
图 1 · 摘自论文原文
  • 用情感控制的中介模块,在不改原意前提下过滤有毒音频。
  • 两种变体分别通过重排序和反事实生成,毒性降低超70%且准确率损失<5%。
  • 适用于需安全可控的语音检索系统,尤其适合部署在开放场景。

文本到音频系统中的有害信息检索处理困难,因语境依赖性强。现有策略(如改写、摘要)易改变意图或丢失细节。本文提出一种后处理因果去偏框架,采用情感控制的中介模块,在保持语义相关性的同时抑制有害言论。该方法与模型无关,可无缝集成至现有检索流程。提出两个变体:Forgive 通过逻辑调整重排序并过滤有毒音频;Forget 生成反事实有毒提示以缓解有害检索。实验表明,毒性显著降低,检索准确率损失极小,安全性与可靠性同步提升。

原文摘要 · Abstract (English)

Handling toxic retrieval in text-to-audio systems is challenging due to contextual dependencies. Existing strategies (e.g., rephrasing, summarization) risk altering intent or omitting details. We propose a post hoc causal debiasing framework with a sentiment-controlled mediator to preserve semantic relevance while suppressing harmful speech. Our approach is model-agnostic and integrates seamlessly with existing retrieval pipelines. We introduce two variants: Forgive, which re-ranks and filters toxic audio via logit adjustment, and Forget, which generates counterfactual toxic prompts to mitigate harmful retrievals. Experiments show consistent toxicity reduction with minimal loss in retrieval accuracy, improving both safety and reliability.

音频检索有害内容因果去偏安全生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。