通过控制生成句子的毒性,模拟人类对无上下文语句的多样理解。
Mimicking How Humans Interpret Out-of-Context Sentences Through Controlled Toxicity Decoding
- 基于输入毒性动态调节生成内容的毒性水平
- 生成结果与人类标注更一致,且不确定性降低
- 适合研究语言歧义与毒性检测的学者
单个句子的含义可能因上下文缺失而产生差异。本文旨在通过生成无上下文句子的多种解释,模拟读者对不同毒性程度内容的理解。通过建模毒性,可预测误解并揭示隐藏的有毒含义。提出的解码策略通过三点实现毒性控制:(i) 使解释的毒性与输入一致,(ii) 对毒性更高的输入放宽约束,(iii) 提升生成解释集合中毒性水平的多样性。实验表明,该方法在句法和语义上均更贴近人工标注的解释,同时降低了模型预测的不确定性。
原文摘要 · Abstract (English)
Interpretations of a single sentence can vary, particularly when its context is lost. This paper aims to simulate how readers perceive content with varying toxicity levels by generating diverse interpretations of out-of-context sentences. By modeling toxicity, we can anticipate misunderstandings and reveal hidden toxic meanings. Our proposed decoding strategy explicitly controls toxicity in the set of generated interpretations by (i) aligning interpretation toxicity with the input, (ii) relaxing toxicity constraints for more toxic input sentences, and (iii) promoting diversity in toxicity levels within the set of generated interpretations. Experimental results show that our method improves alignment with human-written interpretations in both syntax and semantics while reducing model prediction uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。