arXiv:2506.02372cs.CL2025-06被引 7

构建日语大模型安全输出数据集,提升生成内容的恰当性与安全性。

AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output

  • 基于日本社会文化背景人工构建1800组问答对,覆盖多种风险类型。
  • 微调后日语大模型安全性提升,通用响应质量不受影响。
  • 新增英译与标注,助力多语言安全数据集构建。

本文提出 AnswerCarefully,一个用于提升日语大模型输出安全性和适当性的数据集。该数据集包含1,800对问题与参考答案,问题均需特别注意回应。其涵盖先前英文数据集中确立的各类风险类别,但样本为原创,充分反映日本本地大模型使用场景的社会文化背景。实验表明,使用该数据集进行指令微调可显著提升日语大模型输出的安全性,且不损害通用回答的实用性。此外,我们以该数据集为基准评估了12个日语大模型的安全性表现。最后,数据集最新更新提供了问题的英文翻译与标注,旨在促进其他语言和地区的类似数据集开发。

原文摘要 · Abstract (English)

In this paper we present AnswerCarefully, a dataset for promoting the safety and appropriateness of Japanese LLM outputs. The dataset consists of 1,800 pairs of questions and reference answers, where the questions require special attention in answering. It covers a wide range of risk categories established in prior English-language datasets, but the data samples are original in that they are manually created to reflect the socio-cultural context of LLM usage in Japan. We show that using this dataset for instruction to fine-tune a Japanese LLM led to improved output safety without compromising the utility of general responses. We also report the results of a safety evaluation of 12 Japanese LLMs using this dataset as a benchmark. Finally, we describe the latest update on the dataset which provides English translations and annotations of the questions, aimed at facilitating the derivation of similar datasets in different languages and regions.

大模型安全日语数据指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。