arXiv:2508.05938cs.CLcs.AI2025-08被引 1

用人类与AI协作标注,高效识别游戏聊天中的利他行为。

Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale

  • 通过小样本人类标注筛选最优LLM标注策略。
  • 构建人机迭代修正机制,提升标签质量与定义一致性。
  • 轻量级分类器+仅35%疑难样本调用大模型,推理成本降70%。

检测文本中促进他人行为的利他性沟通(如支持、鼓励)是信任与安全系统面临的新挑战。与有害内容检测不同,利他性缺乏明确定义和标注数据,需创新标注与部署方法。本文提出三阶段可扩展管道:首先基于少量人工标注样本筛选最优大模型标注策略;其次引入人机协同修正循环,由标注者审查GPT-4与人工判断不一致案例,持续优化任务定义;最后利用GPT-4合成10,000条高质量标签,训练双阶段推理系统——轻量分类器处理高置信度样本,仅约35%模糊实例交由GPT-4o处理。该架构使推理成本降低约70%,同时保持高精度(约0.90)。结果表明,精准的任务设计、人机协作与部署导向架构,可为新兴负责任AI任务提供可扩展解决方案。

原文摘要 · Abstract (English)

Detecting prosociality in text--communication intended to affirm, support, or improve others' behavior--is a novel and increasingly important challenge for trust and safety systems. Unlike toxic content detection, prosociality lacks well-established definitions and labeled data, requiring new approaches to both annotation and deployment. We present a practical, three-stage pipeline that enables scalable, high-precision prosocial content classification while minimizing human labeling effort and inference costs. First, we identify the best LLM-based labeling strategy using a small seed set of human-labeled examples. We then introduce a human-AI refinement loop, where annotators review high-disagreement cases between GPT-4 and humans to iteratively clarify and expand the task definition-a critical step for emerging annotation tasks like prosociality. This process results in improved label quality and definition alignment. Finally, we synthesize 10k high-quality labels using GPT-4 and train a two-stage inference system: a lightweight classifier handles high-confidence predictions, while only $\sim$35\% of ambiguous instances are escalated to GPT-4o. This architecture reduces inference costs by $\sim$70% while achieving high precision ($\sim$0.90). Our pipeline demonstrates how targeted human-AI interaction, careful task formulation, and deployment-aware architecture design can unlock scalable solutions for novel responsible AI tasks.

人机协作文本检测大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。