arXiv:2608.16030cs.ROcs.CY2026-08中稿 · the Foundation Mod…

研究大模型生成的安防机器人设计受身份提示影响的程度。

Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots

论文配图:Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots
图 1 · 摘自论文原文
  • 用236种身份标签测试提示对生成内容的影响。
  • 不同身份提示导致可读性差异显著,尤其在设计维度间。
  • 可读性可作公平性评估的初步基准,适合政策制定者参考。

大型语言模型(LLMs)在早期机器人开发阶段被用于生成文本形式的机器人设计规范、交互策略和风险评估。这些输出可能影响安防与安保机器人的概念化、文档记录及最终实施。本文评估了身份条件提示是否会导致大模型生成的安防机器人设计描述出现系统性差异。通过在单标签和模型增强提示条件下使用236个种族/性别等身份标签,分析生成内容的可读性,作为可访问性和身份敏感性差异的初始基准。结果表明,在不同提示条件、设计维度和身份标签下,可读性存在显著差异。尽管可读性无法判断输出是否公平或社会适宜,但它为更广泛的评估框架(包含词汇、语义、情感、句法和公平性分析)提供了可解释的基线。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an interpretable baseline within a broader benchmarking framework that also includes lexical, semantic, sentiment, syntactic, and fairness-focused analyses.

大模型身份偏见机器人设计可读性评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。