arXiv:2606.29437cs.HCcs.AI2026-06

用对话记录追踪AI协作全过程,让人类贡献可查可审。

LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators

论文配图:LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators
图 1 · 摘自论文原文
  • 将人机对话转化为可量化的协作指标,记录交互轨迹。
  • 实测学生报告中人类主导得分86.8,输出可追溯性达77.1。
  • 适合教育、工程等需审计协作过程的场景使用。

大型语言模型在教育、软件工程、学术写作和技术文档中的广泛应用引发关键问题:如何评估不仅包括最终输出,还包括生成过程的人机交互?现有讨论多聚焦于检测最终成果是否由AI生成,却忽视了揭示人类引导、AI贡献、修正与验证的对话历史。本文提出LLMography框架,将人机对话转化为可衡量的出处、人类贡献度、AI依赖性、可复现性与审计性指标。类比参考文献与网络资源目录,该框架系统记录人机协同创作的动态过程。我们构建原型分析对话痕迹,生成包含提示质量评分、人类引导分、AI依赖等级、审计性评分、最终输出可追溯性、隐私风险等级及推荐标签的KPI报告。对19份匿名工程学生审计报告的初步探索评估显示,多数交互属人机共创,平均人类引导分86.8/100,提示质量81.9/100,审计性72.8/100,最终输出可追溯性77.1/100。研究还将该框架应用于自身写作过程,判定为人类主导的AI辅助共创。结果表明,AI透明度应从输出检测转向交互过程记录。

原文摘要 · Abstract (English)

The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a key question: how can we evaluate not only AI-assisted outputs, but also the interaction process that produced them? Current debates often focus on detecting whether a final artifact was generated by AI, while overlooking the conversation history that reveals human direction, AI contribution, corrections, validation, and traceability. This paper introduces LLMography, a framework for transforming Human-AI conversations into measurable indicators of provenance, human contribution, AI dependency, reproducibility, and auditability. By analogy with bibliography and webography, LLMography documents the dynamic trajectory of interaction between a human and a Large Language Model as a structured trace of Human-AI co-production. We present a prototype that analyzes Human-AI conversation traces and generates KPI reports including Prompt Quality Score, Human Direction Score, AI Dependency Level, Auditability Score, Final Output Traceability, Privacy Risk Level, and a recommended LLMography label. A preliminary exploratory evaluation was conducted on 19 anonymized audit reports from engineering students. Most interactions were classified as Human-AI co-produced, with average scores of 86.8/100 for Human Direction, 81.9/100 for Prompt Quality, 72.8/100 for Auditability, and 77.1/100 for Final Output Traceability. The paper also applies LLMography to its own writing process, classified as human-originated, human-directed, AI-assisted co-production. The findings suggest that AI transparency should move beyond output detection toward documenting the history of interaction.

人机协作可追溯性审计指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。