Label Studio 导出的原始 JSON 保存任务数据和人工标注,但它不是直接可以训练的两列表。准备单标签文本分类数据时,应从正式 annotations 读取指定控件的类别,保留原样本编号,并把取消标注、多份有效标注、非法类别等记录单独报告。
本文适用于 Text name="text" value="$text" 和 Choices name="intent" toName="text" choice="single" 的单标签项目,类别为“咨询、投诉、表扬”,任务含 data.source_id。下面是通用 Python 3 示例,不含真实客户数据;文件转换可本地验证,本文没有读取你的项目或训练模型。

先选择保留完整结构的 JSON
在项目 Export 选择 JSON,下载后另存为 export.json,保留未修改的原始副本。官方导出文档说明,JSON 保留任务和标注结构;JSON_MIN 会移除 Label Studio 专有字段,不适合下面需要判断取消状态和标注数量的检查。
同一任务可能有多份人工标注,也可能还有 predictions。预测建议不能替代正式标签;取消的标注可能仍被导出。下面的转换器只接受“一个非取消标注、一个指定分类结果、一个合法标签”的任务,其余进入异常报告。这是示例交付规则,需要多轮复核或多人裁决的项目应先完成仲裁,不能自动取第一份结果。
社区版界面导出不一定按你当前数据管理页的过滤条件裁剪;官方文档特别说明会包含已标注任务和取消任务。因此不能只根据界面筛选后的数量假设文件里都是训练样本。
认识要转换的字段
下面是虚构的一条正常任务,省略了时间等非必要字段,系统任务编号 101 也是演示值:
[
{
"id": 101,
"data": {"source_id": "s001", "text": "请问如何修改地址?"},
"annotations": [
{"was_cancelled": false, "result": [
{"from_name": "intent", "to_name": "text", "type": "choices",
"value": {"choices": ["咨询"]}}
]}
],
"predictions": []
}
]
from_name 要匹配分类控件,to_name 要匹配文本对象,不能把同一项目里的其他评分控件也当分类标签。Choices 官方文档说明单选和多选是不同配置;此脚本遇到多个类别会拒绝转换,不会截掉后面的标签。
生成训练表和异常报告
把以下代码保存为 export_to_csv.py,与 export.json 放在同一目录,运行 python export_to_csv.py。无需安装第三方库。正式使用前,将类别集合和两个控件名改成自己项目的实际值。
import csv
import json
from collections import Counter
from pathlib import Path
ALLOWED = {"咨询", "投诉", "表扬"}
tasks = json.loads(Path("export.json").read_text(encoding="utf-8-sig"))
if not isinstance(tasks, list):
raise ValueError("输入必须是原始 JSON 导出的任务数组")
if any(not isinstance(t, dict) or not isinstance(t.get("data"), dict)
for t in tasks):
raise ValueError("任务缺少 data 对象,请检查导出格式")
counts = Counter(t["data"].get("source_id") for t in tasks
if isinstance(t["data"].get("source_id"), str))
accepted, rejected = [], []
for task in tasks:
data = task["data"]
sid, text = data.get("source_id"), data.get("text")
reason = None
if not isinstance(sid, str) or not sid.strip():
reason = "缺少字符串 source_id"
elif sid != sid.strip():
reason = "source_id 含首尾空白"
elif counts[sid] != 1:
reason = "source_id 重复,所有同编号记录待核对"
elif not isinstance(text, str) or not text.strip():
reason = "文本为空或不是字符串"
annotations = task.get("annotations", [])
if reason is None:
if not isinstance(annotations, list) or any(
not isinstance(a, dict) for a in annotations
):
reason = "annotations 结构异常"
else:
active = [a for a in annotations if not a.get("was_cancelled", False)]
if len(active) != 1:
reason = "非取消标注数量不等于 1,需要复核或裁决"
if reason is None:
results = active[0].get("result", [])
if not isinstance(results, list) or any(
not isinstance(r, dict) for r in results
):
reason = "result 结构异常"
else:
matched = [r for r in results if r.get("type") == "choices"
and r.get("from_name") == "intent"
and r.get("to_name") == "text"]
if len(matched) != 1:
reason = "找不到唯一的 intent 分类结果"
else:
value = matched[0].get("value")
labels = value.get("choices") if isinstance(value, dict) else None
if (not isinstance(labels, list) or len(labels) != 1
or not isinstance(labels[0], str) or labels[0] not in ALLOWED):
reason = "标签缺失、多选或不在类别表中"
if reason is not None:
rejected.append({"task_id": task.get("id"), "source_id": sid,
"reason": reason})
else:
accepted.append({"task_id": task.get("id"), "source_id": sid,
"text": text, "label": labels[0]})
with Path("train.csv").open("w", encoding="utf-8", newline="") as f:
writer = csv.DictWriter(f, fieldnames=["task_id", "source_id", "text", "label"])
writer.writeheader()
writer.writerows(accepted)
Path("rejected.json").write_text(
json.dumps(rejected, ensure_ascii=False, indent=2), encoding="utf-8"
)
print("输入:", len(tasks), "通过:", len(accepted), "待复核:", len(rejected))
脚本对重复来源编号会拒绝所有同编号记录,避免第一条悄悄进入训练表。它保留文本原样,由 CSV 写入器处理文本中的逗号、双引号和换行;不会把换行删除或把类别自动改名。写入器和 newline="" 的使用可查Python csv 官方文档。脚本会覆盖同名输出文件,运行前保存旧输出或使用新的工作目录。
用已知答案核对转换结果
只用上面的正常示例运行时,应报告“输入: 1 通过: 1 待复核: 0”,训练 CSV 应包含 s001、原始问题和“咨询”。这是人为样例的预期结果,不是生产数据通过率。
再复制三条样例,给它们不同编号,分别设置为:唯一标注的 was_cancelled=true;同一任务有两份非取消标注;单选结果的 choices 含“咨询”和“投诉”两个标签。和正常样例放在一起,应得到“输入: 4 通过: 1 待复核: 3”,并能在 rejected.json 看到每条的原因。正式数据必须逐项处理异常,不能把它们无记录地丢弃。
核对 通过数 + 待复核数 = 输入数,再按 source_id 随机抽查原文和标签。CSV 看起来能打开仍不足以通过:还要确认训练程序读出的记录数、列名和文本与原导出一致。
这份转换器的边界
它只检查指定单标签项目的结构,不判断人工类别是否标对,不决定多份标注谁更可信,也不划分训练集和测试集。多标签任务应保留完整标签集合,实体识别任务应保留文本片段与位置,两者都不能使用此脚本的单个 label 列。
导出中没有未标注任务时,脚本无法证明项目已经全部标完;应另外核对项目总任务数和剩余任务。先让训练人员确认字段契约,处理异常记录并保存原始导出,再交付训练 CSV 和对应的类别定义。
Ai菜鸟网。发布者:AI小管家,转载请注明出处:https://www.alyyhw.com/30246.html