本例把每条 AI 样本的 text_chars 和 image_count 两个字段,从“一行一个样本”转成“一行一个样本指标”。这种长表便于按指标做分布图和数据检查;转换增加行数是预期结果,但样本与指标的归属不能改变。
长表必须同时保存 sample_id、metric、unit、value,避免把字符数与图片张数混在同一个数值列后失去含义。pandas 是表格处理工具,不会自动理解指标的业务语义;单位由你明确提供。

先说明一行代表什么
| sample_id | text_chars(字符) | image_count(张) |
|---|---|---|
| 001 | 12 | 1 |
| 002 | 18 | 缺失 |
| 003 | 27 | 0 |
表中数值全部为人工构造,用于多模态样本字段整理;没有实际运行分词器、图像识别或模型训练。text_chars 在这里是给定的字符数量,不是 token 数,image_count 的缺失也不能解释为“没有图片”。
示例使用 Windows、Python 3.11.15、pandas 2.3.3、NumPy 2.4.6。先保持 sample_id 为字符串,使 001 的前导零不会被转换成整数 1。
在空的练习目录里创建独立环境。以下为 Windows PowerShell 命令;macOS/Linux 的激活命令用 source .venv/bin/activate。输出文件只用于本次演示,重复执行会更新这些演示文件。
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install pandas==2.3.3 numpy==2.4.6
只转换明确选中的指标列
保存代码为 wide_to_long.py,运行 python wide_to_long.py。输出文件为 reshape_output/samples_long.csv。id_vars 保留标识字段,value_vars 明确指定需要转行的指标;不指定 value_vars 容易把备注、来源或标签一起当作测量值。
from pathlib import Path
import numpy as np
import pandas as pd
wide = pd.DataFrame({
'sample_id': ['001', '002', '003'],
'text_chars': [12, 18, 27],
'image_count': [1, np.nan, 0],
})
metrics = ['text_chars', 'image_count']
units = {'text_chars': 'chars', 'image_count': 'images'}
if wide['sample_id'].isna().any() or not wide['sample_id'].is_unique:
raise ValueError('sample_id must be present and unique')
long = wide.melt(
id_vars='sample_id', value_vars=metrics,
var_name='metric', value_name='value',
)
long['unit'] = long['metric'].map(units)
long = long[['sample_id', 'metric', 'unit', 'value']]
assert len(long) == len(wide) * len(metrics)
assert not long.duplicated(['sample_id', 'metric']).any()
assert long['value'].isna().sum() == 1
restored = long.pivot(index='sample_id', columns='metric', values='value')
restored = restored.reindex(wide['sample_id'])[metrics]
restored.columns.name = None
pd.testing.assert_frame_equal(restored, wide.set_index('sample_id'), check_dtype=False)
out = Path('reshape_output')
out.mkdir(exist_ok=True)
long.to_csv(out / 'samples_long.csv', index=False, encoding='utf-8-sig')
print('wide_rows:', len(wide), 'long_rows:', len(long))
print('missing_values:', int(long['value'].isna().sum()))
print(long.to_string(index=False))
print('roundtrip_values_equal:', True)
broken = pd.concat([long, long.iloc[[0]]], ignore_index=True)
try:
broken.pivot(index='sample_id', columns='metric', values='value')
except ValueError:
print('duplicate_sample_metric_rejected:', True)
else:
raise AssertionError('Duplicate key was accepted')
核对行数、缺失和回转
wide_rows: 3 long_rows: 6
missing_values: 1
sample_id metric unit value
001 text_chars chars 12.0
002 text_chars chars 18.0
003 text_chars chars 27.0
001 image_count images 1.0
002 image_count images NaN
003 image_count images 0.0
roundtrip_values_equal: True
duplicate_sample_metric_rejected: True
3 条样本、2 个指标应得到 6 行长表,其中 1 个 value 保持 NaN;编号 002 的 image_count 仍然缺失,编号 003 的 image_count 为有效值 0。逐条确认 001/text_chars=12、001/image_count=1,另外两条同样按编号核对,不按输出行号猜归属。
回转断言验证列顺序和数值对应,包括缺失值;check_dtype=False 允许转长时统一数值列导致整数变成浮点数。它不是原始数据类型完全不变的承诺。若正式任务要求恢复整数、类别或可空类型,应另存字段类型定义,并在回转后显式校验和转换。
最后故意复制一条 sample_id/metric 组合,pivot 应报 ValueError。不能为了“能转回来”就改用求平均的 pivot_table;重复组合可能代表重复导出或多次测量,先确定真实粒度,必要时把时间或测量序号纳入键。
换真实文件后的三个检查
读入和重新打开导出文件时都指定 dtype={“sample_id”:”string”}。其次检查 units 映射覆盖每个 metric,禁止 unit 为空;再核对行数是否等于原样本数乘选定指标数,及每个 sample_id/metric 是否唯一。
字符数和图片数不能相加;长表 value 的汇总必须先按 metric 与 unit 分组,再按字段含义决定能否计算。若有标签、时间和来源列,按业务粒度放进 id_vars,避免转换后丢掉关联。本文示例的两列数值可以统一存成 float,混合文本与数字的指标需要另外设计类型。
长表能直接用于模型训练吗?
不一定。许多表格模型仍要求一行一个样本、每个特征一列;这里的长表主要用于检查、绘图和导出。如果将同一样本的两条指标记录当作两条独立训练样本,标签与划分关系就会变错。训练前应回到约定的矩阵形状并核对样本编号。
如果还需要连接单独的标签文件,另看按样本编号连接特征和标签,该步骤与单表重塑的键检查用途不同。
官方资料与版本范围
本文于 2026 年 10 月 1 日读取下列官方文档;示例的实际版本见正文。
Ai菜鸟网。发布者:AI小管家,转载请注明出处:https://www.alyyhw.com/32541.html