RAG 资料重复入库时,可以先按原文与嵌入配置查缓存:输入完全一致才复用向量,文本、模型版本或维度变化则重新计算。SQLite 适合验证这个单进程小型流程;缓存命中只说明已保存的同配置结果被取出,不证明检索质量合格。
本文使用 Python 标准库和人工三维、四维向量,不调用真实嵌入 API,不测量费用或模型速度。演示可验证缓存规则;接入实际模型后,还需核对提供方的输出协议、版本含义与真实检索结果。

缓存键必须描述同一项计算
仅按文本生成哈希不够:同一段文字用不同模型、版本、维度或任务类型计算,结果可能不兼容。下面把 model、model_version、dimension、task_type、preprocess_version 和 namespace 全部写入键。文本按原始 UTF-8 字节计算 SHA-256,不自动删空白或改大小写;实际接口的指令、归一化等影响输出的参数也要成为明确的配置字段。
Python hashlib 官方文档提供 sha256 与 hexdigest。这里用哈希识别输入,不能把它当作对资料的加密;向量、配置和哈希也应按来源资料的访问范围保存。
model_version 必须对应你确认的模型版本。如果服务只提供可变别名且没有公开修订标识,应另行维护缓存代次,在供应方变更或重新核验后更新;本地键不能自动发现远端模型已经变化。preprocess_version 也要随实际清洗规则变化,而不是一直写同一个值。
完整示例:相同输入命中,配置变化重新计算
新建练习目录,将代码保存为 embedding_cache.py,执行 python embedding_cache.py。无需额外安装包。内存数据库只用于此次演示,程序关闭后数据消失;真实持久缓存需另选受控数据库路径。
import hashlib
import json
import math
import sqlite3
conn = sqlite3.connect(":memory:")
conn.execute("CREATE TABLE cache (cache_key TEXT PRIMARY KEY, vector_json TEXT NOT NULL)")
calls = 0
def validate_vector(vector, dimension):
if not isinstance(vector, list) or len(vector) != dimension:
raise ValueError("向量维度不符合当前配置")
if any(type(value) not in (int, float) or not math.isfinite(value)
for value in vector):
raise ValueError("向量包含非数值、NaN 或 Inf")
return vector
def get_vector(text, config, embed):
if not isinstance(text, str) or not text.strip():
raise ValueError("输入文本必须为非空字符串")
dimension = config["dimension"]
if type(dimension) is not int or dimension < 1:
raise ValueError("维度必须为正整数")
for field in ("model", "model_version", "task_type", "preprocess_version", "namespace"):
if not isinstance(config[field], str) or not config[field].strip():
raise ValueError("配置缺少已确认的版本或范围")
identity = {**config, "text_sha256": hashlib.sha256(text.encode("utf-8")).hexdigest()}
canonical = json.dumps(identity, sort_keys=True, ensure_ascii=False,
separators=(",", ":"))
key = hashlib.sha256(canonical.encode("utf-8")).hexdigest()
row = conn.execute("SELECT vector_json FROM cache WHERE cache_key=?", (key,)).fetchone()
if row is not None:
return validate_vector(json.loads(row[0]), dimension), True
vector = validate_vector(embed(text, config), dimension)
encoded = json.dumps(vector, allow_nan=False)
with conn:
conn.execute("INSERT INTO cache(cache_key,vector_json) VALUES (?,?)", (key, encoded))
return vector, False
def demo_embed(text, config):
global calls
calls += 1
# 教学假向量:仅验证缓存身份,不表示文本语义。
return [1.0] + [0.0] * (config["dimension"] - 1)
base = {"model": "demo-model", "model_version": "revision-1", "dimension": 3,
"task_type": "document-demo",
"preprocess_version": "raw-v1", "namespace": "public-demo"}
text = "示例设备 A 保修期为 12 个月。"
first, hit1 = get_vector(text, base, demo_embed)
again, hit2 = get_vector(text, base, demo_embed)
assert not hit1 and hit2 and first == again and calls == 1
version_config = {**base, "model_version": "revision-2"}
_, version_hit = get_vector(text, version_config, demo_embed)
dimension_config = {**base, "dimension": 4}
four, dimension_hit = get_vector(text, dimension_config, demo_embed)
_, text_hit = get_vector(text + " 仅限示例。", base, demo_embed)
_, task_hit = get_vector(text, {**base, "task_type": "query-demo"}, demo_embed)
assert not version_hit and not dimension_hit and not text_hit and not task_hit
assert len(four) == 4 and calls == 5
for name, bad in [("dimension", [1.0]), ("NaN", [float("nan"), 0.0, 0.0]),
("Inf", [float("inf"), 0.0, 0.0])]:
try:
get_vector("bad-case-" + name, base, lambda _text, _config: bad)
except ValueError as error:
print("rejected:", name, str(error))
else:
raise AssertionError("非法向量不应写入缓存")
rows = conn.execute("SELECT COUNT(*) FROM cache").fetchone()[0]
assert rows == 5
print({"first_hit": hit1, "repeat_hit": hit2, "version_hit": version_hit,
"dimension_hit": dimension_hit, "text_hit": text_hit,
"task_hit": task_hit,
"demo_callback_calls": calls, "cache_rows": rows})
conn.close()
Python sqlite3 官方文档说明参数绑定与连接事务管理。缓存键与向量用参数传入,写入失败会抛错;这里不把 SQL 执行失败当作成功命中。Python json 官方文档说明 allow_nan=False 会拒绝非标准浮点常量;代码还在写入前及读取后检查维度与有限值。
逐项核对真实运行结果
固定演示中第一次 first_hit 为 false,同一输入第二次 repeat_hit 为 true。模型版本、维度、文本或任务类型分别变化时都为 false,四维配置得到长度为 4 的向量。教学回调共调用 5 次,缓存共 5 条;三条非法向量均被拒绝且未增加缓存条数。
这些数字属于人工回调的运行结果,不能解读成真实 API 节省了多少费用。需要持久化时,再检查同一文件数据库重新打开后是否命中;内存库的一次演示不能证明磁盘备份或跨进程行为。
怎样替换成真实嵌入接口
- 将 demo_embed 替换成已验证的嵌入调用函数,只返回当前文本对应的数值列表,保留错误响应。真实请求中的模型、维度、任务类型与预处理必须和配置一致;document-demo 和 query-demo 是教学值,不能当作真实 API 枚举发送。
- 确认批量接口如何把响应索引映射回输入,不能把第一条向量随意分配给所有文本。本文每次只有一个输入。
- 为实际模型建立固定查询与目标片段,核对缓存前后的向量及检索结果。缓存身份正确不等于模型选型或召回已经通过。
- namespace 由应用根据可信资料范围设置,不能允许请求者随意指定其他用户的范围。它只隔离键,不替代接口鉴权。
什么时候需要扩展或停止
- 缓存内容损坏:读取仍要验证,失败时定位对应条目并重新取得向量,不默默返回空数组。
- 模型或预处理变化:更新配置与相应索引,不能让旧向量和新查询向量在不兼容空间里继续比较。
- 多工作进程同时写:本例未处理并发重复计算、数据库锁等待和冲突后回读;扩量前增加明确的协调及恢复策略。
- 资料删除或过期:同步处理缓存与索引生命周期。缓存存在不能作为仍有访问许可的依据。
读者下一步是先跑固定演示并核对六个命中状态和三条拒绝结果,再接入自己已验证的单条嵌入函数。本文不保证远端模型版本自动可知,也不以人工向量验证任何语义质量。
Ai菜鸟网。发布者:AI小管家,转载请注明出处:https://www.alyyhw.com/31993.html