把 EPUB 接入 AI 阅读工具,先按书籍定义的阅读顺序提取章节,并保留每段文字的章节来源。不能直接按压缩包内文件名排序:chapter-a.xhtml 可能在后面读,目录文件也未必属于正文阅读序列。
本文用 Python 标准库创建并解析一份人为构造的 EPUB,已在 Windows、Python 3.11.15 环境运行。它输出文字章节 JSON,没有调用 AI 阅读服务,也没有测试真实电子书的问答表现。

container、manifest、spine 分别做什么
W3C EPUB 3.3 规范规定容器、包文档与阅读顺序的结构。可以把 EPUB 理解为打包的内容集合;解析时要先找包文档,再按它的引用关系读取内容。
META-INF/container.xml的 rootfile 指出包文档路径,例如 EPUB/package.opf。- 包文档的 manifest 用 item 的 id 对应 href 与 media-type,列出内容资源。
- spine 中的 itemref 按 idref 引用 manifest 项,定义阅读序列。
- 章节 href 相对包文档所在目录解析,不能一律从压缩包根目录拼接。
本例处理默认阅读序列,跳过 linear="no" 的辅助项目。规范把这种项目解释为补充主要内容的材料,可能包括注释、说明和答案;跳过不代表不重要。如果你的问答依赖这些内容,应按业务需要另外提取和关联。
运行文件名与阅读顺序相反的样例
不需要安装第三方包。将下面代码保存为 extract_epub.py,执行 python extract_epub.py。它创建 sample_knowledge.epub:manifest 的 first 指向 chapter-a.xhtml,second 指向 chapter-b.xhtml,但 spine 明确先读 second,再读 first。
from pathlib import Path,PurePosixPath
import json,posixpath,zipfile,xml.etree.ElementTree as ET
from urllib.parse import unquote
container='''<?xml version="1.0"?><container version="1.0" xmlns="urn:oasis:names:tc:opendocument:xmlns:container"><rootfiles><rootfile full-path="EPUB/package.opf" media-type="application/oebps-package+xml"/></rootfiles></container>'''
package='''<?xml version="1.0"?><package xmlns="http://www.idpf.org/2007/opf" version="3.0" unique-identifier="book-id"><metadata xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:identifier id="book-id">urn:example:synthetic-knowledge</dc:identifier><dc:title>人为构造的阅读示例</dc:title><dc:language>zh</dc:language><meta property="dcterms:modified">2026-10-01T00:00:00Z</meta></metadata><manifest><item id="first" href="chapter-a.xhtml" media-type="application/xhtml+xml"/><item id="second" href="chapter-b.xhtml" media-type="application/xhtml+xml"/><item id="nav" href="nav.xhtml" media-type="application/xhtml+xml" properties="nav"/></manifest><spine><itemref idref="second"/><itemref idref="first"/></spine></package>'''
chapter_a='''<html xmlns="http://www.w3.org/1999/xhtml"><head><title>后读章节</title></head><body><h1>第二个阅读位置</h1><p>这个文件名靠前,但 spine 将它放在第二位。</p></body></html>'''
chapter_b='''<html xmlns="http://www.w3.org/1999/xhtml"><head><title>先读章节</title></head><body><h1>第一个阅读位置</h1><p>AI 阅读工具需要按阅读顺序处理章节,再保留来源标识。</p></body></html>'''
nav='''<html xmlns="http://www.w3.org/1999/xhtml" xmlns:epub="http://www.idpf.org/2007/ops"><head><title>目录</title></head><body><nav epub:type="toc"><ol><li><a href="chapter-b.xhtml">先读</a></li><li><a href="chapter-a.xhtml">后读</a></li></ol></nav></body></html>'''
with zipfile.ZipFile('sample_knowledge.epub','w') as archive:
archive.writestr('mimetype','application/epub+zip',compress_type=zipfile.ZIP_STORED)
for name,text in {'META-INF/container.xml':container,'EPUB/package.opf':package,
'EPUB/chapter-a.xhtml':chapter_a,'EPUB/chapter-b.xhtml':chapter_b,
'EPUB/nav.xhtml':nav}.items():archive.writestr(name,text)
def extract_epub(path):
ns={'c':'urn:oasis:names:tc:opendocument:xmlns:container','o':'http://www.idpf.org/2007/opf','h':'http://www.w3.org/1999/xhtml'}
with zipfile.ZipFile(path) as archive:
if 'META-INF/encryption.xml' in archive.namelist():
raise ValueError('This example accepts packages without encryption declarations')
root=ET.fromstring(archive.read('META-INF/container.xml'))
rootfile=root.find('c:rootfiles/c:rootfile',ns)
if rootfile is None:raise ValueError('EPUB has no package rootfile')
opf_path=rootfile.attrib['full-path']
package=ET.fromstring(archive.read(opf_path))
items={x.attrib['id']:x.attrib for x in package.findall('o:manifest/o:item',ns)}
records=[]
for itemref in package.findall('o:spine/o:itemref',ns):
if itemref.get('linear','yes')=='no':continue
item=items[itemref.attrib['idref']]
if item['media-type']!='application/xhtml+xml':
raise ValueError('This parser accepts textual XHTML spine items only')
href=unquote(item['href'].split('#')[0])
entry=posixpath.normpath(posixpath.join(posixpath.dirname(opf_path),href))
if entry.startswith('../') or PurePosixPath(entry).is_absolute():raise ValueError('Invalid package path')
html=ET.fromstring(archive.read(entry))
body=html.find('h:body',ns)
if body is None:raise ValueError('XHTML spine item lacks body')
for element in list(body.iter()):
for child in list(element):
if child.tag.split('}')[-1] in ('script','style'):element.remove(child)
text=' '.join(' '.join(body.itertext()).split())
if not text:raise ValueError('No textual content in spine item')
records.append({'chapter_order':len(records)+1,'item_id':itemref.attrib['idref'],
'source_file':Path(path).name,'package_entry':entry,'text':text})
return records
records=extract_epub('sample_knowledge.epub')
assert [x['item_id'] for x in records]==['second','first']
assert records[0]['package_entry']=='EPUB/chapter-b.xhtml'
assert 'AI 阅读工具' in records[0]['text']
Path('epub_chapters.json').write_text(json.dumps(records,ensure_ascii=False,indent=2),encoding='utf-8')
print('chapters:',len(records))
print('reading_order:',[x['item_id'] for x in records])
print('first_package_entry:',records[0]['package_entry'])
print('saved: epub_chapters.json')
本地运行输出
chapters: 2
reading_order: ['second', 'first']
first_package_entry: EPUB/chapter-b.xhtml
saved: epub_chapters.json
检查章节顺序和来源字段
程序输出 2 个章节,阅读顺序为 second、first;第一个包内路径是 EPUB/chapter-b.xhtml。代码断言这个顺序并检查首章文字,证明演示中读取顺序来自 spine,而不是文件名。
打开 epub_chapters.json,核对 chapter_order、item_id、package_entry 和 text 是否与包文档引用一致。每条记录保留 source_file 与包内章节路径,后续分段时继续带上这些字段,回答引用才能回到具体章节。
本例把 XHTML body 内的文字合并为字符串,并去掉 script、style 节点。得到的是可核对的文字输入,不是原书版式;图片中的文字、图表、公式和脚注关系没有被完整保留。
处理你自己的电子书
- 使用自己有权处理的 EPUB 副本,保留原书。
- 删除创建 sample 文件的演示部分,保留 extract_epub 函数,调用
extract_epub('实际书名.epub')。 - 先检查书中第一章与最后一章,再检查目录、序言和附录是否按照你的任务需要被保留。
- 本例只选择 container 中第一个 rootfile;多种呈现版本的电子书需要先选择正确包文档,不能默认第一个版本就适合任务。
- 将章节分段时保存稳定书籍 ID、版本、章节路径和片段位置,再送入阅读工具支持的导入格式;JSON 本身不代表所有阅读工具都能直接导入。
解析失败怎样判断
| 情况 | 处理方式 |
|---|---|
| container 找不到 rootfile | 确认文件确实是 EPUB,并检查包结构,不猜包文档名称 |
| spine 的 idref 没有对应 manifest 项 | 核对原包的引用关系,保存错误位置再修复来源文件 |
| 章节 media-type 不是 XHTML | 本例明确拒绝,需要为该媒体类型另加处理器 |
| XHTML 不完整、没有 body 或正文为空 | 回读原章节,确认是否为图片型内容或格式异常 |
| 输出顺序正确但关键信息缺失 | 检查辅助章节、图像和脚注,不能把顺序通过当作内容完整 |
本例只接受没有 META-INF/encryption.xml 声明、且 spine 项为文字 XHTML 的输入;发现该声明就停止。encryption.xml 也可能用于字体混淆,并不必然表示整书有 DRM;这里采用的是明确的简化边界,没有识别各种声明后继续处理的实现。
如果输出为空,先检查 spine 是否有主要阅读项目,再决定增加辅助内容处理;没有可用文字就先停在资料转换阶段。本文不绕过电子书保护,也不把无法读取的章节交给模型猜写。
为什么输出没有原书页码?
流式电子书显示的页码会随排版设置变化;本例的来源定位是书籍文件和包内章节路径。若需要引用到原书特定位置,进一步保留元素 ID 或出版方提供的位置标记,再核对阅读工具的定位能力。本例没有生成纸书页码映射。
Ai菜鸟网。发布者:AI小管家,转载请注明出处:https://www.alyyhw.com/31903.html