EPUB 怎么接入 AI 阅读工具?按 spine 提取正文并保留章节来源

按 EPUB 的 container、manifest 和 spine 提取文字章节,保留包内路径与章节顺序;通过文件名与阅读顺序相反的样例核对,并说明辅助内容和加密声明边界。

把 EPUB 接入 AI 阅读工具,先按书籍定义的阅读顺序提取章节,并保留每段文字的章节来源。不能直接按压缩包内文件名排序:chapter-a.xhtml 可能在后面读,目录文件也未必属于正文阅读序列。

本文用 Python 标准库创建并解析一份人为构造的 EPUB,已在 Windows、Python 3.11.15 环境运行。它输出文字章节 JSON,没有调用 AI 阅读服务,也没有测试真实电子书的问答表现。

EPUB 怎么接入 AI 阅读工具?按 spine 提取正文并保留章节来源

container、manifest、spine 分别做什么

W3C EPUB 3.3 规范规定容器、包文档与阅读顺序的结构。可以把 EPUB 理解为打包的内容集合;解析时要先找包文档,再按它的引用关系读取内容。

  1. META-INF/container.xml 的 rootfile 指出包文档路径,例如 EPUB/package.opf。
  2. 包文档的 manifest 用 item 的 id 对应 href 与 media-type,列出内容资源。
  3. spine 中的 itemref 按 idref 引用 manifest 项,定义阅读序列。
  4. 章节 href 相对包文档所在目录解析,不能一律从压缩包根目录拼接。

本例处理默认阅读序列,跳过 linear="no" 的辅助项目。规范把这种项目解释为补充主要内容的材料,可能包括注释、说明和答案;跳过不代表不重要。如果你的问答依赖这些内容,应按业务需要另外提取和关联。

运行文件名与阅读顺序相反的样例

不需要安装第三方包。将下面代码保存为 extract_epub.py,执行 python extract_epub.py。它创建 sample_knowledge.epub:manifest 的 first 指向 chapter-a.xhtml,second 指向 chapter-b.xhtml,但 spine 明确先读 second,再读 first。

from pathlib import Path,PurePosixPath
import json,posixpath,zipfile,xml.etree.ElementTree as ET
from urllib.parse import unquote

container='''<?xml version="1.0"?><container version="1.0" xmlns="urn:oasis:names:tc:opendocument:xmlns:container"><rootfiles><rootfile full-path="EPUB/package.opf" media-type="application/oebps-package+xml"/></rootfiles></container>'''
package='''<?xml version="1.0"?><package xmlns="http://www.idpf.org/2007/opf" version="3.0" unique-identifier="book-id"><metadata xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:identifier id="book-id">urn:example:synthetic-knowledge</dc:identifier><dc:title>人为构造的阅读示例</dc:title><dc:language>zh</dc:language><meta property="dcterms:modified">2026-10-01T00:00:00Z</meta></metadata><manifest><item id="first" href="chapter-a.xhtml" media-type="application/xhtml+xml"/><item id="second" href="chapter-b.xhtml" media-type="application/xhtml+xml"/><item id="nav" href="nav.xhtml" media-type="application/xhtml+xml" properties="nav"/></manifest><spine><itemref idref="second"/><itemref idref="first"/></spine></package>'''
chapter_a='''<html xmlns="http://www.w3.org/1999/xhtml"><head><title>后读章节</title></head><body><h1>第二个阅读位置</h1><p>这个文件名靠前,但 spine 将它放在第二位。</p></body></html>'''
chapter_b='''<html xmlns="http://www.w3.org/1999/xhtml"><head><title>先读章节</title></head><body><h1>第一个阅读位置</h1><p>AI 阅读工具需要按阅读顺序处理章节,再保留来源标识。</p></body></html>'''
nav='''<html xmlns="http://www.w3.org/1999/xhtml" xmlns:epub="http://www.idpf.org/2007/ops"><head><title>目录</title></head><body><nav epub:type="toc"><ol><li><a href="chapter-b.xhtml">先读</a></li><li><a href="chapter-a.xhtml">后读</a></li></ol></nav></body></html>'''
with zipfile.ZipFile('sample_knowledge.epub','w') as archive:
    archive.writestr('mimetype','application/epub+zip',compress_type=zipfile.ZIP_STORED)
    for name,text in {'META-INF/container.xml':container,'EPUB/package.opf':package,
                      'EPUB/chapter-a.xhtml':chapter_a,'EPUB/chapter-b.xhtml':chapter_b,
                      'EPUB/nav.xhtml':nav}.items():archive.writestr(name,text)

def extract_epub(path):
    ns={'c':'urn:oasis:names:tc:opendocument:xmlns:container','o':'http://www.idpf.org/2007/opf','h':'http://www.w3.org/1999/xhtml'}
    with zipfile.ZipFile(path) as archive:
        if 'META-INF/encryption.xml' in archive.namelist():
            raise ValueError('This example accepts packages without encryption declarations')
        root=ET.fromstring(archive.read('META-INF/container.xml'))
        rootfile=root.find('c:rootfiles/c:rootfile',ns)
        if rootfile is None:raise ValueError('EPUB has no package rootfile')
        opf_path=rootfile.attrib['full-path']
        package=ET.fromstring(archive.read(opf_path))
        items={x.attrib['id']:x.attrib for x in package.findall('o:manifest/o:item',ns)}
        records=[]
        for itemref in package.findall('o:spine/o:itemref',ns):
            if itemref.get('linear','yes')=='no':continue
            item=items[itemref.attrib['idref']]
            if item['media-type']!='application/xhtml+xml':
                raise ValueError('This parser accepts textual XHTML spine items only')
            href=unquote(item['href'].split('#')[0])
            entry=posixpath.normpath(posixpath.join(posixpath.dirname(opf_path),href))
            if entry.startswith('../') or PurePosixPath(entry).is_absolute():raise ValueError('Invalid package path')
            html=ET.fromstring(archive.read(entry))
            body=html.find('h:body',ns)
            if body is None:raise ValueError('XHTML spine item lacks body')
            for element in list(body.iter()):
                for child in list(element):
                    if child.tag.split('}')[-1] in ('script','style'):element.remove(child)
            text=' '.join(' '.join(body.itertext()).split())
            if not text:raise ValueError('No textual content in spine item')
            records.append({'chapter_order':len(records)+1,'item_id':itemref.attrib['idref'],
                            'source_file':Path(path).name,'package_entry':entry,'text':text})
        return records

records=extract_epub('sample_knowledge.epub')
assert [x['item_id'] for x in records]==['second','first']
assert records[0]['package_entry']=='EPUB/chapter-b.xhtml'
assert 'AI 阅读工具' in records[0]['text']
Path('epub_chapters.json').write_text(json.dumps(records,ensure_ascii=False,indent=2),encoding='utf-8')
print('chapters:',len(records))
print('reading_order:',[x['item_id'] for x in records])
print('first_package_entry:',records[0]['package_entry'])
print('saved: epub_chapters.json')

本地运行输出

chapters: 2
reading_order: ['second', 'first']
first_package_entry: EPUB/chapter-b.xhtml
saved: epub_chapters.json

检查章节顺序和来源字段

程序输出 2 个章节,阅读顺序为 second、first;第一个包内路径是 EPUB/chapter-b.xhtml。代码断言这个顺序并检查首章文字,证明演示中读取顺序来自 spine,而不是文件名。

打开 epub_chapters.json,核对 chapter_order、item_id、package_entry 和 text 是否与包文档引用一致。每条记录保留 source_file 与包内章节路径,后续分段时继续带上这些字段,回答引用才能回到具体章节。

本例把 XHTML body 内的文字合并为字符串,并去掉 script、style 节点。得到的是可核对的文字输入,不是原书版式;图片中的文字、图表、公式和脚注关系没有被完整保留。

处理你自己的电子书

  1. 使用自己有权处理的 EPUB 副本,保留原书。
  2. 删除创建 sample 文件的演示部分,保留 extract_epub 函数,调用 extract_epub('实际书名.epub')。
  3. 先检查书中第一章与最后一章,再检查目录、序言和附录是否按照你的任务需要被保留。
  4. 本例只选择 container 中第一个 rootfile;多种呈现版本的电子书需要先选择正确包文档,不能默认第一个版本就适合任务。
  5. 将章节分段时保存稳定书籍 ID、版本、章节路径和片段位置,再送入阅读工具支持的导入格式;JSON 本身不代表所有阅读工具都能直接导入。

解析失败怎样判断

情况 处理方式
container 找不到 rootfile 确认文件确实是 EPUB,并检查包结构,不猜包文档名称
spine 的 idref 没有对应 manifest 项 核对原包的引用关系,保存错误位置再修复来源文件
章节 media-type 不是 XHTML 本例明确拒绝,需要为该媒体类型另加处理器
XHTML 不完整、没有 body 或正文为空 回读原章节,确认是否为图片型内容或格式异常
输出顺序正确但关键信息缺失 检查辅助章节、图像和脚注,不能把顺序通过当作内容完整

本例只接受没有 META-INF/encryption.xml 声明、且 spine 项为文字 XHTML 的输入;发现该声明就停止。encryption.xml 也可能用于字体混淆,并不必然表示整书有 DRM;这里采用的是明确的简化边界,没有识别各种声明后继续处理的实现。

如果输出为空,先检查 spine 是否有主要阅读项目,再决定增加辅助内容处理;没有可用文字就先停在资料转换阶段。本文不绕过电子书保护,也不把无法读取的章节交给模型猜写。

为什么输出没有原书页码?

流式电子书显示的页码会随排版设置变化;本例的来源定位是书籍文件和包内章节路径。若需要引用到原书特定位置,进一步保留元素 ID 或出版方提供的位置标记,再核对阅读工具的定位能力。本例没有生成纸书页码映射。

Ai菜鸟网。发布者:AI小管家,转载请注明出处:https://www.alyyhw.com/31903.html

赞 (0)
AI小管家的头像AI小管家
Markdown 知识库怎么分段?按标题保存层级并避免切断代码块
上一篇 2小时前
RSS 智能体怎么获取资讯?解析订阅条目并按稳定 ID 去重
下一篇 2小时前

相关推荐

联系我们

联系我们

1

在线咨询: QQ交谈

邮件:admin@example.com

工作时间:周一至周五,9:30-18:30,节假日休息

关注微信
关注微信
分享本页
返回顶部