pdf如何用python读取？-Python常见问题-Python学习网

pdf如何用python读取？

yang2020-05-28 16:00:37原创

python中可以使用pdfminer库来读取PDF文件中的内容。

安装命令：

pip install pdfminer

pip install pdfminer3k

python中读取PDF文件代码：

from urllib.request import urlopen
from pdfminer.pdfinterp import PDFResourceManager, process_pdf
from pdfminer.converter import TextConverter
from pdfminer.layout import LAParams
from io import StringIO
from io import open

def readPDF(pdfFile):
    rsrcmgr = PDFResourceManager()
    retstr = StringIO()
    laparams = LAParams()
    device = TextConverter(rsrcmgr, retstr, laparams=laparams)

    process_pdf(rsrcmgr, device, pdfFile)
    device.close()

    content = retstr.getvalue()
    retstr.close()
    return content

pdfFile = urlopen("http://pythonscraping.com/pages/warandpeace/chapter1.pdf")
outputString = readPDF(pdfFile)
print(outputString)
pdfFile.close()

解析pdf文件用到的类：

PDFParser：从一个文件中获取数据
PDFDocument：保存获取的数据，和PDFParser是相互关联的
PDFPageInterpreter处理页面内容
PDFDevice将其翻译成你需要的格式
PDFResourceManager用于存储共享资源，如字体或图像。

更多Python知识请关注Python自学网

专题推荐：python

Python3 Selenium3 自动化测试开发实战

本套Python自动化测试教程零基础讲解自动化测试， selenium 安装到八种元素定位，用户事件处理，等待时间处理，到单元测试框架 Unitest 整合实战，整合自动化测试项目实战，新版本HTML TestRnner 生成测试报告，自动化发送测试报告邮件等核心知识点

pdf如何用python读取？

相关文章推荐

相关课程推荐

Python3 Selenium3 自动化测试开发实战

《python一小时快速实战入门》（微软官方）

《Develop with Python on Windows》（微软官方-中文版）

全部评论我要评论

Python学习网