Skip to content

Python PDF guide

Convert PDF to Markdown in Python: A Practical Guide

For a repeatable PDF to Markdown Python workflow, use a document-aware library, write the result as UTF-8, and review the parts that layout extraction cannot guarantee. This guide starts with a short PyMuPDF4LLM example, then covers scanned pages, batches, embedded images, and a simple quality check.

Editorial illustration of a PDF report being converted into structured Markdown with Python
A Python converter can turn a PDF text layer into editable Markdown, but the result still needs a quick structure check.

Choose a Python PDF to Markdown approach

Choose around the PDF layouts and file types you actually have. For PDF-focused Python work, PyMuPDF4LLM is a practical first package to evaluate: it can return Markdown, JSON, or text, and offers layout, page-chunk, OCR, and image options. If your input queue also includes Word, PowerPoint, Excel, and other formats, compare Microsoft's MarkItDown project. Its maintainers describe it as an LLM-oriented converter and caution that the output may not suit high-fidelity human conversion. A browser converter is simpler for an occasional file; a hosted API fits a backend workflow but adds file-transfer and service controls. Compare representative documents before choosing.

MethodGood fitTrade-off to plan for
PyMuPDF4LLMPDF-focused Python extraction with Markdown, JSON/text, layout, page-chunk, OCR, and image optionsReview output on your PDFs; the repository declares AGPL-3.0, so check the current license for your distribution.
Microsoft MarkItDownMixed-format Markdown workflows that include PDF, Office files, images, and other inputsUseful for broad ingestion; its maintainers caution that output may not be ideal for high-fidelity human conversion.
Browser converterA one-off conversion without a local Python environmentLess suitable for scheduled batch jobs or custom pipeline control.
Hosted conversion APIAn application that already runs a backend serviceReview credentials, upload/retention behavior, network errors, and retries.

Convert one PDF to Markdown with Python

PyMuPDF4LLM provides a direct to_markdown() call. The method returns Markdown text, so a small script can decide where the output goes and how to name it. Install the package in the environment that will run the script, then begin with one representative PDF rather than a whole archive. The basic example below uses the package name and method shown in the project documentation.

Minimal conversion example
import pymupdf4llm
from pathlib import Path

source = Path("report.pdf")
markdown = pymupdf4llm.to_markdown(str(source))
output = source.with_suffix(".md")
output.write_text(markdown, encoding="utf-8")

print("Saved", output)

Install the library and save the Markdown file

Use an isolated virtual environment when the conversion script is part of a project, so its PDF dependencies do not unexpectedly affect other Python work. The documented package install command is pip install pymupdf4llm. Once installed, the example writes beside the source PDF and explicitly uses UTF-8, which avoids relying on the machine’s default text encoding. If your application accepts files from users, validate the extension and size before parsing, choose an output directory deliberately, and avoid placing generated files in a public web directory by accident.

A useful first run is intentionally small: convert one file, open the .md output in a plain-text editor, and compare it with the PDF. This tells you whether the library’s defaults are suitable before you build a batch job around them. The exact package options can change, so check the current official API reference before depending on optional arguments in a long-lived pipeline.

Handle scanned PDFs and OCR deliberately

PyMuPDF4LLM checks pages and uses OCR automatically when it detects that recognition is needed. The official documentation checked on 2026-10-02 describes Tesseract and RapidOCR adapters. Make sure an OCR engine is available in the same environment that runs Python; Tesseract also needs the language data for the document. If a scan returns little text, first check the installed package version, engine, and language data. Use force_ocr=True only when the native text layer is corrupt or automatic detection missed a page: it bypasses the normal check and forces OCR, which can slow clean digital pages and reduce their quality. Review names, identifiers, totals, and mixed-language text against the scan.

Conceptual comparison of selectable PDF text and a scanned page that needs OCR before Markdown extraction
Selectable text can be extracted directly; a page image needs OCR first, and recognized text should be reviewed.
Explicit OCR example from the documented API
import pymupdf4llm

# OCR runs automatically when the installed engine detects that a page needs it.
markdown = pymupdf4llm.to_markdown("scanned-report.pdf")

# With Tesseract, choose language data installed in the environment.
markdown = pymupdf4llm.to_markdown(
    "scanned-report.pdf",
    ocr_language="eng+chi_sim",
)

Process multiple PDFs with predictable paths

Batch conversion is mostly a file-management problem. Keep source files separate from output files, create the destination folder before the loop, and derive each Markdown filename from the PDF stem. Catch errors per file so one damaged or password-protected PDF does not silently stop the entire folder. Record failures in a log that an operator can inspect; do not silently mark the batch as complete. For large jobs, add bounded concurrency only after measuring memory use and the converter’s documented behavior.

The example below handles the top level of one input folder. If documents are nested, use rglob("*.pdf") and preserve relative paths so two files named report.pdf do not overwrite each other. If conversion is a recurring backend workload rather than a local script, compare this approach with a job-based API and its authentication, retry, and retention model.

Editorial flow diagram showing a folder of PDFs entering a Python conversion script and producing separate Markdown files
A batch script should keep inputs, Markdown outputs, and per-file errors easy to distinguish.
Simple folder batch with a separate output directory
from pathlib import Path
import logging
import pymupdf4llm

source_dir = Path("pdfs")
output_dir = Path("markdown")
output_dir.mkdir(parents=True, exist_ok=True)
logging.basicConfig(
    filename=output_dir / "conversion.log",
    level=logging.ERROR,
    format="%(asctime)s %(levelname)s %(message)s",
)

for pdf_path in sorted(source_dir.glob("*.pdf")):
    try:
        markdown = pymupdf4llm.to_markdown(str(pdf_path))
        target = output_dir / (pdf_path.stem + ".md")
        target.write_text(markdown, encoding="utf-8")
    except Exception:
        logging.exception("Could not convert %s", pdf_path.name)

Keep images and tables useful in Markdown

Plain Markdown can represent headings, lists, links, and many simple tables. More complex tables, charts, equations, and page artwork may need a different representation or a separate image asset. If your output depends on extracted images, review the current documented options for write_images, image paths, and formats; keep those files alongside the Markdown and verify that generated references resolve from the place where the .md file will be read. Do not assume that every figure can be converted into descriptive text or that visual layout survives as-is.

For a knowledge base, preserve enough context for a reader to trace an important passage back to its source page. Page-separated output or metadata may be useful, but choose based on the retriever and citation behavior you need. For a browser-based workflow focused on exported images and tables, see the PDF to Markdown images and tables guide. The existing PDF to Markdown for RAG guide covers chunking and retrieval preparation; this tutorial stays focused on the Python extraction step.

Save image assets alongside the Markdown output
markdown = pymupdf4llm.to_markdown(
    "report.pdf",
    write_images=True,
    image_path="output_assets",
    image_format="png",
)

Diagnose common Python PDF to Markdown failures

If Python cannot import the package, first check which interpreter runs the script. Run python -m pip show pymupdf4llm from that same environment; a notebook, terminal, or scheduled job may use a different interpreter. A FileNotFoundError often means a relative path is being resolved from an unexpected working directory. Print source.resolve() while debugging or test with an absolute path.

If a scanned PDF produces little text, confirm that an OCR engine and the needed language data are installed in the environment. Keep force_ocr=True for pages whose text layer is missing or unreliable; clean digital pages usually do not need forced OCR. When the Markdown points to missing images, keep the exported asset directory with the .md file and check the relative paths.

For a batch job, record a success or failure for every input and save the exception details to a log. A generated file only proves that the parser returned something. It does not prove that every page, table, or recognized value is correct.

Review the Markdown before you rely on it

A successful function call only means the parser returned output. It does not prove that the document’s structure was reconstructed correctly. Before indexing, publishing, or extracting business facts, compare a few pages against the source PDF. Include an opening page, a page with columns, a table page, a page with footnotes or repeated headers, and a scanned page if your corpus contains scans. Fix the extraction configuration or route difficult files to another method when the output fails your requirements.

CheckWhat to compareWhy it matters
Headings and listsCompare hierarchy, numbering, and list nestingWrong levels make long Markdown hard to scan
Reading orderCheck columns, sidebars, footnotes, and page transitionsMisordered text can change a sentence’s meaning
Tables and figuresVerify rows, captions, image references, and symbolsBroken structure can misstate values or remove evidence
OCR textCheck names, dates, quantities, and uncommon termsRecognition errors can flow into search or summaries
File handlingConfirm output count, filenames, encoding, and logged failuresA partial batch should not look complete

Know what a PDF to Markdown converter cannot guarantee

PDF records where content is drawn on a page; it does not always contain the semantic structure that Markdown needs. A two-column paper, a form, an engineering drawing, and a clean text report are different extraction problems. Layout-aware software can make useful guesses, but output should be judged against the pages and the downstream task. If the document contains tables, formulas, or diagrams that must remain exact, keep the source PDF and use page references or separate assets alongside the Markdown.

Processing locally can keep the file inside the environment you control, but security also depends on dependencies, storage paths, logs, backups, and process isolation. Microsoft's MarkItDown documentation notes that conversion runs with the current process's access; isolate jobs and narrow file access when processing untrusted documents. The PyMuPDF4LLM repository declares AGPL-3.0; review the current license and dependency terms before distributing a product that uses it. A hosted API may reduce the work of managing a parser while requiring you to review upload and retention terms. For a managed workflow, see the PDF to Markdown API guide; for a quick selectable-text preview, return to the online PDF to Markdown converter.

PDF to Markdown Python questions

What is a practical Python library for PDF to Markdown conversion?

PyMuPDF4LLM is one practical option when you want a local Python call that returns Markdown. Start with one representative document, then compare headings, reading order, tables, images, and scans with your requirements. Other libraries and hosted services may fit different layouts or operational constraints.

How do I convert a scanned PDF to Markdown in Python?

PyMuPDF4LLM can run OCR automatically when an OCR engine is available and the page needs recognition. Install the engine and language data for your environment; with Tesseract, set a suitable language when necessary. Use force_ocr only if automatic detection missed a page or its text layer is unreliable, then review the output against the scan.

Can Python keep PDF tables and images in Markdown?

Some tools can rebuild simple tables or write extracted image assets, but results vary by layout. Check the table and image options for the library version you install, keep any asset directory with the Markdown file, and inspect complex pages manually.

How can I batch convert PDFs to Markdown with Python?

Iterate over a known input folder, create a separate output directory, use each PDF filename stem for its Markdown file, and handle exceptions per file. For recursive folders, preserve relative paths so files with the same name do not overwrite one another.

Should I use a Python library or a PDF to Markdown API?

Choose a local library when you need direct control over files and can maintain dependencies. An API may fit a server workflow that already needs authentication, job status, and managed processing. Compare upload, retention, retries, operational cost, and output quality for your documents before choosing.

Does a successful conversion mean the Markdown is accurate?

No. A successful call confirms that output was produced, not that columns, tables, figures, or OCR were reconstructed correctly. Sample pages from each document type and keep the source PDF available for verification.

How do I convert a PDF to Markdown in Python?

Call pymupdf4llm.to_markdown() on the source file, write the returned string with UTF-8 encoding, and compare the result with representative PDF pages before automating the workflow.

A reliable starting workflow

For a local script, install the documented package, convert one file, write the result as UTF-8, and compare it with the source. Add OCR, batch handling, and image extraction only when the document set needs them. Keep failed-file logs, and check output before using it for search or decisions. If you do not need a repeatable Python workflow, use the browser PDF to Markdown converter; for managed jobs, see the PDF conversion API guide.

Official documentation