If you feed documents to a language model, you probably convert them to Markdown first, and you probably don't want a Word file to leave your laptop just to have that done. The convilyn package on PyPI has an offline half that does it locally: no account, no API key, nothing uploaded. It is the same conversion engine the hosted workflows run, published as a library.

This walks through version 4.1.0 as it installs from PyPI today. Every command and output below comes from a fresh virtual environment on my Windows machine, including the parts that didn't go the way the docs suggest.

Install only the formats you need

pip install "convilyn[documents]"

The bare package converts plain text and CSV and nothing else. Each document format is an extra, so you only pull in the parsers you use. pip install convilyn followed by a .docx fails with a message telling you to add the docx extra, which is correct but not a great first minute.

ExtraAdds
pdf, docx, pptx, xlsx, xmlone format each
documentsall five of the above
imagesimage format conversion (Pillow)
alldocuments + images

OpenDocument, legacy .doc/.xls and ebook formats go through LibreOffice and Calibre, which are desktop applications rather than Python packages, so no extra can install them.

One Windows trap I hit while writing this: installing into a virtualenv under a deeply nested folder failed with OSError: [Errno 2] No such file or directory on a license file inside pypdfium2, a dependency of the PDF extra. The path passed Windows' 260-character limit. Either enable long paths in Windows or put the virtualenv somewhere short.

Ask what your machine can do

$ convilyn local doctor
✓ pdfplumber: installed
✓ pypdf: installed
✓ docx: installed
✓ pptx: installed
✓ openpyxl: installed
✓ defusedxml: installed
✓ PIL: installed
✓ libreoffice: installed
! calibre: missing — Install Calibre from https://calibre-ebook.com/download (provides `ebook-convert`).
! ffmpeg: missing — Install FFmpeg from https://ffmpeg.org/download.html (provides `ffmpeg`).
287 of 737 conversions available. Run `convilyn local formats` for the per-format detail.

Your count will differ; this machine happens to have LibreOffice. convilyn local formats lists every route and why an unavailable one is unavailable.

Convert one file

$ convilyn local convert docs/q3-report.docx --to md
▶ Converting q3-report.docx → md
✓ Wrote docs\q3-report.md
Converted q3-report.docx → docs\q3-report.md

The test file had a heading, a paragraph, a bullet list and a table. The output, in full:

# Q3 Support Report

Ticket volume rose in July and fell back in September. Median first response stayed under two hours.

## What changed

- Added a self-serve refund form
- Moved billing questions to a separate queue
- Retired the legacy email alias

## Numbers

| Month | Tickets | Median first response |
| --- | --- | --- |
| July | 1,240 | 1h 50m |
| August | 1,105 | 1h 35m |
| September | 960 | 1h 20m |

Headings, the list and the table came through intact. Two phrases in that paragraph were bold and italic in the Word file, and they came out as plain text. For a model's input that rarely matters, but if you need the emphasis preserved, this isn't the tool for it.

I converted the same file three times and got the same SHA-256 each time, and the PDFs below behaved the same way. The package's published test report also compares offline and hosted output on the 21 test documents both can read, and they are byte-for-byte identical.

Spreadsheets: a formula needs a cached value

Each sheet becomes a ## heading with its own table. I built a two-sheet workbook with openpyxl and put =SUM() formulas in the last row:

## Tickets

| Month | Tickets | Resolved |
| --- | --- | --- |
| July | 1240 | 1198 |
| August | 1105 | 1090 |
| September | 960 | 955 |
| Total |  |  |

The totals are blank, and the result carries a warning that 2 formula cells had no cached value because the workbook had never been recalculated by a spreadsheet application. An .xlsx file stores a formula and, separately, the last value it produced. A file written by a library such as openpyxl has the formula and no value, and the converter reads values rather than evaluating formulas. Passing the file through LibreOffice once fixes it:

soffice --headless --convert-to xlsx --outdir recalced docs/q3-numbers.xlsx

Converting the recalculated copy gave | Total | 3305 | 3243 | and no warning. Files that people saved from Excel already have cached values, so this mostly bites generated reports.

From Python

from pathlib import Path
from convilyn import local

result = local.convert("docs/q3-report.docx", to="md")
print(result.ok, result.output, result.warnings)

results = local.convert_many(sorted(Path("docs").glob("*.pdf")), to="md", out_dir="build")
for r in results:
    if not r.ok:
        print("failed:", r.source, r.error)
    elif any(w.startswith("no text layer") for w in r.warnings):
        print("needs OCR:", r.source)

convert_many returns one ConversionResult per input and doesn't raise on a failed file unless you pass raise_on_error=True, so check ok yourself. The warnings are plain strings today, which is why the scan check matches on the start of the message.

Embedded images, and the ones it drops

Images inside a document are written next to the Markdown and linked from it:

# Site visit

Photo from the loading dock.

![Picture 1](assets/img-0001.jpg)

The report converted above also contained a flat two-colour bar chart, and it disappeared without a warning. The engine treats an image as decorative (a spacer, a rule, a bullet) and skips it if it has fewer than 8 distinct colours, is under 2 KB or 4,096 pixels in area, or is more than 8 times as long as it is wide. A photo passes easily. A simple logo or a two-colour diagram may not. In a batch, each file's images go under assets/<file name>/, so two documents can't overwrite each other's img-0001.jpg.

PDFs: tables need lines, scans need OCR

A PDF has no table markup, only text placed on a page. In my tests the engine picked up a table only when it had ruling lines. The same three-row table converted two ways:

Two rows. Top: a PDF table drawn with grid lines converts to a proper Markdown table with a header row. Bottom: the same table with no lines converts to two lines of plain text, read one column at a time.
With grid lines the table survives. Without them the cells come out as text, one column at a time.

If your PDFs come from a report generator that draws borderless tables, expect the second result and check a sample before trusting a whole folder. Headings in PDFs are inferred from font size, so the result is best effort: the same title came out as a ## heading in one of my files and as plain text in another.

Scanned pages are the other limit. A scan is a picture of text with no text layer, and the engine doesn't do OCR:

$ convilyn local convert scan.pdf --to md
▶ Converting scan.pdf → md
! no text layer found — this PDF is probably a scan
! best_effort: PDF carries no structure; headings inferred from font size
✓ Wrote scan.md
Converted scan.pdf → scan.md

That command exits with status 0 and writes an empty file. A script that only checks the exit code will think it worked. Use --json, which prints one object with a warnings array, or the Python check above. In the package's measurements on 650 real-world PDFs, 145 had no text layer at all, so this isn't a corner case. Those files need OCR, which the hosted API runs as a paid workflow.

HTML is refused outright, with exit status 1:

✗ Unsupported conversion: No route from html to md. This engine cannot read html at all.

That one is a licensing decision rather than a gap nobody got to. The good HTML-to-text libraries are GPL-licensed, and this package is Apache-2.0.

Batch conversion, and where the glob gets expanded

$ convilyn local batch docs/*.pdf --to md --out-dir build/
▶ [1/2] onboarding.pdf
▶ [2/2] pricing-faq.pdf
✓ Converted 2 of 2 files
Converted 2 of 2 files into build.

The command has no glob handling of its own. On macOS and Linux the shell expands docs/*.pdf, so if you quote the pattern the command receives the literal string and reports that the file doesn't exist. The package's own quick-start had exactly that bug until 4.1.0. On Windows it works quoted or unquoted, which surprised me: Click, the library the CLI is built on, expands wildcards itself on Windows because the Windows shells don't. I checked it from cmd.exe and from Git Bash. In Python, convert_many takes paths, so you glob them yourself and the question doesn't come up.

How good is the output?

The package ships a dated test report scored by a separate deterministic evaluator. On its synthetic document set, text fidelity is 0.9664 (normalised edit distance to a golden Markdown) and table structure scores 0.92 on TEDS, a tree-edit similarity measure for tables. The report is also clear about where the engine struggles: dense small-type pages such as dictionaries come out with words in the wrong order, though the words themselves are all there. If your documents look like that, run your own sample first.