Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

PDF to Markdown Converter

Turn a PDF into Markdown or HTML — headings, lists, tables, code and pictures included.

PDF No upload Free to try, no sign-up Pro tool Pro pass from ₹79

Next steps

About the PDF to Markdown Converter

Open a PDF and get clean Markdown — or semantic HTML — that keeps the structure of the document: larger type becomes # headings, bold and italic stay **bold** and *italic*, bullets and numbers become real lists (nested by their indent), web links stay links, monospaced text becomes inline code and fenced code blocks, and tables become GitHub-style pipe tables. Running headers, footers and page numbers are left out, and words split by a hyphen at the end of a line are joined again.

Pictures can come along in a ZIP next to the Markdown file, be embedded in it, or be left out. Scanned pages can be read with OCR, right here. Everything happens in your browser — the PDF is not uploaded.

How to use it

  1. Drop a PDF onto the box above, or tap Choose a PDF. If it is password-protected, enter the password.
  2. Pick Markdown or HTML, choose what happens to Pictures, and untick anything you don’t want — for example tables or the removal of headers and footers. To convert only part of the PDF, type the Pages, such as 1-5 or 3, 7-9.
  3. Press Convert, then check the result: switch between the source and the Preview.
  4. Press Download (a .md or .html file, or a .zip when it has pictures) or Copy the text. If some pages are scans, press Read with OCR first.

Examples

A report page
Input
A 22 pt bold title, 15 pt section titles, body text with a bold word and a link, a bullet list with a sub-item, a ruled table and a logo
Result
# Annual Plan 2027

## 1. Goals

We will grow **revenue** … Read more on [our website](https://example.com/).

- Hire two engineers
- Open an office in Pune
  - near the station

| Item | Cost | Owner |
| --- | ---: | --- |
| Laptops | 4,00,000 | Asha |

![Figure 1: The new office](images/page-1-1.png)
Code set in a monospaced font
Input
function total(items) {
    return items.reduce((a, b) => a + b, 0);
}
Result
A fenced code block (```), with the indentation of each line kept

Common uses

  • Moving a PDF manual, policy or report into a wiki, a docs site, GitHub or a Markdown note app.
  • Pasting a long PDF into an AI assistant as plain text that still shows its headings, lists and tables.
  • Getting clean HTML out of a PDF to publish on a website or in a CMS, without the PDF’s fixed layout.
  • Reusing the tables of a PDF as Markdown or HTML tables instead of retyping them.

How the structure is rebuilt

The page content of a PDF is positioned text, lines and images (ISO 32000-1:2008, §9.10 describes how text is extracted from it); headings, lists and tables are only the way it looks. The converter reads that content with pdf.js — every PDF has it, while structure tags are optional — and rebuilds the rest:

  • Headings come from type sizes. Text clearly larger than the body text is ranked by size: the largest becomes #, the next ##, down to ######. A short bold line that introduces ordinary text, such as “Payment terms” or “1. Definitions”, becomes the next level down.
  • Lists: lines that start with •, –, 1., a), (iv), a tick box or a bullet drawn as a small shape are list items, nested by how far they are indented. Numbered lists keep the number they start at; ☐ and ☑ become GitHub task lists. Markdown has no lettered or Roman-numeral lists, so those keep their labels as text (the HTML output uses real <ol type="a"> lists).
  • Code: lines set entirely in a monospaced font become a fenced code block, with indentation rebuilt from the character positions; monospaced words inside a sentence become inline code.
  • Tables are rebuilt from ruling lines (including merged cells) or from clearly aligned columns. Columns of numbers are right-aligned, and a table that runs onto the next page is joined, with its repeated header row dropped.
  • Pages: text in two or three columns is read column by column, and a paragraph cut by a column or page break is joined again.

Markdown that renders as you expect

The Markdown follows CommonMark, with the pipe tables and task lists of GitHub Flavored Markdown (GFM), which GitHub, GitLab and most Markdown editors understand. Text from the PDF is kept literal: characters that Markdown would read as formatting — *, _, #, [, a number and full stop at the start of a line — are escaped with a backslash, while those that can’t change the meaning where they stand, such as _ inside a word or # in the middle of a line, are left alone, so the result stays readable. A | inside a table cell is written as \|, as GFM requires.

The HTML option writes semantic tags only (h1–h6, p, ul/ol, pre/code, table with thead, figure, strong, em, a) — no classes or scripts, and no styles except text-align: right on columns of numbers — so it drops into a CMS cleanly. The downloaded .html file adds a small stylesheet so it is readable on its own.

Scanned pages and unreadable text

A scanned page is a picture of text. When the PDF has pages like that, the result offers Read with OCR: the pages are recognised on your device with Tesseract, in English, Hindi and other supported languages, and the text is rebuilt into headings, paragraphs, lists and ruled tables like any other page. The OCR engine and the language data are downloaded from this site the first time (a few megabytes per language).

OCR is also offered for pages whose fonts don’t say which letters they draw — common with Hindi and other complex scripts in some PDFs, where copied text comes out with missing letters.

Limitations

  • Layouts with sidebars, text boxes, pull quotes or forms convert only approximately: the text ends up in reading order, not in the same boxes.
  • Markdown tables can’t merge cells or hold several paragraphs per cell; merged cells keep their text in the first cell (the HTML output keeps colspan and rowspan).
  • Equations, footnote links, colours, underlines and fonts are not carried over; charts drawn as vector graphics are lost (pictures stored in the PDF are copied).
  • Text read with OCR can contain mistakes, especially in poor scans, and has no bold or italic. Check names, numbers and amounts.
  • Headings are guessed from type sizes and bold text, so a document that uses large type for other things may get extra headings.

Privacy

Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.

Frequently asked questions

Is my PDF uploaded?

No. The PDF is read and converted inside your browser, on your device. When you use OCR, the OCR engine and language data are downloaded from this site, but your pages never leave your device.

What kind of Markdown does it write?

CommonMark, with the tables and task lists of GitHub Flavored Markdown. It shows correctly on GitHub and GitLab and in most Markdown editors, note apps and static-site generators. Tick Start with front matter to add a YAML block with the title, author and source file, as Jekyll, Hugo, Astro and Obsidian read it.

Where do the pictures go?

With Save in a ZIP you get a .zip holding the Markdown (or HTML) file and an images folder, and the document refers to images/page-3-1.png and so on. Embed in the file puts the pictures inside the document as data URLs — one file, but some sites strip such pictures. Leave out writes text only. A caption next to a picture, such as “Figure 2: …”, becomes its alternative text.

Why did a table come out as plain text?

Tables are found from their ruling lines or from text that lines up in clear columns over several rows. A table with no lines and ragged columns can’t be told apart from ordinary text, so it is kept as paragraphs. Pages that are tables only, like bank statements, convert better with PDF to Excel.

Can I convert a scanned PDF?

Yes. After converting, press Read with OCR for the scanned pages, pick the language of the text and the pages are recognised on your device. You can also make the whole PDF searchable first with OCR PDF.

Can I mark where each page starts?

Yes — tick Mark where each page starts. Each page then begins with an HTML comment such as <!-- Page 3 -->, which Markdown and HTML viewers don’t show but search, scripts and AI assistants can use to cite pages.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.