Word to HTML Converter
Paste from Word or Google Docs, or open a .docx — get clean, semantic HTML or Markdown.
Paste formatted text
Headings, lists, tables, links and bold/italic are read from the clipboard; Office styles are thrown away. Nothing is uploaded.
Paste HTML source (for example from a CMS, an email or Word’s “Save as Web Page”) to clean it. Plain text becomes paragraphs.
Result
Preview
Notes
About the Word to HTML Converter
Copy a document out of Word and paste it into a website, and you usually get hundreds of class="MsoNormal", <span style="font-family:Calibri">, <o:p> tags and empty paragraphs. Google Docs adds a <span> with a dozen styles around every word and wraps the whole thing in a bold tag that is not bold. This converter keeps what the document means — headings, paragraphs, bulleted and numbered lists (nested too), tables, links, bold, italic, strikethrough, superscript, subscript, code and images — and throws the rest away.
Open a .docx file (read with the open-source mammoth converter), paste formatted text, or paste messy HTML code. You see a live preview and the resulting code side by side, and can download HTML, Markdown, or a ZIP with the page and its images. Everything happens in your browser; the document is never uploaded.
How to use it
- Choose where your document is: Paste from Word or Google Docs (click the box and press Ctrl+V / ⌘V, or use Paste), Open a .docx file, or HTML code or text.
- Pick HTML or Markdown, and what to do with images: embed them, save them as separate files in a ZIP, or remove them.
- Adjust the clean-up: keep or remove classes and IDs, keep underlining, replace non-breaking spaces, or wrap the result in a complete HTML page.
- Check the Preview and the code, read any notes (for example images that could not be included), then Copy or Download.
Examples
<p class=MsoNormal><span style='font-family:"Calibri",sans-serif'>Hello <b>world</b><o:p></o:p></span></p>
<p>Hello <strong>world</strong></p>
<span style="font-weight:700;font-family:Arial">Bold</span> and <span style="font-style:italic">italic</span>
<p><strong>Bold</strong> and <em>italic</em></p>
<p class=MsoListParagraphCxSpFirst style="mso-list:l0 level1 lfo1"><span style="mso-list:Ignore">1.</span>Plan</p> <p class=MsoListParagraphCxSpLast style="mso-list:l0 level1 lfo1"><span style="mso-list:Ignore">2.</span>Do</p>
<ol> <li>Plan</li> <li>Do</li> </ol>
<h2>Agenda</h2><ul><li><strong>10:00</strong> Welcome</li></ul>
## Agenda - **10:00** Welcome
Common uses
- Publishing a Word or Google Docs draft in WordPress, a CMS, a blog or a help centre without the hidden formatting that breaks the theme.
- Turning documentation written in Word into Markdown for GitHub, a static site or a wiki.
- Cleaning HTML exported by Word (“Save as Web Page”) or pasted from an email or another website.
- Getting a document’s images out as separate PNG/JPEG files next to the HTML.
What is kept and what is removed
Kept: headings h1–h6, paragraphs, line breaks, bulleted and numbered lists (with their start number and a/i/A/I numbering styles), tables with merged cells, links (only http, https, mailto, tel and in-page links), bold, italic, strikethrough, superscript, subscript, code (text in a monospace font becomes <code>), block quotes, footnote links and images with their alternative text.
Removed: inline styles, fonts, colours and sizes, <span> and <font> wrappers, Office tags such as <o:p>, Word’s conditional comments and XML, empty paragraphs, page breaks, classes and IDs (unless you keep them; IDs that links point to are always kept), and anything unsafe: scripts, event handlers, forms, frames and javascript: links.
How Word and Google Docs lists are rebuilt
When you copy a list out of Word, the clipboard does not contain a list at all: each item is a separate paragraph with an mso-list:l0 level2 style and the bullet or number as text in front. The converter groups those paragraphs, reads the level, and recognises the marker — “1.”, “a)”, “iv.” mean a numbered list, “·”, “o”, “§” a bulleted one — to rebuild properly nested <ul> and <ol> lists, keeping a list’s start number. Google Docs puts nested lists in an invalid place (directly inside the parent list); they are moved into the item they belong to.
Images
From a .docx file, every image is included: embedded in the HTML as a data URI, saved as a separate file in a ZIP (images/image-1.png …), or removed. Images in Windows Metafile format (EMF/WMF) are kept but browsers cannot display them; the tool tells you when this happens.
When you paste from Word, the clipboard only refers to temporary image files on your computer that a web page is not allowed to read, so those images are left out — open the .docx instead. Google Docs pastes links to images on Google’s servers; they are kept as links (the preview does not load them).
Limitations
- Old Word 97–2003
.docfiles cannot be read in a browser. Save them as.docxin Word, LibreOffice or Google Docs first. - Page layout is not converted: headers, footers, page numbers, columns, text colours, fonts, alignment and table borders are dropped on purpose. Text boxes appear as normal paragraphs after the paragraph that held them.
- Equations, charts, SmartArt and drawings are not converted. Comments and tracked-change deletions in a .docx are not included.
- Markdown tables need a header row; when a Word table has none, its first row becomes the header. Tables with merged cells or lists inside cells stay as HTML inside the Markdown.
- Very large documents (hundreds of pages, or many large images) can take a few seconds and a lot of memory, especially with images embedded.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
How do I copy Word content into WordPress without the junk?
Copy the text in Word, paste it into the box here, choose HTML and keep Remove classes and IDs ticked, then Copy HTML and paste it into the WordPress code editor (or a Custom HTML block). Upload the images separately, or use Save as separate files and upload the images folder.
Why are my pasted images missing?
Word puts images on the clipboard as references to temporary files on your computer, which websites are not allowed to read. Open the .docx file instead (Open a .docx file) and every image is included.
Does it keep bold, italic and links from Google Docs?
Yes. Google Docs writes formatting as styles on <span> tags; the converter turns font-weight:700 into <strong>, italic into <em>, strikethrough into <s>, superscript and subscript into <sup>/<sub>, monospace text into <code>, and keeps links (unwrapping Google’s google.com/url?q= redirect links).
Can I convert Word to Markdown?
Yes — choose Markdown. Headings become #, lists - or 1., bold **…**, links [text](url), tables GitHub-style tables, strikethrough ~~…~~, and paragraphs in Word’s code styles fenced code blocks. Superscript and subscript stay as HTML tags, which Markdown allows. Text that would otherwise turn into markup, such as a literal <div> or ©, is escaped with a backslash so it still reads the same.
Is the preview safe if I paste HTML from an unknown source?
Yes. Pasted HTML is read into an inactive document where nothing runs or loads, cleaned against an allowlist, and sanitised again with DOMPurify before it is shown. Scripts, event handlers, frames and javascript: links are removed, and images from other servers are shown as placeholders instead of being loaded.
Is my document uploaded?
No. The .docx file and anything you paste are converted in your browser. Nothing is sent to a server.