# Converting html to Word ?

**URL:** <https://boards.straightdope.com/t/converting-html-to-word/482628>\
**Category:** Factual Questions\
**Created:** [January 23, 2009, 1:30pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628 "2009-01-23T13:30:09Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![C\_K\_Dexter\_Haven](https://avatars.discourse-cdn.com/v4/letter/c/b2d939/32.png) [@C\_K\_Dexter\_Haven](https://boards.straightdope.com/u/C_K_Dexter_Haven)\
**Post date:** [January 23, 2009, 1:30pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/1 "2009-01-23T13:30:09Z")

</div>

I’ve received a fairly long document in html, with all the \<codes\> and such. Is there some easy way to convert this into a normal Word document? That is, replacing all the &146; with ’ etc? It’s long, so going through and fixing by hand seems like a long and tedious process (which is liable to miss stuff anyhow). Is there some simpler way?

Thanks!

---

<div class="post-metadata">

**Author:** ![jjimm](https://avatars.discourse-cdn.com/v4/letter/j/ba8739/32.png) [@jjimm](https://boards.straightdope.com/u/jjimm)\
**Post date:** [January 23, 2009, 1:40pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/2 "2009-01-23T13:40:26Z")

</div>

Have you tried just opening the HTML doc in Word and saving as a .doc?

I just tested this on Word 2003 and it worked fine (or as fine as you’d expect MS Office apps ever manage to handle HTML).

---

<div class="post-metadata">

**Author:** ![C\_K\_Dexter\_Haven](https://avatars.discourse-cdn.com/v4/letter/c/b2d939/32.png) [@C\_K\_Dexter\_Haven](https://boards.straightdope.com/u/C_K_Dexter_Haven)\
**Post date:** [January 23, 2009, 1:43pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/3 "2009-01-23T13:43:28Z")

</div>

That didn’t do it.

I guess I could Find/Replace all the &151; with – and so forth for each ASCII character, but I’m still stuck with how to get \<BR\> to a line break, how to convert text after [noparse]\<i\>[/noparse] to italic, etc.

(Yes, I know the ASCII code actually has a # after the &, but if I type that here, it gets parsed and the no-parse tags don’t seem to avoid it.)

---

<div class="post-metadata">

**Author:** ![Keeve](https://avatars.discourse-cdn.com/v4/letter/k/f07891/32.png) [@Keeve](https://boards.straightdope.com/u/Keeve)\
**Post date:** [January 23, 2009, 1:49pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/4 "2009-01-23T13:49:05Z")

</div>

From your browser, do a Select All and Copy. And then go to Word, and Paste.

I’ve done this lots of times. It’s not perfect, but it is a decent approximation. And it can take quite a long time, especially of it has loads of nested tables and such. But it usually does work, eventually.

---

<div class="post-metadata">

**Author:** ![Keeve](https://avatars.discourse-cdn.com/v4/letter/k/f07891/32.png) [@Keeve](https://boards.straightdope.com/u/Keeve)\
**Post date:** [January 23, 2009, 1:50pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/5 "2009-01-23T13:50:24Z")

</div>

> [@C\_K\_Dexter\_Haven](#):
>
> but I’m still stuck with how to get \<BR\> to a line break.

Find: \<BR\>  
Replace with: ^p

---

<div class="post-metadata">

**Author:** ![Harmonious\_Discord](https://avatars.discourse-cdn.com/v4/letter/h/74df32/32.png) [@Harmonious\_Discord](https://boards.straightdope.com/u/Harmonious_Discord)\
**Post date:** [January 23, 2009, 1:52pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/6 "2009-01-23T13:52:37Z")

</div>

There are converters, but I don’t have a need for them, so can’t tell you how well each one works.

The copy and paste method is what I use.

---

<div class="post-metadata">

**Author:** ![C\_K\_Dexter\_Haven](https://avatars.discourse-cdn.com/v4/letter/c/b2d939/32.png) [@C\_K\_Dexter\_Haven](https://boards.straightdope.com/u/C_K_Dexter_Haven)\
**Post date:** [January 23, 2009, 1:56pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/7 "2009-01-23T13:56:06Z")

</div>

Thanks all, I guess that’s the easiest.

---

<div class="post-metadata">

**Author:** ![jjimm](https://avatars.discourse-cdn.com/v4/letter/j/ba8739/32.png) [@jjimm](https://boards.straightdope.com/u/jjimm)\
**Post date:** [January 23, 2009, 2:29pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/8 "2009-01-23T14:29:33Z")

</div>

I don’t understand why Word doesn’t open this as an approximation of how it is viewable in a browser. Is it opening the raw code, or as a web page?

---

<div class="post-metadata">

**Author:** ![sailor](https://avatars.discourse-cdn.com/v4/letter/s/a587f6/32.png) [@sailor](https://boards.straightdope.com/u/sailor)\
**Post date:** [January 23, 2009, 2:55pm UTC](https://boards.straightdope.com/t/converting-html-to-word/482628/9 "2009-01-23T14:55:47Z")

</div>

Wait a minute. You have HTML code. Just insert that code in an HTML document and then open as HTML, highlight, copy etc.

Same concept explained another way. Does the code have tha header and footer? If not, just add it and save it as text, change extension to .htm and you’re done.

```auto

<HTML>
<HEAD></HEAD>
<BODY>

Your code goes here

</BODY></HTML>

```

Save as TXT, change extension to HTM. Done!

If you have Outlook Express you can use that too. There’s a hundred ways to skin this cat, all very easy.
