# Tagging text strings in MS Word

**URL:** <https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904>\
**Category:** Factual Questions\
**Created:** [December 3, 2009, 10:31pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904 "2009-12-03T22:31:45Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Kizarvexius](https://avatars.discourse-cdn.com/v4/letter/k/da6949/32.png) [@Kizarvexius](https://boards.straightdope.com/u/Kizarvexius)\
**Post date:** [December 3, 2009, 10:31pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/1 "2009-12-03T22:31:45Z")

</div>

Never mind the why and wherefore, here is a quick rundown of what I’m trying to do. Any suggestions that might point me in the right direction would be greatly appreciated.

I’ve got a large block of text in a Word document that I am about to import into a database. Certain segments of the text should be boldfaced, italicized, etc., and are correctly formatted when the process begins. Since the result will be displayed as a webpage, I can insert HTML coding, which is what I normally do. But it’s a tedious process, pasting formatting codes at the beginning and end of each little segment. Is there a way to have Word do this for me automatically, without inserting all of the other crap that normally comes up when Word “webifies” the whole document?

Just as an example, here’s the method I’m currently using.

1. Use search and replace to identify all italicized text in the document and color it red for easier identification.
2. Go through the document manually and type an open bracket at the beginning of each italicized segment and a closed bracket at the end of each. There may be hundreds of these.
3. Use search and replace to swap each open bracket for “\<i\>” and each closed bracket with “\</i\>”.
4. Copy and paste text into database form.

Suggestions?

---

<div class="post-metadata">

**Author:** ![KneadToKnow](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/kneadtoknow/32/3999_2.png) [@KneadToKnow](https://boards.straightdope.com/u/KneadToKnow)\
**Post date:** [December 3, 2009, 10:45pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/2 "2009-12-03T22:45:07Z")

</div>

What version of Word?

---

<div class="post-metadata">

**Author:** ![Kizarvexius](https://avatars.discourse-cdn.com/v4/letter/k/da6949/32.png) [@Kizarvexius](https://boards.straightdope.com/u/Kizarvexius)\
**Post date:** [December 3, 2009, 10:59pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/3 "2009-12-03T22:59:04Z")

</div>

> [@KneadToKnow](#):
>
> What version of Word?

1.

---

<div class="post-metadata">

**Author:** ![KneadToKnow](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/kneadtoknow/32/3999_2.png) [@KneadToKnow](https://boards.straightdope.com/u/KneadToKnow)\
**Post date:** [December 3, 2009, 10:59pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/4 "2009-12-03T22:59:12Z")

</div>

Pending that, here’s one way to do it in Word 2003, which should be pretty easily macro-able for other markups:

Use search and replace. In the “Find what:” box, type _^?_ for any character, then set the format to whatever you’re doing on this iteration (italic, bold, underline, what have you). In the “Replace with:” box, type _\<x\>^&\</x\>_ to replace the text with the same text with the appropriate HTML tags around it (replace the _x_s, obviously with the appropriate HTML tag).

So, where you had **this** , you’ll now have \<b\>t\</b\>\<b\>h\</b\>\<b\>i\</b\>\<b\>s\</b\>

Then do search and replace again searching for any instances of _\</x\>\<x\>_ and replacing it with nothing. This will clear out the intervening tags and leave you with \<b\>this\</b\>.

I do not rule out the possibility of their being simpler ways.

ETA: Should work the same in 2007. 🙂

---

<div class="post-metadata">

**Author:** ![Khadaji](https://avatars.discourse-cdn.com/v4/letter/k/9e8a1a/32.png) [@Khadaji](https://boards.straightdope.com/u/Khadaji)\
**Post date:** [December 3, 2009, 11:21pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/5 "2009-12-03T23:21:03Z")

</div>

\<deleted, I see now why what I wrote was not going to work\>

---

<div class="post-metadata">

**Author:** ![Tim\_T-Bonham.net](https://avatars.discourse-cdn.com/v4/letter/t/46a35a/32.png) [@Tim\_T-Bonham.net](https://boards.straightdope.com/u/Tim_T-Bonham.net)\
**Post date:** [December 3, 2009, 11:27pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/6 "2009-12-03T23:27:24Z")

</div>

Under File menu, Save as Web Page. That will create a .mht file. Now go at that with a plain text editor. The file will start with a couple hundred lines of junk, but look beyond that and you should be able to find the paragraphs that you want, with the appropriate tags already around them. Still a bit of cut-and-paste, but at least the tags are there.

---

<div class="post-metadata">

**Author:** ![Kizarvexius](https://avatars.discourse-cdn.com/v4/letter/k/da6949/32.png) [@Kizarvexius](https://boards.straightdope.com/u/Kizarvexius)\
**Post date:** [December 3, 2009, 11:49pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/7 "2009-12-03T23:49:05Z")

</div>

> [@KneadToKnow](#):
>
> Pending that, here’s one way to do it in Word 2003, which should be pretty easily macro-able for other markups:
> 
> Use search and replace. In the “Find what:” box, type _^?_ for any character, then set the format to whatever you’re doing on this iteration (italic, bold, underline, what have you). In the “Replace with:” box, type _\<x\>^&\</x\>_ to replace the text with the same text with the appropriate HTML tags around it (replace the _x_s, obviously with the appropriate HTML tag).
> 
> So, where you had **this** , you’ll now have \<b\>t\</b\>\<b\>h\</b\>\<b\>i\</b\>\<b\>s\</b\>
> 
> Then do search and replace again searching for any instances of _\</x\>\<x\>_ and replacing it with nothing. This will clear out the intervening tags and leave you with \<b\>this\</b\>.
> 
> I do not rule out the possibility of their being simpler ways.
> 
> ETA: Should work the same in 2007. 🙂

Damn, that’s brilliant. I think it may be exactly what I was looking for. I’ll test it out first thing tomorrow.

---

<div class="post-metadata">

**Author:** ![KneadToKnow](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/kneadtoknow/32/3999_2.png) [@KneadToKnow](https://boards.straightdope.com/u/KneadToKnow)\
**Post date:** [December 3, 2009, 11:49pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/8 "2009-12-03T23:49:41Z")

</div>

> [@Kizarvexius](#):
>
> Damn, that’s brilliant. I think it may be exactly what I was looking for. I’ll test it out first thing tomorrow.

Happy to help. 🙂

---

<div class="post-metadata">

**Author:** ![CookingWithGas](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/cookingwithgas/32/485_2.png) [@CookingWithGas](https://boards.straightdope.com/u/CookingWithGas)\
**Post date:** [December 4, 2009, 1:15am UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/9 "2009-12-04T01:15:09Z")

</div>

> [@KneadToKnow](#):
>
> Pending that, here’s one way to do it in Word 2003, which should be pretty easily macro-able for other markups:
> 
> Use search and replace. In the “Find what:” box, type _^?_ for any character, then set the format to whatever you’re doing on this iteration (italic, bold, underline, what have you). In the “Replace with:” box, type _\<x\>^&\</x\>_ to replace the text with the same text with the appropriate HTML tags around it (replace the _x_s, obviously with the appropriate HTML tag).
> 
> So, where you had **this** , you’ll now have \<b\>t\</b\>\<b\>h\</b\>\<b\>i\</b\>\<b\>s\</b\>
> 
> Then do search and replace again searching for any instances of _\</x\>\<x\>_ and replacing it with nothing. This will clear out the intervening tags and leave you with \<b\>this\</b\>.
> 
> I do not rule out the possibility of their being simpler ways.
> 
> ETA: Should work the same in 2007. 🙂

I haven’t tried it but I think if you search on \* instead of ? it will make the change to the whole block of characters instead of one at a time, eliminating the second step.

---

<div class="post-metadata">

**Author:** ![KneadToKnow](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/kneadtoknow/32/3999_2.png) [@KneadToKnow](https://boards.straightdope.com/u/KneadToKnow)\
**Post date:** [December 4, 2009, 1:23am UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/10 "2009-12-04T01:23:24Z")

</div>

I get “^\* is not a valid special character for the Find What box.”

---

<div class="post-metadata">

**Author:** ![CookingWithGas](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/cookingwithgas/32/485_2.png) [@CookingWithGas](https://boards.straightdope.com/u/CookingWithGas)\
**Post date:** [December 4, 2009, 1:27am UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/11 "2009-12-04T01:27:29Z")

</div>

> [@CookingWithGas](#):
>
> I haven’t tried it but I think if you search on \* instead of ? it will make the change to the whole block of characters instead of one at a time, eliminating the second step.

Oh, you have to check “Use wildcards”. When you use wildcards you don’t use the caret. And I determined that it won’t work (apparently \* is “generous” instead of “greedy”), but if you use

\<\*\>

as the search string it will tag each word. But I haven’t figured out how to make it tag the whole string à la regular expressions.

---

<div class="post-metadata">

**Author:** ![ZipperJJ](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/zipperjj/32/211_2.png) [@ZipperJJ](https://boards.straightdope.com/u/ZipperJJ)\
**Post date:** [December 4, 2009, 1:41am UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/12 "2009-12-04T01:41:33Z")

</div>

Most pre-built HTML editors for Web sites (like one you’d use in a CMS app) have a nice “clean up Word” feature in them.

[This one](http://demos.telerik.com/aspnet-ajax/editor/examples/default/defaultcs.aspx) from Telerik is the best I’ve seen lately.

You can use their demo to strip out Word formatting if you want.

Go to the demo in the link and delete the existing content in the text area. Ctrl-V to paste your Word content in there. A JS confirm will pop up asking if you want to clean up Word formatting, and just click OK. Go to the HTML button underneath the text area to see your cleaned up HTML.

---

<div class="post-metadata">

**Author:** ![Kizarvexius](https://avatars.discourse-cdn.com/v4/letter/k/da6949/32.png) [@Kizarvexius](https://boards.straightdope.com/u/Kizarvexius)\
**Post date:** [December 4, 2009, 4:29pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/13 "2009-12-04T16:29:00Z")

</div>

> [@KneadToKnow](#):
>
> Happy to help. 🙂

**KneadToKnow** , you are a genius. It worked like a charm. I’d upload you a beer, but my USB-powered liquid intake device is on the blink. I hope you will be content to bask in the glow of my admiration, because I cannot tell you how much time your suggestion is going to save me.

🆒

---

<div class="post-metadata">

**Author:** ![KneadToKnow](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/kneadtoknow/32/3999_2.png) [@KneadToKnow](https://boards.straightdope.com/u/KneadToKnow)\
**Post date:** [December 4, 2009, 5:12pm UTC](https://boards.straightdope.com/t/tagging-text-strings-in-ms-word/519904/14 "2009-12-04T17:12:32Z")

</div>

Like I tell people at work, the secret to having a good answer at the ready is to have been asked that question once before. 🙂

Glad it’s a workable solution for you, **Kizarvexius**.
