# HTML I've never seen before

**URL:** <https://boards.straightdope.com/t/html-ive-never-seen-before/289603>\
**Category:** Factual Questions\
**Created:** [February 11, 2005, 4:40pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603 "2005-02-11T16:40:54Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Improv\_Geek](https://avatars.discourse-cdn.com/v4/letter/i/b3f665/32.png) [@Improv\_Geek](https://boards.straightdope.com/u/Improv_Geek)\
**Post date:** [February 11, 2005, 4:40pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/1 "2005-02-11T16:40:54Z")

</div>

I’ve been doing web design for a while and my dad asked me to do some work for him but the code is unlike anything I’ve seen before. Anyone recognize it and know where I can learn more about it?

The code has parts such as \<O:P\> and inside image tags it has v:shapes as an attribute.

Anyone?

---

<div class="post-metadata">

**Author:** ![Duckster](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/duckster/32/1244_2.png) [@Duckster](https://boards.straightdope.com/u/Duckster)\
**Post date:** [February 11, 2005, 4:44pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/2 "2005-02-11T16:44:35Z")

</div>

It may be the page was created in Microsoft Publisher.

---

<div class="post-metadata">

**Author:** ![Bill\_H](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@Bill\_H](https://boards.straightdope.com/u/Bill_H)\
**Post date:** [February 11, 2005, 5:45pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/3 "2005-02-11T17:45:35Z")

</div>

Or with Microsoft Word.

---

<div class="post-metadata">

**Author:** ![ouryL](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/ouryl/32/6067_2.png) [@ouryL](https://boards.straightdope.com/u/ouryL)\
**Post date:** [February 11, 2005, 5:48pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/4 "2005-02-11T17:48:41Z")

</div>

> [@Duckster](#):
>
> It may be the page was created in Microsoft Publisher.

I concur. I have seen programs which clean these extraneous tags off.

---

<div class="post-metadata">

**Author:** ![aldiboronti](https://avatars.discourse-cdn.com/v4/letter/a/9fc348/32.png) [@aldiboronti](https://boards.straightdope.com/u/aldiboronti)\
**Post date:** [February 11, 2005, 6:01pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/5 "2005-02-11T18:01:16Z")

</div>

Yep, it’s Publisher, as you can see from [this page.](http://www.forum4designers.com/message179521.html)

---

<div class="post-metadata">

**Author:** ![ChordedZither](https://avatars.discourse-cdn.com/v4/letter/c/e9c0ed/32.png) [@ChordedZither](https://boards.straightdope.com/u/ChordedZither)\
**Post date:** [February 11, 2005, 6:01pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/6 "2005-02-11T18:01:38Z")

</div>

The part in front of the colon is a namespace prefix. In XML (and newer versions of HTML are incorporating more and more general XML features - some are even well-formed XML) a anemspace is used to prevent clashes between elements that might have the same name but convey different meanings. In your example, \<O:P\> names a \<P\> element that might have nothing to do with the HTML paragraph element.

You can get some clue about the origin of these namespaces by looking for the definition of each prefix. Somewhere up above the place where \<O:P\> and \<v:shapes\> are being used, you should find an element with attributes  
xmlns:O="…some long string…"  
and, not necessarily in the element,  
xmlns:v="…some other long string…"  
The strings are URIs that represent a unique identifier for the set of tags being associated with that prefix (“O” or “v”). The prefixes themselves are arbitrary, but the corresponding URIs uniquely identify the set of tags being used. In many cases, you can figure out the general purpose of that set of tags by examining those URIs. Although it’s not required, some people use “real” URLs for the URIs and place documentation for the tag set at that web location.

If you’re going to work with this style of markup, you should start with some basic tutorials on XML.

---

<div class="post-metadata">

**Author:** ![Cerowyn](https://avatars.discourse-cdn.com/v4/letter/c/82dd89/32.png) [@Cerowyn](https://boards.straightdope.com/u/Cerowyn)\
**Post date:** [February 11, 2005, 6:26pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/7 "2005-02-11T18:26:41Z")

</div>

To add to what **ChordedZither** posted, the tags are often embedded in HTML to enable “round-tripping” documents from other applications. For instance, Word infamously buries a huge volume of extra tags and formatting information when it saves a document in HTML format. However, that allows Word to subsequently re-load the document and retain almost all of the extra information that would otherwise have been lost in a strict HTML page.

---

<div class="post-metadata">

**Author:** ![mhendo](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mhendo/32/3159_2.png) [@mhendo](https://boards.straightdope.com/u/mhendo)\
**Post date:** [February 11, 2005, 8:57pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/8 "2005-02-11T20:57:43Z")

</div>

> [@Cerowyn](#):
>
> To add to what **ChordedZither** posted, the tags are often embedded in HTML to enable “round-tripping” documents from other applications. For instance, Word infamously buries a huge volume of extra tags and formatting information when it saves a document in HTML format. However, that allows Word to subsequently re-load the document and retain almost all of the extra information that would otherwise have been lost in a strict HTML page.

A problem, of course, is that it is often a fucking nightmare to view an HTML page created by a Microsoft program in anything except Microsoft’s IE browser.

---

<div class="post-metadata">

**Author:** ![gotpasswords](https://avatars.discourse-cdn.com/v4/letter/g/c57346/32.png) [@gotpasswords](https://boards.straightdope.com/u/gotpasswords)\
**Post date:** [February 11, 2005, 10:07pm UTC](https://boards.straightdope.com/t/html-ive-never-seen-before/289603/9 "2005-02-11T22:07:05Z")

</div>

> [@mhendo](#):
>
> A problem, of course, is that it is often a fucking nightmare to view an HTML page created by a Microsoft program in anything except Microsoft’s IE browser.

Dreamweaver does a fine job of stripping out all that crap and making normal HTML out of it.

I’m sure other apps do this as well - you know it’s bad when an app has a built in “Clean up Word HTML” item on menu - not even as an option or plugin.
