# Why Is The White House Hiding All Search References To Iraq on Its Web Site?

**URL:** <https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023>\
**Category:** Factual Questions\
**Created:** [January 25, 2004, 6:30pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023 "2004-01-25T18:30:52Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Duckster](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/duckster/32/1244_2.png) [@Duckster](https://boards.straightdope.com/u/Duckster)\
**Post date:** [January 25, 2004, 6:30pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/1 "2004-01-25T18:30:52Z")

</div>

Mods: This thread is not intended as a GD. However, feel free to move it if the Doper responses are more attuned to debate than answering the question.

Background:

> [@](#):
>
> Search engines will look in your root domain for a special file named “robots.txt” ([http://www.mydomain.com/robots.txt](http://www.mydomain.com/robots.txt)). The file tells the robot (spider) which files it may spider (download). This system is called, The Robots Exclusion Standard.

Source: [http://www.searchengineworld.com/robots/robots\_tutorial.htm](http://www.searchengineworld.com/robots/robots_tutorial.htm)

There is nothing nefarious in the use of a robot.txt file on a web site. On the contrary, a robots.txt file is used to assist search engines so that they do not have to collect information that is irrelevant to users.

However, in viewing the [White House robots. txt file](http://www.whitehouse.gov/robots.txt) one notes that _all references to Iraq and only references to Iraq_ are unavailable to searches by search engines.

Why is this?

---

<div class="post-metadata">

**Author:** ![Ringo](https://avatars.discourse-cdn.com/v4/letter/r/779978/32.png) [@Ringo](https://boards.straightdope.com/u/Ringo)\
**Post date:** [January 25, 2004, 6:47pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/2 "2004-01-25T18:47:15Z")

</div>

I don’t know how the robot file works, but a search on the White House site just now for “Iraq” turned up 1,923 results.

---

<div class="post-metadata">

**Author:** ![Duckster](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/duckster/32/1244_2.png) [@Duckster](https://boards.straightdope.com/u/Duckster)\
**Post date:** [January 25, 2004, 6:50pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/3 "2004-01-25T18:50:31Z")

</div>

> [@Ringo](#):
>
> I don’t know how the robot file works, but a search on the White House site just now for “Iraq” turned up 1,923 results.

No, no. The robots.txt file prevents searches by _external_ search engines. An _internal_ search engine can be manipulated by the site owner to only locate what the owner wants you to see.

---

<div class="post-metadata">

**Author:** ![Ringo](https://avatars.discourse-cdn.com/v4/letter/r/779978/32.png) [@Ringo](https://boards.straightdope.com/u/Ringo)\
**Post date:** [January 25, 2004, 6:53pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/4 "2004-01-25T18:53:32Z")

</div>

I don’t think it’s disallowing _all_ references to Iraq; it looks more like a list of specific files and/or subdirectories.

---

<div class="post-metadata">

**Author:** ![Fear\_Itself](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/fear_itself/32/19637_2.png) [@Fear\_Itself](https://boards.straightdope.com/u/Fear_Itself)\
**Post date:** [January 25, 2004, 6:54pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/5 "2004-01-25T18:54:43Z")

</div>

[Google found 19,800](http://www.google.com/search?num=100&hl=en&lr=&ie=ISO-8859-1&safe=off&q=site%3Awww.whitehouse.gov+iraq&btnG=Google+Search)

---

<div class="post-metadata">

**Author:** ![Shiva](https://avatars.discourse-cdn.com/v4/letter/s/85e7bf/32.png) [@Shiva](https://boards.straightdope.com/u/Shiva)\
**Post date:** [January 25, 2004, 6:55pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/6 "2004-01-25T18:55:39Z")

</div>

A Google search confined to the [whitehouse.gov](http://whitehouse.gov) domain found about 19,900 hits for “Iraq”

[http://www.google.com/search?as\_q=iraq&num=10&hl=en&ie=UTF-8&oe=UTF-8&btnG=Google+Search&as\_epq=&as\_oq=&as\_eq=&lr=&as\_ft=i&as\_filetype=&as\_qdr=all&as\_occt=any&as\_dt=i&as\_sitesearch=whitehouse.gov&safe=off](http://www.google.com/search?as_q=iraq&num=10&hl=en&ie=UTF-8&oe=UTF-8&btnG=Google+Search&as_epq=&as_oq=&as_eq=&lr=&as_ft=i&as_filetype=&as_qdr=all&as_occt=any&as_dt=i&as_sitesearch=whitehouse.gov&safe=off)

---

<div class="post-metadata">

**Author:** ![Tapioca\_Dextrin](https://avatars.discourse-cdn.com/v4/letter/t/3d9bf3/32.png) [@Tapioca\_Dextrin](https://boards.straightdope.com/u/Tapioca_Dextrin)\
**Post date:** [January 25, 2004, 6:58pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/7 "2004-01-25T18:58:46Z")

</div>

> [@Duckster](#):
>
> However, in viewing the [White House robots. txt file](http://www.whitehouse.gov/robots.txt) one notes that _all references to Iraq and only references to Iraq_ are unavailable to searches by search engines.
> 
> Why is this?

Just a WAG, but [www.whitehouse.gov](http://www.whitehouse.gov) is a popular site for people looking for information on iraq. It might make sense to disallow searches for searches on iraq on those pages which don’t contain the term iraq. Even Mr. Bush’s hamsters have their breaking point.

e.g the page [http://www.whitehouse.gov/firstlady/recipes](http://www.whitehouse.gov/firstlady/recipes) can safely be assumed to be excluded from an iraq based search.

---

<div class="post-metadata">

**Author:** ![Mops](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mops/32/16494_2.png) [@Mops](https://boards.straightdope.com/u/Mops)\
**Post date:** [January 25, 2004, 6:59pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/8 "2004-01-25T18:59:43Z")

</div>

This became a topic of public discussion [in October of last year](http://yro.slashdot.org/article.pl?sid=03/10/27/2052228).

robots.txt disallows crawling by file name, not by file content. The robots.txt change excludes many file paths that obviously don’t exist:

Disallow: /infocus/everglades/iraq  
…  
Disallow: /infocus/rx-medicare/iraq  
…  
Disallow: /infocus/teacherquality/iraq

What the person who ordered this intended to do is anyone’s guess. My bet is managerial stupidity.

---

<div class="post-metadata">

**Author:** ![typhoon](https://avatars.discourse-cdn.com/v4/letter/t/ed8c4c/32.png) [@typhoon](https://boards.straightdope.com/u/typhoon)\
**Post date:** [January 25, 2004, 7:16pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/9 "2004-01-25T19:16:45Z")

</div>

The question is why would the White House site let you search with the internal search engine but not the external one?

Because the external ones cache pages.

---

<div class="post-metadata">

**Author:** ![robo99](https://avatars.discourse-cdn.com/v4/letter/r/a8b319/32.png) [@robo99](https://boards.straightdope.com/u/robo99)\
**Post date:** [January 25, 2004, 9:04pm UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/10 "2004-01-25T21:04:17Z")

</div>

Whoever is managing the [whitehouse.gov](http://whitehouse.gov) web page could be doing a better job. For instance:

[http://www.whitehouse.gov/index2.html](http://www.whitehouse.gov/index2.html)

is an old page from June 2003. If I were running their web page I would clean this stuff up fairly regularly.

---

<div class="post-metadata">

**Author:** ![friedo](https://avatars.discourse-cdn.com/v4/letter/f/8edcca/32.png) [@friedo](https://boards.straightdope.com/u/friedo)\
**Post date:** [January 26, 2004, 12:25am UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/11 "2004-01-26T00:25:27Z")

</div>

> [@typhoon](#):
>
> The question is why would the White House site let you search with the internal search engine but not the external one?
> 
> Because the external ones cache pages.

The external ones also obey robots.txt as a matter of convention. The file doesn’t “force” any search engine to do anything.

---

<div class="post-metadata">

**Author:** ![Squink](https://avatars.discourse-cdn.com/v4/letter/s/b5e925/32.png) [@Squink](https://boards.straightdope.com/u/Squink)\
**Post date:** [January 26, 2004, 5:28am UTC](https://boards.straightdope.com/t/why-is-the-white-house-hiding-all-search-references-to-iraq-on-its-web-site/226023/12 "2004-01-26T05:28:06Z")

</div>

I did a bit of looking at the robot.txt files of various government organizations.  
The CIA, FBI, Senate, DOE, Air Force, NASA, Secret Service, Supreme court, Federal Election Commission, Federal Reserve, Homeland Security, and FirstGove sites have no robot.txt files at all.  
The House, FDA, NSA, DOJ, USDA, Army, Joint Chiefs, FDIC, and OSHA sites have small restriction files, from a few lines to ~25 in length.  
The only other site that approaches the whitehouse in the size of its robot file is the EPA.
