# searching the web

**URL:** <https://boards.straightdope.com/t/searching-the-web/30289>\
**Category:** Factual Questions\
**Created:** [August 22, 2000, 10:28am UTC](https://boards.straightdope.com/t/searching-the-web/30289 "2000-08-22T10:28:07Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Typo\_Negative](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/typo_negative/32/484_2.png) [@Typo\_Negative](https://boards.straightdope.com/u/Typo_Negative)\
**Post date:** [August 22, 2000, 10:28am UTC](https://boards.straightdope.com/t/searching-the-web/30289/1 "2000-08-22T10:28:07Z")

</div>

Please help an idiot.

Is there a way to do a search for a multi-word term (i.e. double jeapordy) and NOT get sites containing only one of the words?

---

<div class="post-metadata">

**Author:** ![JonF](https://avatars.discourse-cdn.com/v4/letter/j/cab0a1/32.png) [@JonF](https://boards.straightdope.com/u/JonF)\
**Post date:** [August 22, 2000, 10:42am UTC](https://boards.straightdope.com/t/searching-the-web/30289/2 "2000-08-22T10:42:09Z")

</div>

Yes.

Oh, you want to know what it _is_?

It depends on the search engine. For example, in Alrtavista, to get sites containing double and jeapordy but not requiring the words to be together:

+double +jeapordy

In most if not all search engines, to search for the phrase “double jeapordy”:

“double jeapordy”

That is, enclose the phrase with double quotes.

Many search engines have boolean capabilities (sometimes listed in “advanced search”) where you can do queries like:

“double jeapordy” near (legal or law) not “Trebeck”

---

<div class="post-metadata">

**Author:** ![TheThill](https://avatars.discourse-cdn.com/v4/letter/t/35a633/32.png) [@TheThill](https://boards.straightdope.com/u/TheThill)\
**Post date:** [August 22, 2000, 11:48am UTC](https://boards.straightdope.com/t/searching-the-web/30289/3 "2000-08-22T11:48:04Z")

</div>

You might also try “double jeopardy”.

---

<div class="post-metadata">

**Author:** ![Typo\_Negative](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/typo_negative/32/484_2.png) [@Typo\_Negative](https://boards.straightdope.com/u/Typo_Negative)\
**Post date:** [August 22, 2000, 12:34pm UTC](https://boards.straightdope.com/t/searching-the-web/30289/4 "2000-08-22T12:34:08Z")

</div>

Thanks JonF.

---

<div class="post-metadata">

**Author:** ![lee](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/lee/32/7455_2.png) [@lee](https://boards.straightdope.com/u/lee)\
**Post date:** [August 22, 2000, 1:21pm UTC](https://boards.straightdope.com/t/searching-the-web/30289/5 "2000-08-22T13:21:04Z")

</div>

go to [http://www.google.com](http://www.google.com) and simply type in all the trems. It is much better than any onther search engine i have ever used.

---

<div class="post-metadata">

**Author:** ![G.B.H.Hornswoggler](https://avatars.discourse-cdn.com/v4/letter/g/c67d28/32.png) [@G.B.H.Hornswoggler](https://boards.straightdope.com/u/G.B.H.Hornswoggler)\
**Post date:** [August 22, 2000, 1:29pm UTC](https://boards.straightdope.com/t/searching-the-web/30289/6 "2000-08-22T13:29:13Z")

</div>

I’ll second the recommendation of google, but do watch out for those trems. They can turn vicious when you type them. (I find the yellow-bellied trem is the most common type, myself, but that could just be this area.)  
[triple checks for typos in _this_ post…]

---

<div class="post-metadata">

**Author:** ![scr4](https://avatars.discourse-cdn.com/v4/letter/s/59ef9b/32.png) [@scr4](https://boards.straightdope.com/u/scr4)\
**Post date:** [August 22, 2000, 1:49pm UTC](https://boards.straightdope.com/t/searching-the-web/30289/7 "2000-08-22T13:49:44Z")

</div>

I like [google](http://www.google.com/) too, especially when doing searches that return a large number of hits. It does a good job of guessing which one is the most relevant - one method it uses is to prefer pages where the keywords occur close together in the page.

However, it does **not** do phrase searches. If you search for “to be or not to be”, google won’t find many Shakespearean sites. I think [http://www.altavista.com/](http://www.altavista.com/) is a better bet if you want to search for a phrase.

---

<div class="post-metadata">

**Author:** ![Crusoe](https://avatars.discourse-cdn.com/v4/letter/c/49beb7/32.png) [@Crusoe](https://boards.straightdope.com/u/Crusoe)\
**Post date:** [August 22, 2000, 8:42pm UTC](https://boards.straightdope.com/t/searching-the-web/30289/8 "2000-08-22T20:42:01Z")

</div>

AltaVista, and their cut-down version at [raging.com](http://www.raging.com), are very good at “advanced” searches like this, and they offer a very thorough tutorial of their services. They also let you use the following commands (from [here](http://doc.altavista.com/adv_search/syntax.html)):

> [@](#):
>
> **AND** Finds documents containing all of the specified words or phrases. Peanut AND butter finds documents with both the word peanut and the word butter.
> 
> **OR** Finds documents containing at least one of the specified words or phrases. Peanut OR butter finds documents containing either peanut or butter. The found documents could contain both items, but not necessarily.
> 
> **AND NOT** Excludes documents containing the specified word or phrase. Peanut AND NOT butter finds documents with peanut but not containing butter. NOT must be used with another operator, like AND. AltaVista does not accept ‘peanut NOT butter’; instead, specify peanut AND NOT butter.
> 
> **NEAR** Finds documents containing both specified words or phrases within 10 words of each other. Peanut NEAR butter would find documents with peanut butter, but probably not any other kind of butter.
> 
> **( )** Use parentheses to group complex Boolean phrases. For example, (peanut AND butter) AND (jelly OR jam) finds documents with the words ‘peanut butter and jelly’ or ‘peanut butter and jam’ or both.
> 
> **anchor:text** Finds pages that contain the specified word or phrase in the text of a hyperlink. anchor:“Click here to visit [garden.com](http://garden.com)” would find pages with “Click here to visit [garden.com](http://garden.com)” as a link.
> 
> **applet:class** Finds pages that contain a specified Java applet. Use applet:morph to find pages using applets called morph.
> 
> **domain:domainname** Finds pages within the specified domain. Use domain:uk to find pages from the United Kingdom, or use domain:com to find pages from commercial sites.
> 
> **host:hostname** Finds pages on a specific computer. The search host:www.shopping.com would find pages on the [Shopping.com](http://Shopping.com) computer, and host:dilbert.unitedmedia.com would find pages on the computer called dilbert at [unitedmedia.com](http://unitedmedia.com).
> 
> **image:filename** Finds pages with images having a specific filename. Use image:beaches to find pages with images called beaches.
> 
> **like:URLtext** Finds pages similar to or related to the specified URL. For example, like:www.abebooks.com finds Web sites that sell used and rare books, similar to the [http://www.abebooks](http://www.abebooks) site. like:sfpl.lib.ca.us/ finds public and university library sites. like:[http://www.indiaxs.com/](http://www.indiaxs.com/) finds sites about culture on the Indian subcontinent.
> 
> **link:URLtext** Finds pages with a link to a page with the specified URL text. Use link:www.myway.com to find all pages linking to [myway.com](http://myway.com).
> 
> **text:text** Finds pages that contain the specified text in any part of the page other than an image tag, link, or URL. The search text:graduation would find all pages with the term graduation in them.
> 
> **title:text** Finds pages that contain the specified word or phrase in the page title (which appears in the title bar of most browsers). The search title:sunset would find pages with sunset in the title.
> 
> **url:text** Finds pages with a specific word or phrase in the URL. Use url:myway.com to find all pages on all servers that have the word myway in the host name, path, or filename–the complete URL, in other words.

---

<div class="post-metadata">

**Author:** ![Arnold\_Winkelried](https://avatars.discourse-cdn.com/v4/letter/a/3d9bf3/32.png) [@Arnold\_Winkelried](https://boards.straightdope.com/u/Arnold_Winkelried)\
**Post date:** [August 22, 2000, 9:42pm UTC](https://boards.straightdope.com/t/searching-the-web/30289/9 "2000-08-22T21:42:52Z")

</div>

Let me embellish a little bit on what is being suggested here:

If you are looking for an “official” site, e.g.

- USA Treasury Department
- Swiss Federal Government
- Encyclopaedia Britannica
- Britney Spears Fan Club  
etc…

Your best bet is a site that groups websites in categories, such as [Yahoo!](http://www.yahoo.com) or [Excite](http://www.excite.com).

If on the other hand you are looking for a website on an obscure subject or to answer a question, e.g.  
“What colour triangle did the gypsies wear in Concentration Camps in WWII?”  
There are several ways you can approach the search:

a) Go to a “categorized” web site like Yahoo! or Excite, look for the WWII Concentration Camp category, and then go read those sites that may contain your answer.  
b) go to a site like [Google](http://www.google.com) or [AlltheWeb](http://www.alltheweb.com) that searches for text in web pages, and type in as many keywords as you think would probably be on the page, e.g. “gypsy patch concentration camp” with the hope that you will find your page.

Also realize that any web page you find does not necessarily contain reliable information. I could write a 20-page thesis on Mongolian history, but if I have no expertise on the subject, all the information contained therein could be unmitigated crap. Caveat lector!

---

<div class="post-metadata">

**Author:** ![JonF](https://avatars.discourse-cdn.com/v4/letter/j/cab0a1/32.png) [@JonF](https://boards.straightdope.com/u/JonF)\
**Post date:** [August 22, 2000, 10:25pm UTC](https://boards.straightdope.com/t/searching-the-web/30289/10 "2000-08-22T22:25:55Z")

</div>

> [@](#):
>
> AltaVista, and their cut-down version at [raging.com](http://raging.com), are very good at “advanced” searches like this

Which can be used in non-obvious ways. There’s a short article in the print version of [Fast Company](http://www.fastcompany.com/homepage/) this month, claiming that you can find all sorts of interesting and sometimes hidden things on a company’s web site using the “host:” keyword and key words and phrases like “business plan”.
