# Web Search Coverage?

**URL:** <https://boards.straightdope.com/t/web-search-coverage/43846>\
**Category:** Factual Questions\
**Created:** [December 3, 2000, 6:58pm UTC](https://boards.straightdope.com/t/web-search-coverage/43846 "2000-12-03T18:58:06Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![toonerama](https://avatars.discourse-cdn.com/v4/letter/t/ecae2f/32.png) [@toonerama](https://boards.straightdope.com/u/toonerama)\
**Post date:** [December 3, 2000, 6:58pm UTC](https://boards.straightdope.com/t/web-search-coverage/43846/1 "2000-12-03T18:58:06Z")

</div>

I’ve read that typical search engines cover maybe 20% of the web. If that’s true, what is the missing 80%, and how can you get at it?

Perhaps it’s mostly small independant pages or private stuff? And maybe often when we follow links to links to links we’re off the beaten path?

Lastly, any good suggestions for finding sites that are away from the mainstream? I’ve been really enjoying weblogs/blogs recently, and that’s one way.

---

<div class="post-metadata">

**Author:** ![Crusoe](https://avatars.discourse-cdn.com/v4/letter/c/49beb7/32.png) [@Crusoe](https://boards.straightdope.com/u/Crusoe)\
**Post date:** [December 3, 2000, 7:10pm UTC](https://boards.straightdope.com/t/web-search-coverage/43846/2 "2000-12-03T19:10:13Z")

</div>

I don’t have exact figures, but most search engines only index pages that are submitted to them. I believe there are projects intended to manually index as many pages as possible, submitted or not, but that’s a Herculean task and perhaps impossible given the ever-increasing number of sites and pages. As for how to find these other pages…well, you’re right in that following links from other pages is the best way. Randomly entering URLs would be another, but you’re not likely to get much from an awful lot of typing!

For non-mainstream pages, weblogs are probably a good choice (I often use [Robot Wisdom](http://www.robotwisdom.com)). I also check out Yahoo’s picks of the week every Monday – I would post an address, but I only have the England one to hand.

---

<div class="post-metadata">

**Author:** ![Crusoe](https://avatars.discourse-cdn.com/v4/letter/c/49beb7/32.png) [@Crusoe](https://boards.straightdope.com/u/Crusoe)\
**Post date:** [December 3, 2000, 7:17pm UTC](https://boards.straightdope.com/t/web-search-coverage/43846/3 "2000-12-03T19:17:02Z")

</div>

From [Search Engine Watch](http://www.searchenginewatch.com/):

_(from [this article](http://www.searchenginewatch.com/reports/sizes.html))_

> [@](#):
>
> Let’s make it clear. None of the search engines – none of them – index everything on the web. No search engine can claim to have a perfect record of everything out there.
> 
> There are some physical reasons why they miss things, such as problems with frames, image maps or the inability to index dynamically-created web pages.
> 
> There are also hardware limitations: it takes a lot of space to store everything on the web, and a lot of processing power to sort through the material quickly enough to respond. That means spending more money…

They also quote a study by a company called BrightPlanet that estimates that there are around 500 billion web pages, with only 1/500 accessible to search engines.
