# Google search on massive database?

**URL:** <https://boards.straightdope.com/t/google-search-on-massive-database/292273>\
**Category:** Factual Questions\
**Created:** [February 28, 2005, 10:00am UTC](https://boards.straightdope.com/t/google-search-on-massive-database/292273 "2005-02-28T10:00:36Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Rune](https://avatars.discourse-cdn.com/v4/letter/r/e68b1a/32.png) [@Rune](https://boards.straightdope.com/u/Rune)\
**Post date:** [February 28, 2005, 10:00am UTC](https://boards.straightdope.com/t/google-search-on-massive-database/292273/1 "2005-02-28T10:00:36Z")

</div>

Google must have a gigantic database with hundreds of million, if not billions, of text rows. When I type in a search for a string for which they clearly cannot have a pre-stored result, how do they manage to search trough all those rows in less than a second?

---

<div class="post-metadata">

**Author:** ![Futile\_Gesture](https://avatars.discourse-cdn.com/v4/letter/f/f05b48/32.png) [@Futile\_Gesture](https://boards.straightdope.com/u/Futile_Gesture)\
**Post date:** [February 28, 2005, 11:01am UTC](https://boards.straightdope.com/t/google-search-on-massive-database/292273/2 "2005-02-28T11:01:27Z")

</div>

Masses of computing power and indexing techniques that unsurprisingly they keep a trade secret.

[Or maybe it’s pigeons.](http://www.google.com/technology/pigeonrank.html)

---

<div class="post-metadata">

**Author:** ![NillyWilly](https://avatars.discourse-cdn.com/v4/letter/n/3da27b/32.png) [@NillyWilly](https://boards.straightdope.com/u/NillyWilly)\
**Post date:** [February 28, 2005, 11:23am UTC](https://boards.straightdope.com/t/google-search-on-massive-database/292273/3 "2005-02-28T11:23:53Z")

</div>

> [@Futile Gesture](#):
>
> Masses of computing power and indexing techniques that unsurprisingly they keep a trade secret.
> 
> [Or maybe it’s pigeons.](http://www.google.com/technology/pigeonrank.html)

They’re willing to sell it to you: [Google appliance](http://www.google.com/enterprise/)

---

<div class="post-metadata">

**Author:** ![hammos1](https://avatars.discourse-cdn.com/v4/letter/h/ba8739/32.png) [@hammos1](https://boards.straightdope.com/u/hammos1)\
**Post date:** [February 28, 2005, 1:27pm UTC](https://boards.straightdope.com/t/google-search-on-massive-database/292273/4 "2005-02-28T13:27:39Z")

</div>

> [@Futile Gesture](#):
>
> Masses of computing power…

I’m sure the indexing is very clever, but the sheer computing grunt is the real explanation of Google’s speed.

From [here](http://www.answers.com/topic/google)

> [@](#):
>
> Based on the Google IPO S-1 form released in April 2004, Tristan Louis, the Vice President of application development for the Internet unit of a large financial firm, estimated the current server farm to contain something like the following [4] ([http://www.tnl.net/blog/entry/How\_many\_Google\_machines](http://www.tnl.net/blog/entry/How_many_Google_machines)):
> 
> ```
> * 719 racks
> * 63,272 machines
> * 126,544 CPUs
> * 253,088 GHz of processing power
> * 126,544 GB of RAM
> * 5,062 TB of hard drive space
> 
> ```
> 
> According to this estimate, the Google server farm constitutes one of the most powerful supercomputers in the world, at 126-316 teraflops, being able to perform at least three times as many calculations per second as the Earth Simulator.

---

<div class="post-metadata">

**Author:** ![carterba](https://avatars.discourse-cdn.com/v4/letter/c/85e7bf/32.png) [@carterba](https://boards.straightdope.com/u/carterba)\
**Post date:** [February 28, 2005, 1:49pm UTC](https://boards.straightdope.com/t/google-search-on-massive-database/292273/5 "2005-02-28T13:49:48Z")

</div>

In general, search engines work by keeping an inverted list: a list of words, and for each word, a list of all the documents that word appears in. When you enter a query, the engine gets the lists for each of your query words and combines the lists according to some formula to get the ranked list of results. That method can be very fast, even with millions of documents.

Since Google indexes _billions_ of documents, a standard inverted list would still be pretty slow. They probably use the basic inverted list, but also do some pruning, some sampling, and some approximation to speed things up. The nice thing about having a huge heterogeneous corpus like the web is that you can be less than exact and still do very well.
