# Googlewhacking puzzle

**URL:** https://boards.straightdope.com/t/googlewhacking-puzzle/246646
**Category:** Factual Questions
**Created:** [May 23, 2004, 1:38pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646 "2004-05-23T13:38:57Z")
**Posts on this page:** 11
**Page:** 1

<div class="post-metadata">

### Author: ![InvidiousCourgette](https://avatars.discourse-cdn.com/v4/letter/i/e47c2d/32.png) [@InvidiousCourgette](https://boards.straightdope.com/u/InvidiousCourgette)
#### Post date: [May 23, 2004, 1:38pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/1 "2004-05-23T13:38:57Z")

</div>

A googlewhack is a two word query, that when submitted to [www.google.com](http://www.google.com) returns exactly one result. Both of the words must be real words (and listed on [www.dictionary.com](http://www.dictionary.com)); quotation marks are not allowed; and finally if the one result returned is a word list then it does not count. More information about this curious hobby is found on [www.googlewhack.com](http://www.googlewhack.com), a site which also accepts contributors googlewhack discoveries. As I write I notice the latest find to be _simoniac moose_.

Now I was wondering, with google currently referencing 4 point something billion web pages, is the total number of possible googlewhacks out there increasing or decreasing with time? More interestingly (?!), can anyone see a way of estimating the number of referenced (English language) web pages for which the number of possible googlewhacks is a maxima?

---

<div class="post-metadata">

### Author: ![Andy](https://avatars.discourse-cdn.com/v4/letter/a/aeb1de/32.png) [@Andy](https://boards.straightdope.com/u/Andy)
#### Post date: [May 23, 2004, 2:44pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/2 "2004-05-23T14:44:11Z")

</div>

You’re aware that “Invidious courgette” returns…just a single result?

---

<div class="post-metadata">

### Author: ![InvidiousCourgette](https://avatars.discourse-cdn.com/v4/letter/i/e47c2d/32.png) [@InvidiousCourgette](https://boards.straightdope.com/u/InvidiousCourgette)
#### Post date: [May 23, 2004, 3:27pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/3 "2004-05-23T15:27:20Z")

</div>

Yes 🙂

I should explain: I went to see Dave Gorman’s Googlewhack Adventure at the theatre. Of course as soon as I got home, I had a go - and after an hour of trying I stumbled on “Invidious Courgette”.

… and of course that got me wondering as to whether the World Wide Web was becoming richer or poorer with respect to googlewhacks. Hence the post!

Of course, next time google updates its database we can tick off two googlewhacks - “simoniac moose” and “invidious courgette” as they will now appear on these pages as well ☹

---

<div class="post-metadata">

### Author: ![Pasta](https://avatars.discourse-cdn.com/v4/letter/p/ecccb3/32.png) [@Pasta](https://boards.straightdope.com/u/Pasta)
#### Post date: [May 23, 2004, 6:26pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/4 "2004-05-23T18:26:22Z")

</div>

> [@InvidiousCourgette](#):
>
> Now I was wondering, with google currently referencing 4 point something billion web pages, is the total number of possible googlewhacks out there increasing or decreasing with time? More interestingly (?!), can anyone see a way of estimating the number of referenced (English language) web pages for which the number of possible googlewhacks is a maxima?

I cannot believe I got the blue screen of death (in XP!) just before I finished this post! Starting over…

Here’s my take… There are several numbers involved in the scaling law of interest:

_V_ = the size of the vocabulary of the language, in number of words. The FAQ at [dictionary.com](http://dictionary.com) put the number of words in English (including scientific terms) at two million, give or take. I’ll use _V_ = 2x10[sup]6[/sup].  
_L_ = the length of a typical web page, in number of words. The top story at [cnn.com](http://cnn.com) right now has 788 words, so I’ll use _L_ = 1,000.  
_N_ = The number of web pages.

A rule of thumb known as Zipf’s law tells us that the _r_[sup]th[/sup] most common word in a language occurs at a frequency [symbol]k[/symbol]/_r_[sup]_a_[/sup], with _a_ near 1. I’ll take _a_=1. For V=2x10[sup]6[/sup], [symbol]k[/symbol]=0.066 (obtained by requiring the sum of probabilities to be 1.)

Now let’s take two specific uncommon words – words whose ranks _r_ are near _V_. The probability that those two words occur together on a web page can be approximated by

_P_[sub]combo/sub = _L_[sup]2[/sup][symbol]k[/symbol][sup]2[/sup]/(_r_[sub]1[/sub]_r_[sub]2[/sub]),

for _L_/_r_[sub]_i_[/sub]\<\<1. The probability that these two words form a googlewhack is simply the probability that they occur on exactly one web page:

_P_[sub]gw[/sub] = _N_ _P_[sub]combo[/sub](1 - _P_[sub]combo[/sub])[sup]_N_ - 1[/sup].

We can now ask what _N_ needs to be to maximize _P_[sub]gw[/sub]. Since _f_(_x_) and log(_f_(_x_)) are maximized for the same _x_, I’ll work with

log(_P_[sub]gw[/sub]) = log(_N_) + _N_ log(1 - _P_[sub]combo[/sub]) + (terms not dependent on _N_).

The derivative of this with respect to _N_ shows that the maximum _P_[sub]gw[/sub] occurs at

_N_[sub]max gw[/sub] = -1/log(1 - _P_[sub]combo[/sub]).

Using the numbers above, the rarest words have a maximized googlewhack probability for _N_ = 2.1 billion.

---

<div class="post-metadata">

### Author: ![alterego](https://avatars.discourse-cdn.com/v4/letter/a/6bbea6/32.png) [@alterego](https://boards.straightdope.com/u/alterego)
#### Post date: [May 23, 2004, 7:06pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/5 "2004-05-23T19:06:21Z")

</div>

> [@InvidiousCourgette](#):
>
> Yes 🙂
> 
> I should explain: I went to see Dave Gorman’s Googlewhack Adventure at the theatre. Of course as soon as I got home, I had a go - and after an hour of trying I stumbled on “Invidious Courgette”.
> 
> … and of course that got me wondering as to whether the World Wide Web was becoming richer or poorer with respect to googlewhacks. Hence the post!
> 
> Of course, next time google updates its database we can tick off two googlewhacks - “simoniac moose” and “invidious courgette” as they will now appear on these pages as well ☹

Google doesn’t index the board. See [http://boards.straightdope.com/robots.txt](http://boards.straightdope.com/robots.txt)

---

<div class="post-metadata">

### Author: ![Shade](https://avatars.discourse-cdn.com/v4/letter/s/2bfe46/32.png) [@Shade](https://boards.straightdope.com/u/Shade)
#### Post date: [May 23, 2004, 10:37pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/6 "2004-05-23T22:37:29Z")

</div>

> [@Pasta](#):
>
> Using the numbers above, the rarest words have a maximized googlewhack probability for _N_ = 2.1 billion.

I followed your maths and got N=1biln, but I was sloppy and expect I got a rounding error. Right order of magnitude, anyway.

Wow, that’s a really interesting calculation. I feel obliged to nitpick the assumptions, but aren’t coming up with anything yet.

---

<div class="post-metadata">

### Author: ![Shade](https://avatars.discourse-cdn.com/v4/letter/s/2bfe46/32.png) [@Shade](https://boards.straightdope.com/u/Shade)
#### Post date: [May 23, 2004, 10:46pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/7 "2004-05-23T22:46:41Z")

</div>

Aha! How about if we take into account a distribution for L? I can’t be bothered to do the maths but my gut tells me that for a large N there’s going to be a long tail of webpages with only a few words on, making googlewhacks quite likely. But there’s probably still going to be some max…

---

<div class="post-metadata">

### Author: ![Pasta](https://avatars.discourse-cdn.com/v4/letter/p/ecccb3/32.png) [@Pasta](https://boards.straightdope.com/u/Pasta)
#### Post date: [May 24, 2004, 1:42am UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/8 "2004-05-24T01:42:28Z")

</div>

> [@Shade](#):
>
> I feel obliged to nitpick the assumptions, but aren’t coming up with anything yet.

Oh, there’s plenty of handwaving to nitpick, if you wanted. As you point out, _L_ is not a fixed number, and the tails will play a role. Actually, in trying out a few googlewhacks of my own, I found that a successful whack tended to return a very long, disjointed document.

Something else to play with would be to perform the sum over all possible doublets (rather than just taking a representative one) and then maximizing. You’d have to first complicate my expression for _P_[sub]combo[/sub], though, since my approximation won’t work for all doublets. As soon as you write anything of the sort, though, it becomes impossible to calculate anything analytically.

If there were a “typical set” of doublets (in the statistical mechanics sense), maybe one could concoct a more robust calculation, but I don’t think such a thing will work here. Perhaps if I find time later I’ll play with this idea.

Also, there’s the issue of a dynamic language. There’s really no reason to fix the number of words available, and one could fold in the frequency at which new words are added to the lexicon.

But ignoring all of that, it’s still a neat question.

---

<div class="post-metadata">

### Author: ![TJdude825](https://avatars.discourse-cdn.com/v4/letter/t/dec6dc/32.png) [@TJdude825](https://boards.straightdope.com/u/TJdude825)
#### Post date: [May 24, 2004, 4:37am UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/9 "2004-05-24T04:37:45Z")

</div>

> [@alterego](#):
>
> Google doesn’t index the board. See [http://boards.straightdope.com/robots.txt](http://boards.straightdope.com/robots.txt)

True, but if you report your googlewhack on some other site that Google _does_ index, it’ll make it no longer a googlewhack. You’d think that most googlewhackers would know not to do this, but if even a few do, the number of googlewhacks will decrease. OTOH, more are discovered all the time (how frequently? I don’t know…)

---

<div class="post-metadata">

### Author: ![InvidiousCourgette](https://avatars.discourse-cdn.com/v4/letter/i/e47c2d/32.png) [@InvidiousCourgette](https://boards.straightdope.com/u/InvidiousCourgette)
#### Post date: [May 24, 2004, 6:18am UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/10 "2004-05-24T06:18:52Z")

</div>

> [@](#):
>
> OTOH, more are discovered all the time (how frequently? I don’t know…)

The Whack Stack on [www.googlewack.com](http://www.googlewack.com) has had 358000 googlewhacks posted in the last 2 and a bit years… that is 400 Whacks per day. To that one would have to factor in unreported Whacks. Fortunately googlewhack is not indexed by google.

A very nice bit of maths/stats Pasta. I am sure a lot of googlewhackers will be able to sleep easier now, knowing that they probably are living in a golden age for whacking. Maybe in 20 years time, a googlewhack will be a much rarer find indeed.

When I first considered the problem, it occured to me that the frequency of the individual words that constitute a googlewhack would be important. For example google lists the following frequencies:

F\_simoniac = 1510  
F\_moose = 1380000  
F\_invidious = 76800  
F\_courgette = 31300

Multiplying these numbers to produce a googlewhack score gives the following results:

simoniac moose = 2.1 billion  
invidious courgette = 2.4 billion

It would seem likely that when choosing words to googlewhack, a score of approximately this magnitude may maximise your chance of a successful Whack.

It is probably coincidence, but it is interesting that _Pasta’s_ approximation for Nmax gw is of a similar order of magnitude.

---

<div class="post-metadata">

### Author: ![InvidiousCourgette](https://avatars.discourse-cdn.com/v4/letter/i/e47c2d/32.png) [@InvidiousCourgette](https://boards.straightdope.com/u/InvidiousCourgette)
#### Post date: [May 25, 2004, 2:53pm UTC](https://boards.straightdope.com/t/googlewhacking-puzzle/246646/11 "2004-05-25T14:53:56Z")

</div>

I was thinking a little more about this problem :o

I don’t think that Zipf’s law can apply for the population we are considering, in the area that we are considering - i.e. fairly uncommon words.

Using Pasta’s numbers of  
V = 2 million  
k = 0.066  
L = 1000

Then the chance of \*the most \* uncommon word appearing in any given web page would = L . k / V = 0.000033

If we assume that of the 4 billion pages that google indexes, half are English language, then the number of web pages that we might expect to contain the most uncommon word in the English language would be = 2000000000 x 0.000033 = 66000

Hence it appears that Zipf’s law overestimates the frequency of uncommon words appearing on the world wide web, with somoniac, for example, only mustering 1500 apperances. Maybe the constant “a” needs a serious tweek, or perhaps this is pushing the envelope too far for the law.

Can anyone see a way to rescue the calculation? Notwithstanding these comments, it remains a canny piece of work anyway.
