# Why doesn't gmail know how many emails I have?

**URL:** <https://boards.straightdope.com/t/why-doesnt-gmail-know-how-many-emails-i-have/540496>\
**Category:** Factual Questions\
**Created:** [May 23, 2010, 8:24pm UTC](https://boards.straightdope.com/t/why-doesnt-gmail-know-how-many-emails-i-have/540496 "2010-05-23T20:24:45Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![wonky](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/wonky/32/393_2.png) [@wonky](https://boards.straightdope.com/u/wonky)\
**Post date:** [May 23, 2010, 8:24pm UTC](https://boards.straightdope.com/t/why-doesnt-gmail-know-how-many-emails-i-have/540496/1 "2010-05-23T20:24:45Z")

</div>

If I run a search in gmail for emails from 2006, on the first page of results it will say “1-20 of about 80.”

First, “about” 80? Why not an exact number?

If I delete that first page of emails, it will say “1-20 of about 80” still. Why?

If I click on the “older” link, I can scroll through multiple pages of emails until it is finally revealed that I have 282 emails from 2006. Why couldn’t it tell me that up front?

If I click back to go to the first page, the count reverts to “1-20 of about 80.”

All of my searches work like this. I don’t get it.

---

<div class="post-metadata">

**Author:** ![Superfluous\_Parentheses](https://avatars.discourse-cdn.com/v4/letter/s/8edcca/32.png) [@Superfluous\_Parentheses](https://boards.straightdope.com/u/Superfluous_Parentheses)\
**Post date:** [May 23, 2010, 8:49pm UTC](https://boards.straightdope.com/t/why-doesnt-gmail-know-how-many-emails-i-have/540496/2 "2010-05-23T20:49:50Z")

</div>

The google web search results are the same, and the estimates are sometimes way off (but you can only really see that for searches that return a hand full of results\*).

AFAICT, the reason the results are estimated is that the mechanism google uses for searches is based on [MapReduce](http://en.wikipedia.org/wiki/MapReduce). MapReduce is interesting since any tasks built on it can be automatically distributed over many machines, but it also has some limitations, one being that it’s much quicker to return a subset of results than it is to count all the results.

Related: as you may have noted, you only get your web search results in pages, and with only 10 links to further pages of results (for a total of 100 results). That is _probably_ because the system uses the results it’s already found to quickly find the “nearby” ones.

I don’t know the ins and outs of Google, but that’s more or less how some other MapReduce type systems work.

- ETA: I did some checking just now, and it seems that overly-optimistic estimates are quickly scaled down if you work your way to the last page and then search the same term again.

---

<div class="post-metadata">

**Author:** ![wonky](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/wonky/32/393_2.png) [@wonky](https://boards.straightdope.com/u/wonky)\
**Post date:** [May 23, 2010, 9:01pm UTC](https://boards.straightdope.com/t/why-doesnt-gmail-know-how-many-emails-i-have/540496/3 "2010-05-23T21:01:31Z")

</div>

I knew there had to be a techy reason for it. Thanks!
