# What's the deal with Captchas?

**URL:** <https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746>\
**Category:** Factual Questions\
**Created:** [April 9, 2011, 2:36am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746 "2011-04-09T02:36:43Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![randwill](https://avatars.discourse-cdn.com/v4/letter/r/a3d4f5/32.png) [@randwill](https://boards.straightdope.com/u/randwill)\
**Post date:** [April 9, 2011, 2:36am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/1 "2011-04-09T02:36:43Z")

</div>

On NPR today someone was talking about a Captcha project. He said that the smeary, wavy words we are seeing in the Captcha box are words from old scanned books that computers can’t recognize. By typing in the word, we are contributing to the digitization of these old books. Millions of people type in these Captchas per day, so whole books are being converted from old printed pages to digital form in this way.

This makes no sense to me. The purpose of the Captcha is to make sure it is a real person and not a bot trying to access whatever. So the website must compare what I have typed (my interpretation of the smeary words) to what the words actually are. If I got it right, then I’ve passed the human test and I get access. Which means that the correct words of the smeary image already exist in the programs memory. Which means somebody had to type it there. Which means, me typing it again doesn’t add up to anything in terms of providing digitized words to convert old books.

I’m missing something - but what?

---

<div class="post-metadata">

**Author:** ![Terminus\_Est](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/terminus_est/32/3087_2.png) [@Terminus\_Est](https://boards.straightdope.com/u/Terminus_Est)\
**Post date:** [April 9, 2011, 2:43am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/2 "2011-04-09T02:43:03Z")

</div>

No ordinary captcha, but [reCAPTCHA](http://www.google.com/recaptcha). You get two words, one which corresponds to a word being digitized and the other for which the answer is already known. The primary purpose is still to catch autologins, but as long as they’re making a legitimate user recognize digitized words, they might as well get something useful out of it.

---

<div class="post-metadata">

**Author:** ![friedo](https://avatars.discourse-cdn.com/v4/letter/f/8edcca/32.png) [@friedo](https://boards.straightdope.com/u/friedo)\
**Post date:** [April 9, 2011, 2:55am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/3 "2011-04-09T02:55:06Z")

</div>

And the way they know the known word is that they have previously given it to many people who have interpreted it the same way, giving a high degree of confidence as to what it is.

The person doing the reCAPTCHA doesn’t know which word is known and which is unknown. By entering the known word correctly, the system can gain some confidence that his guess for the unknown word is correct. If many people enter the same interpretation for the unknown word, then that word is considered solved and is thereafter used as a known word.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [April 9, 2011, 7:07pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/4 "2011-04-09T19:07:19Z")

</div>

And, of course, if the spammers manage to come up with an algorithm that can defeat them, then Google can appropriate that algorithm and use it to improve their own automated OCR methods. It’s a win-win.

---

<div class="post-metadata">

**Author:** ![Cugel](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/cugel/32/1199_2.png) [@Cugel](https://boards.straightdope.com/u/Cugel)\
**Post date:** [April 10, 2011, 1:29am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/5 "2011-04-10T01:29:52Z")

</div>

> [@friedo](#):
>
> The person doing the reCAPTCHA doesn’t know which word is known and which is unknown.

Sometimes you do. And if the other word is too hard you can just keyboard mash to save time. Rather than mucking about squinting and guessing.

---

<div class="post-metadata">

**Author:** ![Aversin](https://avatars.discourse-cdn.com/v4/letter/a/53a042/32.png) [@Aversin](https://boards.straightdope.com/u/Aversin)\
**Post date:** [April 10, 2011, 2:38am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/6 "2011-04-10T02:38:50Z")

</div>

Actually you can tell the difference between the known word and the unknown one, at least unless it has been changed. The known word always looks a certain way, it has a particular (for lack of a better word) font.

On a side note, there have been movements by groups of people on the internet to replace the unknown word with a certain racial slur.

---

<div class="post-metadata">

**Author:** ![friedo](https://avatars.discourse-cdn.com/v4/letter/f/8edcca/32.png) [@friedo](https://boards.straightdope.com/u/friedo)\
**Post date:** [April 10, 2011, 3:10am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/7 "2011-04-10T03:10:10Z")

</div>

> [@Cugel](#):
>
> Sometimes you do. And if the other word is too hard you can just keyboard mash to save time. Rather than mucking about squinting and guessing.

You can also just press the recycle button to get a new one, if you can’t read it.

> [@Aversin](#):
>
> Actually you can tell the difference between the known word and the unknown one, at least unless it has been changed. The known word always looks a certain way, it has a particular (for lack of a better word) font.

That seems implausible, given the way the system is designed to work. The known and unknown words come from the same sources. If you can provide an example of this I’d like to see it.

---

<div class="post-metadata">

**Author:** ![BigT](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/bigt/32/12044_2.png) [@BigT](https://boards.straightdope.com/u/BigT)\
**Post date:** [April 10, 2011, 6:36am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/8 "2011-04-10T06:36:19Z")

</div>

> [@friedo](#):
>
> You can also just press the recycle button to get a new one, if you can’t read it.
> 
> That seems implausible, given the way the system is designed to work. The known and unknown words come from the same sources. If you can provide an example of this I’d like to see it.

It’s not as implausible as you make it sound. It makes sense that the algorithm might have more trouble with certain fonts.

My observation, though, is that other differences are more common. One word always looks a lot more smeared than the other, for example. The letters are less distinct. Also, one word often isn’t one that’s in a normal dictionary (in fact, I’ve seen quite a few that were obvious typos). I think it is quite likely that these are the words that are hard to determine by OCR software. I know the software I’ve used before has problems with such text.

---

<div class="post-metadata">

**Author:** ![Rigamarole](https://avatars.discourse-cdn.com/v4/letter/r/77aa72/32.png) [@Rigamarole](https://boards.straightdope.com/u/Rigamarole)\
**Post date:** [April 10, 2011, 7:29am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/9 "2011-04-10T07:29:59Z")

</div>

Interesting - that explains why sometimes I get by the captcha even though I’m pretty sure I typed one of the words wrong (hey, they can be hard).

---

<div class="post-metadata">

**Author:** ![psychonaut](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/psychonaut/32/4655_2.png) [@psychonaut](https://boards.straightdope.com/u/psychonaut)\
**Post date:** [April 10, 2011, 8:46am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/10 "2011-04-10T08:46:12Z")

</div>

> [@Chronos](#):
>
> And, of course, if the spammers manage to come up with an algorithm that can defeat them, then Google can appropriate that algorithm and use it to improve their own automated OCR methods. It’s a win-win.

What makes you think the spammers are just going to hand over their source code to Google?

---

<div class="post-metadata">

**Author:** ![ZenBeam](https://avatars.discourse-cdn.com/v4/letter/z/3ab097/32.png) [@ZenBeam](https://boards.straightdope.com/u/ZenBeam)\
**Post date:** [April 10, 2011, 1:26pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/11 "2011-04-10T13:26:40Z")

</div>

I had to do a recaptcha after reading this thread, and it did seem I could pick the real word. It was similar to the example in the link, where the one was a dictionary word, and was much straighter than the other.

---

<div class="post-metadata">

**Author:** ![Dewey\_Finn](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dewey_finn/32/4222_2.png) [@Dewey\_Finn](https://boards.straightdope.com/u/Dewey_Finn)\
**Post date:** [April 10, 2011, 2:26pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/12 "2011-04-10T14:26:45Z")

</div>

> [@Aversin](#):
>
> Actually you can tell the difference between the known word and the unknown one, at least unless it has been changed. The known word always looks a certain way, it has a particular (for lack of a better word) font.

You seem to be suggesting that the known word is one that’s been created by the organization, while the unknown word is one that’s from a genuine scan. My understanding is similar to **friedo** ’s; that the known word is also from a genuine scan but is one that’s already been identified by several users.

---

<div class="post-metadata">

**Author:** ![John\_Bredin](https://avatars.discourse-cdn.com/v4/letter/j/96bed5/32.png) [@John\_Bredin](https://boards.straightdope.com/u/John_Bredin)\
**Post date:** [April 10, 2011, 2:50pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/13 "2011-04-10T14:50:39Z")

</div>

> [@Aversin](#):
>
> On a side note, there have been movements by groups of people on the internet to replace the unknown word with a certain racial slur.

Is there a reason for this, other than some peoples’ endless capacity to be assholes? A “principled” objection to being “forced” to do a couple of seconds of “work” solving a puzzle not of their own choosing? :rolleyes:

---

<div class="post-metadata">

**Author:** ![Interconnected\_Series\_of\_Tubes](https://avatars.discourse-cdn.com/v4/letter/i/13edae/32.png) [@Interconnected\_Series\_of\_Tubes](https://boards.straightdope.com/u/Interconnected_Series_of_Tubes)\
**Post date:** [April 10, 2011, 4:21pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/14 "2011-04-10T16:21:01Z")

</div>

> [@John\_Bredin](#):
>
> Is there a reason for this, other than some peoples’ endless capacity to be assholes? A “principled” objection to being “forced” to do a couple of seconds of “work” solving a puzzle not of their own choosing? :rolleyes:

You’ve clearly never been to 4chan.

---

<div class="post-metadata">

**Author:** ![John\_Bredin](https://avatars.discourse-cdn.com/v4/letter/j/96bed5/32.png) [@John\_Bredin](https://boards.straightdope.com/u/John_Bredin)\
**Post date:** [April 10, 2011, 5:04pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/15 "2011-04-10T17:04:40Z")

</div>

> [@Interconnected\_Series\_of\_Tubes](#):
>
> You’ve clearly never been to 4chan.

That was subsumed into “some peoples’ endless capacity to be assholes.” 😃

---

<div class="post-metadata">

**Author:** ![Derleth](https://avatars.discourse-cdn.com/v4/letter/d/b9e5f3/32.png) [@Derleth](https://boards.straightdope.com/u/Derleth)\
**Post date:** [April 10, 2011, 5:20pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/16 "2011-04-10T17:20:20Z")

</div>

> [@Chronos](#):
>
> And, of course, if the spammers manage to come up with an algorithm that can defeat them, then Google can appropriate that algorithm and use it to improve their own automated OCR methods. It’s a win-win.

Aside from **psychonaut** ’s point, the algorithm currently used by spammers is to hire people in the Third World to solve them at a few pennies per word.

---

<div class="post-metadata">

**Author:** ![Nunzio\_Tavulari](https://avatars.discourse-cdn.com/v4/letter/n/b9bd4f/32.png) [@Nunzio\_Tavulari](https://boards.straightdope.com/u/Nunzio_Tavulari)\
**Post date:** [April 11, 2011, 3:26am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/17 "2011-04-11T03:26:57Z")

</div>

They’re mainly there as a first defense against spambots. I encounter many captchas that only require a vague approximation of what’s in the box. The same number of characters within one, and at least two matching characters. Some are very strict, others are not. You may see some that have an algebraic equation in them. That’s a sure sign that it’s looking for a loose interpretation, just the numbers will do.

I’ve never heard of this project to clarify digitized content but it seems feasible.

---

<div class="post-metadata">

**Author:** ![tellyworth](https://avatars.discourse-cdn.com/v4/letter/t/977dab/32.png) [@tellyworth](https://boards.straightdope.com/u/tellyworth)\
**Post date:** [April 11, 2011, 5:11am UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/18 "2011-04-11T05:11:26Z")

</div>

> [@Derleth](#):
>
> Aside from **psychonaut** ’s point, the algorithm currently used by spammers is to hire people in the Third World to solve them at a few pennies per word.

This is true. Also, several major spambots advertise the ability to break CAPTCHAs (including recaptcha) using OCR techniques.

---

<div class="post-metadata">

**Author:** ![Aversin](https://avatars.discourse-cdn.com/v4/letter/a/53a042/32.png) [@Aversin](https://boards.straightdope.com/u/Aversin)\
**Post date:** [April 17, 2011, 7:25pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/19 "2011-04-17T19:25:19Z")

</div>

Did a short bit of searching and found the picture I was referring to. [Link\*](http://funnyjunk.com/funny_pictures/1019304/Operation+ReNigger/) it contains an image which describes what I posted.

> [@](#):
>
> You seem to be suggesting that the known word is one that’s been created by the organization, while the unknown word is one that’s from a genuine scan. My understanding is similar to friedo’s; that the known word is also from a genuine scan but is one that’s already been identified by several users.

> [@](#):
>
> That seems implausible, given the way the system is designed to work. The known and unknown words come from the same sources. If you can provide an example of this I’d like to see it.

> [@](#):
>
> Is there a reason for this, other than some peoples’ endless capacity to be assholes? A “principled” objection to being “forced” to do a couple of seconds of “work” solving a puzzle not of their own choosing?

The image I just linked provides an example. A few times I entered incorrect words to see if I could distinguish between the fonts used, but I didn’t do it enough to rule out the possibility that I got lucky.

---

<div class="post-metadata">

**Author:** ![control-z](https://avatars.discourse-cdn.com/v4/letter/c/eada6e/32.png) [@control-z](https://boards.straightdope.com/u/control-z)\
**Post date:** [April 18, 2011, 4:56pm UTC](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746/20 "2011-04-18T16:56:42Z")

</div>

> [@friedo](#):
>
> And the way they know the known word is that they have previously given it to many people who have interpreted it the same way, giving a high degree of confidence as to what it is.
> 
> The person doing the reCAPTCHA doesn’t know which word is known and which is unknown. By entering the known word correctly, the system can gain some confidence that his guess for the unknown word is correct. If many people enter the same interpretation for the unknown word, then that word is considered solved and is thereafter used as a known word.

Interesting. That explains why many times I don’t have high confidence in my interpretation of one of the words but it still accepts my answer most of the time. I would think I would be wrong more in trying to interpret the very distorted letters.

[Next page](https://boards.straightdope.com/t/whats-the-deal-with-captchas/577746.md?page=2)
