# SDMB not crawled by search engines?

**URL:** <https://boards.straightdope.com/t/sdmb-not-crawled-by-search-engines/112019>\
**Category:** About This Message Board\
**Created:** [June 2, 2002, 4:05pm UTC](https://boards.straightdope.com/t/sdmb-not-crawled-by-search-engines/112019 "2002-06-02T16:05:54Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Johanna](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/johanna/32/8318_2.png) [@Johanna](https://boards.straightdope.com/u/Johanna)\
**Post date:** [June 2, 2002, 4:05pm UTC](https://boards.straightdope.com/t/sdmb-not-crawled-by-search-engines/112019/1 "2002-06-02T16:05:54Z")

</div>

Ever notice that no matter how many times you Google something, it never picks up any pages from SDMB? Is there some kind of firewall here to keep search engines out?

---

<div class="post-metadata">

**Author:** ![yabob](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/yabob/32/2821_2.png) [@yabob](https://boards.straightdope.com/u/yabob)\
**Post date:** [June 2, 2002, 4:41pm UTC](https://boards.straightdope.com/t/sdmb-not-crawled-by-search-engines/112019/2 "2002-06-02T16:41:36Z")

</div>

I HAVE seen SDMB pages in google occasionally. Not often, I will admit.

There is a recognized standard to preclude crawling by webrobots:

[http://www.searchengineworld.com/robots/robots\_tutorial.htm](http://www.searchengineworld.com/robots/robots_tutorial.htm)

[http://boards.straightdope.com/robots.txt](http://boards.straightdope.com/robots.txt) does not exist, but [http://www.straightdope.com/robots.txt](http://www.straightdope.com/robots.txt) does, and contains:

```auto

User-agent: *
Disallow: /bonus/

```

This suffices to keep well-behaved robots from crawling past the published straight dope “front door”, which might be how they would normally reach the message boards.

There could be some blocks placed against known search engines at other levels, of course, either at a firewall, or simply by IP blocking in vBulletin.

BTW, that “Disallow:” line advertises another path under [www.straightdope.com](http://www.straightdope.com), which produces the straightdope banner and footer with the text “Hey! You’re not supposed to be rooting around in here!” for the content. Does any admin care to comment on what’s in /bonus/?

---

<div class="post-metadata">

**Author:** ![yabob](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/yabob/32/2821_2.png) [@yabob](https://boards.straightdope.com/u/yabob)\
**Post date:** [June 2, 2002, 4:47pm UTC](https://boards.straightdope.com/t/sdmb-not-crawled-by-search-engines/112019/3 "2002-06-02T16:47:18Z")

</div>

DUH!

Excuse me. I’m an idiot. I stated that backwards. The [www.straightdope.com/robots.txt](http://www.straightdope.com/robots.txt) file allows all robots in, except that it disallows them from crawling “/bonus”. So, it DOESN’T stop robots from crawling the links from the front page, only down that mysterious /bonus path.

---

<div class="post-metadata">

**Author:** ![yabob](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/yabob/32/2821_2.png) [@yabob](https://boards.straightdope.com/u/yabob)\
**Post date:** [June 2, 2002, 5:04pm UTC](https://boards.straightdope.com/t/sdmb-not-crawled-by-search-engines/112019/4 "2002-06-02T17:04:29Z")

</div>

More on this. There is also a robots \<META\> tag which is supposed to be honored:

[http://searchengineworld.com/metatag/robots.htm](http://searchengineworld.com/metatag/robots.htm)

The tag doesn’t seem to be present in the SDMB pages. IP’s of known search engines could still be blocked by other mechanisms, as I said.

---

<div class="post-metadata">

**Author:** ![femtosecond](https://avatars.discourse-cdn.com/v4/letter/f/f08c70/32.png) [@femtosecond](https://boards.straightdope.com/u/femtosecond)\
**Post date:** [June 2, 2002, 7:13pm UTC](https://boards.straightdope.com/t/sdmb-not-crawled-by-search-engines/112019/5 "2002-06-02T19:13:28Z")

</div>

> [@](#):
>
> _Originally posted by yabob_  
> **"Hey! You’re not supposed to be rooting around in here!" for the content. Does any admin care to comment on what’s in /bonus/?**

Hehe, I found that, too. But you _dare_ to ask? :eek:

This thread may be interesting to you: [Why isn’t the SDMB indexed on Google?](http://boards.straightdope.com/sdmb/showthread.php?threadid=107972)

No limiting on our side, it seems. At least Google limits its crawling on dynamically generated sites, but we know by now of at least two archiving sites ([www.archive.org](http://www.archive.org) and [www.boardreader.com](http://www.boardreader.com)) which didn’t encounter any form of resistance when spidering our site.

I’m still eager to hear from our ‘board-officials’ if they think losing some bandwidth to traffic generated by crawlers is an issue.

And about the /bonus/ thing. 🙂
