# Borderline surface web sites

**URL:** <https://boards.straightdope.com/t/borderline-surface-web-sites/772211>\
**Category:** Factual Questions\
**Created:** [November 17, 2016, 12:42am UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211 "2016-11-17T00:42:34Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![correlophus](https://avatars.discourse-cdn.com/v4/letter/c/58956e/32.png) [@correlophus](https://boards.straightdope.com/u/correlophus)\
**Post date:** [November 17, 2016, 12:42am UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/1 "2016-11-17T00:42:34Z")

</div>

Now I am reading about the Deep Web, how it is accessed, and if the Surface Web has places that have some characteristics of the Deep Web. The author of the blog characterizes some sites as borderline surface web, that is, Google can index them with difficulty and anonimity is prevalent. He names as examples reddit and 4chan. Are there any sites that are officially called borderline websites? Can we compare some places of the surface web to the deep web? I know reddit is a well-known website. Is it considered deep?

---

<div class="post-metadata">

**Author:** ![snfaulkner](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/snfaulkner/32/433_2.png) [@snfaulkner](https://boards.straightdope.com/u/snfaulkner)\
**Post date:** [November 17, 2016, 1:00am UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/2 "2016-11-17T01:00:35Z")

</div>

> [@correlophus](#):
>
> Now I am reading about the Deep Web, how it is accessed, and if the Surface Web has places that have some characteristics of the Deep Web. The author of the blog characterizes some sites as borderline surface web, that is, Google can index them with difficulty and anonimity is prevalent. He names as examples reddit and 4chan. Are there any sites that are officially called borderline websites? Can we compare some places of the surface web to the deep web? I know reddit is a well-known website. Is it considered deep?

What blog are you reading?

---

<div class="post-metadata">

**Author:** ![LSLGuy](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/lslguy/32/5813_2.png) [@LSLGuy](https://boards.straightdope.com/u/LSLGuy)\
**Post date:** [November 17, 2016, 1:47am UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/3 "2016-11-17T01:47:10Z")

</div>

Deep, surface, and borderline means whatever the blog author wants them to mean. They’re buzzwords, not internet engineering terms.

---

<div class="post-metadata">

**Author:** ![TruCelt](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/trucelt/32/523_2.png) [@TruCelt](https://boards.straightdope.com/u/TruCelt)\
**Post date:** [November 17, 2016, 3:42am UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/4 "2016-11-17T03:42:37Z")

</div>

reddit and 4chan are more analogous to gateway drugs. They are where you would gain the right knowledge, meet the right people, and get into the wrong crowd, all of which could end in an invitation to try some site you’d never find on your own.

Or you could just use a Deep Web or Dark Web search engine . . .

> **[Deep Web Search Engines to Explore the Hidden Internet](https://thehackernews.com/2016/02/deep-web-search-engine.html)**
>
> These are the top Deep Web Search Engines help to Explore the Hidden Internet behind the Tor Network

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [November 17, 2016, 2:41pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/5 "2016-11-17T14:41:46Z")

</div>

You already use parts of the Deep Web, every day. When you check your bank account balance, or read your e-mail, or edit your Google Docs, you’re viewing information over the Web which is not available to the general public. And yet, your bank makes no secret of the fact that they have accounts, and has information prominently available on their public web page about how to get one of your own.

“Deep” does not in any way mean that there’s anything shady going on. There will be some shady stuff in the Deep Web, of course, but then, there will be some in the public Web, too. It might even be safer on the public web, because not requiring login credentials might make it easier to maintain anonymity.

---

<div class="post-metadata">

**Author:** ![aldiboronti](https://avatars.discourse-cdn.com/v4/letter/a/9fc348/32.png) [@aldiboronti](https://boards.straightdope.com/u/aldiboronti)\
**Post date:** [November 17, 2016, 2:52pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/6 "2016-11-17T14:52:10Z")

</div>

Yes, [according to a recent survey](http://opensources.info/tor-browser-news-majority-of-dark-web-content-is-perfectly-legal-study-finds/) much of the content on the Deep Web is quite innocuous.

> [@](#):
>
> If you think the dark web is nothing more than a wretched hive of scum and villainy, think again – research has shown that the majority of content hosted on it is perfectly legal.
> 
> A new report from security firm Terbian Labs reveals that while most people associate the dark web with questionable pornography, exotic narcotics and unlicensed arms deals, the reality is actually quite dull, with over 50% of all domains and URLs in the survey’s sample comprised of legal content.
> 
> “These Tor Hidden Services play host to Facebook, European graphic design firms, Scandinavian political parties, personal blogs about security, and forums to discuss privacy, technology, even erectile dysfunction,” the report explains. “Anonymity does not equate criminality, merely a desire for privacy.”

---

<div class="post-metadata">

**Author:** ![watchwolf49](https://avatars.discourse-cdn.com/v4/letter/w/e9c0ed/32.png) [@watchwolf49](https://boards.straightdope.com/u/watchwolf49)\
**Post date:** [November 17, 2016, 2:53pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/7 "2016-11-17T14:53:21Z")

</div>

Maybe we’re thinking of the Dark Web … bitcoins, hackers, those Guy Fawkes people … all those nasty places on the web.

Deep Web is stuff that’s just buried in layer after layer of directories … think of a specific latitude and longitude on Mars and we’d have to be digging through NASA’s web site to find a photo of that specific place … it’s there … but it’s `deep` in the site … many government agencies post their scientific data they’ve collected, it’s just difficult to find in some cases.

---

<div class="post-metadata">

**Author:** ![Carryon](https://avatars.discourse-cdn.com/v4/letter/c/9fc29f/32.png) [@Carryon](https://boards.straightdope.com/u/Carryon)\
**Post date:** [November 17, 2016, 3:26pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/8 "2016-11-17T15:26:53Z")

</div>

This would never have happened with Gopher 🙂

---

<div class="post-metadata">

**Author:** ![correlophus](https://avatars.discourse-cdn.com/v4/letter/c/58956e/32.png) [@correlophus](https://boards.straightdope.com/u/correlophus)\
**Post date:** [November 17, 2016, 6:27pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/9 "2016-11-17T18:27:14Z")

</div>

Ok, I conflated deep web and dark web, as usual. I was specifically meaning the tor network, and if some surface sites are somewhat similar to that. The specific article in a blog I read that is below:

> **[TOR and The Dark Net Learn To Avoid NSA Spying And Become Anonymous Online](https://1earthunite.wordpress.com/2016/10/31/tor-and-the-dark-net-learn-to-avoid-nsa-spying-and-become-anonymous-online/)**
>
> Introduction Chapter 1 What is the Deep Web and Why Is It Worth Exploring? Chapter 2 The Only Ways to Surf Anonymously Chapter 3 Pros and Cons of Using Tor Chapter 4 Pros and Cons of VPN Chapter 5 …

---

<div class="post-metadata">

**Author:** ![74westy](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/74westy/32/3950_2.png) [@74westy](https://boards.straightdope.com/u/74westy)\
**Post date:** [November 17, 2016, 7:42pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/10 "2016-11-17T19:42:51Z")

</div>

> [@correlophus](#):
>
> Ok, I conflated deep web and dark web, as usual. I was specifically meaning the tor network, and if some surface sites are somewhat similar to that. The specific article in a blog I read that is below:  
> [TOR and The Dark Net Learn To Avoid NSA Spying And Become Anonymous Online | 1EarthUnited](https://1earthunite.wordpress.com/2016/10/31/tor-and-the-dark-net-learn-to-avoid-nsa-spying-and-become-anonymous-online/)

Over 14,000 words and not a single paragraph break.

---

<div class="post-metadata">

**Author:** ![snfaulkner](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/snfaulkner/32/433_2.png) [@snfaulkner](https://boards.straightdope.com/u/snfaulkner)\
**Post date:** [November 17, 2016, 10:24pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/11 "2016-11-17T22:24:17Z")

</div>

I got as far as “1EarthUnited” and dismissed the entire thing as rubbish.

---

<div class="post-metadata">

**Author:** ![aldiboronti](https://avatars.discourse-cdn.com/v4/letter/a/9fc348/32.png) [@aldiboronti](https://boards.straightdope.com/u/aldiboronti)\
**Post date:** [November 17, 2016, 10:26pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/12 "2016-11-17T22:26:17Z")

</div>

Right, I confused the Deep Web and the Dark Web in my post above. Here’s a definition from PC Advisor:

> [@](#):
>
> Although all of these terms tend to be used interchangeably, they don’t refer to exactly the same thing. An element of nuance is required. The ‘Deep Web’ refers to all web pages that search engines cannot find. Thus the ‘Deep Web’ includes the ‘Dark Web’, but also includes all user databases, webmail pages, registration-required web forums, and pages behind paywalls. There are huge numbers of such pages, and most exist for mundane reasons.

The survey I linked covered the Dark Web.

---

<div class="post-metadata">

**Author:** ![quimper](https://avatars.discourse-cdn.com/v4/letter/q/7c8e57/32.png) [@quimper](https://boards.straightdope.com/u/quimper)\
**Post date:** [November 17, 2016, 11:22pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/13 "2016-11-17T23:22:30Z")

</div>

> [@Carryon](#):
>
> This would never have happened with Gopher 🙂

Cmon man, no tables, frames or inline image support? Mosaic kicked Gopher’s ass fair and square.

---

<div class="post-metadata">

**Author:** ![dracoi](https://avatars.discourse-cdn.com/v4/letter/d/90db22/32.png) [@dracoi](https://boards.straightdope.com/u/dracoi)\
**Post date:** [November 18, 2016, 1:23am UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/14 "2016-11-18T01:23:37Z")

</div>

I don’t know if this is a helpful thought, but the SDMB is a perfect example of surface and deep content side by side. This post is now public information and searchable by the likes of Google. But if I sent you a private message, that is part of the deep web, only accessible with your username and password. And if I use the Tor browser to read the SDMB… well, I haven’t changed the SDMB at all, but I am using dark web technology to encrypt data and mask IPs.

---

<div class="post-metadata">

**Author:** ![correlophus](https://avatars.discourse-cdn.com/v4/letter/c/58956e/32.png) [@correlophus](https://boards.straightdope.com/u/correlophus)\
**Post date:** [November 18, 2016, 10:16am UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/15 "2016-11-18T10:16:57Z")

</div>

> [@74westy](#):
>
> Over 14,000 words and not a single paragraph break.

And I read it whole!

---

<div class="post-metadata">

**Author:** ![Terminus\_Est](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/terminus_est/32/3087_2.png) [@Terminus\_Est](https://boards.straightdope.com/u/Terminus_Est)\
**Post date:** [November 18, 2016, 10:32am UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/16 "2016-11-18T10:32:59Z")

</div>

Characterizing Reddit and even 4Chan as “borderline” speaks of a misunderstanding of the nature of those forums. If the primary characteristic of borderline is that Google can’t index them, then the SDMB would have been borderline before Google was allowed to index us. It’s really to block Google and other legitimate webcrawlers as that just requires a couple of lines in the robots.txt file.

---

<div class="post-metadata">

**Author:** ![LSLGuy](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/lslguy/32/5813_2.png) [@LSLGuy](https://boards.straightdope.com/u/LSLGuy)\
**Post date:** [November 18, 2016, 1:44pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/17 "2016-11-18T13:44:35Z")

</div>

> [@Terminus\_Est](#):
>
> …  
> It’s really [easy] to block Google and other legitimate webcrawlers as that just requires a couple of lines in the robots.txt file.

[Bracketing] inserted by me to clarify.

Agree with all you’ve said. Which makes me think of a question …

Ref the snippet above, compliance with robots.txt is 100% voluntary. An interesting question is whether there are any publicly available search engines that advertise they don’t abide by robots.txt?

Sure, any given webmaster could try to IP-block such an unfriendly search engine. But that’s a futile game of whack-a-mole versus any good-sized crawler infrastructure.

---

<div class="post-metadata">

**Author:** ![Derleth](https://avatars.discourse-cdn.com/v4/letter/d/b9e5f3/32.png) [@Derleth](https://boards.straightdope.com/u/Derleth)\
**Post date:** [November 18, 2016, 2:14pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/18 "2016-11-18T14:14:01Z")

</div>

> [@quimper](#):
>
> Cmon man, no tables, frames or inline image support? Mosaic kicked Gopher’s ass fair and square.

Eh, you can serve HTML over Gopher just as easily as anything else. Images, too.

---

<div class="post-metadata">

**Author:** ![watchwolf49](https://avatars.discourse-cdn.com/v4/letter/w/e9c0ed/32.png) [@watchwolf49](https://boards.straightdope.com/u/watchwolf49)\
**Post date:** [November 18, 2016, 2:15pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/19 "2016-11-18T14:15:56Z")

</div>

> [@quimper](#):
>
> Cmon man, no tables, frames or inline image support? Mosaic kicked Gopher’s ass fair and square.

Wow … Mosaic … seeing that word makes me feel very very old … do you Alta Vista ???

---

<div class="post-metadata">

**Author:** ![Derleth](https://avatars.discourse-cdn.com/v4/letter/d/b9e5f3/32.png) [@Derleth](https://boards.straightdope.com/u/Derleth)\
**Post date:** [November 18, 2016, 2:25pm UTC](https://boards.straightdope.com/t/borderline-surface-web-sites/772211/20 "2016-11-18T14:25:14Z")

</div>

> [@LSLGuy](#):
>
> Ref the snippet above, compliance with robots.txt is 100% voluntary. An interesting question is whether there are any publicly available search engines that advertise they don’t abide by robots.txt?

No. That would be not only blatantly antisocial, but monumentally _stupid_ from the perspective of the robot’s operator.

A lot of what robots.txt does these days is protect website backends from robots and, therefore, robots from themselves, in the form of notifying robots about dynamically-generated content which can be effectively infinite, generated programmatically from whatever internal database the website draws from; unless the robot’s owner wants to be on the wrong end of a combinatorial explosion, it programs the robot to respect robots.txt and avoid some infinite tarpits that machines really cannot navigate.

> [@](#):
>
> Sure, any given webmaster could try to IP-block such an unfriendly search engine. But that’s a futile game of whack-a-mole versus any good-sized crawler infrastructure.

Well, here you get into the difference between what semi-legitimate but assholish people do and what spammers with hordes of zombies do. Sure, a good-sized search engine company might own a lot of different computers sitting behind a lot of different IP addresses, but since it will have leased those computers legally from one or two other companies, or will own them themselves, all of those IP addresses will be in a few specific netblocks, owned by the relevant companies, as recorded in the information associated with the Autonomous Systems which advertise those netblocks as their own. In short, all of those IP addresses will be coming from “the same place”, in a networking sense, and it will be easy to block all of them with a few commands.

The spammers don’t do that. They own zombies, created through foul magicks involving unpatched Windows XP machines sitting behind cable modems, and therefore their IP addresses could come from anywhere on Earth. Blocking them is more of a game of whack-a-mole, but programming your server software to rate-limit any specific IP address which tries to go too fast, or tries to grab the wrong things, is a lot easier.
