# What exactly is the arrangement between search engines and news providers?

**URL:** <https://boards.straightdope.com/t/what-exactly-is-the-arrangement-between-search-engines-and-news-providers/713397>\
**Category:** Factual Questions\
**Created:** [February 23, 2015, 10:18am UTC](https://boards.straightdope.com/t/what-exactly-is-the-arrangement-between-search-engines-and-news-providers/713397 "2015-02-23T10:18:58Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![davidmich](https://avatars.discourse-cdn.com/v4/letter/d/e56c9b/32.png) [@davidmich](https://boards.straightdope.com/u/davidmich)\
**Post date:** [February 23, 2015, 10:18am UTC](https://boards.straightdope.com/t/what-exactly-is-the-arrangement-between-search-engines-and-news-providers/713397/1 "2015-02-23T10:18:58Z")

</div>

Hi  
What exactly is the arrangement between search engines and news providers regarding the distribution of news content? I remember reading of an agreement between the New York Times and Google some years ago on this very topic.

As of now, if you haven’t subscribed to the New Yorker magazine, you will not be able to access their articles. Does the New Yorker make agreements with individual search engines not to distribute their content to non-subscribers? What if subscribers were to copy content and make it available on blogs, would search engines block it? I look forward to your feedback.  
davidmich

---

<div class="post-metadata">

**Author:** ![PastTense](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pasttense/32/14550_2.png) [@PastTense](https://boards.straightdope.com/u/PastTense)\
**Post date:** [February 23, 2015, 5:30pm UTC](https://boards.straightdope.com/t/what-exactly-is-the-arrangement-between-search-engines-and-news-providers/713397/2 "2015-02-23T17:30:02Z")

</div>

Any site can set up a robots.txt file on their site:

> [@](#):
>
> A robots.txt file is a file at the root of your site that indicates those parts of your site you don’t want accessed by search engine crawlers. The file uses the Robots Exclusion Standard, which is a protocol with a small set of commands that can be used to indicate access to your site by section and by specific kinds of web crawlers (such as mobile crawlers vs desktop crawlers).  
> You only need a robots.txt file if your site includes content that you don’t want Google or other search engines to index.
> 
> To test which URLs Google can and cannot access on your website, try using the robots.txt Tester.

> **[Robots.txt Introduction and Guide | Google Search Central  | ...](https://developers.google.com/search/docs/crawling-indexing/robots/intro?hl=en&visit_id=638687021343820679-4275454275&rd=1)**
>
> Robots.txt is used to manage crawler traffic. Explore this robots.txt introduction guide to learn what robot.txt files are and how to use them.

Copyright holders can request that Google and other search engines remove copyrighted materials from the search results–and Google does this tens of millions of times a year  
[http://www.google.com/transparencyreport/removals/](http://www.google.com/transparencyreport/removals/)

---

<div class="post-metadata">

**Author:** ![RealityChuck](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/realitychuck/32/195_2.png) [@RealityChuck](https://boards.straightdope.com/u/RealityChuck)\
**Post date:** [February 23, 2015, 5:34pm UTC](https://boards.straightdope.com/t/what-exactly-is-the-arrangement-between-search-engines-and-news-providers/713397/3 "2015-02-23T17:34:24Z")

</div>

The crawlers can only find information on publicly available web pages. If there’s a log in, they have no access.

The robot.txt files are only for public pages.

Now, if you copied text from a protected site and put it on your blog, Google would find it. I suspect that would be against the terms of service of the original site (and maybe even the blog), so you’d lose your access if it’s egregious enough.

---

<div class="post-metadata">

**Author:** ![CC](https://avatars.discourse-cdn.com/v4/letter/c/f1d935/32.png) [@CC](https://boards.straightdope.com/u/CC)\
**Post date:** [February 23, 2015, 5:35pm UTC](https://boards.straightdope.com/t/what-exactly-is-the-arrangement-between-search-engines-and-news-providers/713397/4 "2015-02-23T17:35:17Z")

</div>

Whatever the arrangement is, it seems that some sites have permission to use articles from, say, the New Yorker, even if you normally have to have a subscription in order to access such an article. For example, _Arts & Letters Daily_ frequently does this.

---

<div class="post-metadata">

**Author:** ![davidmich](https://avatars.discourse-cdn.com/v4/letter/d/e56c9b/32.png) [@davidmich](https://boards.straightdope.com/u/davidmich)\
**Post date:** [February 24, 2015, 2:57am UTC](https://boards.straightdope.com/t/what-exactly-is-the-arrangement-between-search-engines-and-news-providers/713397/5 "2015-02-24T02:57:54Z")

</div>

So that’s how “gating” news content is done. I had never heard of the robots.txt file. before. Thank you all very much.  
davidmcih

---

<div class="post-metadata">

**Author:** ![Derleth](https://avatars.discourse-cdn.com/v4/letter/d/b9e5f3/32.png) [@Derleth](https://boards.straightdope.com/u/Derleth)\
**Post date:** [February 24, 2015, 3:18am UTC](https://boards.straightdope.com/t/what-exactly-is-the-arrangement-between-search-engines-and-news-providers/713397/6 "2015-02-24T03:18:02Z")

</div>

> [@PastTense](#):
>
> Copyright holders can request that Google and other search engines remove copyrighted materials from the search results–and Google does this tens of millions of times a year  
> [Google Transparency Report](http://www.google.com/transparencyreport/removals/)

Would that it were only the legitimate copyright holders making such requests. Anyone can request Google stop linking to anything and there are apparently no repercussions, beyond Google ignoring the most egregious asshats.

[Case in point](http://torrentfreak.com/the-worlds-most-idiotic-copyright-complaint-150222/), taken directly from the requests, which Google makes public:

> [@](#):
>
> At least once a month TorrentFreak reports on the often crazy world of DMCA takedown notices. Google is kind enough to publish thousands of them in its Transparency Report and we’re only too happy to spend hours trawling through them.
> 
> Every now and again a real gem comes to light, often featuring mistakes that show why making these notices public is not only a great idea but also in the public interest. The ones we found this week not only underline that assertion in bold, but are actually the worst examples of incompetence we’ve ever seen.
> 
> German-based Total Wipes Music Group have made these pages before after trying to censor entirely legal content published by Walmart, Ikea, Fair Trade USA and Dunkin Donuts.
> 
> [snip]
> 
> Going after alleged pirates of the album “In To The Wild – Vol.7″ on Aborigeno Music, Total Wipes offer their pièce de résistance, the veritable jewel in their crown. The notice, which covers 95 URLs, targets no music whatsoever. Instead it tries to ruin the Internet by targeting the download pages of some of the most famous online companies around.
> 
> In no particular order, here is a larger selection of some of the download pages the notice attacks.
> 
> ICQ, RedHat, SQLite, Vuze, LinuxMint, WineHQ, Foxit, Calibre, Kodi/XBMC, Skype, Java, OpenOffice, Gimp, Ubuntu, Python, TeamViewer, MySQL, VLC, Joomla, Z-Zip, RaspberryPI, Unity3D, Apache, MalwareBytes, Pidgin, LibreOffice, VMWare, uTorrent, WinSCP, WhatsApp, Evernote, AMD, AVG, Origin, TorProject, PHPMyAdmin, Nginx, FFmpeg, phpbb, Plex, GNU, WireShark, Dropbox and Opera.

[Here’s the whole thing.](https://www.chillingeffects.org/notices/10416081)
