# Are Google's archives a copyright violation

**URL:** https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917
**Category:** Factual Questions
**Created:** [August 4, 2003, 4:11pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917 "2003-08-04T16:11:27Z")
**Posts on this page:** 11
**Page:** 1

<div class="post-metadata">

### Author: ![Walloon](https://avatars.discourse-cdn.com/v4/letter/w/fbc32d/32.png) [@Walloon](https://boards.straightdope.com/u/Walloon)
#### Post date: [August 4, 2003, 4:11pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/1 "2003-08-04T16:11:27Z")

</div>

The search engine Google archives virtually all of the Web pages it indexes, meaning that even if the Web page is taken down by its owner, a cached copy of the Web page (a “snapshot”, as they put it) will remain available to the public on Google’s servers.

How is this not a violation of copyright?

---

<div class="post-metadata">

### Author: ![RealityChuck](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/realitychuck/32/195_2.png) [@RealityChuck](https://boards.straightdope.com/u/RealityChuck)
#### Post date: [August 4, 2003, 4:51pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/2 "2003-08-04T16:51:19Z")

</div>

Technically, it is. But no copyright holder has taken them to court over it, so they continue to do it. If the copyright holder contacts Google, Google probably just takes it down to avoid the hassles.

[http://www.archive.org](http://www.archive.org) is in the same boat, and has removed pages when asked.

---

<div class="post-metadata">

### Author: ![micco](https://avatars.discourse-cdn.com/v4/letter/m/5f8ce5/32.png) [@micco](https://boards.straightdope.com/u/micco)
#### Post date: [August 4, 2003, 5:07pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/3 "2003-08-04T17:07:46Z")

</div>

[http://boards.straightdope.com/sdmb/showthread.php?threadid=191895](http://boards.straightdope.com/sdmb/showthread.php?threadid=191895)

---

<div class="post-metadata">

### Author: ![citybadger](https://avatars.discourse-cdn.com/v4/letter/c/45deac/32.png) [@citybadger](https://boards.straightdope.com/u/citybadger)
#### Post date: [August 4, 2003, 5:32pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/4 "2003-08-04T17:32:59Z")

</div>

I visit a web page. The information downloaded goes to my browser’s cache. I can view that page offline later. Am I violating copyright? Of course not, if you put up something for people to download, you can’t complain when they download it.

Likewise, Google could make a case that they are just downloading and caching content like every other visitor. A reasonable position would be that there is an implied consent for a use like Google’s. Especially since I believe a web page owner can prevent Google indexing if they want by using HTML tags.

The big difference is that Google is redistributing their cache while mine sits on my hard drive, accessable only to people who use my computer. Of course my ISP’s router, or my employer’s proxy server, is also “redistributing” the web page is some sense.

And no one is really sure what that means. Wait for more legislation followed by 20 years of case law to sort it all out.

---

<div class="post-metadata">

### Author: ![Whack-a-Mole](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/whack-a-mole/32/141_2.png) [@Whack-a-Mole](https://boards.straightdope.com/u/Whack-a-Mole)
#### Post date: [August 4, 2003, 5:37pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/5 "2003-08-04T17:37:01Z")

</div>

If they were then wouldn’t _any_ archive be a copyright violation? Such a libraries?

---

<div class="post-metadata">

### Author: ![Walloon](https://avatars.discourse-cdn.com/v4/letter/w/fbc32d/32.png) [@Walloon](https://boards.straightdope.com/u/Walloon)
#### Post date: [August 4, 2003, 6:15pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/6 "2003-08-04T18:15:46Z")

</div>

A library holds _authorized_ copies of materials.

---

<div class="post-metadata">

### Author: ![Jeff\_Lichtman](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/jeff_lichtman/32/1531_2.png) [@Jeff\_Lichtman](https://boards.straightdope.com/u/Jeff_Lichtman)
#### Post date: [August 4, 2003, 6:28pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/7 "2003-08-04T18:28:29Z")

</div>

> [@](#):
>
> \*Originally posted by Whack-a-Mole \*  
> \*\*If they were then wouldn’t _any_ archive be a copyright violation? Such a libraries? \*\*

I can’t let this one go by. Copyright is exactly what it sounds like: the right to copy. It is not a violation of copyright merely to archive materials that one has paid for, which is what libraries do. It would be a violation if a library kept materials that had been copied illegally, but libraries generally don’t do this.

Copyright law is complicated, and depends on precedent as well as on written law. The question of caching has not been tested in court. We can only speculate as to how the courts would rule.

An important difference between a browser cache and the Google cache is that Google makes their cache available to others. It is possible that the courts could rule that local caching is a form of fair use, while making cached copies available to others is not.

---

<div class="post-metadata">

### Author: ![micco](https://avatars.discourse-cdn.com/v4/letter/m/5f8ce5/32.png) [@micco](https://boards.straightdope.com/u/micco)
#### Post date: [August 4, 2003, 7:22pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/8 "2003-08-04T19:22:45Z")

</div>

> [@](#):
>
> \*Originally posted by Walloon \*  
> \*\*A library holds _authorized_ copies of materials. \*\*

Google’s archive is comprised of authorized copies as well. Google’s spider creates a copy in exactly the same way your browser creates a copy in its cache when you view a page. Putting content on a webserver would seem to be tacit approval for creation of those working copies. If Google’s copies constitute infringement, then so does every hit by a browser and the entire web is just an instrument for infringment (except for those rare sites that post PD content).

As **Jeff Lichtman** points out, the difference between Google’s archive and your browser cache is that Google is redistributing. This is analogous to a library that makes photocopies of their books so nothing is ever unavailable.

IANAL and I’m not going to speculate on how a court would rule on this, but I wanted to point out that there is nothing illegitimate about Google making a copy in the first place. It’s what they do with it afterward which would bear on whether there is copyright infringement or not.

---

<div class="post-metadata">

### Author: ![sailor](https://avatars.discourse-cdn.com/v4/letter/s/a587f6/32.png) [@sailor](https://boards.straightdope.com/u/sailor)
#### Post date: [August 4, 2003, 7:45pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/9 "2003-08-04T19:45:25Z")

</div>

> [@](#):
>
> [http://www.google.com/webmasters/3.html#B1](http://www.google.com/webmasters/3.html#B1)
> 
> If you do not want your content to be accessible through Google’s cache, you can use the NOARCHIVE meta-tag. Place this in the \<HEAD\> section of your documents:  
> \<META NAME=“ROBOTS” CONTENT=“NOARCHIVE”\>
> 
> This tag will tell robots not to archive the page. Google will continue to index and follow links from the page, but will not present cached material to users.

---

<div class="post-metadata">

### Author: ![handy](https://avatars.discourse-cdn.com/v4/letter/h/b5a626/32.png) [@handy](https://boards.straightdope.com/u/handy)
#### Post date: [August 4, 2003, 9:15pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/10 "2003-08-04T21:15:43Z")

</div>

If you have a TradeMark you have to write a letter to people who are using it without permission to protect your right. But if its stored in millions of computers & you don’t write all those people, doesn’t that mean you no longer have the copyright?

---

<div class="post-metadata">

### Author: ![micco](https://avatars.discourse-cdn.com/v4/letter/m/5f8ce5/32.png) [@micco](https://boards.straightdope.com/u/micco)
#### Post date: [August 4, 2003, 9:33pm UTC](https://boards.straightdope.com/t/are-googles-archives-a-copyright-violation/192917/11 "2003-08-04T21:33:40Z")

</div>

> [@](#):
>
> \*Originally posted by handy \*  
> \*\*If you have a TradeMark you have to write a letter to people who are using it without permission to protect your right. But if its stored in millions of computers & you don’t write all those people, doesn’t that mean you no longer have the copyright? \*\*

Trademarks and Copyrights are completely different. Trademarks must be defended. Copyrights need not be. That means that every instance of trademark infringment can contribute to a [dilution](http://www.ipwatchdog.com/dilution.html#n3) of the mark which might eventually mean the owner of the trademark is unable to protect it because it has entered common use (e.g. the owners of Band-Aid and Xerox constantly battle against their generic use to avoid losing their trademark rights). Copyrights are completely different and infringement has no bearing on continued ownership of the rights.
