# Anyone experienced with Archive.org?

**URL:** https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480
**Category:** Miscellaneous and Personal Stuff I Must Share
**Created:** [October 5, 2024, 6:40pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480 "2024-10-05T18:40:34Z")
**Posts on this page:** 20
**Page:** 1

<div class="post-metadata">

### Author: ![GailForce](https://avatars.discourse-cdn.com/v4/letter/g/dfb087/32.png) [@GailForce](https://boards.straightdope.com/u/GailForce)
#### Post date: [October 5, 2024, 6:40pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/1 "2024-10-05T18:40:34Z")

</div>

Otherwise known as the Wayback Machine? The idea is fabulous, a way to check out anything (in theory) that’s ever appeared online, even if the site itself is defunct, although I’ve found it non-intuitive to use in practice, and I’ve been hit or miss in actually tracking down defunct sites or just things I’ve read online that resist a thorough Google search. Has anyone used it successfully, and do you have any tips you’d like to share? It’s frustrating, knowing that the material is out there somewhere, and there’s a way to search for it, but a search for specific things often yields no usable results. I suspect there’s a better way to use [archive.org](http://archive.org) than I’ve been doing.

---

<div class="post-metadata">

### Author: ![filmore](https://avatars.discourse-cdn.com/v4/letter/f/7993a0/32.png) [@filmore](https://boards.straightdope.com/u/filmore)
#### Post date: [October 5, 2024, 6:51pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/2 "2024-10-05T18:51:19Z")

</div>

It is somewhat hit-or-miss as to the websites they have. I’m not sure that they archive every website. You can ask them to archive a site by using the “Save Page Now” link on the main page.

If you’re wondering what the SDMB was like in the olden times, they have some archives. Here’s the earliest link I could find from 8/15/2000:

[https://web.archive.org/web/20000815063701/http://boards.straightdope.com/sdmb/](https://web.archive.org/web/20000815063701/http://boards.straightdope.com/sdmb/)

---

<div class="post-metadata">

### Author: ![blondebear](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/blondebear/32/1022_2.png) [@blondebear](https://boards.straightdope.com/u/blondebear)
#### Post date: [October 5, 2024, 7:13pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/3 "2024-10-05T19:13:43Z")

</div>

I often go there as sort of last resort looking for videos. The image quality isn’t always great but they’re still watchable and that’s good enough for me. One example: [The Compleat Beatles](https://archive.org/details/the-compleat-beatles-1982/COMPRESSED+FILES/The+Compleat+Beatles+(1982).mkv) – it was surpassed (and kind of killed off) by Anthology but it stands on it’s own as a worthy documentary.

They also have a pretty good collection of books that you can check out.

---

<div class="post-metadata">

### Author: ![GailForce](https://avatars.discourse-cdn.com/v4/letter/g/dfb087/32.png) [@GailForce](https://boards.straightdope.com/u/GailForce)
#### Post date: [October 5, 2024, 7:35pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/4 "2024-10-05T19:35:21Z")

</div>

Kind of klunky, though, innit? Slow, quirky, counter-intuitive. And very incomplete. That SDMB link for example features all sorts of topics in various forums here that I’d LOVE to read, but the links are to nothing. Click on them and you get “Sorry–this page not archived by [archive.com](http://archive.com)” which is kind of going into a cool restaurant, reading the fascinating menu and then finding out “Sorry–we don’t serve actual food here.”

---

<div class="post-metadata">

### Author: ![AHunter3](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/ahunter3/32/368_2.png) [@AHunter3](https://boards.straightdope.com/u/AHunter3)
#### Post date: [October 5, 2024, 8:54pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/5 "2024-10-05T20:54:10Z")

</div>

It’s most useful when you have a saved link that goes to a no longer extant site/page. So you know what the old URL was and can attempt to fetch it from the Wayback Machine.

What they _need_ is a bloody search engine.

God, would it ever be cool if you could go [here](http://web.archive.org/web/20010622115932/http://www.altavista.com/sites/search/adv) and could actually use it!!

---

<div class="post-metadata">

### Author: ![Mops](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mops/32/16494_2.png) [@Mops](https://boards.straightdope.com/u/Mops)
#### Post date: [October 5, 2024, 9:03pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/6 "2024-10-05T21:03:59Z")

</div>

> [@AHunter3](#):
>
> It’s most useful when you have a saved link that goes to a no longer extant site/page. So you know what the old URL was and can attempt to fetch it from the Wayback Machine.
> 
> …

What I often use them is the old versions of a page that still exists.

For example recently I found the URL where the food safety inspectors of our district (_Landkreis_) publish temporary restaurant closures for hygiene violations, verbally and in great detail (also with details of remediation; the restaurants are usually allowed to open again after passing an inspection a few days later).  
By relevant regulations the listings on that page are purged after three months; by using the Wayback Machine I could access earlier entries.

---

<div class="post-metadata">

### Author: ![ParallelLines](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/parallellines/32/3575_2.png) [@ParallelLines](https://boards.straightdope.com/u/ParallelLines)
#### Post date: [October 5, 2024, 9:26pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/7 "2024-10-05T21:26:53Z")

</div>

I’ve had qualified success using it. The complaints are totally fair, it’s slow to load, incomplete, and while it preserves snapshots of individual sites, it doesn’t preserve the entirety of what they link too.

But it’s great for finding otherwise lost little bits of information. In fact, one of my first posts here (rather than just lurking sans account) was because of memories I eventually recovered with the Wayback Machine.

> [@Which part is Cecil's reply and which part is C K Dexter Haven's?](https://boards.straightdope.com/t/which-part-is-cecils-reply-and-which-part-is-c-k-dexter-havens/854878/10):
>
> @Max_S This had been sticking in my brain for whatever reason after I sent my best recollections, and since I was home sick (not with COVID I’m pretty sure) I figured I’d dig it up for you. Hail the Glory of the WaybackMachine! It was indeed as I remembered, the part you have is the response from DEX, the missing information can be found in the link below. Since it is super slow to load, I’ve copied Cecil’s response in it’s entirety. Cecil Adams replies: Normally I don’t comment on Mail…

NOTE that if you go to the SD site these days, it’s still missing the last section, and attributes the whole thing to Cecil despite being only Dex’s segment!

Anyway, I also wanted to say [archive.org](http://archive.org) is far more than just the wayback machine, with audio, visual, “print” and software stored from times otherwise lost. I’ve played Oregan Trail, found old radio dramas, and picked up audiobooks of classic works.

YES, the interface is positively reminiscent of the 80s, and the search is in desperate need of updates, but I consider the whole thing like going to a _good_ used book store. The organization may be haphazard, and they are likely to be missing much of what you want, but what you can find you likely wouldn’t find anywhere else!

---

<div class="post-metadata">

### Author: ![GailForce](https://avatars.discourse-cdn.com/v4/letter/g/dfb087/32.png) [@GailForce](https://boards.straightdope.com/u/GailForce)
#### Post date: [October 7, 2024, 6:08pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/8 "2024-10-07T18:08:40Z")

</div>

Any tips for navigating the counter-intuitive site? It baffles me. I’m sure there are some ways to find things that I’m simply unaware of.

---

<div class="post-metadata">

### Author: ![pulykamell](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pulykamell/32/3166_2.png) [@pulykamell](https://boards.straightdope.com/u/pulykamell)
#### Post date: [October 7, 2024, 6:22pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/9 "2024-10-07T18:22:13Z")

</div>

I love the site, but I usually have a URL that I already know I’m looking for. I just type it in the Wayback machine, see what page snapshots it has, click on one that’s about the timeframe I’m looking for, then possibly click on a time if there’s multiple snapshots per day, and just hope for the best. I don’t expect it to be a complete record of the web; I just can’t imagine what kind of resources it would take to store all that, plus there are simply internal databases they can’t get access to, anyway. It does very well for what it does.

---

<div class="post-metadata">

### Author: ![motu](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/motu/32/259_2.png) [@motu](https://boards.straightdope.com/u/motu)
#### Post date: [October 7, 2024, 6:23pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/10 "2024-10-07T18:23:38Z")

</div>

I mainly use [Archive.org](http://Archive.org) for the Live Music Archive

[https://archive.org/browse.php?collection=etree&field=creator](https://archive.org/browse.php?collection=etree&field=creator)

---

<div class="post-metadata">

### Author: ![Exapno\_Mapcase](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/exapno_mapcase/32/1051_2.png) [@Exapno\_Mapcase](https://boards.straightdope.com/u/Exapno_Mapcase)
#### Post date: [October 7, 2024, 6:37pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/11 "2024-10-07T18:37:41Z")

</div>

As the posts here indicate, the site contains not just huge amounts of material but huge differences in the types of material that they save. What you want to search for will determine whether any search strategy is optimal.

I’ve done hundreds of searches for material that not in Google Books and also before the internet, mostly magazines and newspapers. Somebody has to scan those in one by one from real paper copies. Finding any specific issue is chancy. The more you know about the name and year and date, the better. Keyword searches don’t work well.

Many organizations don’t want you to go to the site to find stuff, preferring to sell the ability to search their archives. Or they want to force you to see all their delicious advertisements on the current page.

The Dope has had times when material was lost, especially in the olden days. Without knowing how Archive ran their spiders in times past I can’t say whether they would ever have captured the pages in the first place.

In short, no one search strategy will work. Oh, and remember there are two search boxes on the front page. URLs don’t work very well in the second box. Keywords do. You also get to choose between metadata and text, an important difference, plus archived web sites.

---

<div class="post-metadata">

### Author: ![hogarth](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/hogarth/32/1773_2.png) [@hogarth](https://boards.straightdope.com/u/hogarth)
#### Post date: [October 7, 2024, 6:43pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/12 "2024-10-07T18:43:30Z")

</div>

I think the most recent thing I’ve used it for was downloading old MP3s of “Fibber McGee and Molly” and “The Great Gildersleeve” (pre-pandemic).

There was an article on the BBC recently explaining the challenges behind it:

> **[We're losing our digital history. Can the Internet Archive save it?](https://www.bbc.com/future/article/20240912-the-archivists-battling-to-save-the-internet)**
>
> Research shows 25% of web pages posted between 2013 and 2023 have vanished. A few organisations are racing to save the echoes of the web, but new risks threaten their very existence.

---

<div class="post-metadata">

### Author: ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)
#### Post date: [October 7, 2024, 6:59pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/13 "2024-10-07T18:59:14Z")

</div>

> [@Exapno\_Mapcase](#):
>
> Somebody has to scan those in one by one from real paper copies.

I know someone who sent the Internet Archive boxes of historically important documents to be properly scanned and digitized. They are supposed to provide that service.

---

<div class="post-metadata">

### Author: ![Pardel-Lux](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pardel-lux/32/20182_2.png) [@Pardel-Lux](https://boards.straightdope.com/u/Pardel-Lux)
#### Post date: [October 7, 2024, 7:30pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/14 "2024-10-07T19:30:34Z")

</div>

Ooooh! I had not thought of looking up my old, long defunct webpages, but they are there! Still functioning! Amazing. That deserves a donation.  
Now if I could only retrieve my old compuserve(dot)com e-mail… but when I send an e-mail to my former self, I get a delivery status notification. I guess that one is gone for good after decades of not using it. Still I miss it sometimes.

---

<div class="post-metadata">

### Author: ![Exapno\_Mapcase](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/exapno_mapcase/32/1051_2.png) [@Exapno\_Mapcase](https://boards.straightdope.com/u/Exapno_Mapcase)
#### Post date: [October 7, 2024, 11:26pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/15 "2024-10-07T23:26:52Z")

</div>

Like most volunteer organizations, they can’t handle the full load. Others provide needed help. People can register to [upload almost any kind of file](https://help.archive.org/help/uploading-a-basic-guide/).

I also donate to help keep them going since I use it so much.

---

<div class="post-metadata">

### Author: ![blondebear](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/blondebear/32/1022_2.png) [@blondebear](https://boards.straightdope.com/u/blondebear)
#### Post date: [October 11, 2024, 2:36am UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/16 "2024-10-11T02:36:15Z")

</div>

They had to take the site offline today due to a DDOS attack. ☹

---

<div class="post-metadata">

### Author: ![Mangetout](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mangetout/32/19_2.png) [@Mangetout](https://boards.straightdope.com/u/Mangetout)
#### Post date: [October 11, 2024, 7:55am UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/17 "2024-10-11T07:55:13Z")

</div>

They got hacked - the database of (31 million) user accounts was breached and published. Then they got DDOSed - possibly by a flood of people trying to hijack accounts using the stolen details, or maybe, I suppose, a flood of genuine users trying to log in and change their passwords (having heard of the attack). Or maybe just a different attack on a different day.

The performance of the site was never all that good or stable at the best of times, so I imagine it didn’t take a lot of extra traffic to tip it over.

---

<div class="post-metadata">

### Author: ![Pardel-Lux](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pardel-lux/32/20182_2.png) [@Pardel-Lux](https://boards.straightdope.com/u/Pardel-Lux)
#### Post date: [October 11, 2024, 8:07am UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/18 "2024-10-11T08:07:46Z")

</div>

That’s why we can’t have nice things ☹ Vandals come and break them just for the lulz.

---

<div class="post-metadata">

### Author: ![Smapti](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/smapti/32/17938_2.png) [@Smapti](https://boards.straightdope.com/u/Smapti)
#### Post date: [October 11, 2024, 8:48am UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/19 "2024-10-11T08:48:34Z")

</div>

> [@Pardel-Lux](#):
>
> Vandals come and break them just for the lulz.

According to Youtuber SomeOrdinaryGamers (who is an absolute expert when it comes to cybersecurity matters), the responsibility for the attack is being claimed by a Russian group who state they’re doing it in protest of America’s support for Israel.

How they expect to change anything about that by destroying the entire history of the internet is beyond me.

[![](https://img.youtube.com/vi/uMGcUZQmDmA/maxresdefault.jpg "Hackers Took Down The Internet Archive...") ](https://www.youtube.com/watch?v=uMGcUZQmDmA)

---

<div class="post-metadata">

### Author: ![Exapno\_Mapcase](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/exapno_mapcase/32/1051_2.png) [@Exapno\_Mapcase](https://boards.straightdope.com/u/Exapno_Mapcase)
#### Post date: [October 11, 2024, 6:15pm UTC](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480/20 "2024-10-11T18:15:40Z")

</div>

It’s apparently believed by many that the Archive is a product of the government rather than a totally separate separate non-profit. You know, “they” are saving everything that you do to use against you in the show trials.

[Next page](https://boards.straightdope.com/t/anyone-experienced-with-archive-org/1008480.md?page=2)
