# Whatever happened to Mega Data Compression

**URL:** <https://boards.straightdope.com/t/whatever-happened-to-mega-data-compression/235918>\
**Category:** Factual Questions\
**Created:** [March 23, 2004, 3:54pm UTC](https://boards.straightdope.com/t/whatever-happened-to-mega-data-compression/235918 "2004-03-23T15:54:37Z")\
**Posts on this page:** 4\
**Page:** 3

<div class="post-metadata">

**Author:** ![ftg](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/ftg/32/2801_2.png) [@ftg](https://boards.straightdope.com/u/ftg)\
**Post date:** [March 25, 2004, 4:44am UTC](https://boards.straightdope.com/t/whatever-happened-to-mega-data-compression/235918/41 "2004-03-25T04:44:18Z")

</div>

I had posted a lengthy response in this thread early on, mostly in reply to **Bytegeist** ’s “interesting” first post. But it went the way of the hamsters.

I dreaded typing it all back in, etc. But it appears the signal to noise ratio has finally recovered (ignoring the ternary hijack).

At this point I want to add just a couple of things:

[Here](http://www.faqs.org/faqs/compression-faq/)’s the comp.compression FAQ. It explains a lot of the basics, makes fun of earlier “we can compress everything!” companies, etc. Some of the posters to this thread _really_ need to read it.

While _some_ compression methods tack on headers:

1. Not all do, e.g., many forms of RLE. You can do LZW without headers if you choose.
2. Even without headers, some files will get larger.

Note that you can’t iterate a lossy compression system. If you feed a jpeg into a jpeg compressor (dressing it up as an image first), the losses incurred will make recovery of the (single compressed) image impossible. _All_ bits in the compressed image have become vital.

For things like jpeg compression, there is a quality setting (mistakenly taken as a “percentage” value by too many people). Set it to 75 and you get a nice approximation of the original. Set it to 20 and it’s probably going to look crappy. Set it to 5, you get a tiny file that is not going to look like the original. (Setting it to 100 means you don’t have a clue.) For audio files, you can tweak the sample rates and ranges. Etc. But it’s one pass.

These kind of companies have been popping up for a _long_ time. A friend of mine in graduate school was brought in to consult for such a company. He knew it was a joke inside of 5 minutes. They were very unhappy with his analysis and adamant that they really could do it. That was the early 70s folks.

Why does the press give free advertising for these clowns? There is no journalism in American Big Media anymore. No reporter actually questions anything they are told.

---

<div class="post-metadata">

**Author:** ![Mangetout](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mangetout/32/19_2.png) [@Mangetout](https://boards.straightdope.com/u/Mangetout)\
**Post date:** [March 25, 2004, 11:28am UTC](https://boards.straightdope.com/t/whatever-happened-to-mega-data-compression/235918/42 "2004-03-25T11:28:37Z")

</div>

> [@r\_k](#):
>
> Surely more of a problem is that with only one control code, you can’t tell where the end of the compressed section is?  
> eg:  
> aaabbbrrrracadabra would become Y3a3b4rNacadabra
> 
> Incidentally, how are you distinguishing between the control character Y and the Y in the data? (I know this is only a simplification, but what characters would a real compression tool use? Presumably it would have to be something that _never_ appears in the files you’re compressing?)

A real compression tool wouldn’t use characters, probably, and might treat the data as a contiuous binary stream, or something else; the control ‘character’ in my example, had to be the first one in the file (i.e. it is part of the header, not the data), which is why you need Y or N. Also, I assumed (for the purposes of the example) that the entire data would be either compressed or uncompressed; if you want mixed segments in a single file, you’d have to do something like:

-Use segments of a fixed size (which means that the break points probably won’t be where you’d find them most convenient in terms of compressibility/or non)) - in this case, you won’t achieve the best possible compression simply because your compressible segments will contain uncompressible sequences and vice versa.

-Use a control ‘character’ that does not otherwise occur in your compressed data (this won’t happen naturally, so you have to make it happen by encoding all of the compressed and uncompressed data in such a way as that the control character doesn’t occur accidentally) - in this case, you’re essentially ‘wasting’ the potential of using the control sequence at every step.

-Give the file a header that explicitly maps the size, position and compession methodology of each segment - probably the best solution all round, but the map takes up space even if the entire content of the file is an uncompressible sequence.

The advantage of the map method is that you can analyse your segments and use a different compression methodology for each one, as appropriate - I think this is what the zip format actually does.

---

<div class="post-metadata">

**Author:** ![ultrafilter](https://avatars.discourse-cdn.com/v4/letter/u/3d9bf3/32.png) [@ultrafilter](https://boards.straightdope.com/u/ultrafilter)\
**Post date:** [March 25, 2004, 4:22pm UTC](https://boards.straightdope.com/t/whatever-happened-to-mega-data-compression/235918/43 "2004-03-25T16:22:15Z")

</div>

> [@Mangetout](#):
>
> -Give the file a header that explicitly maps the size, position and compession methodology of each segment - probably the best solution all round, but the map takes up space even if the entire content of the file is an uncompressible sequence.

Yes, but if the entire file is uncompressed, the map is rather short.

---

<div class="post-metadata">

**Author:** ![Mangetout](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mangetout/32/19_2.png) [@Mangetout](https://boards.straightdope.com/u/Mangetout)\
**Post date:** [March 25, 2004, 11:02pm UTC](https://boards.straightdope.com/t/whatever-happened-to-mega-data-compression/235918/44 "2004-03-25T23:02:24Z")

</div>

> [@ultrafilter](#):
>
> Yes, but if the entire file is uncompressed, the map is rather short.

Indeed, but still longer than the original raw data.

[Previous page](https://boards.straightdope.com/t/whatever-happened-to-mega-data-compression/235918.md?page=2)
