# Source images for AI images

**URL:** <https://boards.straightdope.com/t/source-images-for-ai-images/971897>\
**Category:** Factual Questions\
**Created:** [September 20, 2022, 9:02am UTC](https://boards.straightdope.com/t/source-images-for-ai-images/971897 "2022-09-20T09:02:37Z")\
**Posts on this page:** 7\
**Page:** 2

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [September 23, 2022, 12:31am UTC](https://boards.straightdope.com/t/source-images-for-ai-images/971897/21 "2022-09-23T00:31:48Z")

</div>

It’s also possible that the program “remembers” more of some images than others, because they have more detail in them. I’m thinking, specifically, of broccoli: One of the published images from one of the AIs was a sequence of “an apple fighting a piece of broccoli”, in photorealistic style. I maintain that it is impossible to create a photorealistic image of a piece of broccoli, without either having a 3D model of the world with a level of detail impossible to achieve with any plausible number of still photographs, or copying exactly an already-existing piece of broccoli. Given that the AI was able to produce photorealistic broccoli, I believe that there must, in fact, have been a picture of that exact piece of broccoli in that exact pose somewhere in its training data, and for some reason (possibly the high amount of fine detail in the picture) it decided it was worth remembering that picture exactly (at the expense of remembering less detail in a bunch of other pictures, to maintain its average).

---

<div class="post-metadata">

**Author:** ![Dr.Strangelove](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dr.strangelove/32/6613_2.png) [@Dr.Strangelove](https://boards.straightdope.com/u/Dr.Strangelove)\
**Post date:** [September 23, 2022, 12:55am UTC](https://boards.straightdope.com/t/source-images-for-ai-images/971897/22 "2022-09-23T00:55:59Z")

</div>

Doubtful. Broccoli has a highly repetitive and fairly simple structure. There’s no reason to think it can’t generate a new picture based on the things it’s learned–not just about broccoli, but about similar-looking objects like trees.

It probably doesn’t have a 3D model as such. Again, this is one of those things where it probably has picked up some pattern, but it is essentially alien to us, and doesn’t map exactly to how we think of a rigid model rotating in space.

It knows about depth occlusion, at the very least (nearby objects occlude far off ones), and that far-off objects may appear hazy compared to nearby ones.

The AI can be thought of as an extremely advanced image compressor. Early compression tried to replicate images exactly, and did not perform very well. Later, JPEG and MPEG adopted lossy compression, trying to use the characteristics of human vision to throw away differences you wouldn’t notice. But this is many steps beyond that: you’d never notice if every tiny bud on a piece of broccoli moved around, only that they’re colored and arranged sensibly. That allows a far greater compression rate than you’d otherwise get. The compressor needs to understand something about how the buds are colored, sized, and positioned relative to one another, but with enough “rules” it can generate something indistinguishable from being real. Some of these rules can be shared with other images; for instance, a random distribution that keeps a minimum separation, along the lines of Poisson disk sampling. “Random points that never get too close” is a pattern that will show up in many places.

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [September 23, 2022, 3:05am UTC](https://boards.straightdope.com/t/source-images-for-ai-images/971897/23 "2022-09-23T03:05:37Z")

</div>

I’m fairly confident that google image search uses AI to compare images. Likely what they do is rescale everything to a given size, calculate the embedding tensor, and then only return N images within some radius in embedding space. Or use some other heuristic to ensure diversity in search results.

---

<div class="post-metadata">

**Author:** ![Half\_Man\_Half\_Wit](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/half_man_half_wit/32/21766_2.png) [@Half\_Man\_Half\_Wit](https://boards.straightdope.com/u/Half_Man_Half_Wit)\
**Post date:** [September 23, 2022, 3:56am UTC](https://boards.straightdope.com/t/source-images-for-ai-images/971897/24 "2022-09-23T03:56:57Z")

</div>

You can image-search the data used for training many of these AIs here: [https://haveibeentrained.com/](https://haveibeentrained.com/)

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [September 23, 2022, 5:42am UTC](https://boards.straightdope.com/t/source-images-for-ai-images/971897/25 "2022-09-23T05:42:24Z")

</div>

> [@Chronos](#):
>
> One of the published images from one of the AIs was a sequence of “an apple fighting a piece of broccoli”

Look at “nuclear explosion broccoli” here (on the second page of examples). Generate your own. Each time, it is different broccoli.

[https://huggingface.co/spaces/kuprel/min-dalle](https://huggingface.co/spaces/kuprel/min-dalle)

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [September 23, 2022, 9:11pm UTC](https://boards.straightdope.com/t/source-images-for-ai-images/971897/26 "2022-09-23T21:11:50Z")

</div>

I’m not sure that is different broccoli… Several of those images look like the same slightly-mangled piece of broccoli, differing only in the ways they’re mangled. See especially images 2 and 16, for instance, which look almost identical.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 14, 2023, 4:17am UTC](https://boards.straightdope.com/t/source-images-for-ai-images/971897/27 "2023-02-14T04:17:17Z")

</div>

> [@Darren\_Garrison](#):
>
> Something did occur to me, though. It probably doesn’t heavily weight one copy of one specific image, but (as anyone who has used Google Images knows) some images are found on multiple websites, but often with some alterations, such as changing the resolution, cropping it differently, adding text, tweaking the contrast, etc. Google Images knows that they are similar enough that they are probably “the same”, but I don’t know how the image scrapers and AI training deal with it. There could be some specific press release image that has more influence because it is in the dataset multiple times and each one is treated as a different image.

Ha!

> If, while training an image synthesis model, the same image is present many times in the dataset, it can result in “overfitting,” which can result in generations of a recognizable interpretation of the original image. For example, the _[Mona Lisa](https://twitter.com/ai_curio/status/1564899484819197953)_ has been found to have this property in Stable Diffusion. That property allowed researchers to target known-duplicate images in the dataset while looking for memorization, which dramatically amplified their chances of finding a memorized match.

> **[Paper: Stable Diffusion “memorizes” some images, sparking privacy concerns](https://arstechnica.com/information-technology/2023/02/researchers-extract-training-images-from-stable-diffusion-but-its-difficult/)**
>
> But out of 300,000 high-probability images tested, researchers found a 0.03% memorization rate.

[Previous page](https://boards.straightdope.com/t/source-images-for-ai-images/971897.md?page=1)
