# I'm missing something about AI training

**URL:** https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781
**Category:** Cafe Society
**Tags:** ai
**Created:** [February 25, 2025, 1:22pm UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781 "2025-02-25T13:22:17Z")
**Posts on this page:** 20
**Page:** 9

<div class="post-metadata">

### Author: ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)
#### Post date: [February 28, 2025, 10:19pm UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/161 "2025-02-28T22:19:53Z")

</div>

> [@Babale](#):
>
> By the way is his name really “Jaws the Shark”? I thought it was Bruce, like the shark from Finding Nemo

Just going by what works. At the time just putting “Jaws” wasn’t enough and DE3 got confused, making a giant gremlin in the background. Modifying it to “Jaws the shark” worked. I just tried five new runs based on one of the old prompts: Jaws, Jaws the shark, Bruce the shark, Bruce, and shark. Just “Jaws” continued to be mostly insufficient to tell DE3 what I wanted. “Jaws the shark” works. “Bruce the shark” works, too. Just “Bruce” does not. Just “shark” makes very similar images to the other successful ones. And this time none of the results were particularly close to the Jaws poster.

[[Album] imgur.com ![](https://i.imgur.com/229r9Lc.jpeg?fb "imgur.com") ](https://imgur.com/a/OZn7lmC)

---

<div class="post-metadata">

### Author: ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)
#### Post date: [February 28, 2025, 10:44pm UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/162 "2025-02-28T22:44:16Z")

</div>

Now I explicitly ask for the Jaws poster

[[Album] imgur.com ![](https://i.imgur.com/BCaPAB7.jpeg?fb "imgur.com") ](https://imgur.com/a/CIK5tSW)

---

<div class="post-metadata">

### Author: ![griffin1977](https://avatars.discourse-cdn.com/v4/letter/g/977dab/32.png) [@griffin1977](https://boards.straightdope.com/u/griffin1977)
#### Post date: [February 28, 2025, 11:00pm UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/163 "2025-02-28T23:00:15Z")

</div>

> [@Babale](#):
>
> No more than they are “stored” in my brain. What is stored are rules and associations for all kinds of concepts, and how to generate images that match them

> [@Jophiel](#):
>
> Is this because I have MonaLisa.png stored in my brain for retrieval or because of magical brain fairies? Are those the only two options?

We don’t actually know exactly what is happening in the brain when we create some piece of art, but we can say for sure what is NOT happening. Your brain is absolutely not a deterministic automaton that predictably produces an output based on the inputs it receives. That “tabula rasa” theory has been completely debunked for decades.

But that is 100% definitely what is happening inside all computer software, your AI model included.

If we are going to assign human qualities to your AI model and the ability to create original work based on its inputs (which are other people’s work). Then what makes other software like my mixing software or the image processing kernel I just wrote unable to have those qualities?

---

<div class="post-metadata">

### Author: ![Jophiel](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/jophiel/32/66_2.png) [@Jophiel](https://boards.straightdope.com/u/Jophiel)
#### Post date: [February 28, 2025, 11:18pm UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/164 "2025-02-28T23:18:44Z")

</div>

> [@griffin1977](#):
>
> We don’t actually know exactly what is happening in the brain when we create some piece of art

That didn’t answer the question. Regardless of whether or not I create art, is MonaLisa.png stored in my brain or is it all brain fairies bringing it up when I want it?

---

<div class="post-metadata">

### Author: ![griffin1977](https://avatars.discourse-cdn.com/v4/letter/g/977dab/32.png) [@griffin1977](https://boards.straightdope.com/u/griffin1977)
#### Post date: [February 28, 2025, 11:32pm UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/165 "2025-02-28T23:32:23Z")

</div>

> [@Jophiel](#):
>
> MonaLisa.png stored in my brain or is it all brain fairies bringing it up when I want it?

We literally don’t know the answer to that question for the human brain

We absolutely do for any bit of computer software, it is mathematically provable. In order to produce a close approximation of the Mona Lisa then then a digital representation of the Mona Lisa must be encoded in the data that the computer program reads. We can even say, at a minimum, how much information must be stored for any given approximation

Saying your computer program produced that image without encoding an image of the Mona Lisa from the training image is as plausible as saying your machine is a perpetual motion machine. It’s fundamentally impossible

---

<div class="post-metadata">

### Author: ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)
#### Post date: [February 28, 2025, 11:48pm UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/166 "2025-02-28T23:48:43Z")

</div>

> [@griffin1977](#):
>
> Your brain is absolutely not a deterministic automaton that predictably produces an output based on the inputs it receives.

So brains are magic?

---

<div class="post-metadata">

### Author: ![griffin1977](https://avatars.discourse-cdn.com/v4/letter/g/977dab/32.png) [@griffin1977](https://boards.straightdope.com/u/griffin1977)
#### Post date: [February 28, 2025, 11:52pm UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/167 "2025-02-28T23:52:15Z")

</div>

> [@Darren\_Garrison](#):
>
> So brains are magic?

Probably not but they are absolutely not deterministic automata that predictably produces an output based only on the input the input they receive. That much we do know.

Computer programs (AI models included) are definitely exactly that and no more.

---

<div class="post-metadata">

### Author: ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)
#### Post date: [March 1, 2025, 12:00am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/168 "2025-03-01T00:00:29Z")

</div>

> [@griffin1977](#):
>
> Probably not but they are absolutely not deterministic automata that predictably produces an output based only on the input the input they receive.

Of course they are. The only reason we can’t predict how a person will react to any given input is because we have a vastly insufficient model of the very precise physical and chemical state of any given brain. The only non-deterministic factor is quantum randomness. Anything other than that is believing that human minds are magical things not subject to the laws of physics.

---

<div class="post-metadata">

### Author: ![griffin1977](https://avatars.discourse-cdn.com/v4/letter/g/977dab/32.png) [@griffin1977](https://boards.straightdope.com/u/griffin1977)
#### Post date: [March 1, 2025, 12:06am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/169 "2025-03-01T00:06:56Z")

</div>

> [@Darren\_Garrison](#):
>
> Of course they are. The only reason we can’t predict how a person will react to any given input is because we have a vastly insufficient model of the very precise physical and chemical state of any given brain.

Yeah if we can completely know the entire state of the brain AND everything that effects it, which is everything in the universe, then we can deterministically predict how it will react. Except no…

> [@Darren\_Garrison](#):
>
> The only non-deterministic factor is quantum randomness.

And fortunately none of the chemistry or physical processes of the brain involve any quantum physics 🙄

---

<div class="post-metadata">

### Author: ![Jophiel](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/jophiel/32/66_2.png) [@Jophiel](https://boards.straightdope.com/u/Jophiel)
#### Post date: [March 1, 2025, 12:13am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/170 "2025-03-01T00:13:18Z")

</div>

> [@griffin1977](#):
>
> We literally don’t know the answer to that question for the human brain

That’s not true. We don’t know everything about how the brain stores memory and experiences but we know a good bit and we do know that it doesn’t store MonaLisa.png. I guess that only leaves brain fairies. After all, those are the only two ways to recall information, right?

---

<div class="post-metadata">

### Author: ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)
#### Post date: [March 1, 2025, 12:30am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/171 "2025-03-01T00:30:51Z")

</div>

> [@griffin1977](#):
>
> And fortunately none of the chemistry or physical processes of the brain involve any quantum physics 🙄

So what makes humans special is an insufficiently effective error-correcting code?

---

<div class="post-metadata">

### Author: ![Jophiel](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/jophiel/32/66_2.png) [@Jophiel](https://boards.straightdope.com/u/Jophiel)
#### Post date: [March 1, 2025, 12:46am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/172 "2025-03-01T00:46:26Z")

</div>

I’d say a better example is less “Look at that Jaws poster” and more what it _can’t_ accomplish.

For example, take the painting by Pierre-Auguste Renoir, “Two Sisters (on the Terrace)”. For reference, here it is:  
[![](https://upload.wikimedia.org/wikipedia/commons/thumb/f/f6/Two_Sisters_%28On_the_Terrace%29.jpg/800px-Two_Sisters_%28On_the_Terrace%29.jpg) ](https://upload.wikimedia.org/wikipedia/commons/thumb/f/f6/Two_Sisters_%28On_the_Terrace%29.jpg/800px-Two_Sisters_%28On_the_Terrace%29.jpg)

This is a famous enough painting that it’s certainly in the models. No one trained Stable Diffusion or Midjourney or Dall-E and left this out when they were scraping up all the art in the world. It’s also not especially obscure; any intro art student or amateur interest in French Impressionism is going to be familiar with the work. So, knowing that the work is included in the model, it should be trivial to prompt is back out, right? Here’s Midjourney’s guess at what I mean when I ask it for Renoir’s Two Sisters (On the Terrace):

[![](https://i.imgur.com/slQsTzq.jpeg) ](https://i.imgur.com/slQsTzq.jpeg)

That’s not a “lossy” interpretation. That’s just wrong. That’s something that is saying “Ok, so I know who Renoir was… And I know Impressionism… and I know what two sisters would look like… in an era/style appropriate way… and it’s a terrace so there’s a railing… and I guess I heard that there’s flowers?..”

The dresses are wrong, the background is wrong, the positioning is wrong, even the strokes are way off… it’s something that understood the basics of what the painting would entail but was left to its own devices to take it from there. But why? After all, we know that Renoir-TwoSisters.png is in the training, right? So why couldn’t it just look at Renoir-TwoSister.png and come up with a much more accurate depiction? I wasn’t asking for something LIKE Two Sisters, I literally asked it to give me Two Sisters and it failed. And this is a public domain image so it’s not a rights protection issue either.

The obvious answer, of course, is magic AI fairies. The other obvious answer is that there is no Renoir-TwoSister.png in the model to extract but rather a collection of tokenized bits telling it what generally makes up a Renoir painting and what they look like thematically and perhaps some general information about the painting itself (red hat, flowers, etc) so it starts with noise and tries to work its way towards a French impressionist style painting of two sisters in era appropriate garb with flowers in a way that looks like a Renoir work, etc.

Things like Jaws or R2D2 or the Mona Lisa come out “cleaner” because they’re so popular that the appropriate tags for them have been reinforced a ton of times. Something like Two Sisters is popular enough to be in the set with some weak information. That shouldn’t matter if Renoir-TwoSisters.png was in the model as an image but, well, it ain’t because that’s not how generative AI models work. They work via fairy magic, of course.

Edit: I noticed that I went with a thinner aspect ratio on my prompt so I ran it again closer to the original ratio to give the AI the best chance at getting it right. It came out [considerably worse](https://imgur.com/a/MKoFu01) instead.

---

<div class="post-metadata">

### Author: ![Babale](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/babale/32/15666_2.png) [@Babale](https://boards.straightdope.com/u/Babale)
#### Post date: [March 1, 2025, 1:08am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/173 "2025-03-01T01:08:25Z")

</div>

> [@griffin1977](#):
>
> Your brain is absolutely not a deterministic automaton that predictably produces an output based on the inputs it receives. That “tabula rasa” theory has been completely debunked for decades.

What on Earth are you talking about? These two concepts are entirely disconnected. The fact that you are not a blank slate has nothing to do with whether or not you are deterministic. Everyone is deterministic, but since everyone’s mental models are weighted differently, the result is that different people think differently in different situations.

> [@griffin1977](#):
>
> Probably not but they are absolutely not deterministic automata that predictably produces an output based only on the input the input they receive. That much we do know.

How do we “know” that?

> [@griffin1977](#):
>
> And fortunately none of the chemistry or physical processes of the brain involve any quantum physics 🙄

_Everything_ involves quantum physics, but the fact that quantum randomness makes determinism fuzzy at a fine enough scale doesn’t really leave room for that randomness to be behind consciousness or free will. It all averages out over any meaningful scale anyhow. Pretending otherwise is just the God of the Gaps fallacy, and that gap is constantly shrinking as our understanding of quantum mechanics grows.

---

<div class="post-metadata">

### Author: ![Zyada](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/zyada/32/3028_2.png) [@Zyada](https://boards.straightdope.com/u/Zyada)
#### Post date: [March 1, 2025, 1:32am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/174 "2025-03-01T01:32:21Z")

</div>

Have you actually looked at the code?

---

<div class="post-metadata">

### Author: ![Babale](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/babale/32/15666_2.png) [@Babale](https://boards.straightdope.com/u/Babale)
#### Post date: [March 1, 2025, 1:47am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/175 "2025-03-01T01:47:33Z")

</div>

> [@Zyada](#):
>
> Have you actually looked at the code?.

Would you like to? Here, StableDiffusion is open source and you run it locally, so you can tell exactly what data is available to the model:

> **[GitHub - AUTOMATIC1111/stable-diffusion-webui: Stable Diffusion web UI](https://github.com/AUTOMATIC1111/stable-diffusion-webui)**
>
> Stable Diffusion web UI

---

<div class="post-metadata">

### Author: ![Jophiel](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/jophiel/32/66_2.png) [@Jophiel](https://boards.straightdope.com/u/Jophiel)
#### Post date: [March 1, 2025, 1:55am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/176 "2025-03-01T01:55:47Z")

</div>

> [@Zyada](#):
>
> Have you actually looked at the code?

Sure. These are based off the LAION image set [which is searchable](https://haveibeentrained.com/search/TEXT?search_text=renoir%20two%20sisters).

(The linked site is, I believe, actually an abbreviated version without all the porn and whatnot. There’s more complete search sites but they tend to be jankier and this makes my point anyway)

---

<div class="post-metadata">

### Author: ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)
#### Post date: [March 1, 2025, 2:02am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/177 "2025-03-01T02:02:49Z")

</div>

> [@Jophiel](#):
>
> This is a famous enough painting that it’s certainly in the models. No one trained Stable Diffusion or Midjourney or Dall-E and left this out when they were scraping up all the art in the world.

Dall-E 3 appears to have no clue at all about that painting. I quite like Renoir’s Two Sisters On the Terrace with a shark and R2-D2, though.

[![](https://i.imgur.com/CCl2ECA.jpeg) ](https://i.imgur.com/CCl2ECA.jpeg)

---

<div class="post-metadata">

### Author: ![Babale](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/babale/32/15666_2.png) [@Babale](https://boards.straightdope.com/u/Babale)
#### Post date: [March 1, 2025, 2:04am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/178 "2025-03-01T02:04:19Z")

</div>

The sisters on the first terrace made a huge fucking mess. They deserve that shark attack.

---

<div class="post-metadata">

### Author: ![Zyada](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/zyada/32/3028_2.png) [@Zyada](https://boards.straightdope.com/u/Zyada)
#### Post date: [March 1, 2025, 2:12am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/179 "2025-03-01T02:12:23Z")

</div>

Your git link goes to a user interface program. Here is the Stable Diffusion project:

> **[GitHub - Stability-AI/stablediffusion: High-Resolution Image Synthesis with Latent...](https://github.com/Stability-AI/stablediffusion?tab=readme-ov-file)**
>
> High-Resolution Image Synthesis with Latent Diffusion Models

---

<div class="post-metadata">

### Author: ![Zyada](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/zyada/32/3028_2.png) [@Zyada](https://boards.straightdope.com/u/Zyada)
#### Post date: [March 1, 2025, 2:13am UTC](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781/180 "2025-03-01T02:13:02Z")

</div>

That is not code. That is a database

[Previous page](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781.md?page=8)

[Next page](https://boards.straightdope.com/t/im-missing-something-about-ai-training/1014781.md?page=10)
