# Another AI images question

**URL:** <https://boards.straightdope.com/t/another-ai-images-question/979951>\
**Category:** Factual Questions\
**Tags:** ai\
**Created:** [February 17, 2023, 4:59pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951 "2023-02-17T16:59:10Z")\
**Posts on this page:** 20\
**Page:** 2

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [February 18, 2023, 1:10am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/21 "2023-02-18T01:10:42Z")

</div>

> [@Darren\_Garrison](#):
>
> 2.) DE2 has a clear idea of the front of an elephant (there are no images of an elephant’s ass attached to a truck) but no idea of where best to attach it to the truck–sometimes it _is_ in the front, but also can be at the side or back.

And a lot of those images, humans would interpret as the elephant’s head being at the front, but only because that’s where the elephant’s head is: The truck part doesn’t have any particularly recognizable front or back.

And I’m not clear that it entirely does understand the “frontness” of an elephant, given that one of those has tusks coming out of both ends, and several are missing the most quintessential feature of an elephant’s front, the trunk.

Also, I can’t help but remember one of the experiments from the long NightCafe thread, where someone asked it for a painting of the prompt “Facing the Charging Elephant”, and the AI dutifully created a painting of an elephant plugged into a charging station.

---

<div class="post-metadata">

**Author:** ![Sam\_Stone](https://avatars.discourse-cdn.com/v4/letter/s/ecccb3/32.png) [@Sam\_Stone](https://boards.straightdope.com/u/Sam_Stone)\
**Post date:** [February 18, 2023, 1:26am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/22 "2023-02-18T01:26:00Z")

</div>

> [@Chronos](#):
>
> Also, I can’t help but remember one of the experiments from the long NightCafe thread, where someone asked it for a painting of the prompt “Facing the Charging Elephant”, and the AI dutifully created a painting of an elephant plugged into a charging station.

That would be a failure of ‘word in context’ determination. Word in Context is an emergent capability that just appeared at a certain point in training.

---

<div class="post-metadata">

**Author:** ![Riemann](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/riemann/32/3133_2.png) [@Riemann](https://boards.straightdope.com/u/Riemann)\
**Post date:** [February 18, 2023, 1:32am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/23 "2023-02-18T01:32:38Z")

</div>

> [@Dr.Strangelove](#):
>
> A significant difference here is the idea of _purpose_. It’s almost impossible to talk about any biological system without bringing in purpose: lungs evolved for breathing, legs evolved for walking, etc. Those are post hoc explanations for a purposeless process. No one makes the same mistake for geology. Formations are what they are.

But it’s not a “mistake” to say that lungs have a purpose in a sense that geologic formations do not.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [February 18, 2023, 1:45am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/24 "2023-02-18T01:45:46Z")

</div>

> [@Sam\_Stone](#):
>
> That would be a failure of ‘word in context’ determination.

Who says it’s a failure?

---

<div class="post-metadata">

**Author:** ![Riemann](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/riemann/32/3133_2.png) [@Riemann](https://boards.straightdope.com/u/Riemann)\
**Post date:** [February 18, 2023, 1:54am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/25 "2023-02-18T01:54:59Z")

</div>

It’s a better “charging elephant” joke than ChatGPT is giving me.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 18, 2023, 1:56am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/26 "2023-02-18T01:56:27Z")

</div>

> [@Riemann](#):
>
> What does _post hoc_ have to do with anything? We understand the general principles of evolution. Of course any application of those general principles to explain an example of evolution will be _post hoc_, just as any explanation of a geologic formation is _post hoc_.

But the majority of the current explanations for why a trait evolved boil down to “because it made the organism more fit” which isn’t much of an explanation at all.

A more direct analogy to explaining how an AI recognizes specific things is if you can explain how a trait evolved by saying that a switch from adenine to thymine in the 46th codon in the gene for the 3rd step in a five step process resulted in an enzyme that was 17 percent more efficient. We have _some_ explanations like that, but they aren’t the norm and they aren’t cheap or easy answers to get.

Explaining how the trained set for an AI model “understands” a specific concept would have a similar level of fine detail, difficulty to tease out, and gibberish-soundingness laymen. The real answer for how an AI recognizes the front of the truck would be something like “because values b724 and b725 in node 17 of layer 8 are set as ‘1’”.

---

<div class="post-metadata">

**Author:** ![Riemann](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/riemann/32/3133_2.png) [@Riemann](https://boards.straightdope.com/u/Riemann)\
**Post date:** [February 18, 2023, 1:58am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/27 "2023-02-18T01:58:55Z")

</div>

> [@Darren\_Garrison](#):
>
> A more direct analogy to explaining how an AI recognizes specific things is if you can explain how a trait evolved by saying that a switch from adenine to thymine in the 46th codon in the gene for the 3rd step in a five step process resulted in an enzyme that was 17 percent more efficient. We have _some_ explanations like that, but they aren’t the norm and they aren’t cheap or easy answers to get.
> 
> Explaining how the trained set for an AI model “understands” a specific concept would have a similar level of fine detail, difficulty to tease out, and gibberish-soundingness laymen. The real answer for how an AI recognizes the front of the truck would be something like “because values b724 and b725 in node 17 of layer 8 are set as ‘1’”.

But who cares about that, for either evolution or for AI? Aren’t the _general underlying principles_ by which it is operating the important explanation?

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 18, 2023, 2:02am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/28 "2023-02-18T02:02:06Z")

</div>

> [@Riemann](#):
>
> But who cares about that, for either evolution or for AI? Aren’t the _general underlying principles_ by which it is operating the important explanation?

No? Not when the question is “how does Midjourney recognize the front of an elephant”, or “why is a St. Benard so big”. General principles are not answers to specific questions.

---

<div class="post-metadata">

**Author:** ![Danger\_Man](https://avatars.discourse-cdn.com/v4/letter/d/6de8d8/32.png) [@Danger\_Man](https://boards.straightdope.com/u/Danger_Man)\
**Post date:** [February 18, 2023, 6:07am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/29 "2023-02-18T06:07:08Z")

</div>

Looking at the Midjourney pictures, it also has both the truck and the elephant rotated in the same way in 3d space. It does look like two 3d models merged together.

The examples using other AI services were not as good as the ones from Midjourney. I’m not sure why, but it seems like Midjourney has its own style.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 18, 2023, 6:20am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/30 "2023-02-18T06:20:02Z")

</div>

Which is weird, because Midjourney is Stable Diffusion with some custom tweaks.

---

<div class="post-metadata">

**Author:** ![Jophiel](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/jophiel/32/66_2.png) [@Jophiel](https://boards.straightdope.com/u/Jophiel)\
**Post date:** [February 18, 2023, 6:23am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/31 "2023-02-18T06:23:21Z")

</div>

I believe Midjourney has a _much_ larger model than what you get from Stable Diffusion. At least, locally run Stable Diffusion where the models are maybe 2-3GB. So Midjourney probably has a lot more data to work with and extrapolate what an elephant-shaped truck might look like.

For that matter, I suppose you could custom train a SD model on elephants and garbage trucks if you really wanted.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 18, 2023, 7:04am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/32 "2023-02-18T07:04:22Z")

</div>

Okay, apparently the latest Midjourney is not SD-based.

> **[r/StableDiffusion - If Midjourney runs Stable Diffusion, why is its output...](https://www.reddit.com/r/StableDiffusion/comments/10liqip/if_midjourney_runs_stable_diffusion_why_is_its/)**
>
> 204 votes and 167 comments so far on Reddit

---

<div class="post-metadata">

**Author:** ![Danger\_Man](https://avatars.discourse-cdn.com/v4/letter/d/6de8d8/32.png) [@Danger\_Man](https://boards.straightdope.com/u/Danger_Man)\
**Post date:** [February 18, 2023, 8:11am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/33 "2023-02-18T08:11:40Z")

</div>

From that link:

“Yeah, all of my Midjourney results seem to be a pastiche high quality 3D render of the prompt, instead of mimicking the style asked for.”

---

<div class="post-metadata">

**Author:** ![MrDibble](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mrdibble/32/114_2.png) [@MrDibble](https://boards.straightdope.com/u/MrDibble)\
**Post date:** [February 18, 2023, 11:54am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/34 "2023-02-18T11:54:04Z")

</div>

Given that learning systems ultimately involve human feedback, isn’t the answer “because humans keep picking the versions that have the right front bits”?

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 18, 2023, 5:21pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/35 "2023-02-18T17:21:46Z")

</div>

> [@MrDibble](#):
>
> Given that learning systems ultimately involve human feedback

This is not true.

As to the OP, the best answer you’re likely to get is that the images in the training data tend to show more of the front of both trucks and elephants since those are the interesting parts.

---

<div class="post-metadata">

**Author:** ![jjakucyk](https://avatars.discourse-cdn.com/v4/letter/j/3d9bf3/32.png) [@jjakucyk](https://boards.straightdope.com/u/jjakucyk)\
**Post date:** [February 18, 2023, 6:03pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/36 "2023-02-18T18:03:12Z")

</div>

This video is five years old now but still entirely relevant to why we don’t really know how AI works.

[![](https://img.youtube.com/vi/R9OHn5ZF4Uo/maxresdefault.jpg "How AIs, like ChatGPT, Learn") ](https://www.youtube.com/watch?v=R9OHn5ZF4Uo)

---

<div class="post-metadata">

**Author:** ![MrDibble](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mrdibble/32/114_2.png) [@MrDibble](https://boards.straightdope.com/u/MrDibble)\
**Post date:** [February 18, 2023, 9:45pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/37 "2023-02-18T21:45:25Z")

</div>

> [@Snarky\_Kong](#):
>
> This is not true.

OpenAI’s [own discussions of their research](https://openai.com/blog/tags/research/) seems to mention using human feedback quite a bit.

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 18, 2023, 9:48pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/38 "2023-02-18T21:48:34Z")

</div>

Having the ability to incorporate human feedback is not the same as human feedback being necessary. A large part of recent success of deep learning is that there are techniques (masked language or image modeling are most common) that allow algorithms to learn relevant features without annotations or feedback.

---

<div class="post-metadata">

**Author:** ![MrDibble](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mrdibble/32/114_2.png) [@MrDibble](https://boards.straightdope.com/u/MrDibble)\
**Post date:** [February 18, 2023, 9:50pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/39 "2023-02-18T21:50:22Z")

</div>

> [@Snarky\_Kong](#):
>
> Having the ability to incorporate human feedback is not the same as human feedback being necessary.

I didn’t say it was _necessary_. I just said it did involve it. Because human feedback is all over the existing datasets these systems are trained on.

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 18, 2023, 9:55pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/40 "2023-02-18T21:55:51Z")

</div>

You said they “ultimately involve human feedback” (which suggests necessary to me…) and suggested human feedback is the reason why these systems can reason about relevant part of images. Both are straight up incorrect.

[Previous page](https://boards.straightdope.com/t/another-ai-images-question/979951.md?page=1)

[Next page](https://boards.straightdope.com/t/another-ai-images-question/979951.md?page=3)
