# Another AI images question

**URL:** <https://boards.straightdope.com/t/another-ai-images-question/979951>\
**Category:** Factual Questions\
**Tags:** ai\
**Created:** [February 17, 2023, 4:59pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951 "2023-02-17T16:59:10Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Danger\_Man](https://avatars.discourse-cdn.com/v4/letter/d/6de8d8/32.png) [@Danger\_Man](https://boards.straightdope.com/u/Danger_Man)\
**Post date:** [February 17, 2023, 4:59pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/1 "2023-02-17T16:59:10Z")

</div>

If I try to generate a picture of “a garbage truck shaped like an elephant”, how does it know to place the head of the elephant at the front of the vehicle?

Sometimes pictures are combined more like a collage, but many times they are combined in a seamingly logical way. I understand that two human-like figures can be combined because it somehow finds the eyes/heads etc.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 17, 2023, 5:11pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/2 "2023-02-17T17:11:54Z")

</div>

If it manages to create that image, it is because it taught itself about the “frontness” of both a truck and an elephant.

---

<div class="post-metadata">

**Author:** ![Riemann](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/riemann/32/3133_2.png) [@Riemann](https://boards.straightdope.com/u/Riemann)\
**Post date:** [February 17, 2023, 5:19pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/3 "2023-02-17T17:19:57Z")

</div>

> [@Danger\_Man](#):
>
> how does it know to place the head of the elephant at the front of the vehicle?

I’m not sure what you’re asking.

If you asked people to do this, wouldn’t you expect them all to place the front of the animal at the front of the vehicle, unless there were some compelling shape correspondence that dictated otherwise?

An A.I. does it for the same reason that people do it. It’s sensible to put the front end of the animal at the front end of the vehicle.

---

<div class="post-metadata">

**Author:** ![Danger\_Man](https://avatars.discourse-cdn.com/v4/letter/d/6de8d8/32.png) [@Danger\_Man](https://boards.straightdope.com/u/Danger_Man)\
**Post date:** [February 17, 2023, 6:11pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/4 "2023-02-17T18:11:49Z")

</div>

Maybe my question is how it knows it is the front

---

<div class="post-metadata">

**Author:** ![Riemann](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/riemann/32/3133_2.png) [@Riemann](https://boards.straightdope.com/u/Riemann)\
**Post date:** [February 17, 2023, 6:53pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/5 "2023-02-17T18:53:03Z")

</div>

We have AI that can _drive_ cars, and you’re surprised that an AI knows which end of a vehicle\* or an elephant\*\* is the front?

\*\* the end with the trunk  
\* the other end

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 17, 2023, 7:01pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/6 "2023-02-17T19:01:36Z")

</div>

> [@Danger\_Man](#):
>
> Maybe my question is how it knows it is the front

Nobody knows how or why the AIs reach the conclusions that they have.

---

<div class="post-metadata">

**Author:** ![Johnny\_Bravo](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/johnny_bravo/32/493_2.png) [@Johnny\_Bravo](https://boards.straightdope.com/u/Johnny_Bravo)\
**Post date:** [February 17, 2023, 7:07pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/7 "2023-02-17T19:07:43Z")

</div>

For giggles, I went ahead and entered the prompt into DALL-E 2 just as the OP wrote it.

[![](https://i.ibb.co/qFH39Qx/elephunts.png) ](https://i.ibb.co/qFH39Qx/elephunts.png)

I quite like the second one.

---

<div class="post-metadata">

**Author:** ![Riemann](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/riemann/32/3133_2.png) [@Riemann](https://boards.straightdope.com/u/Riemann)\
**Post date:** [February 17, 2023, 7:17pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/8 "2023-02-17T19:17:45Z")

</div>

> [@Darren\_Garrison](#):
>
> Nobody knows how or why the AIs reach the conclusions that they have.

Sure we do. Here’s the original 2014 paper by the Google researchers that developed Google’s image captioning AI.

> **[Show and Tell: A Neural Image Caption Generator](https://arxiv.org/abs/1411.4555)**
>
> Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent architecture...

> **[Deep learning](https://en.wikipedia.org/wiki/Deep_learning)**
>
> Deep learning is the subset of machine learning methods based on artificial neural networks with representation learning. The adjective "deep" refers to the use of multiple layers in the network. Methods used can be either supervised, semi-supervised or unsupervised.
> Deep-learning architectures such as deep neural networks, deep belief networks, recurrent neural networks, convolutional neural networks and transformers have been applied to fields including computer vision, speech recognition, nat...

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 17, 2023, 8:58pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/9 "2023-02-17T20:58:49Z")

</div>

> [@Riemann](#):
>
> Sure we do.

No, we do not. I don’t mean nobody knows how to write the software, obviously people know how to do that. 1.) Write an AI that learns from patterns. 2.) Feed it a large dataset. But notbody knows why it reaches any particular conclusion. And I don’t mean the mechanics of it–I obviously don’t mean nobody knows what a neural network is–I mean that _nobody_, not you, not me, not the programmers, knows what output that AI will provide until they test it. _Nobody_ knows the specific visual clues that generate the specific neural net weightings that make an image generating AI reach a specific understanding. That data file is a black box.

---

<div class="post-metadata">

**Author:** ![Riemann](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/riemann/32/3133_2.png) [@Riemann](https://boards.straightdope.com/u/Riemann)\
**Post date:** [February 17, 2023, 9:27pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/10 "2023-02-17T21:27:55Z")

</div>

Ok, but that means we know more about current AI behavior than we know about the behavior of biological intelligences.

I would prefer to reserve hyperbole like “we have no idea how or why the AI is doing what it’s doing” to the post-singularity scenario where AI gets better at developing AI than humans, so that humans are no longer required and each generation of superintelligent AI programs the next generation.

---

<div class="post-metadata">

**Author:** ![Dr.Strangelove](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dr.strangelove/32/6613_2.png) [@Dr.Strangelove](https://boards.straightdope.com/u/Dr.Strangelove)\
**Post date:** [February 17, 2023, 9:46pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/11 "2023-02-17T21:46:29Z")

</div>

> [@Darren\_Garrison](#):
>
> _Nobody_ knows the specific visual clues that generate the specific neural net weightings that make an image generating AI reach a specific understanding. That data file is a black box.

As a kind of “exception that proves the rule”, here is a study of GPT-2 about discovering a neuron that predicted whether the net would choose “a” vs. “an” as the next word:  
[https://clementneo.com/posts/2023/02/11/we-found-an-neuron](https://clementneo.com/posts/2023/02/11/we-found-an-neuron)

So, after a great deal of research, they found _one_ trivial example of how the net is choosing one word over another. And they still don’t know much about what goes into the decision; they can just identify how it is correlated with the output.

In comparison, figuring out things like how an image generation net knows what the front of a vehicle is, or that an elephant’s head should go on the front, etc. is totally hopeless. It’s not even clear that there _is_ an answer; the decision is likely emergent across the set of weights, and there’s nothing that could be explicitly said to correspond to those decisions.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 17, 2023, 10:03pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/12 "2023-02-17T22:03:10Z")

</div>

> [@Danger\_Man](#):
>
> If I try to generate a picture of “a garbage truck shaped like an elephant”, how does it know to place the head of the elephant at the front of the vehicle?

Okay, I tried a number of images in SD and DE2 using that exact prompt.

Stable Diffusion:

[![](https://i.imgur.com/hiVhXyE.jpeg) ](https://i.imgur.com/hiVhXyE.jpeg)

Dall-E 2:

[![](https://i.imgur.com/uy5202L.jpeg) ](https://i.imgur.com/uy5202L.jpeg)

So we see

1.) SD has _no_ ability to generate this type of image _at all_ (and outside these images I tried different combinations of words and arrangements with no better result)

2.) DE2 has a clear idea of the front of an elephant (there are no images of an elephant’s ass attached to a truck) but no idea of where best to attach it to the truck–sometimes it _is_ in the front, but also can be at the side or back.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 17, 2023, 10:13pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/13 "2023-02-17T22:13:27Z")

</div>

BTW, _this_ is the most elephanty garbage truck I could get with any prompt in SD:

[![](https://i.imgur.com/PxfcxqM.jpeg) ](https://i.imgur.com/PxfcxqM.jpeg)

And I quite like this image I got when I weighted the prompt too far into “elephant”.

[![](https://i.imgur.com/OGZtjFG.jpeg) ](https://i.imgur.com/OGZtjFG.jpeg)

---

<div class="post-metadata">

**Author:** ![Danger\_Man](https://avatars.discourse-cdn.com/v4/letter/d/6de8d8/32.png) [@Danger\_Man](https://boards.straightdope.com/u/Danger_Man)\
**Post date:** [February 17, 2023, 10:50pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/14 "2023-02-17T22:50:47Z")

</div>

Midjourney got me this

[![](https://dl.dropboxusercontent.com/s/yzvn5a6w3m9opn0/Bilde%2016.02.2023%2C%2018%2045%2007.png?dl=0) ](https://dl.dropboxusercontent.com/s/yzvn5a6w3m9opn0/Bilde%2016.02.2023%2C%2018%2045%2007.png?dl=0)

---

<div class="post-metadata">

**Author:** ![Danger\_Man](https://avatars.discourse-cdn.com/v4/letter/d/6de8d8/32.png) [@Danger\_Man](https://boards.straightdope.com/u/Danger_Man)\
**Post date:** [February 17, 2023, 10:54pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/15 "2023-02-17T22:54:04Z")

</div>

> [@Darren\_Garrison](#):
>
> > [@Danger\_Man](#):
> >
> > Maybe my question is how it knows it is the front
> 
> Nobody knows how or why the AIs reach the conclusions that they have.

I have trouble accepting this. Evolution is very complex, but we still are able to identify lots of reasons for why things turn out like they do.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 17, 2023, 11:41pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/16 "2023-02-17T23:41:19Z")

</div>

> [@Danger\_Man](#):
>
> Evolution is very complex, but we still are able to identify lots of reasons for why things turn out like they do.

Not really. We have sometimes post-hoc explainations for why it was useful for a trait to evolve, but there are countless traits that would be useful that haven’t evolved, and vast numbers of variations in what did evolve. Sharks, cuttlefish, copepods, and scallops are all solutions to the problem of “mobile organisms in a marine environment” but none of those specific forms could have been predicted 600 million years ago and none if them are very similar to each other.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 17, 2023, 11:44pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/17 "2023-02-17T23:44:12Z")

</div>

> [@Danger\_Man](#):
>
> I have trouble accepting this.

This is from 2017, but is equally valid today. Or much more valid, given the vast rise in complexity in the past five years.

> **[The Dark Secret at the Heart of AI](https://www.technologyreview.com/2017/04/11/5113/the-dark-secret-at-the-heart-of-ai/)**
>
> No one really knows how the most advanced algorithms do what they do. That could be a problem.

---

<div class="post-metadata">

**Author:** ![pulykamell](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pulykamell/32/3166_2.png) [@pulykamell](https://boards.straightdope.com/u/pulykamell)\
**Post date:** [February 18, 2023, 12:49am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/18 "2023-02-18T00:49:09Z")

</div>

My wife is in this field – specifically NLP (natural language processing). From what I’ve understood her talking about this, yes, it gets to the point where we, as humans, don’t really understand the connections learning models are making. I mean, there is some overall sense of it, of course, but when you get into the weeds, it’s hairy.

---

<div class="post-metadata">

**Author:** ![Riemann](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/riemann/32/3133_2.png) [@Riemann](https://boards.straightdope.com/u/Riemann)\
**Post date:** [February 18, 2023, 12:59am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/19 "2023-02-18T00:59:18Z")

</div>

> [@Darren\_Garrison](#):
>
> We have sometimes post-hoc explainations for why it was useful for a trait to evolve, but there are countless traits that would be useful that haven’t evolved, and vast numbers of variations in what did evolve.

What does _post hoc_ have to do with anything? We understand the general principles of evolution. Of course any application of those general principles to explain an example of evolution will be _post hoc_, just as any explanation of a geologic formation is _post hoc_.

The general principles here are not validated by predicting a future that will take millions of years to unfold. They are validated as predictions about future _data_ - that we will not discover precambrian rabbits.

---

<div class="post-metadata">

**Author:** ![Dr.Strangelove](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dr.strangelove/32/6613_2.png) [@Dr.Strangelove](https://boards.straightdope.com/u/Dr.Strangelove)\
**Post date:** [February 18, 2023, 1:06am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/20 "2023-02-18T01:06:31Z")

</div>

> [@Riemann](#):
>
> just as any explanation of a geologic formation is _post hoc_

A significant difference here is the idea of _purpose_. It’s almost impossible to talk about any biological system without bringing in purpose: lungs evolved for breathing, legs evolved for walking, etc. Those are post hoc explanations for a purposeless process. No one makes the same mistake for geology. Formations are what they are.

Trying to figure out what the neural net is doing is possibly an invitation to the idea of purpose, even when it does not exist.

[Next page](https://boards.straightdope.com/t/another-ai-images-question/979951.md?page=2)
