# Another AI images question

**URL:** <https://boards.straightdope.com/t/another-ai-images-question/979951>\
**Category:** Factual Questions\
**Tags:** ai\
**Created:** [February 17, 2023, 4:59pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951 "2023-02-17T16:59:10Z")\
**Posts on this page:** 20\
**Page:** 3

<div class="post-metadata">

**Author:** ![MrDibble](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mrdibble/32/114_2.png) [@MrDibble](https://boards.straightdope.com/u/MrDibble)\
**Post date:** [February 18, 2023, 10:02pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/41 "2023-02-18T22:02:41Z")

</div>

> [@Snarky\_Kong](#):
>
> You said they “ultimately involve human feedback” (which suggests necessary to me…)

No, I made a statement about what is, not what ought to be. That you interpreted it that way is, however, understandable.

That you continue to harp on it, after I’ve corrected you, is not. I’m outright _telling_ you I was _not_ talking about “necessary”.

I’m talking about the pre-existing training datasets and the initial pre-training as discussed by OpenAI themselves, not the subsequent algorithms used.

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 18, 2023, 10:15pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/42 "2023-02-18T22:15:26Z")

</div>

> [@MrDibble](#):
>
> No, I made a statement about what is, not what ought to be. That you interpreted it that way is, however, understandable.

No, I was talking about what is. I don’t really have any opinion on what ought to be in generative modeling.

Leaving that aside, you’re factually incorrect about human feedback being the cause of this phenomenon. CLIP and Dall-E don’t use human feedback. They don’t even use labels.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [February 19, 2023, 12:11am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/43 "2023-02-19T00:11:46Z")

</div>

> [@Snarky\_Kong](#):
>
> They don’t even use labels.

…How could they not even use labels? They have to have some way of translating the human-language prompt into images.

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 19, 2023, 12:29am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/44 "2023-02-19T00:29:08Z")

</div>

CLIP stands for “contrastive language-image pretraining.” Basically they scraped a bunch of image/caption pairs from the internet and then trained two models, one each for the image and the caption, to create embeddings that are close to each other when a given caption is from that image and far if it’s from a different image (this is the contrastive part). This creates a shared embedding space for both visual and language information. So when you pass a prompt to Dall-E it’s converting that prompt into this embedding that makes sense as an image and the diffusion model uses it as a guide.

You might have the thought that the caption is just a label. Not really. You can have incredibly large amount of captions for an image. Pandas at a zoo might be “pandas at a zoo” or “people looking at animals in cages” or “a sunny day outside” or whatever. It’s also very noisy. You might scrape things and use text on a webpage instead of a direct caption. Does the text directly comment on the image? At the internet scale it doesn’t really matter since it will tend to correlate with the image.

My main point being is knowing how to create an image of “red ball on top of a green cube” does not require a human to look at the output of a model and give it feedback (this is very common now in NLP although AI feedback instead is becoming a thing: [https://www.anthropic.com/constitutional.pdf](https://www.anthropic.com/constitutional.pdf)) or even to label images of red balls and green cubes.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [February 19, 2023, 3:16am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/45 "2023-02-19T03:16:00Z")

</div>

> [@Snarky\_Kong](#):
>
> You might have the thought that the caption is just a label. Not really. You can have incredibly large amount of captions for an image.

That still just sounds like labels, to me. The word “label” doesn’t seem to imply that every image has one and only one. Indeed, it’s hard to imagine any meaningful way of associating text with an image that would only connect one text string to each image.

And yeah, with the current generation of AIs, all of the human intervention (in attaching the text strings to the images) happened before the AI was growprammed.

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 19, 2023, 4:02am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/46 "2023-02-19T04:02:10Z")

</div>

Labels would be like a description of what’s in the image, which pixels go with which thing, etc. For caption you’re going to have an image where the only text is something like the date when it was taken.

You can do pure image generation without labels for sure ([thispersondoesnotexist.com](http://thispersondoesnotexist.com) looks like it’s been taken down, but you may recall face generation), I’m not aware of it research done without paired examples.

Still, no human feedback on the model.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 19, 2023, 4:05am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/47 "2023-02-19T04:05:42Z")

</div>

> [@Chronos](#):
>
> growprammed

Is that a typo or a phrase coining?

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 19, 2023, 4:07am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/48 "2023-02-19T04:07:24Z")

</div>

> [@Snarky\_Kong](#):
>
> [thispersondoesnotexist.com](http://thispersondoesnotexist.com) looks like it’s been taken down, but you may recall face generation

Try this site.

> **[This X Does Not Exist](https://thisxdoesnotexist.com/)**
>
> Using generative adversarial networks (GAN), we can learn how to create realistic-looking fake versions of almost anything, as shown by this collection of sites that have sprung up in the past month.

---

<div class="post-metadata">

**Author:** ![Dr.Strangelove](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dr.strangelove/32/6613_2.png) [@Dr.Strangelove](https://boards.straightdope.com/u/Dr.Strangelove)\
**Post date:** [February 19, 2023, 5:05am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/49 "2023-02-19T05:05:32Z")

</div>

> [@Darren\_Garrison](#):
>
> Is that a typo or a phrase coining

Borrowed from _Schlock Mercenary_.

---

<div class="post-metadata">

**Author:** ![MrDibble](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mrdibble/32/114_2.png) [@MrDibble](https://boards.straightdope.com/u/MrDibble)\
**Post date:** [February 19, 2023, 8:11am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/50 "2023-02-19T08:11:55Z")

</div>

> [@Snarky\_Kong](#):
>
> CLIP and Dall-E don’t use human feedback. They don’t even use labels.

So [DALL-E](https://en.wikipedia.org/wiki/DALL-E) is _not_ based off GPT3?

> DALL-E was revealed by OpenAI in a blog post in January 2021, and **uses a version of GPT-3** modified to generate images.

Do [the GPT-3 authors know](https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf) they weren’t using labels? Because they certainly seem to think they were:

> Our training procedure consists of two stages. The first stage is learning a high-capacity language  
> model on a large corpus of text. **This is followed by a fine-tuning stage, where we adapt the model to a discriminative task with labeled data**

My emphases.

> [@Chronos](#):
>
> And yeah, with the current generation of AIs, all of the human intervention (in attaching the text strings to the images) happened before the AI was growprammed.

That’s my understanding - but that doesn’t mean _no_ human intervention at all.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [February 19, 2023, 1:00pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/51 "2023-02-19T13:00:19Z")

</div>

> [@Darren\_Garrison](#):
>
> Is that a typo or a phrase coining?

Not my phrase coining. Howard Tayler first used it [in 2009](https://www.schlockmercenary.com/2009-04-19).

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 19, 2023, 4:15pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/52 "2023-02-19T16:15:40Z")

</div>

Finetuning is not necessary for functionality, it improves performance.

Ok, so when you said image generation models “ultimately use human feedback” you didn’t mean a human was giving feedback to that model and when you said “humans keep picking the right front bits” you meant that a human taught GPT3 (a different model) how to speak better English?

---

<div class="post-metadata">

**Author:** ![MrDibble](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/mrdibble/32/114_2.png) [@MrDibble](https://boards.straightdope.com/u/MrDibble)\
**Post date:** [February 19, 2023, 4:50pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/53 "2023-02-19T16:50:58Z")

</div>

Again with the “necessary” - that nobody but yourself has introduced into this. 🙄

> [@Snarky\_Kong](#):
>
> Ok, so when you said image generation models “ultimately use human feedback” you didn’t mean a human was giving feedback to that model

No, which is why I said “ultimately”, not “constantly” or “immediately” or any other word that might be misunderstood that way.

> [@Snarky\_Kong](#):
>
> when you said “humans keep picking the right front bits” you meant that a human taught GPT3 (a different model) how to speak better English?

It’s more than “speak better English”.

You clearly have no interest in responding to what I _actually_ wrote, and have repeatedly elaborated on, but only the strawman version of it you initially jumped on.

When you then show no understanding that GPT is not, in fact “a different model”, but underlies DALL-E, I don’t see the point of responding to you any more.

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 19, 2023, 5:03pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/54 "2023-02-19T17:03:20Z")

</div>

> [@MrDibble](#):
>
> No, which is why I said “ultimately”, not “constantly” or “immediately” or any other word that might be misunderstood that way.

Constantly or immediately wouldn’t be necessary. MW gives a definition of eventually. Which is how I understood you. That feedback happens at some point. Which is false. You’re

The current version of Dall-E isn’t even the autoregressive model anymore. It’s diffusion based.

It’s not really more than speaking better English. Base GPT3 trained without human labeling at all, does next token prediction. The finetuning you’re referring to is to get it to perform better at tasks like question answering, or “speaking better English.”

The entire point of my responses have been that your assertion that compositionality is due to humans labeling images to show front is false. That’s directly responding to what you actually wrote

---

<div class="post-metadata">

**Author:** ![Snarky\_Kong](https://avatars.discourse-cdn.com/v4/letter/s/a183cd/32.png) [@Snarky\_Kong](https://boards.straightdope.com/u/Snarky_Kong)\
**Post date:** [February 19, 2023, 6:08pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/55 "2023-02-19T18:08:15Z")

</div>

I notice you cited the GPT3 paper and not the actual [Dall-E](https://arxiv.org/pdf/2102.12092.pdf) (or Dall-E 2) paper.

The actual Dall-E paper does not use humans to finetune.

Section of the paper relevant to this thread/discussion:

> [@OpenAI](#):
>
> We found that our model has the ability to generalize in  
> ways that we did not originally anticipate. When given the  
> caption “a tapir made of accordion…” (Figure 2a), the model  
> appears to draw a tapir with an accordion for a body, or an  
> accordion whose keyboard or bass are in the shape of a  
> tapir’s trunk or legs. This suggests that it has developed a  
> rudimentary ability to compose unusual concepts at high  
> levels of abstraction.

and

> [@OpenAI](#):
>
> We speculate that our zero-shot approach is  
> less likely to compare favorably on specialized distributions  
> such as CUB. We believe that fine-tuning is a promising  
> direction for improvement, and leave this investigation to  
> future work

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 19, 2023, 7:19pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/56 "2023-02-19T19:19:45Z")

</div>

> [@OpenAI](#):
>
> When given the  
> caption “a tapir made of accordion…” (Figure 2a), the model  
> appears to draw a tapir with an accordion for a body, or an  
> accordion whose keyboard or bass are in the shape of a  
> tapir’s trunk or legs.

FWIW, here is how Stable Diffusion handles that.

[![](https://i.imgur.com/dlkzulX.jpeg) ](https://i.imgur.com/dlkzulX.jpeg)

---

<div class="post-metadata">

**Author:** ![jjakucyk](https://avatars.discourse-cdn.com/v4/letter/j/3d9bf3/32.png) [@jjakucyk](https://boards.straightdope.com/u/jjakucyk)\
**Post date:** [February 20, 2023, 12:36am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/57 "2023-02-20T00:36:38Z")

</div>

The two at the bottom of the righthand column made my day, those are so hilariously bad.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [February 20, 2023, 1:30am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/58 "2023-02-20T01:30:48Z")

</div>

Hilarious? I’ll have you know that wild populations are being decimated by Tapir Accordianbutt Disease.

---

<div class="post-metadata">

**Author:** ![jjakucyk](https://avatars.discourse-cdn.com/v4/letter/j/3d9bf3/32.png) [@jjakucyk](https://boards.straightdope.com/u/jjakucyk)\
**Post date:** [February 20, 2023, 2:55am UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/59 "2023-02-20T02:55:36Z")

</div>

I’m afraid to ask what sort of sounds that produces.

---

<div class="post-metadata">

**Author:** ![Johnny\_Bravo](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/johnny_bravo/32/493_2.png) [@Johnny\_Bravo](https://boards.straightdope.com/u/Johnny_Bravo)\
**Post date:** [February 20, 2023, 12:44pm UTC](https://boards.straightdope.com/t/another-ai-images-question/979951/60 "2023-02-20T12:44:31Z")

</div>

A pretty strange one, but it tapers off quickly enough.

[Previous page](https://boards.straightdope.com/t/another-ai-images-question/979951.md?page=2)

[Next page](https://boards.straightdope.com/t/another-ai-images-question/979951.md?page=4)
