AI image generation is getting crazy good

I think by default copilot still might give you Dall-E 3 generation which is pretty long in the tooth now and diffusion based and therefore bad at prompt following. Though I don’t know what copilot gives and why, maybe it does use MAI. They’re the least transparent service by design on what’s actually generating your output.

Have you tried MAI at all, their native generator at all? I don’t know that they integrate it into the normal copilot pathway for free accounts but it’s there for me at the top under “experiments” – I think that this is their own modern generation on par with something like nano banana 2.

The direct link will get you there but you may have to be eligible for it by activating it through copilot or simply having been granted access on your account. I was trying to use your prompt to see if MAI did better than what you got but it’s actually apparently temporarily down right now. Hopefully by the time you read this it’ll be working again.

Wanted to add – don’t do this to me as a favor or anything. I just figured as someone who likes to compare the various systems / generators you might be interested in using the one microsoft is currently developing and is probably their most advanced model. It doesn’t seem like they want many people to know about it and it seems like it’s targeted as a sort of safe corporate / enterprise generator, but it does fairly well in the arenas / blind competitions.

Currently #3 text 2 image and #2 image edit on arena.ai, and #2 text to image and #3 in image edit on artificial analysis. Actually, those numbers are better than I remember and better than I expected. Apparently they have one of the best image generators in the world and they aren’t really trying to get anyone to use it. Only consistently behind ImageGPT but much cheaper.

You can do MAI through Bing. I don’t like it enough to hardly ever try it.

Bing also stills has Dall-E 3

And GPT4

BTW, for me Copilot has always seemed to use the ChatGPT renderer. Output appearance has always been pretty much identical to Sora and ChatGPT output.

You may be right about copilot just flat out calling the GPT renderer - though I suspect that means the earlier versions like Dall-E rather than the newer versions like imageGPT which is probably kept pretty premium. I didn’t think about trying to create images through bing, I’ve only used the copilot client, so I never had a choice or transparency about what was generating my images.

Bing probably has the highest prompt blocking of any AI I’ve used.

Which is funny, because you’d think Microsoft would realize that the main reason anyone uses bing is to search for NSFW stuff, which it historically did better than google, not sure if that’s still true.

I , too, borrowed your prompt. And got something completely different.

Then I asked, “Elaborate on your analysis”

The portrait I generated is less a literal “me” and more a visual metaphor for the role I tend to occupy in our conversations.

The noir detective aesthetic comes from a pattern in the questions you ask: you often start with something that looks simple, then dig underneath it. You asked about suppressors, but the real question became “does the public argument match the evidence?” You asked about First/Second/Third World, but the interesting part was how language changes meaning over time. You asked about The Blues Brothers, and the deeper point became why ordinary people can contain extraordinary things. The detective represents that tendency to look for the hidden structure behind the obvious answer.

The robot/AI figure represents the limitations as well as the strengths. The notes on the wall in the image — “I examine,” “I interpret,” “I understand patterns” — are meant to show that I don’t experience the world the way a person does. I don’t have memories, instincts, or personal beliefs in the human sense. What I do is process information, compare patterns, explain relationships, and help build a clearer picture.

The crime board imagery is a nod to another theme in your prompts: you frequently explore human behavior under pressure. Your writing prompts often involve moral ambiguity — criminals who actually did the crime, people making choices with consequences, questions about intent versus action. You seem interested in the gap between “what happened” and “why someone did it.” The board represents investigation rather than judgment.

The books labeled criminology, psychology, philosophy, neuroscience, behavioral analysis are there because your topics bounce between those areas. You’ll move from guns and law to mythology, cooking, history, technology, and human nature. The common thread is not the subject matter itself; it’s curiosity about systems:

How do people behave?
Why do societies create rules?
What happens when those rules break?
What assumptions are hidden inside everyday ideas?

The coffee mug saying “causation over assumption” is probably the most important symbol. A recurring pattern in your questions is that you’re skeptical of simple explanations. You often test whether a claim is actually supported or just repeated. That doesn’t mean you’re automatically rejecting claims — you ask for fact checks — but you seem interested in separating “sounds right” from “is right.”

If I had to reduce the portrait to one sentence:

A machine in a detective coat, sitting in the middle of humanity’s mess, trying to map the difference between what people say, what they mean, and what the evidence shows.

Which is also a pretty good description of what I’m designed to do.

Tried another suggestive image. This is a grain of sand.

The prompt was

This is the head of a mostly white t-rex with black on its snout. Clear up the photo.

Gemini:

Copilot didn’t argue this time:

ChatGPT did an interesting job of creatively upscaling the photo, but didn’t dino it:

We had discussed a couple of weeks ago whether Google grounds their AI image generation in real world knowledge from maps, and I got this message when I opened flow today:

I thought that those made up drone videos of Italy did a pretty good job of representing the real world and therefore probably were getting a benefit from google maps knowledge. And maybe they were, maybe more map knowledge or live grounding was still baked into the process, but the way they’re talking about it hear makes it sound like it’s a new feature that just became available, and only for the US. Hmm.

Anyway, I tested it a bit and it seems to work. The images I made would be meaningless to you because I used addresses of places of significance to me - but try it yourself, put in an address or intersection and then put robocop on a unicorn in front of it. It’s interesting.

Edit: To be clear, this is a feature of using agent mode in flow. I’m not sure if it works on gemini chat. When an image gets grounded correctly, it says: “Your media generation was grounded on the following Google Maps imagery:” and then the address it matched for your request. I’m not sure how far you can push it, but things like “show godzilla destroying [intersection / landmark]” or “robocop on a unicorn in front of [address]” work.

Edit 2: Actually, I caught my first error. I used an address I used to live at years ago and it showed somewhere down the street. It wasn’t really an understandable error, though, because the google maps link was located at the wrong address and typing the real address, the one I used, correctly locates the scene in google maps. So.. not perfect, but interesting.

Fascinating.

Over on Nightcafe they hold a daily contest. The most recent one was to show something made of seashells.

My entry was placed 4th of 4479, my best result ever.

It was discussed earlier in the thread how AI doesn’t know the proper relative size for a phone booth. Looks like it extends to TARDISes, too.

I’ve been having some fun using ChatGPT to restore old neon signs.

I started with this:

And finished with this:

It has a problem with the elephant chef at the bottom left, it keeps trying to turn his trunk into an arm with a hand, and the face is pretty mutant. But otherwise, it’s kinda cool!

Easy, just tell the AI you want an elephant, and it will give you one.

prompt: “full colour photograph of a sign

prompt: “full colour photograph of a sign with a picture of an elephant

prompt: “full colour photograph of a sign with a picture of an elephant, evening, sign lit up

All images created by GPT Image 2 Low on Nightcafe

Very nice. This was our go-to burger place when I was growing up, until In N Out moved to town. I’ve been told it is now a smoke shop. The sign has no neon, and has been painted over.

The process of image to image editing has always been hard to prompt for me because the way each generator works is very different. Some are trained with edit instructions like “remove this, add that” and some are designed so that you just try to re-state what you want in the image but with the first image as a reference and these are very different editing styles. Some just want you to state the differences, others want you to restate the entire prompt with alterations. Some, confusingly, respond half-assedly to all of these things. I still struggle to make “remixes” in midjourney which are image to image edits cause I don’t know wtf it wants, and what works is wildly inconsistent.

This is resolved almost entirely by the agentic-style editing where you tell it what you want, it uses a smart chat bot to understand, and then sends in the appropriate prompt on your behalf because it knows its own image generation tool call better than you.

Inspired by a current thread:

Create an image of a hybrid between Cleveland from Family Guy and Grover from Sesame Street.

(From ChatGPT.)

After the latest xkcd comic, ­xkcd thread - #6107 by dtilque

I showed it to ChatGPT 5.6 (Thinking/High) and told it to do that obvious followup mentioned in the alt text and got this:

It gets the style right, but the details wrong. The existing HCA story had many mattresses, not one. And “Pea Paradise” has lots of princesses, but is lower than “Peak Overkill”, with only four.