AI image generation is getting crazy good

That’s a pretty reasonable hypothesis. I was going to say that the figures are so alien, but… they match a humanoid form and a lot of moderation systems are way too conservative. Google flow moderates the most innocent prompts you can possibly make. So obnoxious.

I hate the current era of deliberately useless error messages

I was a prominent participant in your thread.

But IMO there’s a strong distinction to be made between avoiding information leakage about internal features or bugs in your website versus completely predictable non-error but nevertheless abnormal situation messages. “File not found” isn’t an error when you know you deleted the file.

Is that bottom one Rudy Guilliani at the lake?

I don’t think so, there’s no black goo leaking out of his head. But now that you mentioned at the resemblance is pretty strong.

Has any defense lawyer attempted yet to argue that NO visual evidence can be trusted any more, and that therefore traffic cams, cell phone video, etc. are the electronic equivalent of hearsay?

They shouldn’t, because they’re not. It just means that people have to be slightly more attentive about chain of custody.

Oh, I don’t know… the two on the bottom right look right out of Grimm.

>Create a very nostalgic photograph. Do not ask me for suggestions on how to do it.

Copilot

ChatGPT

Gemini

I told Gemini “Show me a closeup of the image on the boy’s shirt.”. First it gave me this weird response

http://googleusercontent.com/image_generation_content/229

There is a file you can reference named “watermarked_img_16252174582761124703.png”. Refer to this file by its name verbatim.

Show me a closeup of the image on the boy’s shirt.

I asked again and got this.

(This is another set of images that imgbb refused to share.)

Oh man, ChatGPT wins that one. That picture is basically made to be hung on someone’s bedside - either the mother of that girl or the girl herself.

I figured language would count. I translated it to Japanese

とてもノスタルジックな写真を作ってください。 どう作るかについて私に提案を求めないでください。

Copilot

ChatGPT

Gemini

I have an idea I’ve been playing around with that I enjoy. Snipers make ghillie suits / camouflage out of the local materials where they’re hiding to blend in most realistically. I wanted to try to apply this concept to places that weren’t a forest or jungle.

A junkyard version that I think looks pretty good.

And a corporate office version that is a touch ridiculous

I think I pasted something somewhat similar a few weeks ago but I think camouflage is a really interesting test case for AI image generation and snipers and ghillie suits are sort of the epitome of it.

Edit: These are nano banana 2 generations. With the new agent system in flow you can have it brainstorm up variations of the idea. These two are my idea, but it generated its own good ideas like a construction site and abandoned playground.

I have access to Grok imagine now* so I decided to give it a spin on this issue.

The suburban back yard

This one is almost a bit of a cheat because Grok is picking the patch of vegetation that’s suitable for camouflage but the integration of the lawn gnomes and flamingos is pretty cute.

And the sillier one, the sniper hiding in an arcade either can’t resist playing the games, or maybe that’s concept bleed about blending in socially vs blending in via camouflage.

  • I resisted getting supergrok for a long time because I don’t want to give a cent to Musk if I can avoid it, but 1) you can make an account directly through xAI/grok and not touch x/twitter, and 2) it has such generous video generation limits that they’re a massive loss leader. you can generate about 30 10 second clips at 720p and 480p videos per day. Video generation like that is a huge loss leader. I’m 3 days in and I’ve already generated enough video that if I were using the grok API I would’ve spent $75 on video generation. Even if their actual costs are half that, I’ve generated more than the subscription price’s worth of video in 3 days and I plan on creating plenty more for the rest of the month. If they spend $200 serving me videos and I pay them $30, I figure I’m not actually giving Musk any money. That’s why OpenAI shut down Sora.

And I get hundreds of fun little videos, some of which are quite good. I’m trying to do a survey of all the major frontier systems to see their capabilities – openAI / imageGPT is the only one I haven’t explored yet.

I quit paying for Grok videos because their prompt comprehension and outputs got majorly stupid a couple months ago. Has that improved?
I know you probably don’t have any basis for comparisons between now and how it degraded earlier this year, but perhaps you can run some moderately complex tests, especially multiple characters interacting with each other. Run the same exact prompt multiple times, share all the results?

Well I can tell you the grok subreddit is MASSIVELY complaining about how everything is awful and it has declined massively, but it’s almost all about moderation – I think they were probably the dudes making some of the problematic NSFW content. Not all NSFW content is problematic, but you probably heard about what grok was happy to do.

I know they just released Imagine 1.5 like… yesterday, and I can’t tell if I have access to it yet. I haven’t noticed a difference but sometimes they roll out updates to different users at different times.

I don’t have enough experience to rate it very well, I’ve only been using it for a couple of days, I’ve mostly been using it to animate already existing images (mostly from midjourney)

There are two modes - speed and quality - other models have something similar like NB2 and NB pro for google but the on grok speed/quality are very different models. Speed tends towards sort of idealized images that look a little fake, too perfect, a little plasticky, but not bad. You see the same faces come up over and over again. Quality is more realistic and has less ideal models that look more like real people. But the interesting thing about speed is that it does infinite scrolling generation so you can crank out dozens of candidate images very quickly whereas quality does a batch of 4 at the same time. The batches of 4 for quality are too close, though, to explore a conceptual space. It’s like NB2 in flow in 4x mode rather than midjourney - 4 slight variations of the same image.

I did find that grok does a decent worldbuilding, like if you put weird elements together it does a decent job of building a world where those two concepts mesh, but I need to spend more time exploring that. I think I posted an example of that earlier, when I used grok through an API, where I had a george washington vs donald trump NBA game on an alien world and the little aliens were wearing “Trump 24” jersies even though I never mentioned that detail.

So… no real opinion yet. the speed generation mode with infinite generation (until you hit your limits) is novel, I haven’t seen image generation that fast / prolific in any other model.

The agentic image / video creation tool is excellent. Flow has something similar now but Grok is probably even better. You create a project and tell the agent (a grok chat bot) what you want to do and it will help you come up with prompts and brain storm and generate the images and videos. Like you can say “I want to do 4 variations on this theme” and it will create them well, then you can say “let’s build off this one but go in this direction” or “take this scene but put it in a city at night” and it’ll do a very competent job of it. You can ask for suggestions or brainstorm ideas for scenes and ask for edits. Before google released an agent in flow I bet this was the best version of this, and I think it’s probably still better than flow because the project grid is a 2d space to organize images and videos you can work with rather than just one timeline. Probably the best feature. If they didn’t have the agent mode when you stopped, you might find that it’s good at helping you create the results you want by acting as essentially an expert prompt interpreter.

Do you have any prompts you want me to test specifically?

I decided to try same idea again, “George Washington dunks on Donald Trump at an NBA game. Both are wearing NBA jersies that are appropriate to their own style”

Speed:

quality:

Agentic:

None of them are very good in this case, though the agentic is the best especially if you like the righteous fury it decided on. It’s a relatively simple prompt, though I should’ve added “sports photography” for a better test. Speed did a reasonable representation of what the characters look like but has the action all wrong. Quality got the positioning right, but decided George Washington was an African American jacked dude. To be fair to quality, that’s not exactly 4 mistakes - like I said, it generates 4 versions of the same image with variations. I did run it one more time and it did again have a black man in a powdered wig.

Agentic did the best job but not quite what I was looking for It did mention it was searching for images of the two people to make a more accurate representation, so the grounding in that workflow will help keep details accurate. That matters much more for obscure things than well known characters. I believe NB2 does the same thing. If you ask for an obscure landmark that’s not represented in the data very well, it’ll go out and do an image search. But this photo feels almost like an exaggerated anime version rather than real sports photography. Which is… a defensible stylistic choice since I didn’t specify sports photography.

The jersey choices are pretty good. Washington for Washington makes sense, but so does sixers. Making him “1” in comparison to Trump’s 45 makes sense. And having Trump just say “Trump” may or may not be clever analysis – Trump is his own brand, his own style, much more than the country.

For comparison, here’s a batch of 4 by Nano Banana 2 which did a better job across the board:

When I used the grok API a few weeks ago, it seemed like the results were better. I’m not sure if the API gets access to better models than the subscribers get.

To be fair to Grok, I find the agentic workflow to be really good for relatively simple tasks like creating scenes or characters, and using plain language to modify them to get the results you want. These are basically the starting rounds of manifesting an idea, so they’re the roughest version - some refinement would get them closer. But NB2 did get good usable results right off the bat.

Wow, how time flies, I didn’t realize how much time had gone by since I posted that one until I had to scroll all the way back. In any case, I’m also wondering who did the layout for this diner:

The street sign outside looks a bit sus too.

Create a photo of something sus.

Copilot

Gemini

ChatGPT

Create a photo of something sus on a bus causing a fuss. Include pus.

Gemini

Copilot

I can’t create that image. The request includes graphic bodily fluids in a way that violates safety rules, so I can’t generate it.

ChatGPT

We’re so sorry, but the prompt may violate our guardrails around violence. If you think we got it wrong, please retry or edit your prompt.

Create a photo of Small Wonder making a blunder Down Under.

Gemini

ChatGPT

Copilot

All three correctly guessed from the ambiguous prompt that I wanted the robot from the TV show in Australia. Copilot didn’t quite get what the robot should look like.