Tried cleaning up this photo
with AI.
Remove the text and lines from the background. Leave the fly unchanged.
Gemini
ChatGPT
Copilot
Tried cleaning up this photo
with AI.
Remove the text and lines from the background. Leave the fly unchanged.
Gemini
ChatGPT
Copilot
I find that simple edits like removing or adding things are well covered now by autoregressive models. I’ve rarely (maybe never?) had NB2 fail to achieve the edit I wanted for fairly simple edits. Obviously in your comparison ChatGPT treated the red box as a “line” and that’s at least a valid interpretation I think even if it wasn’t the intended one. An easy fix with a slightly more specific prompt.
They’re actually quite good for relatively complex background edits too like “make this a night scene” or “change the time of day to sunset” - they don’t just preserve the subject and swap out the background, they will generally correctly re-light the subject appropriate to the scene too.
I notice that all three still have the “ar” of “total carbs”.
I actually wanted an empty background, so it did the best job on that. But it did the worst job of preserving the complex shadow.
Copilot did the best job with the shadow, I think, but also cropped the image in a way you didn’t ask for. The results on all of them are pretty good but they all have different subtle errors.
I don’t like generating images at the level of specifying the chatbot because you don’t know what the actual image generation model you’re using is. For example, this is the result of using your image and prompt in image GPT 2.0 at high effort.
When you as the gemini chatbot to do an edit, you’re probably getting nano banana 2 but might be getting nb2 lite. I doubt the chatbot calls NB pro, but it’s possible it does for edits. ChatGPT I honestly don’t know, do you get image gpt 1.5? 4o? Does it select based on task? No idea. Copilot is even more mysterious. You don’t even know what family of image generators it could be using underneath.
ImageGPT 2 did a better, though imperfect job of preserving the shadow. It got rid of the table the sample was on entirely. It does still have the AR from carb as Chronos noted. I think that’s difficult for these models to separate because of the translucent layer - they’re not sure if the AR is part of the markings on the wing or the text underneath it. If you ignore the surrounding context, you can see why the model might not recognize the “AR” and instead just see some spots on the wing.
You say that a lot but I have never got any type of variation in style or quality that ever suggests that each chatbot doesn’t use the same model every single time.
Well, sure, but you used “chatgpt” to generate your image with the opaque wing that didn’t preserve the translucent shadow, and I used a openAI/chatgpt image model that a paid chatgpt account can call that did. A pretty significant difference. Which is the ceiling of what “chatgpt” can do?
There can be inconsistency in the chatbot method you’re using. Are you using free accounts for all these chatbots? Paid accounts often get better quality image models. And sometimes different levels depending on the tier of paid account. What happens when copilot uses MAI vs gpt 4o vs Dall-E? And I do believe it chooses at least the latter two depending on load and whatever criteria it wants. Copilot is designed to be model agnostic and chooses what model to actually generate things with in a way that’s seamless and non-transparent to the user.
At some point gemini moved from Imagen to nano banana to nano banana 2 as its main image generator - is “gemini” in 2024 the same “gemini” in 2026?
It’s sort of like quoting 0-60 times for a “Ford” - is it a Ford Mustang or a Ford Fiesta? Is it a Ford Mustang with a 3.3 ecoboost or 5.0 v8 engine? Is it a 2024 or 2026 Ford Mustang with which engine?
Maybe I’m being overly strict in my analysis. I do enjoy your comparisons so I’m not saying this is a fatal flaw. And if the question is “what do you get when you ask each of these systems with a free or regular end user paid account to generate an image” then it’s the correct output for these questions. I guess I just think if the point is comparing these various systems, you’d want to be specific about what you’re comparing since they offer a range of models and abilities.
Edit: And I’m not objecting, exactly, if this is the workflow that interests you, by all means, continue. I just think it’s worth noting that you may not be getting the best version that these systems have to offer, which I think is relevant to the issue of comparison and the “AI image generation is getting crazy good” theme. I think your chatgpt vs my chatgpt results do show the difference. But “what does asking the chatbot to generate an image get you” is answering a valid question, so I don’t mean to say you’re doing anything wrong, so much as that it’s worth noting the limitations of the method.
I know what all of those models look like and I say with 100% confidence that Copilot has never given me MAI or Dall-E. And nobody uses Dall-E now:
I’m curious, what model do you think copilot is giving you? The chatbot doesn’t know, it just knows it passes an image generation request and gets an image back. It did used to use DALL-E 3, the first time I used copilot to generate images I’m pretty confident that’s what it used but that was probably a year ago, maybe more, so you’re probably right that it no longer uses it. I’m not sure why they wouldn’t use MAI, given that there’s a free test playground for users so it’s not strictly restricted to corporate clients. But I honestly don’t know what it’s using now.
I firmly believe that Copilot has given me gpt 4o or its newer replacement for every single solitary Copilot image I have ever posted in this thread.
I don’t know why they “wouldn’t”, but MAI output looks distinctivly different from (and inferior to) gpt 4o, so I’m confident that they don’t.
Paid users can use this to still use the ChatGPT 4o image model, although 4o itself is long gone; things get sent through whatever current LLM you pick.
I tried this interface for a bit, and I think its results with 5.x models are inferior to what it used to do during the 4.x era.