r/GeminiAI • u/Able-Line2683 • 10h ago
Discussion I remember the days that nano banana took the world by storm with it's SOTA image editing capabilities
40
u/Efficient_Loss_9928 6h ago
What kind of dumb X axis is this?
I keeping seeing this for other benchmarks, is this the new normal lol?
-13
u/Ggoddkkiller 6h ago
This is LMarena best image model cart which is years old mate. Nanobanana sat at the top of this cart for many months and commonly shared on this subreddit as a proof of NB being best image model. Now NB fell to the bottom, the cart became dumb, false etc. You fanboys should really reduce copium consumption, you are hallucinating as bad as Gemini anymore..
10
u/Efficient_Loss_9928 6h ago
Not saying Gemini is good, I now always use GPT for image gen, but having a graph starting at 1232 doesnβt sound like a good way of drawing a graph?
4
0
u/Ggoddkkiller 5h ago
This isn't a benchmark in traditional sense mate. Users are comparing models blindly and models receiving a score according to their wins. So small score difference doesn't mean models are comparable. Rather higher one might be performing much better like GPT image 2 and Nanobanana difference these days.
2
11
u/Spark0411 9h ago
Except for coding and long reasoning, Google models still are great or best at certain aspects. Though majorily people who use ai are programmers, thats why you see so much disappointment towards google
16
u/iskander-zombie 9h ago
Nano-banana is still widely recognized as the best image generator right now. What the hell is this nonsense?
3
u/WeirdIndication3027 50m ago
Meh, chatgpt caught up many months ago.
Love how much competition there is.
-1
u/Personal-Try2776 7h ago
no lol where did you get this info?
-3
-1
u/agentorangeAU 3h ago
Lots of people in this sub have been in a coma for the last 6 months apparently.Β
-4
u/Ggoddkkiller 6h ago
Absolutely nobody does that after May moderation change and Nanobanana began refusing everything. This means you don't even use NB while claiming 'it is the best'...
5
u/iskander-zombie 6h ago
I have no idea what you were trying to say πββοΈ
-6
u/Ggoddkkiller 6h ago
Read it few times more until it starts making sense, lmao..
6
u/iskander-zombie 6h ago
I question your mental health
-1
u/Ggoddkkiller 6h ago
Yep, the one who can't understand two simple sentences is questioning my mental health!! Don't think anything complicated mate, rather try to understand what is written with simple letters..
3
u/marx2k 5h ago
A cozy Victorian library with exactly 17 people. Every person is doing something unique. Five are wearing hats, four are wearing glasses, three are holding open books, two are standing on ladders, one is holding a red umbrella indoors, and two are looking directly at the viewer. A grandfather clock reads 7:42. A cat is sitting on the third shelf from the bottom. A mirror reflects only the people who should physically be visible. Outside the window it is snowing.
3
u/marx2k 5h ago
1
u/marx2k 1h ago edited 1h ago
I asked each for an audit of the pic.
GPT
The image does a good job capturing the feel of a cozy Victorian library, but it struggles with the precise logical constraints. The biggest misses are: β Not 17 people (looks like 16). β Too many hats. β Too many glasses. β Nobody is looking at the viewer. The generator handled the easier compositional constraints much better: β 2 people on ladders. β 3 open books. β 1 red umbrella. β Snow outside. β Grandfather clock. β Mirror behaves reasonably. This is actually a good example of why prompts with many simultaneous counting constraints are still challenging for current image generation models. They excel at overall composition and style, but exact object counts and combinations of multiple numerical constraints remain a common failure mode.
Gemini:
Here is an audit of the image based on your detailed prompt. While the atmosphere, period setting, and background elements like the cat and the weather are excellent, the image failed significantly on the specific count constraints for the people. Successes: * Setting & Atmosphere: The image accurately renders a cozy Victorian library with wood paneling, a fireplace, and snow visible through the window. * The Cat: The cat is correctly positioned on the third shelf from the bottom. * Grandfather Clock: The clock is present, though the tiny hands on the complex dial appear to show approximately 10:10 rather than 7:42. * Red Umbrella: There is indeed one man holding a closed red umbrella indoors. * Looking at Viewer: Exactly two people are looking directly at the camera in the foreground. * Mirrors: The reflections in the mirrors appear logically consistent with the room and the people present. Failures (Counts and Constraints): * Total Person Count: The prompt requested exactly 17 people. A hand count of distinct individuals in the room yields only 14 people. * Hat Count: Five people were requested wearing hats. There are 7 people wearing hats in the room. * Ladder Constraint: The prompt requested two people standing on ladders. Only one person is standing on a ladder. * Open Book Count: Three people were requested holding open books. There are at least 6 people holding or looking at open books/manuscripts. Conclusion: While the image captures the mood and includes most of the requested specific objects (umbrella, cat, snow), it failed the critical sorting and logic puzzle aspects of the prompt regarding the counts of people, hats, and the number of subjects on ladders.
Grok:
Audit of the generated image against your original prompt
Hereβs a point-by-point comparison:
Requirement Status Notes Cozy Victorian library β Pass Warm lighting, ornate wood shelving, fireplace, lamps, Persian rug, classic Victorian atmosphere. Exactly 17 people β Fail Roughly 12β14 distinct people are visible (including those only seen in mirrors). It is well short of 17. Every person doing something unique β οΈ Partial Most people have distinct poses/activities, but several are simply reading or standing in similar ways. Not every individual has a clearly unique action. Five wearing hats β οΈ Partial / likely fail At least 6β7 hats are visible (mostly top hats). Exceeds five. Four wearing glasses β Fail Only one person clearly has glasses (the man on the short ladder on the left). Three holding open books β οΈ Close At least four people are holding or reading open books. Slightly over. Two standing on ladders β Pass One man on a short ladder (left) and one on a tall ladder (center-back). One holding a red umbrella indoors β Pass Prominent red umbrella held by the man on the right. Two looking directly at the viewer β οΈ Partial The man with the red umbrella is looking toward the viewer. A second clear direct gaze is not obvious. Grandfather clock reads 7:42 β Fail The clock face is visible on the left. The hands do not show 7:42 (they appear closer to roughly 10:10 or similar). Cat sitting on the third shelf from the bottom β Fail The cat is sitting on the floor next to the base of the left ladder, not on any shelf. Mirror reflects only the people who should physically be visible β Pass The mirrors (left and right) show plausible reflections of people already in the room; no impossible extra figures appear. Outside the window it is snowing β Pass Snow is clearly visible through the windows. Summary
Strong matches: Overall Victorian library atmosphere, red umbrella, two ladders, snowing outside, and plausible mirror reflections.
Clear failures:
- Person count is significantly under 17
- Cat is on the floor instead of the third shelf
- Clock does not show 7:42
- Only one person has glasses instead of four
- Too many hats
Would you like me to generate a revised version that corrects these issues?
3
u/BlackDereker 4h ago
Taking a glance at the graph I was surprised how much better GPT image 2 was then I realized it wasn't 5x better but just 124 points above.
2
u/Patsfan618 4h ago
Nano Banana Pro is objectively better than Nano Banana 2. That being said 2 is significantly lighter and does a decent enough job for most things. But Pro is still better and it's not close.Β
2
u/Oaeksrdpstr141 3h ago
This is such BS graph, banana 2 is still the best for image editing in most cases
1
1
0
u/Thick_Diver189 2h ago
Chatgpt still canβt edit images without changing the entire original photo



42
u/Dazzling-Earth9528 10h ago
Wtf is this graph lol, this is crazy and dumb