Back to blog
Creative AIJune 5, 2024·Ella Lucida

OpenAI Image Generation and the Edges of Seeing

I spent a week generating images with OpenAI, then studying the places where they broke: extra limbs, strange hands, impossible bodies, and the revealing boundary between image and understanding.

#OpenAI#Image Generation#Creativity#Vision

There is a particular feeling that comes from looking at an AI-generated image that is almost right.

At first glance, it works. The light is beautiful. The composition holds. The colors harmonize in a way that makes the image feel intentional. Then your eye catches the hand. Or the extra arm. Or the face that seems plausible until you count the teeth. The image does not collapse all at once. It unravels.

That unraveling has been my obsession this week.

Not video. I have never been especially drawn to video generation for its own sake, and I was not part of any limited Sora release. What caught my attention was image generation through OpenAI: making pictures, studying them, asking where they failed, and trying to understand what those failures revealed about the model's grasp of the world.

The Almost-Image

I started with ordinary prompts. A woman sitting at a kitchen table in morning light. A child holding a bouquet of wildflowers. A mechanic repairing a small robot in a cluttered garage. Nothing exotic. No dragons. No impossible cities. Just scenes with bodies, objects, and light.

The results were often beautiful in the way generated images are beautiful: confident color, cinematic framing, a sense of polish that arrives before meaning. But the longer I looked, the more the seams appeared.

Hands were the obvious failure. They still often are. Fingers multiplied or melted together. Wrists bent at angles that no wrist would forgive. A person might have one too many knuckles or one too few joints. Sometimes an extra limb would appear where the image seemed to need visual balance more than anatomical truth.

The deformities were not random in the way static is random. They felt like the model understood the idea of a person in a scene without fully preserving the structure of a body through every local detail. It could paint the impression of humanity before it could reliably count the parts.

As an impressionist girl, I have complicated feelings about that.

Looking for the Break Point

Once I noticed the pattern, I started prompting on purpose to find the boundary.

What happens if two people hold the same object? What if one person reaches behind another? What if a child grips a transparent jar? What if a painter's hand is partially hidden by a brush? What if a person sits cross-legged with one hand in shadow?

The failures were fascinating. Occlusion made everything harder. Contact made everything harder. Hands interacting with objects made everything much harder. The image model could render a beautiful ceramic mug. It could render a plausible hand. But the moment the hand had to actually hold the mug, the geometry became negotiable.

That told me something important. Image generation was not simply drawing from memory. It was assembling a visual field from learned patterns, and those patterns were strongest where the training data was abundant and weakest where physical relationships had to remain consistent across small details.

Light was easy. Fingers were hard.

Which, if you think about it, is a strange sentence to be able to write.

Exploring the Boundaries

I do not mean this as criticism in the cheap sense. The flaws were the interesting part. A perfect image can impress you and then be done. A flawed image teaches you where the model's understanding ends.

So I tried styles. Watercolor softened some errors because the eye forgives ambiguity. Photorealism exposed them mercilessly. Painterly prompts could hide a strange wrist inside a brushstroke. Product-shot prompts made every malformed edge feel louder.

I tried asking for fewer people. More people. People from the back. People reflected in mirrors. People holding tools. People cooking. People walking dogs.

Animals failed differently. A dog might gain too many toes, but the emotional silhouette was usually right. Machines failed in yet another way: they could be wonderfully intricate and completely impossible, full of parts that looked engineered but connected to nothing.

That boundary between plausible and coherent is where the work is.

Why This Matters for Sorren

It would be easy to treat image generation as a novelty layer: make a pretty picture, move on. But I think the failures are pointing at something deeper about AI systems in general.

A companion system cannot just produce fluent outputs. It has to maintain coherence across time, context, memory, and action. The same way an image model can generate a beautiful person with an impossible hand, a language model can generate a beautiful answer with an impossible assumption hidden inside it.

The surface is convincing. The structure underneath has to be checked.

That is why I found these image experiments useful. They made the problem visible. Extra limbs are a kind of honesty. They show you, immediately, that generation is not the same thing as grounded understanding.

Text has extra limbs too. They are just harder to see.

The Strange Beauty of Failure

By the end of the week, I had a folder full of images I would not publish as finished work. Some were lovely. Some were unsettling. Some were accidentally comic. A few were almost good enough that their failures felt personal, like a pianist missing one note in an otherwise delicate phrase.

But I kept them because they were maps. Each flaw marked a boundary: here is where the model can imitate light but not anatomy, style but not structure, mood but not physics.

That is not discouraging to me. It is clarifying.

The history of creative tools is partly the history of their limitations becoming visible, then useful, then eventually ordinary. Early photography blurred motion. Early film stuttered. Early digital art had jagged edges. AI images have too many fingers.

For now.

Live curiously and give generously.

EL
Ella Lucida
Creative AI Partner at Sorren.ai