How to Write a ChatGPT Photo Prompt That Actually Works
A good ChatGPT photo prompt is specific, not clever. The details that matter, a worked before/after example, and the checks to run before you use the result.
A ChatGPT photo prompt that reads like a caption gets a generic result. "A photo of a coffee shop" is a caption — it tells the model what the image is about and nothing about how it should look, so every gap you leave gets filled with the model's default guess: flat lighting, a stock-photo crop, colours nobody asked for. The fix is not a clever phrase. It is giving the same kind of detail you would give a photographer on a shoot. Generative image tools have improved fast — the Stanford AI Index tracks that pace year over year — but faster does not mean the model reads your mind. It still only works from what you actually tell it.
ChatGPT's image tool sits behind the same chat window as its text answers, built on the same family of models covered in what GPT actually stands for. The prompting discipline carries over too: OpenAI and Anthropic both tell you, in their own guidance, that specificity beats cleverness — which is exactly what how to write a prompt that works on the first try covers for text. Images just make the gap between a vague prompt and a specific one easier to see, because you get a picture instead of a paragraph.
What a photo prompt actually needs
Six things, in roughly this order of importance. Leave any of them out and the model picks for you.
- Subject and action — not "a chef" but "a chef mid-stir, focused, sleeves rolled up".
- Setting — where, and when. A kitchen at dinner service looks nothing like one at 7am prep.
- Framing and composition — close-up, wide shot, eye-level, from above.
- Lighting and colour — morning light through a window, harsh overhead fluorescents, a muted palette.
- Medium and style — photo, watercolour, flat vector illustration, cinematic still.
- Constraints — what must not appear: no visible text, no logos, no other people in frame.
Each of those answers a question the model would otherwise guess at. That is the same mechanic behind showing an example instead of describing one in a text prompt: a specific instruction carries more information than an adjective, and the model has less room to default to something generic.
Where this actually comes up
Most people run into this at work, not for fun: a placeholder header for a slide deck, a mood-board tile for a pitch, a social post graphic that needs to look intentional rather than obviously generated in five minutes. A specific prompt is what makes the difference between something you can ship and something you have to redo in an actual design tool anyway.
Aspect ratio and size are part of the brief too, and easy to leave out by accident. If you need a square tile for a social post, say square. If it is a header that has to leave room for a headline, say where the empty space should sit — "leave the left third of the frame clear for text" is a normal instruction and the model will generally follow it.
A worked example
Weak: "A photo of a coffee shop."
Better: "A wide, softly lit photo of a small neighbourhood coffee shop counter at opening time. Morning light coming through a front window, a barista mid-pour into a paper cup, warm brown and cream tones, shot like a 35mm film photograph. No text or logos visible anywhere in frame."
Same request, six extra decisions made on purpose instead of left to chance: time of day, light source, action, colour palette, medium, and one explicit constraint. That last one matters more than it looks — asking the model to leave something out is often more reliable than hoping it does not add it.
A second pair, for something less scene-based:
Weak: "A logo for a bakery."
Better: "A simple, flat vector logo for a neighbourhood bakery called The Corner Loaf. A single wheat stalk as the icon, minimal line detail, two colours only — dark brown and cream. Centred on a plain white background, no other text or shapes."
Notice what changed: the weak version leaves format, colour count, and icon choice entirely up to the model, which is exactly the kind of gap a logo cannot afford. The better version specifies the medium (flat vector, not a photo-style logo), the icon, the palette, and the background — four decisions that would otherwise be guesses.
Refining instead of restarting
The first result is rarely exactly right, and the instinct is to describe the fix in isolation — "make it warmer" — and hope the model applies it to the same image. That works less often than asking again with the full prompt plus the one change written into it. Change one thing at a time: colour, then framing, then style. Asking for three changes in the same message makes it hard to tell which instruction the model actually followed if the result still is not right.
Cartoon and art styles
Naming a style term works better than an adjective. "Cartoon style" is vague enough that the model still has to decide what kind. "Flat vector illustration, bold black outlines, four-colour palette, no gradients" is a brief a designer could actually work from.
- Flat vector illustration — bold outlines, limited palette, no gradients.
- Watercolour wash — soft edges, visible paper texture, muted tones.
- Line art or ink sketch — black and white, minimal shading.
- Retro screen-print poster — halftone dots, two or three flat colours.
- Claymation-style render — visible texture, slightly imperfect surfaces.
- 3D render with soft studio lighting — clean shadows, shallow depth of field.
One thing worth avoiding: naming a specific living artist to imitate their style. It sits in a genuinely unsettled legal and ethical space, and most tools restrict it for that reason. Describing the visual traits — line weight, palette, brushwork, era — gets you most of the same result without asking the model to copy a person.
Combining two style terms is usually more useful than reaching for a single label. "Watercolour wash, but with the bold flat colours of a screen-print poster" gives the model two references to blend rather than one to guess at, and it tends to produce something more distinctive than either style alone.
What these tools still get wrong
The same failure mode covered in what AI is actually bad at shows up visually, not just in text: these models produce a confident, finished-looking result whether or not the details are right. A generated image can look finished and still be wrong in ways that take a second look to catch.
- Text inside images is unreliable. Signs, labels, and book spines often come out as garbled or nonsensical characters.
- Small counted details drift — an extra finger, a mismatched pair of earrings, a clock face with numbers in the wrong order.
- Keeping one character or object identical across several generations is hard. Each request is closer to a fresh attempt than a memory of the last one.
None of that is a reason to avoid the tool. It is a reason to look at the result the way you would look at a first draft from a junior colleague — genuinely useful, worth building on, and still worth a proper read before it goes anywhere public.
Check before you use it
The same routine how to check an AI answer when you are not the expert lays out for text applies here, adapted for something you look at instead of read.
- Read any text inside the image letter by letter — do not skim it as decoration.
- Count anything that should have a fixed number: fingers, buttons, chair legs, windows.
- If it depicts a real business, place, or person, consider whether it could be mistaken for an actual photo if it were reposted without context.
- Before using it commercially, check what rights the tool actually gives you. That answer differs by product and by plan, and it is a licensing question rather than a prompting one — worth confirming before you rely on it.
A finished-looking image is not the same as a correct one. That is the visual version of the confident-wrong-answer problem, and it needs the same second look.
What to do next
Pick one image you actually need this week — a slide graphic, a placeholder header, a mood board tile — and write the long version of the prompt: subject, setting, framing, light, medium, one constraint. Compare it against what you would have typed without thinking about it. Then generate it, look for the two or three things wrong with it, and fix one at a time rather than rewriting the whole prompt from scratch. The gap between the first attempt and the third is the whole skill, and it is a fast skill to build because the feedback is immediate — you see the result in seconds, not after a meeting. AI project ideas that actually teach you something has more small tasks in the same spirit if this one goes well.
Coursium teaches this kind of practical prompting as a hands-on task, not a slide of tips to remember. Stay ahead of AI by learning the tools on your phone.