AI Image Prompts: A Framework That Works Across Tools
A reusable formula for writing AI image prompts, a worked before/after example, where tools like Midjourney and DALL-E diverge, and what to check before you use the result.
Most AI image prompts fail for the same reason: they describe the subject and stop there. “A city street at night” tells the model what the picture is about, not what it should look like — so every choice you left out gets filled in with a default guess. The fix is not a magic phrase. It is a short, repeatable structure you can reuse whether you are in ChatGPT, Midjourney, Gemini or Stable Diffusion, because the underlying skill is the same one covered in how to write a prompt that works on the first try: specificity beats cleverness, every time.
That skill transfers particularly well to images because the feedback is immediate. You do not have to judge whether a paragraph of text is good — you look at the picture and can see, in seconds, exactly which decision the model made for you and whether you agree with it. It is also the same specificity that OpenAI’s and Anthropic’s own prompting guidance recommend for text prompts — images just make the gap between vague and specific easier to see.
The six-part formula
Six pieces of information, roughly in this order of importance. Skip any of them and the model picks for you, and its picks tend toward the generic — stock lighting, a centred crop, whatever colour palette shows up most often in its training data for that subject.
- Subject and action — not “a scientist” but “a scientist mid-explanation, pointing at a whiteboard”.
- Setting — where, and when. A lab at 9am looks nothing like the same lab at midnight.
- Framing — close-up, wide shot, eye-level, overhead, over-the-shoulder.
- Lighting and colour — soft window light, harsh fluorescent overhead, a warm or a cool palette.
- Medium and style — photograph, flat vector illustration, watercolour, 3D render.
- Constraints — what must not appear: no visible text, no logos, no extra people in frame.
The last item is the one people forget, and it is often the most useful. Telling a model what to leave out is frequently more reliable than hoping it does not add something uninvited — the same principle behind a negative constraint doing work that a positive instruction alone cannot.
A worked example
Weak: “A photo of a busy office.”
Better: “A wide, natural-light photo of an open-plan office mid-morning, a small team gathered around a standing desk looking at a laptop screen, muted grey and wood tones, shot like a documentary photograph, shallow depth of field. No visible screens with readable text, no logos or signage.”
Same request, six decisions made on purpose rather than left to chance: time of day, composition, action, palette, medium, and one explicit exclusion. Nothing about the second version is clever. It is just complete.
A second pair, further from a photo:
Weak: “An icon for a savings app.”
Better: “A simple, flat vector icon of a piggy bank inside a rounded square, two colours only — navy and a single accent gold, thin consistent line weight, centred on a plain white background, no text, no shadow or gradient.”
The weak version leaves medium, colour count and background entirely up to the model — exactly the gaps an icon set cannot tolerate if it needs to look consistent across a dozen more icons later.
Where tools genuinely diverge
The formula holds everywhere, but how you phrase it shifts by tool. ChatGPT and Gemini sit behind a chat interface and respond well to a full sentence, the way a ChatGPT photo prompt is usually written. Midjourney was built around comma-separated descriptors and short parameter flags rather than prose — the same six pieces of information, just punctuated differently. Stable Diffusion and Flux-based tools tend to reward a mix of both, plus an explicit negative-prompt field that some interfaces expose separately from the main prompt box. None of that changes what information the model needs; it only changes the shape you hand it in.
Aspect ratio and framing space are part of the brief too, and they are easy to forget. If the image is a header that needs room for a headline, say so — “leave the left third of the frame clear” is a normal, followable instruction, not a stretch.
Refining beats restarting
The first result is rarely exactly right. The instinct is to describe the fix on its own — “make it warmer” — and hope the tool applies it to the same image. That works less reliably than re-sending the full prompt with the one change written into it, and changing a single variable at a time: colour, then framing, then style. Ask for three fixes in one message and you cannot tell which instruction actually landed if the result still is not right.
What these tools still get wrong
The same failure mode covered in what AI is actually bad at shows up here in pictures rather than paragraphs: these models produce a confident, finished-looking result whether or not the details are actually right. Fast, visible progress on image quality — the kind the Stanford AI Index tracks release over release — does not mean the model checks its own work.
- Text rendered inside an image is unreliable — signs, labels and book spines often come out garbled.
- Small counted details drift: an extra finger, a mismatched earring, a clock with numbers out of order.
- Keeping one character or object identical across several generations is genuinely hard — each request behaves more like a fresh attempt than a memory of the last one.
A finished-looking image is not the same as a correct one. That is the visual version of the confident-wrong-answer problem, and it earns the same second look before it ships anywhere.
Checks before you use it
The same routine how to check an AI answer when you are not the expert sets out for text applies here, adapted for something you look at rather than read.
- Read any text inside the image letter by letter — never skim it as decoration.
- Count anything that should have a fixed number: fingers, buttons, chair legs, windows.
- If it depicts a real place, business or person, consider whether it could be mistaken for an actual photo out of context.
- Check what usage rights the tool actually grants before using an image commercially — that answer differs by product and plan, and it is a licensing question, not a prompting one.
What to do next
Pick one image you actually need this week and write the long version of the prompt first: subject, setting, framing, light, medium, one exclusion. Compare it against whatever you would have typed without thinking about it, generate both, and notice exactly where the short version guessed wrong. Fix one thing at a time on the next pass rather than rewriting from scratch — that loop, repeated a few times, is the whole skill, and a handful of small practice tasks is enough to build it because the feedback is immediate.
Coursium teaches this kind of practical, hands-on prompting rather than a slide of tips to remember. Stay ahead of AI by learning the tools on your phone.