The one idea
Generation is not a slot machine. It is art direction with a slow, literal collaborator.
You give a direction. You look at what came back. You identify the single thing that is wrong. You change only that. Repeat.
This is exactly the discipline from when-it-breaks — change one variable,
observe the effect — applied to pictures instead of code.
Change one thing per attempt and keep notes. Otherwise you cannot tell what caused what.
Establish a base description that is close, then vary one control at a time. Save the exact text of anything that worked.
Hold the description and, where available, the seed fixed while varying one parameter. Controlled variation is the only way to attribute an effect to a cause in a stochastic system.
The loop
| Step | What you do | Why |
|---|---|---|
| 1 | Write the full direction — all seven controls | A vague start cannot be debugged |
| 2 | Generate a small batch | See the range, not one sample |
| 3 | Pick the closest, name the one biggest problem | Naming it forces precision |
| 4 | Change only the words governing that problem | Attribution |
| 5 | Regenerate; keep or revert | Now you know what that word does |
| 6 | Repeat until only small details are wrong | — |
| 7 | Fix the details by editing, not regenerating | Regenerating loses everything that worked |
Step 3 is the one people skip. "It's not quite right" is not actionable. "The light is coming from the front and it should be from the side" is one change with a predictable effect.
Diagnosing what is wrong
Most complaints map onto one of the controls from the previous lesson. Find the control, change that.
| What is wrong | The control | What to change |
|---|---|---|
| Flat, lifeless, characterless | Light | Specify direction and hardness |
| Feels staged, artificial | Angle and setting | Eye level, specific real place, specific materials |
| Too busy, nothing stands out | Composition and depth | Fewer elements, dissolve the background |
| Looks like every other AI image | Everything unspecified | Specify all seven controls |
| Wrong emotional register | Light and colour | Change hardness and narrow the palette |
| Subject is right, frame is wrong | Shot type | Wide, medium, close-up — state it |
You need an image for a report on rural connectivity. First attempt reads as a stock advertisement: bright, smiling, front-lit.
One change: hard, low afternoon light from the side, and available light only. It now reads as documentary rather than promotional — which was the actual problem all along, and it was one clause.
Consistency across a set
Getting one good image is a skill. Getting eight that look like they belong together is a different and harder one, and it is what real work usually needs.
Making a set hold together
1 of 5Write a style block and never change it.
One paragraph fixing everything that must stay constant: light quality and direction, colour palette, lens feel, level of detail, materials.
This block is copied identically into every prompt in the set. Only the subject line changes. Most inconsistency comes from people rewriting the whole description each time and unconsciously varying five things.
Six images for a product page. Written as six separate prompts, they arrive with six different lighting setups and read as a collage.
Written as one style block plus six subject lines, they read as one shoot. The work of building the style block once is the whole job; the six generations after it are mechanical.
Edit the detail, do not regenerate
When the image is 90% right and one thing is wrong — an extra finger, a wrong object, a distracting element in the corner — regenerating throws away the 90% that took you six attempts to get.
Instead: use the tool's editing or inpainting feature, which regenerates only a region you select. Or fix it in an image editor. A five-minute edit beats twenty more generations, and it keeps everything that worked.
This is the habit that separates people who finish images from people who accumulate near-misses.
Try this
You have generated an image that is right in every way except the subject is looking at the camera and you wanted them absorbed in their work. What do you do?
Your challenge
Level 3 · IndependentProduce a set of four images that clearly belong together and serve one stated purpose — a small campaign, a set of article headers, a product's supporting images.
You have succeeded when: you have a written style block you used unchanged in all four; you can show your iteration history with one change per step and say what each change did; at least one image was finished by editing a detail rather than regenerating; and you can state the rights position for using these where you intend to use them.
That last one is not paperwork. It is the difference between a portfolio piece and a liability.
What people usually get wrong
- Changing everything after a bad result. You learn nothing and you lose what was working.
- Not saving prompts that worked. You will want that light again next month and you will not remember it.
- Regenerating to fix a small detail. Edit the region instead. You are throwing away good work.
- Judging at full size only. Most people will see it small. Check it at thumbnail size, where hierarchy either survives or does not.
- Asking for text in an image. Generators are unreliable at text. Add text afterwards in an editor, where you control the font and it is spelled correctly.
- Assuming the tool's terms match another tool's. They differ, and they change. Check the one you are using.
- Publishing a generated image as if it were a photograph of a real event. This is the one with reputational consequences.
How someone experienced does it
Experienced people keep a small library of prompts that produced good results, labelled by what they were good for — "soft window light, muted, documentary feel". Over a few months this becomes a personal palette worth more than any prompt guide, because it is calibrated to their taste and their tools.
They also generate in small batches and look at the spread rather than the best one. If four generations of the same description differ wildly, the description is underspecified and the next fix is to add specificity, not to keep rolling. If they are all similar and all wrong, the description is specific and wrong, which is a different fix entirely. That diagnostic saves hours.
And they treat the generator as one tool in a chain. Generate the base, edit in an image editor, add real text, adjust colour. The output of a generator is rarely the finished piece, and people who expect it to be spend twenty generations chasing something a two-minute edit would have fixed.
When not to use this
Do not use a generated image when:
- It must depict something real. Your actual office, your actual product, your actual team. Use a photograph. A generated approximation is a small lie that people detect.
- Accuracy matters. Technical diagrams, anatomy, product mechanics, maps. Generators produce plausible-looking wrong detail with total confidence.
- The subject is a real, identifiable person. Unless you have their consent and the depiction is honest.
- A photograph would be faster. For a desk, a book, a plate of food, your phone and a window will beat twenty generations and will be true.
Prove it
Document one image from brief to final: the brief, every prompt version, what you changed at each step and why, and the final edit.
Six to ten steps. This document is more valuable than the image — it shows you can direct rather than roll, and it is the thing you can repeat next time on a different subject.
Keep learning this
Paste this into any AI assistant. It turns the assistant into a tutor that tests you instead of just answering you.
Act as an experienced practitioner who is good at teaching. I have just learned directing image generation models through controlled iteration. Assume I am intelligent but relatively new to this — treat me as intermediate level. Work through this in order, and wait for my reply at each step: 1. Ask me 5 questions that test whether I actually understood directing image generation models through controlled iteration. Do not reveal the answers yet. 2. After I answer, tell me which parts I got right, which I got wrong, and which I only half-understand. Explain only what I misunderstood — do not re-teach what I already know. 3. Give me one practical challenge based on something I could genuinely encounter at work or in daily life. Do not solve it for me. 4. Evaluate my solution the way an experienced person would judge it, including what a professional would have done differently. 5. Tell me what to learn next, and why that comes next. 6. Give me trustworthy sources for deeper study — prefer official documentation, primary research or standards bodies over blogs and videos. Rules for you: no buzzwords. No motivational filler. Say "I'm not certain" when you are not certain, and tell me which parts of your answer I should verify myself. Clearly separate facts from your recommendations and your opinions.
Become independent at this
Use this when you want a path from where you are to actually good, with checkpoints you can test yourself against.
I want to become independently capable at getting the image you actually intended — not permanently dependent on AI, tutorials or step-by-step guides. Design a progression for me with five stages: Beginner, Guided practice, Independent practice, Real-world application, Professional level. For each stage tell me: - what I must know - what I must be able to do without help - the mistakes people make at this stage - one practical challenge - one real project that would prove I reached this stage - one way I can test myself honestly Then tell me the signals that I am ready to move to the next stage, and the signals that I have skipped ahead too early. Keep the theory to the minimum I actually need. Focus on ability I can transfer to situations you and I have not discussed.