Back to BlogDevlog

Every Fix Broke the Next Thing: What 52 AI Character Portraits Actually Took

I needed 52 portraits for a card deck in Epic Words. I assumed it was one good prompt and a loop. Three failed takes and about 43 discarded images later, I had a much better understanding of what these tools actually do.

Every Fix Broke the Next Thing: What 52 AI Character Portraits Actually Took

The art in Epic Words is generated. I want to say that plainly at the top rather than let anyone discover it as a gotcha, because the interesting part of this post is not that I use these tools — it's how much work is left over once you do.

Yesterday I needed a deck of 52 characters. Roaming tradespeople for Epic Words: a baker, a blacksmith, a weaver, a falconer, fifty-two of them, each with a name and a blurb, each needing a portrait in the same painted style so they read as one set. Alongside that, sixteen named characters for a second deck.

I budgeted an afternoon. My mental model was: write one really good prompt, put the character description in the middle of it, run it fifty-two times, drink coffee.

That is not what happened. What happened is that I fixed six problems in a row and five of those fixes directly caused the next problem. By the end I had a working pipeline and about 43 discarded images, and a much clearer idea of why "just use AI" is a sentence that hides an enormous amount of labour.

Here is the whole chain, in order, with the actual pictures.

Failure one: fifty-two people from the same family

The first thing you want from a character set is a consistent style. Same painter, same light, same world. The obvious way to get that is to hand the model a reference image — here is a finished portrait I like, make more like this.

So that is what I did. And this is what came back.

A montage of sixteen generated character faces where the middle-aged women are nearly identical to each other and the elderly men are nearly identical to each other, despite having different names

Look at the women. Averil, Kerensa, Nerys, Angharad, Dilys. Same oval face, same dark arched brows, same nose, same set of the eyes. They are not similar — they are the same person in different clothes. The old men have the same problem: one white-bearded gentleman, painted four times.

I want to flag the thing I assumed first, because I think most people assume it and it is wrong. I thought this was a seed problem. Same random seed, same face, vary the seed and you vary the person. That is a very natural theory and it is not what was happening.

What was actually happening is that I had handed the reference image to an edit endpoint. Edit endpoints are built to preserve the content of the picture you give them. I thought I was passing a style. The model thought I was passing a face — and dutifully kept it, fifty-two times.

That distinction matters far more than seeds do, and it is the first thing I would check now.

Failure two: the model kept signing its own work

While I was untangling the faces, a second problem surfaced that had nothing to do with them.

The engine I was using kept painting an artist's signature into the bottom-right corner. Sometimes a monogram, sometimes a little cursive scrawl. Not always — which is worse, because it meant a sample of three or four images could look completely clean.

The reason is almost charming once you see it. The model learned to paint from paintings, and a very high proportion of finished oil paintings are signed. It had correctly concluded that a signature is part of what a finished painting looks like. It was not glitching. It was being consistent with its training data, which happened to include a convention I did not want.

We eventually measured this properly rather than eyeballing it, because eyeballing it is how you get burned:

Assume every engine stamps a corner mark until measured at n≥15. One engine wrote a monogram on roughly 94% of portraits; another painted faux cursive signatures on roughly 40%. At n=3, either can look clean.

Ninety-four percent. And at a three-image sample, plausibly invisible.

Failure three: the cure was worse than the disease

Here is the first place a fix caused the next failure.

To escape the signatures, I switched image engines. Different model, different training data, no ingrained habit of signing things. Reasonable move.

The new engine gave me genuinely distinct faces — it solved the sibling problem almost for free — and the paint quality was nowhere near as good. Not subtly worse. Obviously, immediately worse, in a way that would have dragged the whole deck down. I rejected the entire take.

Three portraits from the rejected engine — a baker, a blacksmith and a cooper, each clearly a different person, but painted flatter and muddier than the house style, with a faint cursive signature visible at the lower right of the third

Three different people — so the sibling problem really was solved. But look at the surfaces. It is flatter, muddier and more literal than the warm luminous oil I had already settled on, and dropped into a deck alongside the others it would have read as a different artist. Look at the lower right of the cooper, too: a faint cursive squiggle. That is the signature habit I switched engines to escape, still there.

Because it still signed about 40% of them.

So I had traded a great painter with an annoying tic for a worse painter with a slightly less frequent version of the same tic. That is the shape of the mistake and I want to name it clearly, because I have now written it into our own guidelines as a rule:

Never switch engines for a prompt-layer problem.

A corner signature is not a reason to change your painter. It is a reason to write a better instruction and then verify the corners. Changing engines to dodge a prompt issue means re-litigating every single quality decision you had already settled — style, palette, lighting, brushwork — to fix one small thing you were going to have to handle anyway.

Failure four: describing a type gets you the average of that type

Back on the good painter, with the reference image removed and the face described in words instead. Surely this is the answer: stop showing it a face, tell it a face.

Good paint. Still repeated people.

Not identical this time — just drawn from a very small pool that I recognised as the model's faces rather than mine. And this is the subtlest failure of the whole day, so it is the one I would most want another developer to take away:

A brief that describes a type is averaged into the nearest prior.

If I write "a weathered blacksmith in his forties," I have not specified a person. I have specified a category. And the model resolves a category by going to the middle of everything it has ever seen labelled that way. Fifty-two different category descriptions will still collapse onto a handful of faces, because the categories overlap and the middle of each one is close to the middle of the others.

The output is repetitive even though the input genuinely is not. That is why "they all used the same description" feels true when you look at the results, and is not quite what happened. The descriptions were different. They were just not specific enough to escape the average.

Failure five: forcing them apart made everyone a cartoon

The fix for the averaging is to give each character something the average does not have. Not "weathered" — an actual named, asymmetric, unmistakable feature. A nose broken and set badly to the left. A split nostril healed into a notch. One corner of the mouth pulled down by an old palsy.

This works. It is the single highest-leverage change in the whole pipeline.

It also immediately produced two new problems.

The first is that a small pool of distinguishing features is just a new kind of sameness.

Give the batch "a distinctive mark" and let the model choose which, and it will choose from a very short list — the same short list every time. I got a run where nearly everybody had a mole. Corrected that, and got a run where nearly everybody had a scar across the mouth. Not one or two. Most of them, in the same place, at roughly the same angle, as though the entire guild had been in the same bar fight.

This is the averaging problem from failure four wearing a disguise. "A distinctive mark" is a type of thing, so the model resolves it the way it resolves any type — by going to the middle of it. I had not created individuals. I had created a family with a shared injury, which is arguably a worse result than the siblings, because at least the siblings looked like people.

The fix is that the mark has to be named per person, and the pool has to be genuinely wide. The roster that finally worked draws on nineteen distinct kinds of feature — a split nostril, a nose broken and set to one side, a missing eye tooth, an old palsy, a milky blind eye, a burn, a notched ear, teeth turned inward. Across forty characters, "scar" appears six times and "mole" exactly once. That distribution did not happen on its own. It happened because it was written down.

The second is caricature drift, and this is where the whole approach nearly came apart.

My grocer is a small, neat, fussy man who sells vegetables and disapproves of you. Here is what came back.

A generated portrait of a grocer in a green apron holding a brass balance scale, with enormous pointed ears, an elongated tapering skull and oversized eyes — an elf rather than a human

That is an elf.

He was supposed to be a man who sells vegetables. Nobody asked for a fantasy creature — the world these characters live in is deliberately mundane, and the surviving character brief for him describes an ordinary fifty-one-year-old with a waxed moustache and never mentions his ears at all.

I would love to show you the exact prompt that produced him. I can't: this is from one of the morning takes, before I started saving the prompt alongside every image, so the wording is gone. That omission comes back at the end of this post, and this is the moment it cost me something.

What I can show you is the phrasing class that does this, because it survives elsewhere in the roster. My town crier's brief asks for "large ears standing well out." Perfectly ordinary words. But "large" is a direction, not a destination — and stacked with the other amplifiers in his description (an unusually large head, a jaw like a shovel, a mouth that opens enormously, shot from a low angle, mid-bellow) it walks steadily away from a person and toward an ogre.

That is the trap in one image. Every feature I asked for was individually reasonable. The model does not know where the human range ends. It knows which direction "larger" points, and it will keep going.

The same drift, less dramatically, gave me a juggler with a bulbous red nose and ears like jug handles:

A generated portrait of a juggler in patchwork motley with greatly oversized protruding ears and a bulbous red nose, painted in a warm storybook oil style

Not a person with distinctive features — a caricature of a village idiot, which would have sat in the deck next to fifty-one painted portraits looking like it wandered in from a different product.

One more thing about the grocer, and it is the corner-signature problem coming back to bite. This elf is one of two candidates generated for him in the same run. The other one is signed — a rust-red cursive scrawl on the shop wall behind him. This one is completely clean.

Same character, same engine, same minute, and a fifty-fifty split on whether the model signed its work. That is why the measured rates in failure two matter so much, and why checking three images and declaring the corners clean is worth nothing at all.

The lesson I took from the elf is that a distinguishing feature needs a ceiling, not just a direction. "Large ears standing out from the head" invites escalation. "Ears that stand out enough to notice, in normal human proportion" does not. And where a feature really does need to be extreme, it is safer to reach it through a specific cause — an ear notched by an old injury, a nose broken and badly set — because a cause has a natural limit and an adjective does not.

There is a second failure hiding in that same image, incidentally. Bottom-right corner. The monogram from failure two is still there — this was generated after I had added anti-signature language to the prompt. One image, two unresolved bugs, and I only noticed the second one because I had built a habit of checking corners.

Here is that whole batch, and you can watch both problems at once:

A montage of sixteen regenerated faces that are now genuinely distinct from each other, but where two have drifted into caricature with exaggerated ears and noses, and one has a broken composition with the head cropped out of frame

The faces are fixed. Genuinely — compare it to the first montage, that is real progress. But two have gone cartoonish, one composition is broken outright with the head sliding out of frame, and one character came back looking like he belonged to a completely different part of the world than the one I had described.

That last one is worth a paragraph on its own.

The setting problem, which is really a specification problem

One character came back visibly not matching the world I had asked for.

The instinct is to read that as the model making a mistake about geography. It is not, and the framing matters. The model does not hold a setting. It holds an average. If you write "medieval merchant" and leave it there, you get whatever the middle of its training data thinks a medieval merchant looks like — which is an aggregate of centuries of illustration, film, book covers and other people's fantasy art, and which will sometimes match the world you are building and sometimes not.

The fix is not to police the output. The fix is to specify each character's origin deliberately, in the prompt, as a decision you actually made.

And here is the part I did not expect: doing that produced a wider and more considered range than leaving it to chance ever did. The finished roster spans thirty distinct stated origins — Norwegian, Levantine, Yoruba, Sinti Roma, Amazigh, Kazakh, Basque, Igbo, Sicilian, Somali and more — because a trading world full of roaming tradespeople should look like that, and because once you are writing each person down you make each of those calls on purpose. Left to the average, I would have got a much narrower and much blander deck.

Specification did not constrain the range. Vagueness did.

Failure six: the finger fix specified a pose

This is my favourite one, because I would never have predicted it and it is invisible in any single image.

Hands are the classic AI art problem. Six fingers, fused fingers, a thumb where no thumb belongs. So the prompt carried a hand-safety instruction, phrased roughly as: both hands held well apart from each other and clear of the clothing.

Sensible. Hands that are separated and unobstructed are much less likely to fuse into each other or melt into a sleeve.

Then I rendered sixteen finished cards and looked at them together.

Sixteen finished character cards laid out in a grid, in which nearly every character stands with both arms held out from the body and palms open, producing an identical welcoming pose across the whole set

Every single one of them is doing the same thing with their arms.

There is only one way to satisfy "both hands held well apart from each other and clear of the clothing," and that way is arms extended, palms open. I had not written a safety constraint. I had written a pose, without realising it, and the model obeyed exactly. Seven of sixteen came back in an identical welcome; the rest are close variants of it.

The real anti-fusion rule turns out to be much narrower than what I wrote: what causes finger disasters is gripping an articulated prop. Hands clasped, folded, in pockets, or flat on a surface are all completely safe, and none of them dictate a stance.

A comparison strip of regenerated portraits in which two characters now stand with their hands clasped or folded at the waist rather than arms extended, in varied settings

Same safety, no chorus line.

The thing that actually fixed it

The working pipeline, once the dust settled, has four properties. None of them are clever. All of them are just discipline.

It runs in two stages. First a cheap square head-and-shoulders study on a plain background, where the only thing being judged is the face. Then, once a head is approved, a full portrait — using that approved head as the identity input. The face-copying behaviour of the edit endpoint, which was the entire bug in failure one, is the mechanism that holds identity here. The same behaviour is a disaster in one position and load-bearing in the other.

Every character is written down properly before anything is generated. Ancestry, age, build, a specific face, one named asymmetric feature, camera angle, and expression. Not a type — a person. Roughly 700 to 900 characters of specification each, and the shared house-style preamble is the only text any two of them have in common.

Constraints are stated positively. This API has no negative prompt field at all. You cannot pass six fingers, extra limbs and call it handled. Every constraint has to be phrased as something the image does contain — "all four corners are clean painted material" rather than "no watermark." Honestly this is better discipline anyway, and I would keep it even where a negative field exists.

The prompt is saved next to the image. Every single one. This is the least glamorous item on the list and it is the reason this post exists at all — I could go back and read exactly what produced each face. For the early failed takes I did not do this, which is why the worst prompts of the day are gone forever and I am describing them from memory.

Here is where it ended up.

A montage of forty finished character head studies, all in the same painted style but every face clearly a different individual, spanning a wide range of ages, builds, origins and expressions

Forty of the fifty-two, in one style, and every one of them a different person.

What I actually learned

The failures were invisible one image at a time. This is the real lesson and everything else is detail. Every single problem in this post — the sibling faces, the mole run, the identical pose, the corner monogram — looks fine when you review one picture. You cannot see a repeated face until you see it next to the other repeated face. You cannot see a chorus line in a single portrait. Reviewing images as they arrive, one by one, will pass all of it.

So the pipeline now ends in montages, not images. A contact sheet of everything. A face-only crop montage, which is what catches siblings. A corner montage — just the bottom-right of every candidate, tiled — which is what catches signatures. That corner sheet takes about four seconds to scan and it found things I had already approved.

Fixes travel in chains. Style reference caused siblings. Escaping signatures cost me the good painter. Forcing divergence caused caricature. The hand-safety line caused the pose. If you change one thing in a generation pipeline, the thing to do is not admire the fix — it is to go and look at what else moved.

Version numbers are not quality rankings. Two of the engines I had available were, by name, version "2" and version "pro." The one with the higher number is the cheaper, faster tier. It is not an upgrade. This bit me in an unusually modern way: I had two AI agents working on this, both given the instruction "use the latest engine," and they resolved it differently — one took the newer-sounding version number and got the worse painter. Which is a completely defensible reading of what I said. The instruction was ambiguous and I was the one who wrote it.

And the lesson was already written down. This is the part I find genuinely funny. We keep an internal document of image-generation traps. It already contained this line, from a previous project:

Rotate it for a set — a fixed anchor leaks one face into every card.

That is failure one. Precisely. Written down, in advance, by me, and I walked directly into it anyway — because knowing a trap exists and verifying that you did not fall into it are two entirely different activities, and only one of them takes discipline.

I have form here, for what it's worth. While fact-checking the post about how Pairdle got built I discovered that the board generator had been silently throwing away seven of every eight words for six months, behind a comment that said it was doing the opposite. Same shape of mistake: the thing was measurable the whole time, and nobody measured it.

The tools are extraordinary. Genuinely — the finished heads in that last montage are better than anything I could commission at the budget of a small studio, and they exist because of an afternoon and about forty rejected images rather than an afternoon and a hiring process.

But nothing about the job of art direction went away. It moved. It used to be drawing. Now it is specifying precisely, and then inspecting ruthlessly — and of those two, the inspection is the part that is easy to skip and expensive to have skipped.

If you want to see where these end up, they are the tradespeople in Epic Words. And if you would rather read about a puzzle than a pipeline, Pairdle is the game that started all of this.

Share this post