If you have ever generated an illustrated story and watched the main character change species somewhere around page four, the problem is almost never the model. It is that the description was not doing enough work.
Every page in an AI-illustrated book is drawn independently. The picture on page seven has no memory of page one. Whatever the character looks like has to be rebuilt from scratch each time, from words alone. So the question is not "how do I describe this character" but "what is the shortest description that rebuilds the same character every single time".
One: the silhouette
Shape reads before detail. Before a reader registers eye colour they register outline: how tall, how round, what is on the head, what sticks out. A bear cub with tiny rounded ears and a scarf has a silhouette. "A cute bear" does not.
So lead with the shape words. Size, body proportions, hair or fur arrangement, anything that changes the outline — two puff ponytails, long ears, a wide-brimmed hat. These do more for consistency than any amount of careful facial description, because they survive being drawn small, from behind, or at an angle.
Two: a signature object
Give the character one thing that is always with them, and always the same colour. A red knitted scarf. A blue tin pail. Purple slippers. It sounds like a small thing and it is disproportionately effective: it acts as a label the reader can find in every frame, and it gives the image model an anchor that is easy to get right.
The object should be simple enough to draw consistently — a solid colour, a recognizable shape, nothing with fine detail or text on it. And it should ideally be something the story can use. An object that only exists for consistency feels like a prop. An object the character actually carries, drops, looks for, or hands to someone else becomes part of the book.
Three: one repeatable expression
Not a range of expressions — one. The way this particular character looks when they are thinking, or delighted, or unsure. Write it once, concretely: "one paw lifted", "head tilted kindly", "eyebrows up and mouth slightly open".
You will not use it on every page, but having it written down means that when a page calls for the character to be curious, you already know what curious looks like on them. That is what makes a character feel like a person rather than a description that happens to recur.
Write it once, repeat it verbatim
Then the mechanical part, and this is the part most people get wrong: once you have that description, repeat it word for word in every single page prompt where the character appears. Not a shortened version. Not "the bear from before". The same words, every time, however long it makes the prompt.
It feels absurdly redundant. It is also the entire mechanism. Any variation in the words produces variation in the picture, and referring back to a previous page refers to something the model cannot see.