01 The drift
My assistant has a face. Not because it needed one, but because a bot that lives in the corner of the screen and reports on your mail is easier to build when it looks like someone rather than a notification.
The first attempt failed the way these attempts fail. I described the character in each prompt, by hand. By the fifth image the horns had changed their curve, the glasses had gone from round to oval, and the crescent pendant had quietly disappeared. Every frame was fine on its own. Together they were four different people.
The fix turned out to be dull and reliable: stop describing her by hand at all.
02 The canon, in one file
The whole appearance lives between two markers in a single file. The generation scripts splice that block onto the front of every prompt, so the description exists in exactly one place. Change it there and it changes everywhere — including in clips generated six months later.
<!-- canon:start -->
Horns: two curved dark crimson horns with fine ridges,
growing from the top of the head just in front of the bun,
curving up and slightly outward.
Glasses: thin dark round wire-frame glasses.
Jewelry: black choker with a small silver crescent-moon
pendant; silver crescent-moon drop earrings.
Nothing else: no other accessories, no tattoos, no wings,
no tail, no hat, no other characters.
<!-- canon:end -->The last line earns its place. Without an explicit prohibition the model adds things — and the longer the prompt, the more eagerly it adds them. Wings show up. A second character wanders into frame. Naming what must not appear turned out to matter as much as naming what must.
03 The sheet of base poses
Text holds the details but not the proportions. Written description alone gives you the right horns on a face that is subtly the wrong shape. So the canon is only half the input: the other half is a reference frame.
Nine poses were cut from a single sheet, shot in one pass. Because they came out of one generation, there is no drift between them at all — same light, same framing, same face. Any of them can serve as the reference for a new clip.
04 Clips through the neutral
A companion that only holds still is a picture. To move, it needs clips — and clips need to follow each other without a jolt. The trick is one rule: every one-shot clip begins and ends on the same neutral pose.
That single constraint removes the whole transition problem. Any clip can follow any other, in any order, and no in-between frames are needed. Longer states — thinking, working, sleeping — are cut into three pieces through the same frame.
The generation itself is unremarkable and cheap, which was the point. The clips go through OpenRouter's video API on Kling v3.0 std: five seconds, square, 720p, about $0.42 a clip, with the neutral frame handed in as both the first and the last frame. The neutral frame itself was drawn once for about four cents. Only the action goes into the clip prompt — the appearance is already there, spliced in from the canon.
05 What did not go to plan
The spec said the background must be flat bright green, #00FF00, evenly lit, no shadows. The model agreed, and then produced lime. Not the green that was asked for — near it, off by enough that a keyer tuned to #00FF00 left a rim of fringe on every strand of hair.
Arguing with the prompt was the obvious move and the wrong one. The keying script now reads the background colour out of the corner of the first frame and keys against whatever it actually finds. The model is allowed to be approximately green.
That is the shape of most of this work. The interesting decisions were not about prompts at all: put the description in one place, shoot the references in one pass, pick a single frame everything returns to, and measure what came out instead of insisting on what was asked for.