The prompt structure that survives a 40-shot film
One good prompt gets you one good image. A film needs forty of them that agree with each other. This is the five-layer structure we use in Huski Studio, with the real screens and prompts from a 24-shot short.

Most first AI films die the same way. Shot one looks great. By shot six the jacket has changed colour. By shot twelve the kitchen has a second window. By shot twenty you're pasting a 900-word description of everything into every prompt, and the model is treating it as a mood board rather than a set of instructions.
The problem usually isn't the prompts themselves. It's the assumption that a film is a sequence of prompts.
It helps to think of a film as a small database instead. There are a handful of facts (what the world looks like, who's in it, where each scene happens) and forty views of those facts. The prompt for any one shot should be assembled from the facts, not written from scratch each time. Once that structure is in place, consistency stops depending on luck. Without it, no amount of clever prompt-writing will rescue shot thirty.
This article walks through the structure we use, with the actual screens and prompt text from Second Season, a 24-shot short made in Huski Studio with eight characters and five locations. It's one of the demo projects, so you can open it read-only from the Projects page and click through every decision described here. Nothing about the structure is specific to our tool. If you're prompting Kling, Veo or Seedance by hand, you can keep the same ledger in a spreadsheet. It just takes longer.
Why prompt soup fails at shot six
Three things go wrong when every shot re-describes the whole world.
Words are lossy. "A warm brown leather jacket", "a brown jacket" and "a leather coat" are three different garments to an image model. Describe the jacket forty times and you'll describe it forty slightly different ways, and get forty slightly different jackets back.
Attention is finite. Everything in the prompt competes for the model's attention. Two hundred words of wardrobe and set dressing water down the twenty words that say what this particular shot is about.
Contradictions pile up. The style says golden hour, the scene says night, the shot says harsh fluorescent light. The model picks one, more or less at random, and you're rolling that die forty times.
The maths is unforgiving. If each shot independently matches the last one 90% of the time, the odds that all forty match are 0.9⁴⁰, which is about 1.5%. Even at a heroic 98% per shot, the film holds together less than half the time. You can't get there by trying harder on each shot. You get there by taking the coin out of the model's hands.
The structure: five layers, each written once
Every fact about the film lives in exactly one layer. Each layer is written once and then reused, unchanged, by everything below it.
| Layer | Owns | Written | Reused by |
|---|---|---|---|
| 1. Style | Medium, lighting character, palette, lens feel | Once per film | Every keyframe, every clip |
| 2. Cast | Who each character is: face, build, design, wardrobe by state, physical scale | Once per character, plus one per wardrobe change | Every shot they appear in |
| 3. World | Each location: the empty environment and its time of day | Once per location | Every shot set there |
| 4. Shot | What's different about this frame: framing, who's visible, action, expression, props in hand | Once per shot | That shot's keyframe |
| 5. Motion | What moves and what's said | Once per shot | That shot's clip |
Three rules keep the layers honest:
- A fact lives in one layer only. The tomato's bump is a Cast fact. It never shows up in a shot description.
- A lower layer never restates a higher one. The shot says "Main Tomato", not "a plump red tomato with a bump on one side".
- Where a picture can replace words, use the picture. A reference image beats any adjective. The layers above the shot are really images (character sheets and location references) with some text attached, not the other way round.
In Huski Studio these layers are separate steps: Project Settings, Character Design and Character States, Scene Design, Shots Breakdown, then Keyframes and Clip Generation. We built the workflow that way so the structure can't be skipped.

Layer 1 — Style: traits, not names
The style layer is one paragraph about what the camera and the light do. Here is Second Season's, exactly as stored:
Modern 3D CGI animation: rounded stylized proportions, soft subsurface skin shading, big expressive eyes, appealing exaggerated features, polished cinematic lighting, vibrant colorful designs. Not photorealistic, not live-action, not a photograph.
Two things are worth pointing out.
It names traits, not studios. "In the style of" a particular animation studio isn't a trait. It's a pointer into the model's memory, and different models remember different things. It also trips content filters on some providers and gets quietly ignored on others. Traits carry over between models; names don't. Huski Studio strips studio and franchise names out of style directives automatically and keeps the visual traits, because when we measured it, a name added no fidelity and cost us refusals.
It says nothing about time of day. This is the mistake we see most often. A style that mentions golden hour makes every night scene golden. The character of the light (soft, polished, high-contrast) belongs here. The hour belongs to the location, in Layer 3. The pickers in Project Settings preview every style at sunny, cloudy and night for exactly this reason. If a directive only looks right at one of them, a scene description has crept into it.

Layer 2 — Cast: a sheet, and traits you could check
The cast layer starts as text and ends as an image.
The text is an appearance written as a checklist someone could verify on a frame. Not "a charming, expressive heirloom tomato", because nobody can check for charming. Here's Main Tomato's appearance as it stands in Character Design:
A large, plump heirloom tomato with bright, glossy red skin. He has a prominent, irregular, asymmetrical bump on one side of his body in a slightly darker color, making him lopsided and less round than a normal tomato. His face features very large, cartoonish, expressive black eyes with bright white reflections and a small, highly mobile mouth capable of stretching into a broad smile or puckering in worry. A five-pointed, star-shaped green calyx stem sits jauntily on top of his head.
Asymmetrical bump, one side, darker. Five-pointed calyx. Large black eyes with white reflections. Three to six traits like that, things you can point at, are worth more than a paragraph of adjectives. The bump, by the way, came straight from the author's story text ("an irregular bump on the side of his head"). A checkable detail in the source becomes a checkable detail in every frame.
That text is rendered once into a three-view reference sheet, and from that point on the sheet does the describing. The appearance paragraph never gets pasted into a shot prompt again. Instead, the shot prompt attaches the sheet with an instruction to take the character's face, hair, eyewear, build and clothing from this sheet only.

Three refinements help the cast layer survive a long film.
States. A character who changes clothes gets one sheet per look, and each shot is assigned to one of them. Ruth wears a head scarf at the market and an apron in her kitchen. Rather than describing the apron in six shot prompts, the apron look has its own sheet and owns shots 13, 15, 16, 17 and 18.


Scale. Image models can't make sense of "ten centimetres", so numbers do nothing. A tomato-sized character gets drawn person-sized the moment he's alone in a frame, because the model fills the frame with its subject. Huski Studio asks for a size class instead (hand-held, tabletop, child-sized, human-sized, on up to vehicle- and building-sized) and turns it into a comparison the model can act on: far smaller than any person; never enlarge to fill the frame. In our tests those few words took correct-scale renders from 2 out of 12 to 18 out of 20.
Roles. Everyone who appears on screen gets a sheet, extras included. Voice-only characters, like a narrator or a voice on the phone, get nothing visual at all, so they can never accidentally show up in a frame.
Layer 3 — World: one empty reference per location
Each location gets one wide image, rendered empty, with its time of day. Here's the description for Inside Farmer's Crate:
The cramped, dark interior of a rustic wooden vegetable crate. Thin parallel wooden slats form the walls, allowing narrow beams of golden daylight to filter through. Completely empty.
TIME OF DAY: afternoon
"Completely empty" is doing real work there. The set is rendered without the cast so that nobody gets drawn twice, once from the set and once from the sheet. When a location needs more than one angle, the extra views are generated from that first image, so the slats, the posters on the wall and the leaf on the floor stay where they were.

Every shot set in that location then uses its view as the background. The keyframe isn't asked to imagine the crate. It's handed the crate and asked to put the cast in it.

Layer 4 — Shot: two to four sentences about this frame only
With style, cast and world settled above it, the shot layer gets short, and short is the point. Here is the complete keyframe prompt for shot 5, exactly as stored:
Main Tomato sits nestled at frame center, packed closely among several other silent, still tomatoes, facing toward the foreground with his eyes open and directed slightly upward, as a sharp, thin bar of golden light cuts horizontally across his face. Main tomato has a scared facial expression.
Look at what's missing. No red skin. No bump. No crate. No "same as the previous shot". All of that lives upstairs. What's here is only what this frame owns:
- Who's visible, by name. "Main Tomato", never a description.
- Where they are in the frame. "At frame center." Positions are described relative to the frame, not the room; "by the door" makes the model guess where the door is.
- One pose per character. Give one character two actions and the model tends to draw them twice.
- Expression, which is the one thing about a face the shot is allowed to say.
- How the light falls in this frame. A bar across his face. That's different from what the light is, which belongs to the style and the location.
In Huski Studio the shot is a form rather than a paragraph. Framing, camera angle and movement, duration, characters in the shot, props in hand and dialogue are all separate fields, and the "what happens" text is deliberately not allowed to mention physical traits, the environment or camera moves. We added that rule after watching what happens without it: a trait echoed in the shot text competes with the sheet, and the character drifts.


What the model actually receives
Here's the shape of the prompt Huski Studio puts together for that keyframe, abridged. Everything except the last block came from a layer above the shot:
Generate a film keyframe. Render the ENTIRE frame, characters included, in this visual style: <the style directive, Layer 1>
SCENE: Use this environment as the background setting. ← the location view, Layer 3, attached as an image
CHARACTER PLACEMENT RULES:
- Place EXACTLY 1 named character in this shot: Main Tomato
- Each named character appears ONLY ONCE — do NOT duplicate, mirror, or reflect…
SCALE RULES (only when a character is not human-sized) ← Layer 2
HELD PROP RULES / SET PROP RULES (only when props are in the shot)
ATTACHED IMAGES — the input images are attached in this exact order:
1. SCENE …
2. CHARACTER REFERENCE SHEET for Main Tomato — three views of this one
character; take the design, colors, and proportions from THIS sheet only
SHOT DESCRIPTION: <the two-to-four sentences above, Layer 4>
IMPORTANT: The final image must look like a candid still captured mid-scene…
NOT a movie poster, promotional shot, or magazine cover…
The order of the attached images matters more than any sentence in that prompt. They go in from most important for identity to least: the scene first, then one sheet per visible character, then (only when the previous shot shares a character with this one) the previous keyframe for continuity, then props. The continuity frame goes after the sheets on purpose. Put it first and the model will happily give a new character the face of whoever was standing there last shot.

Layer 5 — Motion: what moves and what is said
The keyframe becomes the first frame of the clip, so the clip prompt only describes what moves and what's said. The look is already in the frame, and describing it again just invites the video model to change it.
Shot 5's clip prompt, exactly as stored:
Main Tomato sits nestled at frame center, packed closely among several other silent, still tomatoes. He faces the foreground, looking up and blinking slowly. A single sharp, thin bar of golden light cuts horizontally across his face, leaving his lower body in deep shadow. Camera slowly and smoothly pushes in toward Main Tomato, who remains still, gradually tightening the frame. As the camera moves forward, the sound of creaking wooden crate planks rises, followed by a faint road vibration rattle.
No music.
Every clip prompt follows the same shape: the action in order (camera move, what the characters do, sound effects), a blank line, then the dialogue, with each line marked as on-screen or off-screen and given a delivery cue, and finally the words No music. The score is composed later, and a video model's own music would fight it. Whether a speaker is on screen is decided from the shot's cast list rather than left to the model, and any line the model invents that isn't in the script gets removed before generation.

What twenty-four shots look like when the structure holds
Here's Main Tomato across the film: on the vine, in the crate, on the market table, in Ruth's hand, on the kitchen sill. Six shots, four locations, three kinds of light. Same design, same calyx, same bump wherever he's big enough in frame to show it.

None of that came from re-describing him. It came from attaching the sheet to every shot, rendering the sets empty, and keeping the shot text from restating either.
The structure also gives you something prompt soup never can, which is a place to fix things. When shot 14 comes out wrong, the question isn't "what do I add to the prompt?" but "which layer is wrong?" A wrong outfit is a state assignment. A wrong room is a location reference. A character drawn twice is an action with two poses. Fix the one layer and the fix reaches every shot that shares it. Huski Studio flags the affected shots as out of date automatically; in a spreadsheet you'd flag them yourself.
The checklist
Whatever tool you're using:
- One style directive for the whole film. Traits only. No studio names, no franchise names, no time of day.
- Per character: three to six checkable traits, rendered once into a reference sheet. One extra sheet per wardrobe change, with the shots assigned to it. A size class for anything that isn't human-sized.
- Per location: one empty wide reference image with its time of day. Derive every other angle from it.
- Per shot: two to four sentences. Names, not descriptions. One pose per character. Positions relative to the frame. Expression is allowed; appearance, environment and camera moves aren't.
- Per clip: the keyframe, plus only what moves and what's said. End with No music.
- Never restate a higher layer in a lower one. If you feel the urge to, the higher layer is missing a fact. Add it there instead.
The prompt that survives forty shots isn't really a prompt at all. It's a bill of materials, plus a short note about what's different today.