HUSKI STUDIO

The prompt structure that survives a 40-shot film

One good prompt gets you one good image. A film needs forty of them that agree with each other. This is the five-layer structure we use in Huski Studio, with the real screens and prompts from a 24-shot short.

All 24 shots of the short film Second Season laid out as a contact sheet — the same tomato, the same grandmother, the same kitchen, shot after shot

Most first AI films die the same way. Shot one looks great. By shot six the jacket has changed colour. By shot twelve the kitchen has a second window. By shot twenty you're pasting a 900-word description of everything into every prompt, and the model is treating it as a mood board rather than a set of instructions.

The problem usually isn't the prompts themselves. It's the assumption that a film is a sequence of prompts.

It helps to think of a film as a small database instead. There are a handful of facts (what the world looks like, who's in it, where each scene happens) and forty views of those facts. The prompt for any one shot should be assembled from the facts, not written from scratch each time. Once that structure is in place, consistency stops depending on luck. Without it, no amount of clever prompt-writing will rescue shot thirty.

This article walks through the structure we use, with the actual screens and prompt text from Second Season, a 24-shot short made in Huski Studio with eight characters and five locations. It's one of the demo projects, so you can open it read-only from the Projects page and click through every decision described here. Nothing about the structure is specific to our tool. If you're prompting Kling, Veo or Seedance by hand, you can keep the same ledger in a spreadsheet. It just takes longer.

Why prompt soup fails at shot six

Three things go wrong when every shot re-describes the whole world.

Words are lossy. "A warm brown leather jacket", "a brown jacket" and "a leather coat" are three different garments to an image model. Describe the jacket forty times and you'll describe it forty slightly different ways, and get forty slightly different jackets back.

Attention is finite. Everything in the prompt competes for the model's attention. Two hundred words of wardrobe and set dressing water down the twenty words that say what this particular shot is about.

Contradictions pile up. The style says golden hour, the scene says night, the shot says harsh fluorescent light. The model picks one, more or less at random, and you're rolling that die forty times.

The maths is unforgiving. If each shot independently matches the last one 90% of the time, the odds that all forty match are 0.9⁴⁰, which is about 1.5%. Even at a heroic 98% per shot, the film holds together less than half the time. You can't get there by trying harder on each shot. You get there by taking the coin out of the model's hands.

The structure: five layers, each written once

Every fact about the film lives in exactly one layer. Each layer is written once and then reused, unchanged, by everything below it.

LayerOwnsWrittenReused by
1. StyleMedium, lighting character, palette, lens feelOnce per filmEvery keyframe, every clip
2. CastWho each character is: face, build, design, wardrobe by state, physical scaleOnce per character, plus one per wardrobe changeEvery shot they appear in
3. WorldEach location: the empty environment and its time of dayOnce per locationEvery shot set there
4. ShotWhat's different about this frame: framing, who's visible, action, expression, props in handOnce per shotThat shot's keyframe
5. MotionWhat moves and what's saidOnce per shotThat shot's clip

Three rules keep the layers honest:

  1. A fact lives in one layer only. The tomato's bump is a Cast fact. It never shows up in a shot description.
  2. A lower layer never restates a higher one. The shot says "Main Tomato", not "a plump red tomato with a bump on one side".
  3. Where a picture can replace words, use the picture. A reference image beats any adjective. The layers above the shot are really images (character sheets and location references) with some text attached, not the other way round.

In Huski Studio these layers are separate steps: Project Settings, Character Design and Character States, Scene Design, Shots Breakdown, then Keyframes and Clip Generation. We built the workflow that way so the structure can't be skipped.

The 3-view character sheet for Main Tomato beside the appearance text that generated it
Layer 2. The appearance is written once and turned into a three-view sheet. From then on, the sheet does the describing.

Layer 1 — Style: traits, not names

The style layer is one paragraph about what the camera and the light do. Here is Second Season's, exactly as stored:

prompt
Modern 3D CGI animation: rounded stylized proportions, soft subsurface skin shading, big expressive eyes, appealing exaggerated features, polished cinematic lighting, vibrant colorful designs. Not photorealistic, not live-action, not a photograph.

Two things are worth pointing out.

It names traits, not studios. "In the style of" a particular animation studio isn't a trait. It's a pointer into the model's memory, and different models remember different things. It also trips content filters on some providers and gets quietly ignored on others. Traits carry over between models; names don't. Huski Studio strips studio and franchise names out of style directives automatically and keeps the visual traits, because when we measured it, a name added no fidelity and cost us refusals.

It says nothing about time of day. This is the mistake we see most often. A style that mentions golden hour makes every night scene golden. The character of the light (soft, polished, high-contrast) belongs here. The hour belongs to the location, in Layer 3. The pickers in Project Settings preview every style at sunny, cloudy and night for exactly this reason. If a directive only looks right at one of them, a scene description has crept into it.

The Visual Style picker, showing the same style rendered sunny, cloudy and at night
A style has to survive three times of day. If it only works at one, the hour has leaked into it.

Layer 2 — Cast: a sheet, and traits you could check

The cast layer starts as text and ends as an image.

The text is an appearance written as a checklist someone could verify on a frame. Not "a charming, expressive heirloom tomato", because nobody can check for charming. Here's Main Tomato's appearance as it stands in Character Design:

prompt
A large, plump heirloom tomato with bright, glossy red skin. He has a prominent, irregular, asymmetrical bump on one side of his body in a slightly darker color, making him lopsided and less round than a normal tomato. His face features very large, cartoonish, expressive black eyes with bright white reflections and a small, highly mobile mouth capable of stretching into a broad smile or puckering in worry. A five-pointed, star-shaped green calyx stem sits jauntily on top of his head.

Asymmetrical bump, one side, darker. Five-pointed calyx. Large black eyes with white reflections. Three to six traits like that, things you can point at, are worth more than a paragraph of adjectives. The bump, by the way, came straight from the author's story text ("an irregular bump on the side of his head"). A checkable detail in the source becomes a checkable detail in every frame.

That text is rendered once into a three-view reference sheet, and from that point on the sheet does the describing. The appearance paragraph never gets pasted into a shot prompt again. Instead, the shot prompt attaches the sheet with an instruction to take the character's face, hair, eyewear, build and clothing from this sheet only.

Main Tomato's three-view reference sheet: front, three-quarter and profile
The sheet is the single source of truth for a character. Every shot references this image, not a description of it.

Three refinements help the cast layer survive a long film.

States. A character who changes clothes gets one sheet per look, and each shot is assigned to one of them. Ruth wears a head scarf at the market and an apron in her kitchen. Rather than describing the apron in six shot prompts, the apron look has its own sheet and owns shots 13, 15, 16, 17 and 18.

Ruth's two states — head scarf and apron — each with its own three-view sheet and its own list of shots
Each look is a sheet plus the shots that use it. Nobody types the word 'apron' into a shot.
Ruth in four shots across two locations: head scarf at the market and in the backyard, apron in the kitchen
Shots 10, 13, 18 and 20. The dress is the same in all four; only the layer on top changes.

Scale. Image models can't make sense of "ten centimetres", so numbers do nothing. A tomato-sized character gets drawn person-sized the moment he's alone in a frame, because the model fills the frame with its subject. Huski Studio asks for a size class instead (hand-held, tabletop, child-sized, human-sized, on up to vehicle- and building-sized) and turns it into a comparison the model can act on: far smaller than any person; never enlarge to fill the frame. In our tests those few words took correct-scale renders from 2 out of 12 to 18 out of 20.

Roles. Everyone who appears on screen gets a sheet, extras included. Voice-only characters, like a narrator or a voice on the phone, get nothing visual at all, so they can never accidentally show up in a frame.

Layer 3 — World: one empty reference per location

Each location gets one wide image, rendered empty, with its time of day. Here's the description for Inside Farmer's Crate:

prompt
The cramped, dark interior of a rustic wooden vegetable crate. Thin parallel wooden slats form the walls, allowing narrow beams of golden daylight to filter through. Completely empty.
TIME OF DAY: afternoon

"Completely empty" is doing real work there. The set is rendered without the cast so that nobody gets drawn twice, once from the set and once from the sheet. When a location needs more than one angle, the extra views are generated from that first image, so the slats, the posters on the wall and the leaf on the floor stay where they were.

Scene Design: the empty set views for shots 5 and 6 inside the crate, both derived from the location's reference image
Layer 3. Two shots, one location, one reference image. The views are empty on purpose.

Every shot set in that location then uses its view as the background. The keyframe isn't asked to imagine the crate. It's handed the crate and asked to put the cast in it.

Left: the empty crate reference. Right: the finished keyframe for shot 5, with Main Tomato placed in the same crate
Reference on the left, keyframe on the right. Same slats, same sign, same leaf on the floor.

Layer 4 — Shot: two to four sentences about this frame only

With style, cast and world settled above it, the shot layer gets short, and short is the point. Here is the complete keyframe prompt for shot 5, exactly as stored:

prompt
Main Tomato sits nestled at frame center, packed closely among several other silent, still tomatoes, facing toward the foreground with his eyes open and directed slightly upward, as a sharp, thin bar of golden light cuts horizontally across his face. Main tomato has a scared facial expression.

Look at what's missing. No red skin. No bump. No crate. No "same as the previous shot". All of that lives upstairs. What's here is only what this frame owns:

  • Who's visible, by name. "Main Tomato", never a description.
  • Where they are in the frame. "At frame center." Positions are described relative to the frame, not the room; "by the door" makes the model guess where the door is.
  • One pose per character. Give one character two actions and the model tends to draw them twice.
  • Expression, which is the one thing about a face the shot is allowed to say.
  • How the light falls in this frame. A bar across his face. That's different from what the light is, which belongs to the style and the location.

In Huski Studio the shot is a form rather than a paragraph. Framing, camera angle and movement, duration, characters in the shot, props in hand and dialogue are all separate fields, and the "what happens" text is deliberately not allowed to mention physical traits, the environment or camera moves. We added that rule after watching what happens without it: a trait echoed in the shot text competes with the sheet, and the character drifts.

Shots Breakdown for shot 5: the 'what happens in this shot' text, performance notes and the keyframe
Layer 4. This text drives the keyframe and the clip. It's short because everything else already has a home.
The structured fields under the shot text: location, camera angle and movement, duration, characters in shot, props, dialogue
The rest of the shot is fields, not prose: location, camera, duration, who's in it, what they're holding.

What the model actually receives

Here's the shape of the prompt Huski Studio puts together for that keyframe, abridged. Everything except the last block came from a layer above the shot:

prompt
Generate a film keyframe. Render the ENTIRE frame, characters included, in this visual style: <the style directive, Layer 1>

SCENE: Use this environment as the background setting.          ← the location view, Layer 3, attached as an image

CHARACTER PLACEMENT RULES:
- Place EXACTLY 1 named character in this shot: Main Tomato
- Each named character appears ONLY ONCE — do NOT duplicate, mirror, or reflect…

SCALE RULES (only when a character is not human-sized)         ← Layer 2
HELD PROP RULES / SET PROP RULES (only when props are in the shot)

ATTACHED IMAGES — the input images are attached in this exact order:
1. SCENE …
2. CHARACTER REFERENCE SHEET for Main Tomato — three views of this one
   character; take the design, colors, and proportions from THIS sheet only

SHOT DESCRIPTION: <the two-to-four sentences above, Layer 4>

IMPORTANT: The final image must look like a candid still captured mid-scene…
NOT a movie poster, promotional shot, or magazine cover…

The order of the attached images matters more than any sentence in that prompt. They go in from most important for identity to least: the scene first, then one sheet per visible character, then (only when the previous shot shares a character with this one) the previous keyframe for continuity, then props. The continuity frame goes after the sheets on purpose. Put it first and the model will happily give a new character the face of whoever was standing there last shot.

The Keyframes step showing the elements that went into shot 5: the scene view and Main Tomato's sheet
The reference stack for one keyframe. Add a character to the shot and their sheet joins the stack. Nothing else changes.

Layer 5 — Motion: what moves and what is said

The keyframe becomes the first frame of the clip, so the clip prompt only describes what moves and what's said. The look is already in the frame, and describing it again just invites the video model to change it.

Shot 5's clip prompt, exactly as stored:

prompt
Main Tomato sits nestled at frame center, packed closely among several other silent, still tomatoes. He faces the foreground, looking up and blinking slowly. A single sharp, thin bar of golden light cuts horizontally across his face, leaving his lower body in deep shadow. Camera slowly and smoothly pushes in toward Main Tomato, who remains still, gradually tightening the frame. As the camera moves forward, the sound of creaking wooden crate planks rises, followed by a faint road vibration rattle.

No music.

Every clip prompt follows the same shape: the action in order (camera move, what the characters do, sound effects), a blank line, then the dialogue, with each line marked as on-screen or off-screen and given a delivery cue, and finally the words No music. The score is composed later, and a video model's own music would fight it. Whether a speaker is on screen is decided from the shot's cast list rather than left to the model, and any line the model invents that isn't in the script gets removed before generation.

Clip Generation for shot 5: the generated clip above its keyframe, with the description beside it
Layer 5. The clip inherits the keyframe. The prompt only adds motion and sound.
Shot 5, as delivered. A push-in on a tomato that was designed once and drawn twenty-four times.

What twenty-four shots look like when the structure holds

Here's Main Tomato across the film: on the vine, in the crate, on the market table, in Ruth's hand, on the kitchen sill. Six shots, four locations, three kinds of light. Same design, same calyx, same bump wherever he's big enough in frame to show it.

Main Tomato in shots 2, 5, 8, 11, 13 and 16
Six shots, four locations. The character was described once.

None of that came from re-describing him. It came from attaching the sheet to every shot, rendering the sets empty, and keeping the shot text from restating either.

The structure also gives you something prompt soup never can, which is a place to fix things. When shot 14 comes out wrong, the question isn't "what do I add to the prompt?" but "which layer is wrong?" A wrong outfit is a state assignment. A wrong room is a location reference. A character drawn twice is an action with two poses. Fix the one layer and the fix reaches every shot that shares it. Huski Studio flags the affected shots as out of date automatically; in a spreadsheet you'd flag them yourself.

The checklist

Whatever tool you're using:

  1. One style directive for the whole film. Traits only. No studio names, no franchise names, no time of day.
  2. Per character: three to six checkable traits, rendered once into a reference sheet. One extra sheet per wardrobe change, with the shots assigned to it. A size class for anything that isn't human-sized.
  3. Per location: one empty wide reference image with its time of day. Derive every other angle from it.
  4. Per shot: two to four sentences. Names, not descriptions. One pose per character. Positions relative to the frame. Expression is allowed; appearance, environment and camera moves aren't.
  5. Per clip: the keyframe, plus only what moves and what's said. End with No music.
  6. Never restate a higher layer in a lower one. If you feel the urge to, the higher layer is missing a fact. Add it there instead.

The prompt that survives forty shots isn't really a prompt at all. It's a bill of materials, plus a short note about what's different today.

Written by

Huski Studio

Roll camera

Put the structure to work on your own story.

Every plan runs the full crew — cast, shots, keyframes, clips and the final cut.

Start your film — free