Tool tutorial · Xiaohu's take

Higgsfield shot a 110-minute feature film using real actors' likenesses and AI, then published the entire production manual and prompt methods

Higgsfield Studio's write-up of The Cully Hill Boys covers the full pipeline — assets, prompts, performance, lip-sync — and even makes its three private Claude skills available to download.
The 60-second read
  • A 1-hour-54-minute AI feature film, with the entire production playbook and three custom Claude skills released to the public.
  • The whole manual exists to solve one problem: video models have no memory. "Keeping the same person the same person" is the entire job.
  • The most valuable lessons are the counterintuitive taboos: no "studio" in character sheets, no heads in full-body shots, and never write "don't" in a prompt.
Source material is the studio's own project write-up. Runtime, asset count, and results are all self-reported with no third-party verification. Check notes and gaps are added by this site. The original page is dynamic; scrapers couldn't get the text, so the full transcript was provided by a reader.
Opening

An AI feature film — and the entire production manual, released

For the past two years, AI video has mostly stalled at a few seconds to a few dozen seconds, with everyone competing on image quality, camera moves, and single-shot spectacle. Higgsfield (the AI video generation platform) has brought something different this time: a 110-minute feature film with four real actors who licensed their likenesses and voices, produced in 4 weeks at a cost of $2 million.

They also did something else — they announced the film is 100% open source, with all prompts and assets now public. We checked item by item: what's directly downloadable are three Claude skill files. The rest of the material is scattered between the project page text and that canvas. Where the gaps are is laid out in the final section.

The film is called The Cully Hill Boys, an action comedy about three down-on-their-luck London rappers who stumble into a drug firefight while chasing fame. It premieres in New York on August 5.

What's being broken down here is what shipped with the film: a complete production manual. How to build assets, write prompts, direct performances, lip-sync rap tracks, and chase down fixes in post — written end to end, including the three Claude skill files they use internally.

The Cully Hill Boys film poster
The official poster. The small text at the top reads "A full AI-generated film starring real people," with the four actors' names listed below.
137 scenes
1 hr 54 min runtime, August 5 NYC premiere
600 assets
Approved assets on the canvas: characters, locations, props
~150 locks
Named locks in the prompt library; ~80 end with "= discard"

Every frame is generated. There's no set and no crew. The four actors weren't on any shoot — they contributed digitally scanned likenesses and voices. There's exactly one exception in the entire film, covered later.

The full film is attached to the official announcement post and can be watched in full. All prompts and assets are on the project page. We verified the video file: runtime is 1:54:53, size is 2.08 GB. "110 minutes" is the external rounding.

Tools

What it was built with: five tools, each with its own job

All shots, video, and voices were generated in Seedance. Faces and character sheets went through Soul Cinema. Editing, shot-reverse-shot work, and viewpoint changes used Seedream and Nano Banana.

All prompts were written by Claude — and deliberately in two separate chats, one for image prompts only and one for video prompts only. The reasoning: the rules of one side contaminate the other. Image work needs flat lighting and language that dodges the CG look; video work needs field-of-view in degrees and every light source justified. Mixed together, the two rule sets bleed into each other.

Three skills, in the order they're used

They packaged their common rules into Claude skills — one rulebook, loaded by Claude itself, which then executes accordingly. All three SKILL files are public; the filenames are clickable to download (also listed again at the end of this piece).

1
LIRA — Writes image prompts. You tell it what you need (this character in this state, this location at this time of day), and it produces the prompt from the character sheet or location base map. It has a built-in checklist of each image model's weaknesses, and self-checks the prompt against them before sending: which words summon a film studio, which light sources get baked into the frame, which details a given model always drops.LIRA SKILL.md ↓
2
CINEDANCE — Writes video prompts in three layers. The "writer" breaks down the scene, blocking, optics, physics, and timing; the "reviewer" re-reads each prompt before sending, hunting for empty first frames, outdated tags, weak spatial relations, and self-contradicting lines; the "workbench" saves prompts to files and only patches the failed segment, never rewriting the whole thing — a rewrite throws away the parts that were already right.CINEDANCE HIGGSFIELD SKILL.md ↓
3
ACTING — The performance system prompt. Behavior over emotion, how to write faces and bodies specifically, and the character behavior profile format all live here.ACTING SKILL.md ↓

Four other rules carried over from the previous film

The team had previously made a project called HELL GRIND, learning as they went. These four were already in hand on day one:

  • Scene context: every shot opens by stating what's happening dramatically, who's present, and how long it runs.
  • First-frame space lock: frame one already has everyone in position — no empty establishing shots.
  • One second of silence after every line: a clean tail inside each clip so the editor gets a seam, and the model has no room to invent sound.
  • Expression and performance written beat by beat: face and body written increment by increment, not dismissed with adjectives.
Core problem

Every rule exists to fight one thing: the model has no memory

The manual is this thick, but its starting point is a single sentence: consistency is the entire job.

Video models have no memory. Describe a character incompletely, and the next shot changes his face, his behavior, his voice. Add to that spaces that fall apart when the camera moves, sound that drifts between clips, and scenes that lose spatial relationships — that's everything an AI feature-length film has to fight.

Their foundation is one sentence: every shot comes from "reference + text." Assets handle the visuals; text handles everything else, including geometry: who stands where, which side the camera is on, which direction the light comes from.

Asset What this person looks like Text Who stands where, camera side Lighting, acting One shot Both inputs are required No shot is grown from a still image
Diagram by this site: visuals belong to the asset; everything else belongs to the text — including spatial relationships, which readers overlook most.

One harder constraint runs through the whole film: text-to-video only, no first frame, no image-to-video. It's harder — but that's exactly why the shots cut together.

An official side-by-side demo, 9 seconds. (From Higgsfield's official account; re-encoded and self-hosted by this site.)

The one time they broke this rule was the fight over a gun in the Cadillac: two people in constant contact, both grabbing the same weapon. What got generated was limbs tangled, the gun vanishing from one hand and appearing in the other. Finally they brought in stunt performers to choreograph the fight, spent a day filming on a phone in a real car, and fed that footage in as motion reference while the model layered the characters, interior, and lighting on top. One day of phone footage saved a week of retries.

What an asset is

A pair: a text description that goes verbatim into every prompt, plus a reference image that anchors the model. Characters, locations, and props are each assets.

Think of the studio's look-book photo plus a character bio. The photo handles the face, the bio handles personality and wardrobe, and both are handed to every new cinematographer on set.

Breakdown

And the script gets broken into cards before shooting

Their breakdown turns the script into a pre-visualized shot list that already contains the director's intent — not just a list of things to make. Each shot gets a card, and each card has four groups.

Card groupWhat goes in it
FootageLocation and interior/exterior, and the asset that covers it; time of day (because it determines which asset variant is used); every person in the frame, with tags and state; props and vehicles with tags; action in one to three sentences; dialogue verbatim; duration in seconds; complexity: simple, medium, or complex
DirectorGoal of the shot, in one sentence; the character's task written as a verb — interrogate, expose, humiliate; what changes from start to finish; blocking relative to the camera; what the face and body do in performance, and what the character is hiding
CameraShot size, movement, lens, angle
EditCut style, pace, and how it connects to the next shot

Each card carries a scene number plus a shot letter, like 50B, 50C; the same identifier follows from card to version log to prompt file, so two adjacent shots can never be confused.

That version log ends with 137 entries, showing which shots were reworked repeatedly — the hardest ones went to v15, v10, v9. Without that record, a good shot can't be reproduced, and you can't know whether you've already tried a given fix.

Once cards are filled, the prompts are almost mechanically transcribed. More importantly, gaps appear before a single generation is wasted. A shot without a goal, a character without a task, a scene without an asset — all visible on the card at a glance.

Law 1

So not a single shot is made until assets are locked

The 600 approved assets break down like this: 74 cards for leads and villains, 52 for supporting characters and animals, 90 for background characters appearing in one or two scenes, 159 for props, and 200+ for location base maps.

State itself is an asset. Cal in a clean jacket, Cal soaked after falling in the water, Cal with a split brow and blood in the third act — these are three assets, not one asset with a note appended.