Higgsfield shot a 110-minute feature film using real actors' likenesses and AI, then published the entire production manual and prompt methods
- A 1-hour-54-minute AI feature film, with the entire production playbook and three custom Claude skills released to the public.
- The whole manual exists to solve one problem: video models have no memory. "Keeping the same person the same person" is the entire job.
- The most valuable lessons are the counterintuitive taboos: no "studio" in character sheets, no heads in full-body shots, and never write "don't" in a prompt.
An AI feature film — and the entire production manual, released
For the past two years, AI video has mostly stalled at a few seconds to a few dozen seconds, with everyone competing on image quality, camera moves, and single-shot spectacle. Higgsfield (the AI video generation platform) has brought something different this time: a 110-minute feature film with four real actors who licensed their likenesses and voices, produced in 4 weeks at a cost of $2 million.
They also did something else — they announced the film is 100% open source, with all prompts and assets now public. We checked item by item: what's directly downloadable are three Claude skill files. The rest of the material is scattered between the project page text and that canvas. Where the gaps are is laid out in the final section.
The film is called The Cully Hill Boys, an action comedy about three down-on-their-luck London rappers who stumble into a drug firefight while chasing fame. It premieres in New York on August 5.
What's being broken down here is what shipped with the film: a complete production manual. How to build assets, write prompts, direct performances, lip-sync rap tracks, and chase down fixes in post — written end to end, including the three Claude skill files they use internally.
Every frame is generated. There's no set and no crew. The four actors weren't on any shoot — they contributed digitally scanned likenesses and voices. There's exactly one exception in the entire film, covered later.
The full film is attached to the official announcement post and can be watched in full. All prompts and assets are on the project page. We verified the video file: runtime is 1:54:53, size is 2.08 GB. "110 minutes" is the external rounding.
What it was built with: five tools, each with its own job
All shots, video, and voices were generated in Seedance. Faces and character sheets went through Soul Cinema. Editing, shot-reverse-shot work, and viewpoint changes used Seedream and Nano Banana.
All prompts were written by Claude — and deliberately in two separate chats, one for image prompts only and one for video prompts only. The reasoning: the rules of one side contaminate the other. Image work needs flat lighting and language that dodges the CG look; video work needs field-of-view in degrees and every light source justified. Mixed together, the two rule sets bleed into each other.
Three skills, in the order they're used
They packaged their common rules into Claude skills — one rulebook, loaded by Claude itself, which then executes accordingly. All three SKILL files are public; the filenames are clickable to download (also listed again at the end of this piece).
Four other rules carried over from the previous film
The team had previously made a project called HELL GRIND, learning as they went. These four were already in hand on day one:
- Scene context: every shot opens by stating what's happening dramatically, who's present, and how long it runs.
- First-frame space lock: frame one already has everyone in position — no empty establishing shots.
- One second of silence after every line: a clean tail inside each clip so the editor gets a seam, and the model has no room to invent sound.
- Expression and performance written beat by beat: face and body written increment by increment, not dismissed with adjectives.
Every rule exists to fight one thing: the model has no memory
The manual is this thick, but its starting point is a single sentence: consistency is the entire job.
Video models have no memory. Describe a character incompletely, and the next shot changes his face, his behavior, his voice. Add to that spaces that fall apart when the camera moves, sound that drifts between clips, and scenes that lose spatial relationships — that's everything an AI feature-length film has to fight.
Their foundation is one sentence: every shot comes from "reference + text." Assets handle the visuals; text handles everything else, including geometry: who stands where, which side the camera is on, which direction the light comes from.
One harder constraint runs through the whole film: text-to-video only, no first frame, no image-to-video. It's harder — but that's exactly why the shots cut together.
The one time they broke this rule was the fight over a gun in the Cadillac: two people in constant contact, both grabbing the same weapon. What got generated was limbs tangled, the gun vanishing from one hand and appearing in the other. Finally they brought in stunt performers to choreograph the fight, spent a day filming on a phone in a real car, and fed that footage in as motion reference while the model layered the characters, interior, and lighting on top. One day of phone footage saved a week of retries.
A pair: a text description that goes verbatim into every prompt, plus a reference image that anchors the model. Characters, locations, and props are each assets.
Think of the studio's look-book photo plus a character bio. The photo handles the face, the bio handles personality and wardrobe, and both are handed to every new cinematographer on set.
And the script gets broken into cards before shooting
Their breakdown turns the script into a pre-visualized shot list that already contains the director's intent — not just a list of things to make. Each shot gets a card, and each card has four groups.
| Card group | What goes in it |
|---|---|
| Footage | Location and interior/exterior, and the asset that covers it; time of day (because it determines which asset variant is used); every person in the frame, with tags and state; props and vehicles with tags; action in one to three sentences; dialogue verbatim; duration in seconds; complexity: simple, medium, or complex |
| Director | Goal of the shot, in one sentence; the character's task written as a verb — interrogate, expose, humiliate; what changes from start to finish; blocking relative to the camera; what the face and body do in performance, and what the character is hiding |
| Camera | Shot size, movement, lens, angle |
| Edit | Cut style, pace, and how it connects to the next shot |
Each card carries a scene number plus a shot letter, like 50B, 50C; the same identifier follows from card to version log to prompt file, so two adjacent shots can never be confused.
That version log ends with 137 entries, showing which shots were reworked repeatedly — the hardest ones went to v15, v10, v9. Without that record, a good shot can't be reproduced, and you can't know whether you've already tried a given fix.
Once cards are filled, the prompts are almost mechanically transcribed. More importantly, gaps appear before a single generation is wasted. A shot without a goal, a character without a task, a scene without an asset — all visible on the card at a glance.
So not a single shot is made until assets are locked
The 600 approved assets break down like this: 74 cards for leads and villains, 52 for supporting characters and animals, 90 for background characters appearing in one or two scenes, 159 for props, and 200+ for location base maps.
State itself is an asset. Cal in a clean jacket, Cal soaked after falling in the water, Cal with a split brow and blood in the third act — these are three assets, not one asset with a note appended.
