Prompting Is Not Directing
We asked for one character and got five. A build-in-public report from Amado Studio on why a text prompt cannot carry a spatial decision, what seven hand-staged poses proved, and the two things that still do not work.
We asked for a full resolution frame of one character standing in a forest, and we got five characters. Not five variations. Five people, in one frame, in the same forest, all of them rendered convincingly, all of them wrong.
Nothing had broken. That is the part worth sitting with. No error appeared, no process failed, no piece of the system misfired. The frame had been assembled from several overlapping pieces, and each piece had been handed the entire text description of the scene. Each piece read that description, saw a character in it, and drew a complete scene with a character in it. Five pieces, five characters. Everything worked exactly as instructed. The instruction was the problem.
The fix was not a better description. Giving each piece only its own slice of what the user had staged in 3D cut five figures down to one. A piece whose staging contains no person does not invent a person. It has nothing to invent from. The information that resolved the problem was not linguistic. It was spatial, and it was already sitting in the scene the user had built.
A prompt is a lossy channel for a spatial decision#
Consider what a director actually does when they set up a shot. They do not describe the frame. They place a camera. They walk it in, drop it, tilt it, and look through it. The decision is made in space, with a physical object, by a person standing in the room.
Now consider what a text prompt can carry. It can carry the fact that there is a forest, that there is a character, that the light is low and warm. It cannot carry the eighteen inches. Current tools have no vocabulary for "move the camera eighteen inches left and drop it to waist height," which is an ordinary instruction on any set, given a hundred times a day, understood immediately by everyone who hears it. The sentence exists. There is simply nowhere to put it.
That gap is not a matter of the prompt being too short or insufficiently detailed. A spatial decision compressed into a sentence loses the part that made it a decision. And when several parts of a system each read that same sentence independently, with no shared sense of where anything sits, you get five characters in a forest, and every one of them is a correct reading of the words.
The backflip test#
If the argument is that staging carries information a description cannot, the argument needs to be testable. So we staged seven poses by hand in 3D: flying, crawling, running, a low crouch, arms overhead, a jump, and a mid-air backflip. Same forest, same props, nothing tuned for the run, about 150 seconds each.
Six of the seven came back unmistakable. Running was the one that failed. It reads as standing or walking, because the stride did not carry.
The backflip is the one that settles it. Nothing in the text description said "upside down." The character came out upside down because the 3D staging was upside down. No word did that work. The pose travelled from the scene into the image without passing through language at all.
There was a second finding in that run, and it points the same way. Two of the seven poses had to be re-staged before any generation time was spent, because they had been aimed straight down the lens and foreshortened into an unreadable shape. A pose with no silhouette carries no pose. That is a familiar constraint to anyone who has blocked a shot, and it is a useful one, because it is the kind of problem you can see and fix in the scene, before committing any time to generating anything.
What still does not work#
Across those seven clips, the character wore seven different outfits: dark jacket, blue top, grey shirt, shirtless, white shirt. His hair is red every time, and it is red only because the word red is in the description. That is the whole gap between describing a character and having one. Staging holds the body in space. It does not hold identity, and we do not yet have a solution for that.
The props behave the same way. They read as generic shapes rather than the specific things described. The log lands well. The boulder often becomes a mossy mound. The crate is just a dark box.
We also tried to score the pose matching numerically, and that attempt partly failed. Measured against a clip's own staging we got 1.927, against a different pose 1.432, roughly 35 percent in the right direction. But the per-clip ranking only held for 5 of 7. We report both numbers because reporting only the first would be dishonest. Our reading is that the metric was partly measuring background texture rather than pose, since the airborne poses sit against blown-out sky while the crouch and the crawl sit against textured forest floor. The direction is encouraging. The measurement is not yet trustworthy, and we are not going to present it as though it were.
A workflow problem, not a quality problem#
The five figures incident is often heard as a quality complaint. It is not. Every one of those five characters was well made. The images these systems produce are frequently beautiful. That was never the obstacle.
Think about how notes work on a film set. They are local. Fix the hand. Kill the reflection. That expression on the third beat. The rest of the shot is fine and nobody touches it. A tool whose only answer to a note is to regenerate the whole shot from a new starting point cannot take a note. You do not get a corrected version of what you had. You get a different thing, and you compare the two, and you hope. That is a workflow failure, and no improvement in image quality fixes it, because the images were never the part that was broken.
The alternative is not a cleverer prompt. It is deciding composition before a single frame is generated, because camera, characters, props and lights are real spatial objects rather than words in a prompt. When the scene exists as objects, the camera can move eighteen inches. The character can be upside down without anyone typing the phrase. And a note can be local, because there is something stable to give the note about.
What we are building#
Avalon Partner is a company built on a specific view: that AI filmmaking tools have concentrated on generation quality and largely ignored the workflow that surrounds it, and that this is now the binding constraint.
Amado Studio is Avalon Partner's first product. It puts a simplified 3D production environment in front of the generation step, so that the spatial decisions a director already knows how to make are made in space, and carried through intact.
We are building it in public, and this is the first entry. We will publish what fails alongside what works, at the same level of detail, in the same posts. The running pose did not carry. The identity problem is unsolved. The metric measured the wrong thing. Those belong here as much as the backflip does, and you should hold the rest of what we publish to that standard.
Filed under
Want help putting this into practice?
Avalon Partner helps Fall River and South Coast businesses fix the gaps that cost them leads. Call 774.559.8992 or email Joshua.Amado@AvalonPartner.com.


