We Tested Character Consistency Across Eight AI Worlds
See how one AI character survives eight world changes in a Seedance 2 test of identity, wardrobe, motion, and camera consistency, with prompts and real frames.
Character World-Shift Test — Final Edited Showcase
Final edited showcase · 8 sec · Silent edited showcase. The article separates generation decisions from post-production pacing.
- 0:00The red-coated runner establishes the identity lock
- 0:02Salt, storm, and fire environments replace the world
- 0:04Orbital and underwater spaces stress the same action
- 0:07The closing beat reveals which locks survived
We wanted to know what "keep the same character" meant after the background changed eight times. A similar face would not be enough if the coat changed length, the runner reversed direction, or the camera stopped tracking from the side.
The test used one anchor image and one 15-second Seedance 2.0 generation. We did not retry it. The source was shortened later for the showcase above, but every observation in this report comes from the unedited generation.
This is a stress test, not a model benchmark. It tells us what survived in one demanding prompt and which failures were visible. Another run could distribute the errors differently.
Test setup and scoring criteria
The subject was an androgynous Black woman with close-cropped silver hair, a long crimson coat, matte-black clothing, and black boots. She ran from left to right in side profile.

We scored four consistency locks:
- Identity covered the face, hair, age, and body type.
- Wardrobe covered the crimson coat, black outfit, and boots.
- Spatial consistency covered screen position, scale, and direction.
- Motion covered the running phase and tracking speed.
The environments were free to change. Those four areas were the parts we intended to hold.
Side profile made direction easy to judge. It also exposed leg phase, body lean, and coat motion. A front-facing runner could appear to run in place without making the failure as obvious. The black studio removed most background ambiguity, while the white diagonal and crimson coat gave us strong tracking marks.
The image prompt was:
Premium cinematic editorial still of one androgynous Black woman with close-cropped silver hair in stable full-body side profile, sprinting left to right across a pure black studio crossed by one hard white light slash. She wears a long crimson coat over a matte black fitted outfit and black boots. One foot airborne, coat trailing behind, face clearly visible, fierce focused expression. Eye-level 35mm side view, centered subject, clean silhouette, realistic anatomy and fabric motion. No words, letters, numbers, logo, watermark, interface, border or duplicate person.
The world-shift prompt
The video prompt repeats the locks after describing the worlds:
Use supplied image as strict identity/action anchor. A relentless side-tracking fashion trailer: same runner moves left to right while hard match cuts strike every 0.8-1.2 seconds on footfalls, coat occlusions and whip-pans. Worlds: black studio, mirrored rain tunnel, sun-blasted salt flat, burning baroque opera hall, zero-gravity orbital corridor, underwater cathedral, violent snowstorm brutalist city, final crimson void. No dissolves, no slow morphs, no shot held longer than 1.2 seconds. Lock face, body, crimson coat, screen position, scale, running phase, direction, camera height, 35mm lens and tracking speed. Secondary rain, ash, bubbles and snow react physically. End on a sharp hero stride. No text/signage/logos, face drift, duplicates, extra limbs, flicker or camera jumps.
The prompt changes the world while repeating one action. Footfalls, coat occlusions, and whip-pans provide places for the environment to switch. We did not ask the model to invent a narrative between the locations.
Direction was more stable than anatomy
The runner kept moving left to right through the whole source. She did not reverse, face the camera, or stop for a pose. On the salt flat, the pale environment removed much of the original contrast, yet the coat and black outfit still separated from the ground.

Leg shape and stride phase were less stable. Stride length changed between worlds, and a few frames compressed the legs. In this run, the broad intent to keep running survived better than the exact phase of the run cycle. We would need a controlled follow-up to know how repeatable that difference is.
The coat carried more identity than the face
The burning opera scene changed the color temperature, architecture, floor reflections, and surrounding particles. The crimson coat remained immediately visible.

This does not establish a general rule that wardrobe matters more than a face. The runner often appeared at medium or full-body scale, where coat color was easier to observe than facial detail. In that framing, a generic shirt and ordinary haircut would have made the result harder to judge.
Two or three visible tokens are enough for a useful test: a coat color, hairstyle, accessory, silhouette, or repeated movement. The best tokens depend on the shot scale.
Camera geometry drifted inside the same screen direction
The orbital corridor kept the runner moving left to right, but its perspective lines and visible planet made the set feel wider than the original 35mm studio view.

Direction, position, and lens behavior need separate review. The runner can travel the same way while occupying a different part of the frame. Perspective and field of view can also change without reversing the action. That is what happened here: direction held, while placement and apparent focal length moved.
The underwater world exposed a physics conflict
The underwater cathedral introduced buoyancy, caustic light, bubbles, jellyfish, and an environment where ordinary running does not make physical sense.

The model kept the running graphic instead of changing the action to swimming. The image works within a surreal fashion trailer but would be difficult to justify in a narrative scene.
A production has to choose what happens when the anchor action conflicts with the new environment. It can keep the action and accept surreal physics, adapt the action and accept a visible motion change, or plan a transition that belongs to both spaces. A leap, for example, can continue as zero-gravity drift or underwater suspension more naturally than a running stride.
What the dense-frame audit found
We sampled the unedited source every 0.25 seconds. Fast playback hid several short failures.

The face stayed generally recognizable but softened as the subject became smaller. Coat color held better than coat length and folds. Direction stayed stable while scale and vertical placement shifted. Running continued, though the stride reset near several transitions. The camera remained broadly side-tracking, with changes in apparent focal length and background parallax.
All requested environments appeared, although a few remained longer than the requested 0.8-1.2 seconds. We would not describe the result as perfect character consistency. It shows that a limited set of visible tokens can persist through large scene changes.
Turn "same character" into visible constraints
The phrase "keep the same character" hides several independent requests. A more testable version names what the viewer should be able to compare:
Use the supplied image as the exact identity and action reference. Preserve face, hairstyle, body proportions, signature wardrobe, prop, movement direction, subject scale, screen position, camera height, focal length, and action phase. Only change the environment, weather, secondary particles, and environment lighting. At every transition, the character must continue the same action without posing, reversing direction, changing clothes, duplicating, or resetting the gait.
That list should shrink for simpler shots. A portrait may only need face, expression, and eye line. A dance test will care more about body proportions, clothing, and motion. Product interaction adds the relationship between the hand and the object.
What one reference image cannot show
Our side-profile anchor worked because the camera mostly stayed on that side. It contains little evidence for a close frontal portrait, a back view, or a full orbit. The model would have to invent those unseen angles.
Additional references become useful when the camera crosses to the other side, facial identity must survive both close-up and full-body shots, or the outfit contains asymmetric details. The same applies to a prop that must remain correct from several views and to a recurring character produced across separate clips.
When separately generated stages need to inherit one character, a controlled reference handoff gives each stage a confirmed parent image instead of asking every shot to reconstruct the character independently.
What this single test supports
Direction and coat color were the most stable signals in this source. Stride phase, scale, and apparent lens geometry were weaker. Those observations can guide the next test, but one generation cannot tell us whether the same ranking holds across other characters, motions, or models.
If you want to reproduce the setup, compare image-to-video generators for their reference controls, then start from a side-profile action reference in Motiofy. Three environments will make the first failure easier to isolate than eight. Review the raw output frame by frame before adding more worlds.
