Trending News

Seedance 2.5 vs Veo 3.1 vs Kling 3.0: Which Workflow Fits Indie Film Previsualization?

The location is booked for six hours.

The scene looks simple on paper: a woman enters a convenience store, spots a bag behind the counter and hears something outside just before she reaches it.

Then the questions start.

Do we follow her from the door? Stay wide until she sees the bag? Cut before she turns, or after?

Small decisions — until the crew is standing in a rented location with the clock running.

AI pre-vis can move some of those arguments to an earlier, cheaper moment. Not by making a prettier storyboard, but by giving the director, cinematographer and editor something they can watch, question and change.

Seedance 2.5, Veo 3.1 and Kling 3.0 can all play a part, but I would not use them for exactly the same job.

This isn’t a lab test with identical prompts and assets. It’s a workflow comparison based on the features currently available, and results will vary with the access point, settings, source material and production itself.

Start With What the Art Department Already Has

Before planning every cut, I would want to see the whole convenience-store beat once.

How long does it take her to notice the bag? Does the walk toward the counter build tension, or just feel slow? How much space should there be between spotting the bag and hearing the sound?

The project may already have an original character design, a rough store concept and a prop sketch.

With Seedance 2.5 on XMK, supported image, video and audio references can be brought into the same setup instead of rebuilding all of that material through the prompt. For a team carrying several references into the test, the current interface allows up to 50 files in total, while clip options of up to 30 seconds leave room to explore a longer beat before deciding where to cut.

So the direction can stay focused on the scene:

Wide shot from the back of a quiet convenience store at night. The character enters, pauses and notices a dark bag behind the counter. She walks toward it as the camera slowly moves closer. Near the end, a sound outside makes her turn toward the glass door.

I am not looking for a finished shot here. I want to know whether the beat works before breaking it apart.

If the cereal boxes change between frames, fine. If the character reaches the counter before the tension has had time to build, that is worth noticing.

References are still references, not locks. Clothing, props and backgrounds can shift between results. XMK also states that real human faces, including portraits and celebrity material, are not supported in this workflow, and copyrighted material is restricted. An eligible original or non-real-person character design is therefore a better fit for this kind of test than an actor headshot.

A synthetic character should not be designed to imitate a performer who has not approved that use. If the pre-vis leaves the internal production process, its illustrative or AI-generated nature should be clear where viewers could mistake it for real footage.

Maybe the Problem Isn’t the Picture

Suppose the blocking looks fine, but the turn toward the door still feels wrong.

Change the rough placement of the sound cue.

Should it seem to arrive while she is walking, when she reaches the counter, or shortly after she stops?

Now I would be interested in Veo 3.1.

Google’s current Veo workflow can generate audio with video. A rough version can help compare those timing ideas instead of asking everyone to imagine the sound against a silent clip. Frame-accurate cue placement still belongs in editing software.

Dialogue makes the difference even clearer.

A character asks a question across a table. The other person waits before answering.

The script says the pause lasts two seconds.

Watch it with temporary audio and two seconds might suddenly feel like forever. Or it might be exactly where the tension comes from.

Better to find that out before two actors and a crew are waiting on set.

The audio does not need to sound finished. It needs to make the timing easier to judge.

Generated dialogue needs caution too. Google says natural and consistent speech, especially in shorter segments, is still being improved. For pre-vis, I would treat it as timing material rather than a dependable performance or final soundtrack.

What If It Should Never Have Been One Shot?

After watching the longer version, somebody may ask the obvious question:

“Why are we doing all of this in one shot?”

Try three instead.

Shot 1: Wide as she enters.

Shot 2: Move toward the counter.

Shot 3: Reaction at the door.

Now the cuts are part of the test.

Kling 3.0 becomes more interesting here. Kuaishou’s 3.0 family emphasizes multi-shot storytelling, narrative control and consistency across characters, objects and scenes. The series supports generation of up to 15 seconds and native audio.

But forget the feature list for a moment.

Watch the three shots together.

Do we need Shot 2?

Should the reaction be tighter?

Does the screen direction still make sense after the cut?

Would the reveal work better if we stayed wide for another second?

A rough sequence that starts that conversation has already done its job.

One technical distinction is worth keeping straight: Video 3.0 and Video 3.0 Omni are separate models within the wider Kling 3.0 family. A feature specific to Omni should not automatically be treated as standard Video 3.0 behavior.

Don’t Fix the Cereal Boxes

Generated footage has a habit of making the wrong problems look urgent.

The cereal boxes change.

A sign disappears.

The bag looks slightly different.

Unless one of those details changes the story or the planned shot, leave it alone.

Now imagine the character reaches the counter before the camera settles.

That matters.

Or the turn toward the door happens so quickly that the audience cannot understand what caught her attention.

That matters too.

For pre-vis, I would rather ask:

  • Does the action fit inside the shot?
  • Can we read the reveal?
  • Is the camera moving for a reason?
  • Does the reaction come too early?
  • Would a cut make the beat clearer?
  • Can this camera move actually fit inside the location?

A messy clip that answers those questions can save more time than a polished one.

Know When to Throw a Version Away

Near-miss generations are tempting.

The shot works, but the bag is wrong. Fix the bag. Then the background changes. Fix that. Now the lighting feels different.

Soon the crew is polishing footage nobody plans to use.

XMK presents Seedance 2.5’s local re-draw controls for adjusting a specific visual element, such as a product, background or subject. That makes sense when one visual detail genuinely gets in the way.

Check the full clip again after the change. Motion, lighting, composition or nearby details may shift too.

Timing is different.

If the character turns too early or the camera move simply feels wrong, stop repairing the frame.

Change the direction. Generate again. Or make it two shots.

Pre-vis is cheap partly because you are allowed to throw things away.

Show It to the People Who Have to Shoot It

I would skip the usual AI comparison scorecard.

“Realism: 9/10” is not going to help much when the crew arrives at the store.

Show the clips to the director, cinematographer and editor instead.

The director might prefer one version because the blocking finally makes sense.

The cinematographer may point at the same clip and say:

“We can’t put the camera there.”

The editor might watch another and say:

“Cut before she turns.”

Now the test has produced something useful.

Preventing one unnecessary setup could matter more to a small production than getting the most impressive generation.

A Good Pre-Vis Can Still Be Completely Wrong

The more convincing generated footage becomes, the easier it is to trust too much of it.

That smooth camera move may not fit between two real shelves. Generated lighting may have little connection to what the crew can build in six hours.

A convincing synthetic action beat still does not prove that a camera path, lighting setup, stunt, prop interaction or performer movement is physically achievable or safe.

The clip starts a conversation. It does not replace scouting, blocking, lens tests, budgeting, technical checks or safety planning.

The material going into the generator needs care too. Teams should only use assets they are authorized to process. Identifiable performers and voices, copyrighted works, unreleased footage and confidential production material require particular attention, along with the rules of the platform being used.

Maybe There Isn’t One Model for the Film

The convenience-store scene could start with Seedance 2.5 because the art department already has useful reference material.

If the off-screen sound is still causing trouble, Veo 3.1 might help the team explore rough timing.

Once the director decides the scene needs three shots, Kling 3.0 may be useful for exploring that structure.

The next scene could be completely different.

That is fine.

An indie production does not get anything extra for using the same model from the first rough idea to the last pre-vis.

The director needs to understand the beat.

The cinematographer needs to understand the shot.

The editor needs to see the cut.

Once they have the answer, stop generating.

The location is booked for six hours. You have a film to shoot.

 

Share via: