Back to home
EnglishEN

Three Photos to a Gas Station Dance Video: Our Real 20-Second Test

Watch the unedited output, compare all three input portraits with four scenes, and see what our single 480p generation does and does not establish.

Last updated: 2026-10-05

By GasStationDance.video

We used three fictional adult portraits to generate the video below. The file is 20.04 seconds long at 480 × 854 pixels, with generated audio. No original dance clip was supplied for this test. This is a new scene, not the source reel with three faces replaced.

20.04 seconds, 480 × 854, model-generated audio; results can vary

The player contains the full output. We have not replaced its audio, sped it up, or cut together the best moments from several attempts.

The inputs: three portraits, three jobs

Actual dancer input: fictional adult in a blue overshirt.Actual selfie input: fictional adult in a rust-colored shirt.Actual car-exit input: fictional adult in a teal jacket.

Left to right: dancer, selfie character, car-exit character. These are the actual test inputs, not customer photos.

Follow the blue overshirt for the dancer, the rust-colored shirt for the selfie character, and the teal jacket for the car-exit character.

What happens in the output?

Output frame: blue-shirted dancer beside the pumps.

1. Opening dancer

Output frame: rust-shirted character in selfie view.

2. Camera turns to the selfie

Output frame: teal-jacketed character getting out of a dark car.

3. Car-door reveal

Output frame: rust-shirted selfie character returns.

4. Return to the selfie

Frames from the same output file. Watch the player to judge the transitions and motion.

The sequence includes the intended four parts. The blue overshirt, rust-colored shirt and teal jacket remain useful visual cues across these scenes. The second character returns in the final selfie.

Still frames do not settle every quality question. In the car-exit frame, for example, one moving hand is blurred. To decide whether that moment looks natural enough, watch the movement rather than judging only this screenshot. Face resemblance is also a viewer judgment, not a measured identity score in this test.

How we made and checked it

The real generation was tested locally on October 5, 2026 using Wan, model ID wan/3-0-video, through Kie. The request asked for 20 seconds, 480p and audio, using the three assigned portraits. We submitted one generation, not a batch from which this was selected.

After generation, we played the saved MP4 in a browser, checked its audio and downloaded it through the result interface.

Technical checks and test environment

The file contains H.264 video at 30 fps and AAC stereo audio at 44.1 kHz. The audio decoded successfully. The generation was real; checkout used Stripe test mode. This test did not verify production payment.

What to expect from this example

This set of three portraits produced all four scenes with generated audio. Use the input photos and full video to judge the cast, motion and ending.

This is one 480p generation, not a benchmark or a measured success rate. Your photos can produce different results. Exact reference choreography, the original song, 720p and 1080p were not tested.

Is this result suitable for your clip?

If your main goal is a recognizable three-person scene for a private joke, compare the cast and full sequence with what you have in mind. If you need the original dance, camera path and soundtrack unchanged, do not treat this sample as proof of that capability.

For your own attempt, use the step-by-step guide. For input choices, see the role-photo examples. Questions about this test or a result you received? Contact us.