Parametric Trajectory Distillation

Few-step distillation for video and sound.

Every clip in the film was generated by PTD on MiniMax-H3, picture and sound together: the six-step version throughout, and the four-step version in the comparison rows, next to LightX2V Turbo’s 4-step LoRA. Music and transitions were added in the edit.

Three short films

Three one-minute films made with the six-step version of PTD on MiniMax-H3. In the first two, every clip, picture and sound, comes from the six-step version, which also runs in MiniMax-H3's image and reference modes, and these chain the shots: some continue from the last frame of an earlier shot, others take a crop of an earlier frame as their reference, so faces and places carry through. The third is a test of the ceiling: it reproduces an existing film from that film's own keyframes. Captions, overlays, grading, cuts and music were added in the edit.

  • From text
  • Continues from the last frame of an earlier shot
  • From a reference crop of an earlier frame
  • From keyframes of the original film
  • Card added in the edit

The Open Window

After the short story by Saki (1914) · 64 s · 17 shots

A nervous guest on a rest cure calls at a country house. While he waits, the hostess's calm teenage niece tells him why the French window stays open every evening. Then, at dusk, three figures and a spaniel walk across the lawn toward it.

  • 17 clips, each picked from 8 to 24 seeds: 7 from text, 3 continuing from the last frame of an earlier shot and 7 from a reference crop, so the girl, the guest and the spaniel keep their faces.
  • All eight spoken and sung lines were generated with the picture, and each was checked word for word with Whisper.
  • Added in the edit: a sepia grade for the two stories the girl tells, captions, the cards and music (“Classical 7”, “Fun and Games” and “Little Bells” from Mixkit).
  • Not perfect: the spaniel changes colour between shots, a light, sandy cocker in the sepia stories and a dark red-brown one in the house.

Pebble No. 47

An aquarium at night, seen mostly through its security cameras · 63 s · 16 shots

Every night at 3:12, Dumpling, a small gentoo penguin, sneaks out to the gift shop to court a giant plush penguin the gentoo way: with a pebble. One night the guard catches him.

  • 14 clips, each picked by eye from 8 to 24 seeds, two of them used twice: 7 shots from text, 6 continuing from the last frame of an earlier shot, five of them straight after it and cut on still moments so that the join plays as one take, and 3 from reference crops, so the guard and the penguin stay the same.
  • Both spoken lines were generated with the picture and checked with Whisper.
  • Added in the edit: the camera overlays, the night-vision grade, captions, the cards, a blur over a printed label and music (“Skyline” by Eugenio Mininni, Mixkit).
  • Not perfect: the plush came out larger than scripted, and small text on signs and tags is garbled.

Erlang Shen, reproduced

A reproduction of 《二郎显圣真君》 by 昔年 · 69 s · 25 shots

A test of how far the six-step version can go rather than a story of our own. The keyframes, script and background music come from 《二郎显圣真君》 by 昔年 on Douyin, a film made with LibTV. Every shot is conditioned on keyframes of the original, so the composition is given; what PTD generates is the motion, the effects and the voices in between.

  • Each of the 25 shots was generated as one clip from the first and last frames of the original shot and a prompt for the action in between, in MiniMax-H3's first-and-last-frame mode, and picked from 8 seeds. Shot 23 starts from the last frame of shot 22 instead, and the last shot is split into two clips through one of its middle frames.
  • The Chinese lines were generated with the picture and placed at the original's timing, or in sync with the lips where the speaker is on screen; the music is the original's soundtrack with the vocals removed.
  • Edit: the cuts fall on the original's cut frames, and each shot plays the end of its clip, sped up 1 to 2.75 times so that it ends on the conditioning frame; shot 11 plays a stretch at normal speed instead.
  • One deliberate change: in shot 23 the camera follows the eye beam up to the Heavenly Palace, which the original does not show.

Abstract

Few-step video distillation poses a capacity allocation problem: a student must approximate a teacher's iterative generation with far fewer sequential evaluations while retaining comparable sample quality. Existing trajectory methods ask the student to reproduce complex teacher evolution, which can exceed its predictive capacity and degrade fine detail.

We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize each teacher trajectory segment as a polynomial and learn from teacher guidance along its own predicted path. PTD is designed to let the learned curvature adapt to the backbone's predictive capacity, preserving motion and diversity without forcing the student to fit trajectory details it cannot reliably learn. The interior curve serves as a training scaffold; inference uses only the predicted endpoint update.

Using LoRA and trajectory supervision alone, we distill MiniMax-H3, a 33B joint audio-video model, into a four-step generator. GPT-6 Astra evaluation shows greater diversity and higher naturalness than a four-step LightX2V distribution matching distillation (DMD) baseline, while matching or improving dynamic and static quality. On Wan2.1-14B, PTD sets a new VBench state of the art among trajectory-only video distillation methods.

Teacher
MiniMax-H3, a 33B model that generates video and sound together, sampled in 50 steps
Student
The same model with one LoRA, sampled in four or six steps
Conditions
Text, an image, a first and a last frame, or a reference picture
Sampling time
31 s in four steps and 46 s in six, instead of 343 s

A learned path, taken in a few strokes

A video model turns noise into a clip by following a path, fifty small steps long. PTD teaches the student the shape of that path, one curved segment per step, so it can cover the same ground in four or six strokes.

How it learns

Each step predicts a whole curved segment: a mean velocity and a few curve coefficients. The teacher's velocity is read at points along the student's own predicted path, and one least-squares loss fits the segment to those readings. Only a LoRA is trained, on text-to-video data.

How it samples

The curve is a training scaffold. At inference each step jumps straight to the segment's predicted end point, so a clip with sound takes four or six network evaluations. The previews above are the student's own predictions after each step, for the same prompt and the same starting noise.

Six steps, in motion

Fast, hard motion with its own sound. Six-step version; each clip was picked from eight seeds. Point at a clip to preview it, open it to watch with sound.

Preferred in a blind study

Four-step version against the public four-step LightX2V DMD LoRA for MiniMax-H3. Raters saw both methods as two rows of three seeds in random order and picked the better row or called a tie.

  • PTD preferred
  • Tie
  • LightX2V DMD preferred
PTD was preferred in 66% of the votes that picked a side (95% interval 54 to 76%, two-sided sign test p = 0.01). 82 votes from 10 raters on 74 prompts, counting each rater's last vote per prompt.

Everyday scenes held-out prompts from the MiniMax-H3 validation set

Fast motion prompts written for the study

The ten prompts raters preferred PTD on most. Open one to watch it with sound, or to see both rows the way the raters did.

Different seeds, different videos

PTD follows the teacher's path from each starting noise, so every seed keeps its own video. Distribution-matching students tend to land on one answer: compare the LightX2V rows below with the same prompt.

“A child in a yellow raincoat runs down a rainy street and jumps into a puddle.” Across a row, one starting noise: the six-step and four-step students make nearly the same video, close to the teacher's in layout, light and timing. Down a column, every seed is a different street and framing. Seeds 0, 1, 2, shown without selection.
“Top-down view of a clear koi pond.” Six-step version, seeds 0 to 15, shown without selection.
“A child in a yellow raincoat runs down a rainy street and jumps into a puddle.” Top row PTD, bottom row LightX2V Turbo. Seeds 0, 1, 2 for both, shown without selection or grading.
“A snowboarder carves down a freshly groomed slope.” Top row PTD, bottom row LightX2V Turbo. Seeds 0, 1, 2 for both, shown without selection or grading.

One LoRA, four ways in

The LoRA is distilled on text-to-video data only. Loaded unchanged, it also drives MiniMax-H3's image, first-and-last-frame and reference modes, in the same few steps and with sound.

From text

A paper boat sails down a rainy street at dusk.

From an image

The input picture: a woodblock print of a towering wave

From a first and a last frame

First frame: an oak in winterLast frame: the same oak in summer

From a reference

The reference picture: a woman with a silver bob in a mustard coat

A few steps, measured

Sampling time for one five-second 1344×768 clip with sound, measured in the same pipeline on 8 H100 GPUs (one clip per GPU, weights sharded); VAE decoding excluded.

Citation

@misc{feng2026ptd,
  title  = {Parametric Trajectory Distillation for Few-Step Video Generation},
  author = {Feng, Lan and Karkus, Peter and Igl, Maximilian and Berner, Julius and Chen, Yuxiao and
            Tan, Shuhan and Alahi, Alexandre and Ivanovic, Boris and Pavone, Marco},
  year   = {2026},
  note   = {Under review}
}

Prompt