Wan2.2 VAE — Flow-Weighted LoRA A/B vs control

Does weighting the reconstruction loss by optical-flow magnitude buy sharper moving regions? Two adapters whose configs differ in exactly one field.

Flow weighting did not do what it was built to do

Across 32 held-out val clips, the flow-weighted adapter is -0.28 dB on exactly the moving regions it was supposed to improve, while static regions gain +0.20 dB. That is the intended trade running backwards: capacity moved away from motion, not toward it.

The CutTofu clip disagrees — it shows moving +0.12 dB. That is why both are shown: one clip cannot separate a tenth of a dB from clip-to-clip variation, and picking the clip that agrees with the hypothesis would have manufactured a positive result.

Caveat that cuts the other way: one training run per arm (seed 0). A 0.28 dB gap is small enough that run-to-run variance is not ruled out. The honest reading is no evidence flow weighting helps moving regions — not proof that it hurts.

Overall PSNR was never going to settle this: it averages every pixel, so a pure redistribution nets out. Only the moving column can discriminate. Both adapters do beat the frozen base by ~2.8 dB — the LoRA finetune works; only the flow weighting is a wash.
Source: CutTofu_1_tactile_left.mp4 · Clip: 77 frames · 256×256 · 6fps (stride 5 from 390 native @ 30fps) · Adapters: lora_last.pt · r32 α32 targets=all · step 4000

Reconstruction

originalresized + subsampled input
reference · 77 frames @ 256×256
base Wan2.2frozen VAE, no adapter
PSNR 35.84 dB · MSE 2.61e-04
control · flow_weight 06fps, edge+LPIPS only
PSNR 38.61 dB · MSE 1.38e-04 · +2.77 dB vs base
treatment · flow_weight 16fps, L1 weighted by motion
PSNR 38.64 dB · MSE 1.37e-04 · +2.80 dB vs base

Error maps

diff ×20|original − base Wan2.2|
brighter = larger error
diff ×20|original − control · flow_weight 0|
brighter = larger error
diff ×20|original − treatment · flow_weight 1|
brighter = larger error
motion weightm̂ from optical flow
what flow up-weighted · 5.1% of pixels moving

PSNR split by motion · this clip

adaptermoving (m̂≥0.5) static (m̂≤0.1)all pixels
base (frozen)31.4337.1435.84
control (flow_weight 0)34.3039.8738.61
flow (flow_weight 1)34.4239.8738.64
flow − control+0.12 dB-0.00 dB+0.04 dB
Source: tactile_right.mp4 · Clip: 117 frames · 256×256 · 6fps · first 20s (stride 5 from 600 native @ 30fps) · Adapters: lora_last.pt · r32 α32 targets=all · step 4000

Reconstruction

originalresized + subsampled input
reference · 117 frames @ 256×256
base Wan2.2frozen VAE, no adapter
PSNR 45.76 dB · MSE 2.65e-05
control · flow_weight 06fps, edge+LPIPS only
PSNR 46.16 dB · MSE 2.42e-05 · +0.40 dB vs base
treatment · flow_weight 16fps, L1 weighted by motion
PSNR 46.18 dB · MSE 2.41e-05 · +0.42 dB vs base

Error maps

diff ×20|original − base Wan2.2|
brighter = larger error
diff ×20|original − control · flow_weight 0|
brighter = larger error
diff ×20|original − treatment · flow_weight 1|
brighter = larger error

PSNR split by motion · this clip

No *_flow.npz exists beside this clip, so there is no motion map to split on — only whole-frame PSNR is available here. The motion split for these adapters is on the CutTofu tab and, at population scale, in the val table below.

Source: tactile_right.mp4 · Clip: 117 frames · 256×256 · 6fps · first 20s (stride 5 from 600 native @ 30fps) · Adapters: lora_last.pt · r32 α32 targets=all · step 4000

Reconstruction

originalresized + subsampled input
reference · 117 frames @ 256×256
base Wan2.2frozen VAE, no adapter
PSNR 45.40 dB · MSE 2.88e-05
control · flow_weight 06fps, edge+LPIPS only
PSNR 45.85 dB · MSE 2.60e-05 · +0.45 dB vs base
treatment · flow_weight 16fps, L1 weighted by motion
PSNR 45.86 dB · MSE 2.59e-05 · +0.47 dB vs base

Error maps

diff ×20|original − base Wan2.2|
brighter = larger error
diff ×20|original − control · flow_weight 0|
brighter = larger error
diff ×20|original − treatment · flow_weight 1|
brighter = larger error

PSNR split by motion · this clip

No *_flow.npz exists beside this clip, so there is no motion map to split on — only whole-frame PSNR is available here. The motion split for these adapters is on the CutTofu tab and, at population scale, in the val table below.

Source: tactile_left.mp4 · Clip: 117 frames · 256×256 · 6fps · first 20s (stride 5 from 600 native @ 30fps) · Adapters: lora_last.pt · r32 α32 targets=all · step 4000

Reconstruction

originalresized + subsampled input
reference · 117 frames @ 256×256
base Wan2.2frozen VAE, no adapter
PSNR 42.83 dB · MSE 5.21e-05
control · flow_weight 06fps, edge+LPIPS only
PSNR 43.55 dB · MSE 4.41e-05 · +0.73 dB vs base
treatment · flow_weight 16fps, L1 weighted by motion
PSNR 43.59 dB · MSE 4.37e-05 · +0.77 dB vs base

Error maps

diff ×20|original − base Wan2.2|
brighter = larger error
diff ×20|original − control · flow_weight 0|
brighter = larger error
diff ×20|original − treatment · flow_weight 1|
brighter = larger error

PSNR split by motion · this clip

No *_flow.npz exists beside this clip, so there is no motion map to split on — only whole-frame PSNR is available here. The motion split for these adapters is on the CutTofu tab and, at population scale, in the val table below.

Source: tactile_right.mp4 · Clip: 117 frames · 256×256 · 6fps · 210–230s (stride 5 from 600 native @ 30fps) · Adapters: lora_last.pt · r32 α32 targets=all · step 4000

Reconstruction

originalresized + subsampled input
reference · 117 frames @ 256×256
base Wan2.2frozen VAE, no adapter
PSNR 43.27 dB · MSE 4.71e-05
control · flow_weight 06fps, edge+LPIPS only
PSNR 43.38 dB · MSE 4.59e-05 · +0.12 dB vs base
treatment · flow_weight 16fps, L1 weighted by motion
PSNR 43.39 dB · MSE 4.58e-05 · +0.13 dB vs base

Error maps

diff ×20|original − base Wan2.2|
brighter = larger error
diff ×20|original − control · flow_weight 0|
brighter = larger error
diff ×20|original − treatment · flow_weight 1|
brighter = larger error

PSNR split by motion · this clip

No *_flow.npz exists beside this clip, so there is no motion map to split on — only whole-frame PSNR is available here. The motion split for these adapters is on the CutTofu tab and, at population scale, in the val table below.

Val set · 32 held-out clips · the population read

The tabs above are individual samples; this is the population, and it is what the conclusion rests on. Same motion split, same lora_last.pt adapters, 32 val clips at 6fps (4.8% of pixels moving) · compare_flow_adapters.py

adaptermoving (m̂≥0.5) static (m̂≤0.1)all pixels
base (frozen)34.0436.7736.48
control (flow_weight 0)36.8339.9639.62
flow (flow_weight 1)36.5540.1639.84
flow − control-0.28 dB+0.20 dB+0.22 dB