generative scene extrapolation - benchmark

DL3DV ctx18 - 20-clip extrapolation benchmark

Observe the first 18 frames of a real DL3DV walkthrough with known camera poses, then generate the continuation along the ground-truth trajectory (frames 18-80, 63 extrapolation frames) into partially-unseen regions. All methods run the identical protocol; metrics are scored on the extrapolated region only.

context = 18 framestargets = 63 (18-80) 20 clips - 832x48011 methods incl SEVA, LVSM & CameraAnything

Leaderboard - extrapolation slice, sorted by PSNR, best per column highlighted

MethodPSNRSSIMLPIPSFIDTSEDMEt3Rn
SEVAStable Virtual Camera - generative NVS diffusion [NEW]16.060.4840.50365.90.6490.113020
LVSMLarge View Synthesis Model - regression transformer [NEW]15.620.4850.593121.80.6730.110820
FrameCrafterI2V-14B + LoRA warp-inpaint14.950.4090.46343.10.5160.175120
VACE-14B (ours)warp + inpaint baseline14.280.4220.54081.40.5090.135920
NVS-Solvertraining-free diffusion prior14.150.4190.62178.40.3680.057720
GEN3Cexplicit 3D cache + render14.130.4400.628114.20.2970.120920
FlexWorld3DGS + video diffusion14.120.4320.60790.20.3090.114520
Lyra-2AR + 3D memory13.180.4000.65879.90.3930.099120
trajcrafter_gtpose12.340.3660.64478.30.4160.126920
TrajCrafterdiffusion NVS - 3-DOF orbit11.910.3510.68085.20.3240.142820
CameraAnythingrefilming w/ arbitrary camera control (ECCV'26) [NEW-EXT]11.350.3480.65577.30.4240.107520
MCSDF (ours, prev ckpt)memory-cond. DF - pre-retrain, refreshing11.350.3300.681132.30.4400.206520

PSNR/SSIM/LPIPS over the extrapolated region; FID = generation realism; TSED = 2-view SfM inlier fraction; MEt3R = multi-view consistency. Raw PSNR is low because this is generation, not reconstruction. SEVA leads PSNR; FrameCrafter leads LPIPS & FID. MCSDF is our method at a pre-retrain checkpoint (a new-objective model is training; this row refreshes when it lands).

CameraAnything audit (2026-07-29). Its last-place PSNR with best-tier MEt3R (0.107) and worst rotation error (43.8°) is a task-domain result, not a broken adapter - we verified the pose chain at three levels: (1) code - their pipeline consumes our OpenCV c2w with no inverse/transpose, relativized to observed frame 0; (2) numerical - independently recomputed Plücker rays match their function to 2.2e-02 max abs (half-pixel grid offset; a convention flip would be O(1)); (3) behavioral - two mirrored synthetic yaw trajectories (±40°, no translation) produce cleanly mirrored camera motion in exactly the direction the matrices dictate, and their own 8 shipped examples all obey their trajectories. The failure mode is that CameraAnything is trained for refilming (new viewpoints on content that stays visible); once the target view leaves the observed frustum it has no geometric evidence and invents a plausible different room. Fairness caveat: translations are similarity-scaled to max‖t‖=2.0 (their training range); doubling to 4.0 on a held-out clip gained +0.45 dB and increased motion magnitude without fixing the route, so a full rescaled rerun would move this row to roughly 11.8 - not enough to change the ordering. Panel note: the per-clip comparison videos now show CameraAnything in place of Lyra-2 (Lyra-2's scores remain in the leaderboard above; the video grid holds 9 method panels + GT).

Per-clip comparison - method grid + camera trajectory

12aa6a12145314fb_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
1c2554841888ab34_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
1c2554841888ab34_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
23cf048933459b1f_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
23cf048933459b1f_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
24b70d178abebbb0_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
24b70d178abebbb0_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
377430f9813379c5_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
3ab050170ad955b3_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
3ab050170ad955b3_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
3fa142f449c51c3e_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
3fa142f449c51c3e_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
408f0740fbf4395c_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
408f0740fbf4395c_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
42ac4b19a6ec45ee_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
42ac4b19a6ec45ee_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
4f0eb45686bc4a10_c0
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
4f0eb45686bc4a10_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
53b39a7ec355b8f3_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame
552646ce086e79d5_c1
context 0-17generated 18-80
method comparison - GT + 9 methods (SEVA, LVSM, FrameCrafter, VACE-14B, GEN3C, NVS-Solver, FlexWorld, Lyra-2, TrajCrafter); context 18 frames, extrapolate 18-80. Bar: green = context 0-17 / orange = generated 18-80, marker = playback position.
camera path
camera trajectory - observed 0-17 vs target 18-80, colored by frame