Ten seconds of AI video on a laptop GPU with 8 GB VRAM
One 640 × 480 still in; sixteen seconds of 1440 × 1080 video out, generated, upscaled and interpolated on a laptop. Full 1080p on YouTube.
The clip above is the spaghetti test, the scene that has become an informal benchmark for AI video, and every frame of it was made on my laptop: an RTX 4070 Laptop GPU with 8 GB of VRAM and 32 GB of system RAM. No cloud, no API, no credits. It came out of Forge VFX, the compositor I build in my spare time, which runs a local generation pipeline in three stages.
| Stage | Model | Result | Wall time |
|---|---|---|---|
| Generate | HunyuanVideo 1.5, 480p image-to-video, step-distilled, 12 steps | 241 frames at 720 × 544, 10 s at 24 fps, in one run | ~50 min |
| Upscale | SeedVR2 3B, FP8 weights | 2× to 1440 × 1088, all 241 frames | ~19 min |
| Interpolate | RIFE 4.26 | 24 to 48 fps, original frames untouched | ~2 min |
| Edit | Forge VFX | 48 fps played at 0.625× in a 30 fps scene: 16 s of slow motion |
The interesting part is the memory. HunyuanVideo 1.5 is an 8.3 billion parameter video transformer with a 7 billion parameter text encoder, about 35 GB of weights. Tencent's own README gives its minimum as 14 GB of VRAM with model offloading enabled, for its default five-second clip. Third-party guides put an FP8, CPU-offloaded setup at 8 to 10 GB for 480p and list 12 GB as the smallest workable card. This run generated a clip twice the default length, in one go, using about 4.7 GB of video memory for the job itself (under 5.4 GB counting everything else on the card), with nothing spilled to system memory. This post is about how.
Results at a glance
Every number in this post was measured on the machine, or read from the model's own configuration and documentation.
| What changed | Before | After |
|---|---|---|
| Peak VRAM, 241-frame generation | spilled, 10+ min per step | ~4.7 GB, no spill |
| Transformer weights in memory | 16.6 GB (bf16) | 8.3 GB (FP8 storage) |
| Peak memory commit while loading | 36.8 GB | ~20 GB |
| Prompt encoding on the CPU | 135 s | 37 s |
| Denoising step, 121 frames | 85 s | 56 s |
| Decode joins vs. a whole-clip decode | 31–40 dB PSNR, visible | ≥ 44 dB PSNR, invisible |
| Clip length in one run | 33 frames, ~4 min | 241 frames, ~50 min |
Why clip length is the hard part
Weights are a fixed cost. Activations grow with the clip, and at 241 frames they are the whole problem. The model's VAE compresses 16× in each spatial direction and 4× in time, and the transformer uses a patch size of one, so every latent pixel is one token:
720 × 544 frame → 45 × 34 latent = 1,530 tokens per latent frame 241 frames → 61 latent frames = 93,330 tokens one activation 93,330 × 2,048 × bf16 = 382 MB per tensor feed-forward middle 93,330 × 8,192 × bf16 = 1.5 GB, plus a copy for the activation naive attention 93,330² × 16 heads × 2 B ≈ 280 GB one VAE decoder map 128 ch × 720 × 544 × 241 × 2 B ≈ 24 GB
Against that, the budget was about 7.8 GB of usable VRAM and 32 GB of RAM. One mental model drove every decision that followed: treat VRAM as a cache and system RAM as the store, and keep on the GPU only what it needs this instant.
1. The text encoder gets its own phase
The 7B text encoder runs exactly once, to turn the prompt into a few megabytes of embeddings, yet keeping it resident beside the transformer costs about 15 GB for nothing. So generation runs in two phases. First only the text encoder is loaded, run on the CPU so it never touches VRAM, and deleted once the embeddings exist. Then the transformer loads into the memory just freed.
A side discovery: the tokenizer padded every prompt to 1,108 tokens, and on the CPU the padding was most of the work. The encoder only looks backwards and the padding comes after the prompt, so trimming to the real length changes nothing in the output. Encoding went from 135 s to 37 s. The negative prompt is now only encoded when the model's guidance uses it, and the distilled checkpoint does not.
2. FP8 storage, not FP8 math
bf16 has 8 exponent and 7 mantissa bits; FP8 E4M3 has 4 and 3, holding each weight to about one part in sixteen. Weights tolerate that well, because rounding errors average out over the thousands of multiply-adds behind every output. With layerwise casting the weights sit in FP8 and a hook upcasts one layer to bf16 right before it runs, then drops the copy. The math stays bf16, and small sensitive layers such as norms and embeddings stay bf16 throughout.
That halves the transformer from 16.6 GB to 8.3 GB: half the RAM, and half the bytes streamed to the GPU on every step. It does not make the matrix math any faster, and it is worth being precise about that.
Loading needed care too. The model is built on PyTorch's meta device, shapes only and no memory, and then each tensor is read from disk, converted to FP8 and placed, one at a time. I replaced memory mapping with a small safetensors reader, because on Windows mapping a file charges its whole size against the commit limit, briefly twice over. Peak commit fell from 36.8 GB to about 20 GB.
3. Streaming transformer blocks from RAM
The transformer's 54 blocks run in sequence, so only one block's weights are needed at a time. All 54 live in system RAM. Before each block runs, its roughly 150 MB of FP8 weights are copied to the GPU, upcast, used and released. VRAM holds one block, never the model.
The reason this costs almost nothing here is the clip length again. One denoising step moves about 8 GB across PCIe, well under a second, while computing that step over 93,330 tokens takes minutes. Compute grows faster than transfer, so a long clip turns streaming into a rounding error. Ordinary model CPU offload, by contrast, moves the whole transformer onto the GPU while it runs and needs 8 GB or more at once.
4. The VAE and image encoder leave the GPU while denoising
Image-to-video encodes the start image once, through the VAE into latents and through a SigLIP image encoder into features, and then both models sit idle for all twelve steps while holding VRAM. A hook on the conditioning step now moves both to the CPU the moment it finishes and returns their memory to the driver; the VAE only comes back for the final decode. At 121 frames that took each step from 85 s to 56 s, because the freed memory stopped the denoiser spilling.
5. Slicing per-token work along the sequence
The key observation: the q, k and v projections, the norms, the rotary position encoding and the entire feed-forward network all treat each token independently. Splitting them along the token axis is mathematically exact. Only attention mixes tokens, and PyTorch's fused attention kernel already streams through keys and values in tiles, so the 280 GB attention matrix is never built. The real peaks were the plain tensors around attention: 382 MB each for Q, K and V, and the 1.5 GB feed-forward middle.
The fix preallocates each output and fills it 8,192 tokens at a time. This is the feed-forward wrapper from the generation helper:
def sliced(forward):
def run(hidden_states, *arguments, **keywords):
length = hidden_states.shape[1]
if length <= slice_length:
return forward(hidden_states, *arguments, **keywords)
out = torch.empty_like(hidden_states)
for start in range(0, length, slice_length):
out[:, start:start + slice_length] = forward(
hidden_states[:, start:start + slice_length], *arguments, **keywords)
return out
return run
for block in transformer.transformer_blocks:
block.ff.forward = sliced(block.ff.forward)
Temporaries dropped from gigabytes to about 130 MB, and the output matches the unsliced version to within 1e-7. Only the real, unpadded prompt tokens are appended to the sequence, so no attention mask is needed; a mask would force a slower and hungrier attention path. Before this change the 241-frame run spilled and took more than ten minutes per denoising step. After it, the peak was about 4.7 GB and nothing spilled.
6. Decoding in chunks, with warm-up
The decoder supports spatial tiles but decodes every frame of a tile at once, which at about 24 GB per feature map is impossible. So I added chunking in time: six latent frames, roughly 24 video frames, at a time, each moved to the CPU as soon as it is decoded.
The catch is that this VAE is causal in time. Each frame's decode depends on the frames before it, so a chunk started cold comes out slightly different and the join shows. My first attempt cross-faded overlapping chunks, and the joins measured only 31 to 40 dB PSNR against a whole-clip decode: visible. The final version re-decodes the four latent frames before each chunk as warm-up and throws those frames away. The joins now match a whole-clip decode at 44 dB PSNR or better, which is invisible. Earlier, with the library's default tile size, decoding alone reserved 11.46 GB on a card with 8 GB VRAM and ran for nine minutes while spilling. Smaller tiles, clearing the cache first and chunking fixed it.
The Windows trap: spilling looks exactly like fitting
PyTorch keeps freed GPU memory reserved for reuse, and Windows still counts it as used, so between phases the cache is explicitly returned. The bigger lesson was this: on Windows, running out of VRAM does not raise an error. The driver silently pages the overflow into system RAM over PCIe and the job carries on four or more times slower. It was the single biggest source of mysterious slowness in the whole project. So the helper reports memory continuously, and Forge shows a live line in the Generate panel: GPU memory in use, and how much has spilled to RAM. If it is not visible, it is not measured.
Upscaling: SeedVR2, and two instructive failures
I tried Real-ESRGAN first. It is a single-image model, so it upscales each frame alone and re-invents texture on every frame: skin comes out waxy and painted, and hairlines look cut out and wobble from frame to frame. SeedVR2 restores a batch of neighbouring frames together in one diffusion step, so the detail it adds holds still. On the same clip it measured about 15% sharper, with natural hair, skin pores, stubble and eyebrow strands.
The same frame, upscaled per frame (left) and as a video (right).
Six consecutive frames. Top: Real-ESRGAN re-invents the hairline each frame. Bottom: SeedVR2 holds it steady.
SeedVR2 3B is a 3.4 GB FP8 transformer with 32 blocks plus a 0.5 GB VAE, and its packaging recommends 16 GB. It got the same playbook: blocks streamed from RAM, a tiled VAE with 512 px encode and 384 px decode tiles on 8 GB VRAM, frames in batches, and the models kept in RAM between chunks.
Failure one: the silent spill, again. The first run used all 8 GB of VRAM, spilled about 3.3 GB, and took 94 s for a single five-frame encode batch. The fix was to cap the process at whatever VRAM is free when it starts, with torch.cuda.set_per_process_memory_fraction. Overflow then becomes a real out-of-memory error instead of a silent spill, and the helper retries with smaller tiles or batches. The same 33-frame test then ran in under three minutes, peaking at 6.9 GB with nothing spilled.
Failure two: the cross-fade that caused a stutter. With five-frame batches, detail visibly pulsed at each batch join, about 30% more change from frame to frame across a join than within a batch. The obvious fix was to overlap batches by three frames and cross-fade them, and it measured fine on the detail-change metric. In playback, though, I could see motion-blur hiccups. A per-frame sharpness plot showed why: the two batches restore their shared frame differently, and a 50/50 blend of two diffusion outputs is a ghosted frame at about 70% of its neighbours' sharpness. That happened every six frames, about two and a half times a second after interpolation and slow-down. Averaging two diffusion outputs of the same frame does not give a better frame; it gives a blurred one, and a clean step in detail is far less visible than a regular blur. The metric said success. My eyes, and a better metric, said otherwise.
The final settings use no overlap and the biggest batch that fits. 33-frame batches did not fit. 25 failed with 870 MB free, but fragmented, with no single block for a 476 MB request. Switching PyTorch to CUDA's own pooling allocator, backend:cudaMallocAsync, removed the fragmentation and 17-frame batches fit. All 241 frames upscaled in about 19 minutes at a 6.9 GB peak, with no spill and no soft frames, and only 5 of the 14 batch joins show a sharpness step over 15%, against 62 jumps with the cross-fade.
Interpolation: order matters
RIFE 4.26 estimates optical flow between each pair of frames and warps both towards the in-between moment. The original frames are written untouched, and for 2× one new frame goes between each pair. I checked it on the fastest motion in the clip, the fork going into the mouth and the noodles lifting: tines and single strands stay single, with no ghosting. Upscale first, then interpolate. The upscaler then has half the frames to do, and a per-frame upscaler would otherwise treat the softer synthesised frames differently from the real ones, pulsing every other frame.
How it lives in the app
Forge VFX is written in C++, and the multi-gigabyte Python and CUDA stack stays outside it. Generation runs in Python helper processes that talk to the app over a line protocol on standard output, PROGRESS, MEMORY, DONE and ERROR, so the app builds and runs without any of it installed. Models come from a catalogue grouped by family, each declaring its memory needs, and the Generate panel warns, without blocking, when a machine is below a model's minimum. Each video model's clip imports at the frame rate it was trained for, so motion plays at the speed it was generated, and every result lands on a new layer timed to its source footage, ready to composite.
Forge VFX (C++)
├── Generate panel model catalogue, live VRAM / spill readout
└── Python helper one process per job, line protocol on stdout
├── HunyuanVideo 1.5 image → 241 frames @ 720 × 544
├── SeedVR2 3B 2× video restoration, 17-frame batches
└── RIFE 4.26 24 → 48 fps
What I took away
- VRAM is a cache, RAM is the store. Keep only what the GPU needs this instant, and stream the rest.
- Long sequences make streaming cheap. Compute grows faster than transfer, so the bigger the job, the smaller the cost of moving weights.
- Per-token work slices exactly. Only attention mixes tokens, and fused kernels already handle attention.
- Make the invisible visible. On Windows, fitting and spilling look identical until you measure, so the app measures all the time.
- Metrics can lie. A detail-change measure called the cross-fade a success; a per-frame sharpness plot, and watching the clip, showed a stutter.
The progression tells the story best: 33 frames in about four minutes, then 121 frames in about fifteen, then all 241 frames in one run, about 4.7 GB, on a laptop. Working inside a hard constraint is, honestly, the fun part. You can see what the rest of the app does on the Forge VFX page.