Forge VFX: building a compositor on its own renderer
A final render from Forge VFX. The rain is a particle system, held behind the glass by a matte.
Forge VFX is a desktop motion graphics and compositing application: a multi-layer timeline, a real-time viewport, stackable effects with a node editor, 3D layers and cameras, audio, export to production codecs, and local AI generation of video, images and sound. It is about 52 thousand lines of first-party C++ with 59 shaders, written since early 2024, and it runs on macOS, Windows and Linux on its own renderer with no game engine or UI toolkit underneath. This article is a tour of how it is put together and of the decisions that turned out to matter. There is a product page if you would rather see what it does than how.The shape of the program
The architecture is deliberately flat. There is one window loop, one thread, and a render hardware interface that everything draws through, including the interface itself.Application (single thread, cooperative)
├── UI retained widget tree, drawn every frame through the RHI
├── Project scenes, layers, resources, .vxproj (JSON)
│ └── Scene offscreen composite, nested scenes, cameras
│ └── Layer transform, keyframes, effect node graph
├── Effects 16 effects, pixel / compute / CPU passes
├── Audio miniaudio mixing, spectrum analysis
├── Export state machine, VideoToolbox / Media Foundation
├── Generate Python sidecar, depth warp via ONNX Runtime
└── RHI IRenderDevice + IRenderContext
├── Vulkan default: dynamic rendering, sync2, VMA
└── OpenGL reference backend
One seam, two backends
Forge started life on OpenGL. When it moved to Vulkan I did not want a rewrite, I wanted a seam: a small interface that the rest of the program could draw through without knowing which API sat behind it. The result is two interfaces.IRenderDevice creates things and drives frames: swapchains (one per window, all on one shared device), buffers, textures, render targets and pipelines. IRenderContext records work:
class IRenderContext {
public:
virtual void beginRenderPass(RenderTargetHandle target, const float clearColor[4], bool doClear) = 0;
virtual void endRenderPass() = 0;
virtual void bindPipeline(PipelineHandle pipeline) = 0;
virtual void bindTexture(uint32_t slot, TextureHandle texture) = 0;
virtual void bindVertices(const void* data, size_t bytes) = 0;
virtual void pushConstants(const void* data, uint32_t bytes) = 0;
virtual void draw(uint32_t vertexCount, uint32_t firstVertex = 0) = 0;
virtual void bindComputePipeline(PipelineHandle pipeline) = 0;
virtual void bindComputeImages(TextureHandle input, TextureHandle output) = 0;
virtual void dispatch(uint32_t groupsX, uint32_t groupsY, uint32_t groupsZ) = 0;
};
On OpenGL each call executes immediately. On Vulkan the same calls are recorded into a command buffer and submitted at the end of the frame. Resources are typed 32-bit handles, so nothing outside a backend ever sees a GLuint or a VkImage, and shaders are referenced by logical name. "ui/unlitTexture" resolves to GLSL 330 source on OpenGL and to SPIR-V compiled from GLSL 450 at build time on Vulkan, and the two shader trees mirror each other file for file.A few decisions kept the seam small. Every shader takes at most one 128-byte push constant block, so there is no uniform buffer machinery to abstract. There is a single descriptor set layout with two combined image samplers, which covers every effect in the program; unbound slots fall back to a shared 1×1 white texture, and each frame in flight caches descriptor sets keyed by the pair of textures bound. Pipelines are described without a target format, and the Vulkan backend lazily builds a concrete variant per format and sample count, persisted through a
VkPipelineCache so the second launch skips the compile. The Vulkan side uses dynamic rendering and synchronization2, so there are no VkRenderPass or framebuffer objects at all, and the edges between a producing pass and a consuming one become explicit image barriers.The OpenGL backend is kept as the reference. It was written to wrap the exact GL calls the program already made, so moving the whole application onto the seam changed nothing on screen, and it remains the thing I compare Vulkan against when a composite looks wrong. Today that comparison is done by eye with an environment variable,
FORGE_RENDER_BACKEND=opengl, rather than by automated image diffs, which is the obvious next step for it. If Vulkan device creation fails the application falls back to OpenGL on its own, so a machine without a Vulkan driver still runs.Compositing without nested passes
A composition in Forge is a stack of layers rendered into an offscreen target, and any layer can itself be a whole scene. On OpenGL that nests naturally: bind a framebuffer, draw a child scene into it, restore, continue. Vulkan does not allow a render pass to open inside another one, so a scene renders its frame in two phases:PHASE 1 outside any render pass
bake every layer's animated values for this frame
prepare offscreen content, dependencies first:
effect chains, nested scenes, 3D meshes
PHASE 2 one render pass
draw each layer's prepared texture, bottom to top
Phase 1 walks layers depth first over their dependencies, because a layer's effects can read another layer's picture (a hidden displacement map, a matte, a camera), and cycles are skipped rather than followed. Animated values are baked in a separate pass before anything prepares, since a 3D layer may look through a camera that sits anywhere in the stack. Once a frame is composed it is cached behind a dirty flag, so a paused viewport or a UI redraw costs nothing. The coding standard for the project says it plainly: prefer dirty flags and caches to a render graph. For a program with one author, that trade has held up well.Scene targets are RGBA8 today with premultiplied alpha and four blend modes. A half-float format is already defined through both backends for the HDR work that comes later.
Effects, and why the stack is a graph
The node editor. An input feeds a particle system, which a matte mask confines to the windows.
An effect is a list of typed parameter descriptions plus anapply that runs as a pixel shader, a compute shader, or a CPU pass. The inspector builds its controls from those descriptions, so no effect contains UI code, and the same descriptions drive keyframing and project serialization for free. There are sixteen effects so far, from blur, keying and colour tools to masks, a displacement map, an audio spectrum, a particle system and a crowd generator.Every layer owns a small directed acyclic graph with input, output, effect, animator and layer-source nodes. The familiar effect stack is simply what that graph looks like when it is a straight line, and the inspector only offers reordering while it still is one. Connections that would form a cycle or join mismatched types are rejected at the moment you draw them. This gives beginners a stack and gives complex shots a graph, without two systems to keep in sync.
Animation follows one rule: every value is a pure function of time. Keyframes are baked once per frame, and animators layer on top of them. The wiggle animator, for example, takes an amount, a frequency and a seed and evaluates directly at a time in seconds, with no accumulated state. The particle system works the same way: each particle's position is a closed-form function of its index, the seed and the time, with trajectory parameters sampled at its birth frame. Nothing is simulated step by step, so scrubbing backwards, jumping to a frame, previewing and exporting all produce the same picture. Up to twenty thousand particles are built on the CPU each frame, budgeted to a third of the 16 MB vertex ring each frame in flight owns.
One thread
There is nostd::thread anywhere in Forge. Decoding, compositing, audio scheduling, AI generation and export all cooperate on the main loop. Long jobs are written as small state machines that do a slice of work per tick: the exporter renders sixteen offscreen frames per UI frame, then reads them back and hands them to the encoder after submission, and a generation job is polled for progress between frames. It sounds limiting, and on a larger team it probably would be, but it removed a whole class of bugs. There are no locks, no races between the timeline and the renderer, and a GPU resource is never touched from two places at once.Generating inside the timeline
Region-only generation: only the fireplace is regenerated, then feathered back into the photograph.
Generation runs as a Python sidecar. The first time you open the Generate panel it offers to set things up: it downloadsuv, creates a private Python environment in its own folder, checks nvidia-smi to choose between the CUDA and CPU builds of PyTorch, and installs the requirements. Nothing outside that folder is touched, so deleting it removes everything. Models come from a catalogue that states the memory each one needs, among them Wan 2.2, LTX-Video, HunyuanVideo 1.5 and CogVideoX for video, FLUX.1, SDXL and Stable Diffusion 3.5 for stills, and Stable Audio for sound, all run through Hugging Face diffusers.The sidecar speaks a line protocol on standard output,
PROGRESS <0..1> <message>, DONE or ERROR, which the main loop polls without blocking. The generation backend interface is cooperative, with a step inside the frame and a step outside it, so the in-process depth warp and an out-of-process diffusion model sit behind the same interface without the effect, the cache, the project format or the UI knowing which one ran.The feature I use most is region-only generation. Rather than regenerate a whole frame, you draw a rectangle and only that crop goes to the model; the result comes back as a feathered layer over its source. Video memory limits resolution times frame count, so spending it only on the pixels that change buys considerably longer clips on the same card. The fireplace on the product page was made this way: the flames were generated into a 908 by 642 region of a still photograph, then composited back with a mask, particle snow behind the window glass, and animated candles.
Knowing where generation stops being useful shaped the rest of the tool. Video diffusion models are poor at moving many small objects coherently, which is exactly why the crowd effect exists: it instances real 3D models along a path under a perspective camera, so the part of a shot that must be precise is deterministic and the part that benefits from invention is generated. Results are cached by a hash of everything they were made from, the source, the geometry, the parameters and the backend's version, so the cache is never part of the document and is always safe to delete.
Reading video without seeking
Video is decoded with OpenCV, and the naive approach of seeking to every requested frame is slow, because a seek has to find the previous keyframe and decode forward. Measured on real footage, reading the next frame sequentially cost 0.56 ms while a seek cost 27 ms. So each file keeps up to four decode cursors in least-recently-used order, which stops two layers showing the same clip at different times from fighting over one read position, and a cursor will walk forward through gaps of up to 48 frames instead of seeking, with the budget adapting to measured cost. Image sequences sit in a RAM cache sized to an eighth of system memory, capped at 2 GB.Making export fast
Export was the part of the program that most needed measuring rather than guessing. A benchmark mode,FORGE_EXPORT_BENCHMARK=<frames>, prints a per-phase breakdown of each exported frame, including an explicit remainder so that unaccounted time cannot hide. On an Apple M2 at 4K, sustained throughput went from 2.2 to 51.8 frames per second. None of the fixes were exotic:
- The build had been running unoptimized, because CMake defaults to no build type.
- Vulkan validation layers were enabled by default.
- Readback buffers were reallocated on every frame.
- Export frames were being presented to the window, which tied throughput to the display's refresh rate: a hard 30 fps cap on a 30 Hz monitor, identical at 640×360 and at 4K.
- The copy into the encoder was scalar. A NEON version is 3.8× faster on the encoder's write-combined memory, and waiting on the GPU fell 2.7× as a side effect.
- Video layers were seeking instead of walking forward.
Small things that mattered
The interface is set in a variable font rendered through FreeType, using its weight axis rather than separate font files. An early version built a glyph atlas per label, which cost 128 GPU round trips each time; now one ASCII atlas is built per style and size and simply never evicted. Window layouts are JSON descriptors that a Python script turns into C++ headers at build time, so a panel's structure is data while its behaviour stays in code. The project file is JSON with a format version, currently 21, and a comment above the version constant explains every bump, so an old project can always be traced back to what it was missing.Open source
Forge VFX will be free, and I am preparing the source for release under the GNU GPL v3, the license behind Blender, Krita, GIMP and Kdenlive. Strong copyleft suits a tool like this: anyone who ships a modified version has to publish it, so every fork stays open. Every dependency is compatible, and the optional ONNX Runtime is detected when present rather than bundled. The remaining work is mostly housekeeping, such as SPDX headers, a contribution sign-off policy, and a note on codec patents, and a pre-commit hook already keeps commercial product names out of the source.In the meantime the work continues on keyframe easing, shape tools and the export pipeline. If you want to follow along, the Discord is where builds will show up first.