Shader Variants
dev_opaque.frag is one shader that can draw every surface the engine knows: lightmaps, the probe volume, reflection probes, the detail layer and its decals, matcaps and envmaps, parallax, stochastic tiling, voxel light, water caustics and fog, and every debug view. Most draws use a handful of those. On a desktop GPU the rest costs little, because a branch on a uniform skips it. On a mobile GPU it costs nearly everything: the compiler allocates registers for the busiest point anywhere in the program, so a block a draw never reaches still decides how many waves of that draw fit on the GPU, and what spills goes out to memory.
So the scene pipeline is specialized per draw. Each draw records with a variant of the shader that compiled in only the blocks its material and its frame can reach. This is the same idea as Poiyomi’s and Mochie’s shader “locking”, where a material’s shader is rebuilt with every disabled feature stripped, except that it is done with a Vulkan specialization constant at load time instead of by rewriting the source.
What a variant is
Section titled “What a variant is”The fragment stage declares one uint specialization constant, Features (constant_id 2), and one boolean per optional block derived from it (UsesDetail, UsesWater, …). Every optional block’s runtime gate is ANDed with its boolean, so a cleared bit is a constant false that the driver’s compiler folds away together with everything only that block used. The constant defaults to all bits set, which is the whole shader.
| Bit | Block | Owned by | Set when |
|---|---|---|---|
Detail | detail layer, matcap, envmap, UV animation, channel UV sets | material | the material has a detail row |
Decals | the decal list | material | its row names a decal block |
Height | parallax march and self-shadow | material | a base height field, or a detail height field with depth |
Stochastic | stochastic anti-tiling | material | the base or the detail layer asks for it |
Lightmap | the lightmap read and baked specular | material and frame | a lightmapped material while an atlas is bound |
BlendMask | the detail layer’s blend mask or height blend comparison | material | its row binds a blend mask or authors heightBlend |
ProbeVolume | irradiance volume and light probe group | frame | giBoundsMin.w is live (a bake is loaded and client.render.giVolume is on) |
Reflections | reflection probe gather | frame | a probe table is bound |
VoxelLight | the voxel light field | frame | a voxel light pool exists |
Water | water column daylight, caustics, underwater fog | frame | the eye is submerged, or a body’s column is in the frame |
SunShadows | cascaded shadow sampling | frame | shadows are on and the sun has intensity |
ScreenOcclusion | the AO upsample | frame | client.render.ao is active |
DebugViews | client.render.debugView and the cascade tint | frame | either is on |
Cutaway | the third-person cutaway’s coverage math | frame | client.camera.cutaway.enabled is on and its fade has strength above zero |
PlayerShadow | the grounding shadows under players | frame | client.render.playerShadow is not none and at least one player is shadowed this frame |
ProbeBlend | a light probe group’s per-pixel blend of its cell’s list | frame | a probe source is live and the frame does not light each leaf as one (client.render.giPerPixel on) |
ProbeLattice | the irradiance volume’s lattice boxes: the region walk and the eight-corner blend | frame | a lattice volume is live, which a map lit only by a light probe group never has |
SceneFeatures in Rendering/SceneFeatures.cs is the C# side, bit for bit, pinned by SceneFeaturesTests. The draw-dissolve demote (the session-end dissolve, the proximity fade, the first-person head chop and the loading stand-in’s animation) is never a bit: every variant keeps it. The one exception inside that call is the cutaway’s term, which is off by default and has a bit of its own, so a variant without it skips the math entirely.
The contract: a variant never changes the picture
Section titled “The contract: a variant never changes the picture”A bit may be cleared only where the block’s own runtime gate in the shader is already false for that draw in that frame. That is why a variant draws the same picture as the whole shader: specialization removes code that would have been branched around, never code that would have run.
SceneFeatureKey enforces this by deriving every bit from the same values the shader’s gates read:
OfMaterialreads the draw’s push-constant flags and its detail row (the row flags, the decal block, the detail height depth). A row the table does not hold proves nothing absent, so the whole material half stays in.OfFramereads the frame block exactly as it was written for that view (one result per uniform slot, so a camera preview or an eye gets its own), plus whether a reflection table and a voxel pool are bound.Combinekeeps a material’s blocks within what its frame allows and adds the frame’s own blocks.
Measured on the Steam Frame (Turnip), every Night City and MltnCity capture with MSAA off is byte-identical between the main-branch whole shader and the variants. MSAA captures are not bit-reproducible on that hardware even run to run (see verify-render-determinism.bat), which is why the comparison pins client.render.msaa=1. A driver may still schedule floating-point work differently in a smaller program, so a variant is not promised to be bit-exact on every driver; d1_canals_01 showed one frame of fifteen with two bytes one step apart.
How a draw gets its variant
Section titled “How a draw gets its variant”Renderer.BindSceneVariant computes the key for each scene draw (opaque, blended, the backdrop’s, and the coverage-capture depth pass) and asks ScenePipelines.Resolve for a pipeline. Consecutive draws that resolve to the same pipeline share one bind.
ScenePipelines still builds the whole-shader set up front, one pipeline per kind, backdrop and blend equation, exactly as before. Beside it sits a SceneVariantCache, keyed by kind, backdrop, blend and features (normalized, so an opaque key ignores the blend and a key holding every known bit is the whole shader):
- Live client: compiled in the background. The first draw that asks for a missing variant starts its compile on the thread pool and records with the whole shader. A later frame picks the variant up once it is ready. Because the whole shader draws the same picture, a missing variant costs time and never a wrong frame. This is Apple’s pattern from WWDC23 (draw with the general shader while the specialized one compiles, then switch).
- Captures: compiled inline.
HeadlessRenderersetsRenderer.CompileVariantsInline, so a variant compiles before the first draw that needs it and every frame a--render,--renderUior--gpuTimingsrun writes is the variant’s: the same pipeline a live client settles on, from the same key function. That keeps the determinism and render proofs meaningful.
Every compile is logged on the render channel with its key and its time, inline or background (scene variant Opaque [Height, Water] compiled in the background in 345 ms), so a variant compiled mid-play is never silent. A compile that fails is logged once and its draws stay on the whole shader.
client.render.shaderVariants (on) turns specialization off live: every draw records with the whole shader, and the variants already compiled stay cached. It is the A/B a suspected variant artifact is compared against.
What chooses a variant
Section titled “What chooses a variant”- The content. A map’s materials decide the material bits, and its bake decides
Lightmap,ProbeVolumeandReflections. A map that uses none of them never compiles them. - The quality preferences. Shadows (
client.render.shadowDistanceand the shadow toggle), ambient occlusion (client.render.ao), indirect light (client.render.giVolume), the per-pixel probe blend (client.render.giPerPixel) and the debug views each flip a frame bit. Turning one on compiles the variants that need it the next time they are drawn; turning it off moves the draws onto smaller variants. A quality-presets feature drives exactly these preferences, so a preset is also a choice of variants.
What a variant costs
Section titled “What a variant costs”A variant is one more graphics pipeline. How many exist is bounded by the feature combinations the loaded content and the client’s settings actually produce, not by the 131,072 the seventeen bits could spell: Night City needs one, MltnCity five, d1_canals_01 eight (counting its backdrop and blended kinds). The cache lives as long as the pipeline set, which a sample-count change or a shader reload rebuilds.
Compile time on the Steam Frame (Turnip, cold Mesa cache) is about 200 to 300 ms for a small variant and about 0.6 to 1 s for a heavy one (probe volume, reflections, detail, water and shadows together). Mesa’s on-disk shader cache makes every later launch nearly free (1 to 2 ms a variant). On the live client none of that time is on the frame: it runs on the thread pool while the whole shader draws.
Measured on the Steam Frame
Section titled “Measured on the Steam Frame”GPU time of the opaque pass, --render ... --gpuTimings at 1600x900 with the player defaults (MSAA 4x), Adreno 750 under Turnip:
| Map | Main (whole shader) | Whole shader, array fixes | Variants |
|---|---|---|---|
| Night City | 390 ms | 18.7 ms | 1.8 ms |
| MltnCity | 673 ms | 37.8 ms | 9.0 ms |
| d1_canals_01 | device lost | 129 ms | 66.7 ms |
d1_canals_01’s remaining cost was real work rather than registers: its light probe group (13,170 probes in 2,084 cells) walked a partition up to 24 deep and blended every valid probe a cell lists, per pixel, and every pixel looped every ranked light. Four changes since:
- The walk starts at a voxel grid that names the deepest node, or the cell, each voxel lies wholly under, so the lookup lands in the same cell in a load or two (see the group’s lookup). Exact.
- Each pixel loops only its screen cluster’s lights (clusters). Exact: a light left out is one whose range the pixel is outside.
client.render.giPerPixeloff lights each leaf as one probe, as Source lights a leaf, and clearsProbeBlendso the list loop compiles out. Medium and Low set it; High and Ultra keep the per-pixel blend.- A map with no lattice compiles the lattice out (
ProbeLattice). The lattice’s region walk and eight-corner blend were what held gm_construct’s group variants at 32 and 48 full registers even with nothing to sample; without them they are 19 and 32 (10 and 6 waves).
Measured on the Frame at 1600x900 from spawn1, whole GPU frame (every pass), repeated runs within 0.1 ms. Before is main at 2b9eac07a; High is the defaults with MSAA 4x, Medium and Low the presets; the pallets are the installed ones, whose lightmaps are a version this build no longer reads, so world surfaces take the probes and the live lights:
| Map | High before | High after | Medium before | Medium after | Low before | Low after |
|---|---|---|---|---|---|---|
| gm_construct | 120.0 ms | 57.2 ms | 76.6 ms | 23.3 ms | 17.5 ms | 9.7 ms |
| gm_flatgrass | 27.9 ms | 21.3 ms | 25.6 ms | 8.9 ms | 4.0 ms | 4.0 ms |
| d1_canals_01 | 82.0 ms | 51.8 ms | 66.9 ms | 28.0 ms | 14.9 ms | 13.2 ms |
What gm_construct’s Medium frame still spends: the leaf probe about 7 ms, the lights about 4.7 ms, MSAA 2x about 2.5 ms. A lightmapped surface skips both the probes and the mixed lights, so a re-imported map’s world sheds most of the first two.
The ir3 compiler’s own statistics (IR3_SHADER_DEBUG=disasm) for the opaque program show where it went. Main’s whole shader was 45,000 instructions with 5,127 memory instructions (cat6, nearly all of them scratch loads and stores of spilled arrays), 48 full registers and 4 waves. After the register fixes below the whole shader is about 17,900 instructions with 14 to 34 memory instructions, still at 48 registers and 4 waves. Night City’s variant ([SunShadows]) is 1,352 instructions, no memory instructions at all, 12 full registers and 8 waves with double thread size.
Register rules for scene shaders
Section titled “Register rules for scene shaders”The array fixes were worth a factor of 20 on their own, on every variant, because the scene shader had two habits a tiler’s compiler cannot keep in registers:
- No dynamically indexed local arrays.
fetchUvSetbuilt avec2[4]of the UV sets and indexed it by a runtime set number, and it is called a dozen times a fragment. ir3 keeps such an array as an addressable register range, which pins the whole range for its lifetime, and spills what does not fit. It is now a select chain. On its own (with the old probe gather below) it took Night City from 390 ms to 36 ms. - No arrays of structs carried between loops. The reflection probe gather copied up to four probe structs (64 floats) into a local array for the blend to walk again. The gather now returns the table index of its last candidate, and the blend re-reads the probes from the table, which costs a few cached loads and no registers. On its own (with the old UV-set array) it took Night City to 112 ms; together with the select chain, 19 ms.
- Read what is used, when it is used. The probe volume scattered 27 halves through a
float[28]before assembling nine coefficients, unpacked a second probe’s nine beside the first’s for a set crossfade, and kept eight corners’ slots, weights and validity bytes in arrays across the whole blend.giCoefficientnow reads one coefficient’s halves directly, the crossfade mixes the previous set in coefficient by coefficient, andgiCornerrecomputes a corner where it is needed. On d1_canals_01 that took its heaviest variant from 48 registers to 32 (4 waves to 6).
When a new block goes into dev_opaque.frag, check the program on a tiler: IR3_SHADER_DEBUG=disasm and MESA_SHADER_CACHE_DISABLE=true on a headless --render, then grep "; FRAG prog" the log for instr, cat6, full and max_waves.
Adding a block
Section titled “Adding a block”- Gate it in the shader with a new
Uses*boolean on the next free bit ofFeatures, ANDed into the block’s existing runtime gate. - Add the member to
SceneFeaturesand toMaterialOwnedorFrameOwned. - Derive it in
SceneFeatureKeyfrom the same lanes the gate reads, so the bit is set whenever the gate could be true. SceneFeaturesTestsfails until the shader and the enum agree.
Half precision (not yet)
Section titled “Half precision (not yet)”Running the safe parts of the shader at 16-bit precision would roughly halve their register weight and double their ALU rate on Adreno. It is deliberately not done yet: GLSL precision is fixed when the SPIR-V is generated, so a client.render.halfPrecision preference needs the shader compiled twice (a mediump build beside the full one) and a per-variable argument for where 16 bits is safe. The plan is for it to default on for mobile-class devices and the lower quality presets and off on desktop High and Ultra, with the on/off difference measured and shown before any default ships.