Skip to content

Shader Variants

dev_opaque.frag is one shader that can draw every surface the engine knows: lightmaps, the probe volume, reflection probes, the detail layer and its decals, matcaps and envmaps, parallax, stochastic tiling, voxel light, water caustics and fog, and every debug view. Most draws use a handful of those. On a desktop GPU the rest costs little, because a branch on a uniform skips it. On a mobile GPU it costs nearly everything: the compiler allocates registers for the busiest point anywhere in the program, so a block a draw never reaches still decides how many waves of that draw fit on the GPU, and what spills goes out to memory.

So the scene pipeline is specialized per draw. Each draw records with a variant of the shader that compiled in only the blocks its material and its frame can reach. This is the same idea as Poiyomi’s and Mochie’s shader “locking”, where a material’s shader is rebuilt with every disabled feature stripped, except that it is done with a Vulkan specialization constant at load time instead of by rewriting the source.

The fragment stage declares one uint specialization constant, Features (constant_id 2), and one boolean per optional block derived from it (UsesDetail, UsesWater, …). Every optional block’s runtime gate is ANDed with its boolean, so a cleared bit is a constant false that the driver’s compiler folds away together with everything only that block used. The constant defaults to all bits set, which is the whole shader.

BitBlockOwned bySet when
Detaildetail layer, matcap, envmap, UV animation, channel UV setsmaterialthe material has a detail row
Decalsthe decal listmaterialits row names a decal block
Heightparallax march and self-shadowmateriala base height field, or a detail height field with depth
Stochasticstochastic anti-tilingmaterialthe base or the detail layer asks for it
Lightmapthe lightmap read and baked specularmaterial and framea lightmapped material while an atlas is bound
BlendMaskthe detail layer’s blend mask or height blend comparisonmaterialits row binds a blend mask or authors heightBlend
ProbeVolumeirradiance volume and light probe groupframegiBoundsMin.w is live (a bake is loaded and client.render.giVolume is on)
Reflectionsreflection probe gatherframea probe table is bound
VoxelLightthe voxel light fieldframea voxel light pool exists
Waterwater column daylight, caustics, underwater fogframethe eye is submerged, or a body’s column is in the frame
SunShadowscascaded shadow samplingframeshadows are on and the sun has intensity
ScreenOcclusionthe AO upsampleframeclient.render.ao is active
DebugViewsclient.render.debugView and the cascade tintframeeither is on
Cutawaythe third-person cutaway’s coverage mathframeclient.camera.cutaway.enabled is on and its fade has strength above zero
PlayerShadowthe grounding shadows under playersframeclient.render.playerShadow is not none and at least one player is shadowed this frame
ProbeBlenda light probe group’s per-pixel blend of its cell’s listframea probe source is live and the frame does not light each leaf as one (client.render.giPerPixel on)
ProbeLatticethe irradiance volume’s lattice boxes: the region walk and the eight-corner blendframea lattice volume is live, which a map lit only by a light probe group never has

SceneFeatures in Rendering/SceneFeatures.cs is the C# side, bit for bit, pinned by SceneFeaturesTests. The draw-dissolve demote (the session-end dissolve, the proximity fade, the first-person head chop and the loading stand-in’s animation) is never a bit: every variant keeps it. The one exception inside that call is the cutaway’s term, which is off by default and has a bit of its own, so a variant without it skips the math entirely.

The contract: a variant never changes the picture

Section titled “The contract: a variant never changes the picture”

A bit may be cleared only where the block’s own runtime gate in the shader is already false for that draw in that frame. That is why a variant draws the same picture as the whole shader: specialization removes code that would have been branched around, never code that would have run.

SceneFeatureKey enforces this by deriving every bit from the same values the shader’s gates read:

  • OfMaterial reads the draw’s push-constant flags and its detail row (the row flags, the decal block, the detail height depth). A row the table does not hold proves nothing absent, so the whole material half stays in.
  • OfFrame reads the frame block exactly as it was written for that view (one result per uniform slot, so a camera preview or an eye gets its own), plus whether a reflection table and a voxel pool are bound.
  • Combine keeps a material’s blocks within what its frame allows and adds the frame’s own blocks.

Measured on the Steam Frame (Turnip), every Night City and MltnCity capture with MSAA off is byte-identical between the main-branch whole shader and the variants. MSAA captures are not bit-reproducible on that hardware even run to run (see verify-render-determinism.bat), which is why the comparison pins client.render.msaa=1. A driver may still schedule floating-point work differently in a smaller program, so a variant is not promised to be bit-exact on every driver; d1_canals_01 showed one frame of fifteen with two bytes one step apart.

Renderer.BindSceneVariant computes the key for each scene draw (opaque, blended, the backdrop’s, and the coverage-capture depth pass) and asks ScenePipelines.Resolve for a pipeline. Consecutive draws that resolve to the same pipeline share one bind.

ScenePipelines still builds the whole-shader set up front, one pipeline per kind, backdrop and blend equation, exactly as before. Beside it sits a SceneVariantCache, keyed by kind, backdrop, blend and features (normalized, so an opaque key ignores the blend and a key holding every known bit is the whole shader):

  • Live client: compiled in the background. The first draw that asks for a missing variant starts its compile on the thread pool and records with the whole shader. A later frame picks the variant up once it is ready. Because the whole shader draws the same picture, a missing variant costs time and never a wrong frame. This is Apple’s pattern from WWDC23 (draw with the general shader while the specialized one compiles, then switch).
  • Captures: compiled inline. HeadlessRenderer sets Renderer.CompileVariantsInline, so a variant compiles before the first draw that needs it and every frame a --render, --renderUi or --gpuTimings run writes is the variant’s: the same pipeline a live client settles on, from the same key function. That keeps the determinism and render proofs meaningful.

Every compile is logged on the render channel with its key and its time, inline or background (scene variant Opaque [Height, Water] compiled in the background in 345 ms), so a variant compiled mid-play is never silent. A compile that fails is logged once and its draws stay on the whole shader.

client.render.shaderVariants (on) turns specialization off live: every draw records with the whole shader, and the variants already compiled stay cached. It is the A/B a suspected variant artifact is compared against.

  • The content. A map’s materials decide the material bits, and its bake decides Lightmap, ProbeVolume and Reflections. A map that uses none of them never compiles them.
  • The quality preferences. Shadows (client.render.shadowDistance and the shadow toggle), ambient occlusion (client.render.ao), indirect light (client.render.giVolume), the per-pixel probe blend (client.render.giPerPixel) and the debug views each flip a frame bit. Turning one on compiles the variants that need it the next time they are drawn; turning it off moves the draws onto smaller variants. A quality-presets feature drives exactly these preferences, so a preset is also a choice of variants.

A variant is one more graphics pipeline. How many exist is bounded by the feature combinations the loaded content and the client’s settings actually produce, not by the 131,072 the seventeen bits could spell: Night City needs one, MltnCity five, d1_canals_01 eight (counting its backdrop and blended kinds). The cache lives as long as the pipeline set, which a sample-count change or a shader reload rebuilds.

Compile time on the Steam Frame (Turnip, cold Mesa cache) is about 200 to 300 ms for a small variant and about 0.6 to 1 s for a heavy one (probe volume, reflections, detail, water and shadows together). Mesa’s on-disk shader cache makes every later launch nearly free (1 to 2 ms a variant). On the live client none of that time is on the frame: it runs on the thread pool while the whole shader draws.

GPU time of the opaque pass, --render ... --gpuTimings at 1600x900 with the player defaults (MSAA 4x), Adreno 750 under Turnip:

MapMain (whole shader)Whole shader, array fixesVariants
Night City390 ms18.7 ms1.8 ms
MltnCity673 ms37.8 ms9.0 ms
d1_canals_01device lost129 ms66.7 ms

d1_canals_01’s remaining cost was real work rather than registers: its light probe group (13,170 probes in 2,084 cells) walked a partition up to 24 deep and blended every valid probe a cell lists, per pixel, and every pixel looped every ranked light. Four changes since:

  • The walk starts at a voxel grid that names the deepest node, or the cell, each voxel lies wholly under, so the lookup lands in the same cell in a load or two (see the group’s lookup). Exact.
  • Each pixel loops only its screen cluster’s lights (clusters). Exact: a light left out is one whose range the pixel is outside.
  • client.render.giPerPixel off lights each leaf as one probe, as Source lights a leaf, and clears ProbeBlend so the list loop compiles out. Medium and Low set it; High and Ultra keep the per-pixel blend.
  • A map with no lattice compiles the lattice out (ProbeLattice). The lattice’s region walk and eight-corner blend were what held gm_construct’s group variants at 32 and 48 full registers even with nothing to sample; without them they are 19 and 32 (10 and 6 waves).

Measured on the Frame at 1600x900 from spawn1, whole GPU frame (every pass), repeated runs within 0.1 ms. Before is main at 2b9eac07a; High is the defaults with MSAA 4x, Medium and Low the presets; the pallets are the installed ones, whose lightmaps are a version this build no longer reads, so world surfaces take the probes and the live lights:

MapHigh beforeHigh afterMedium beforeMedium afterLow beforeLow after
gm_construct120.0 ms57.2 ms76.6 ms23.3 ms17.5 ms9.7 ms
gm_flatgrass27.9 ms21.3 ms25.6 ms8.9 ms4.0 ms4.0 ms
d1_canals_0182.0 ms51.8 ms66.9 ms28.0 ms14.9 ms13.2 ms

What gm_construct’s Medium frame still spends: the leaf probe about 7 ms, the lights about 4.7 ms, MSAA 2x about 2.5 ms. A lightmapped surface skips both the probes and the mixed lights, so a re-imported map’s world sheds most of the first two.

The ir3 compiler’s own statistics (IR3_SHADER_DEBUG=disasm) for the opaque program show where it went. Main’s whole shader was 45,000 instructions with 5,127 memory instructions (cat6, nearly all of them scratch loads and stores of spilled arrays), 48 full registers and 4 waves. After the register fixes below the whole shader is about 17,900 instructions with 14 to 34 memory instructions, still at 48 registers and 4 waves. Night City’s variant ([SunShadows]) is 1,352 instructions, no memory instructions at all, 12 full registers and 8 waves with double thread size.

The array fixes were worth a factor of 20 on their own, on every variant, because the scene shader had two habits a tiler’s compiler cannot keep in registers:

  • No dynamically indexed local arrays. fetchUvSet built a vec2[4] of the UV sets and indexed it by a runtime set number, and it is called a dozen times a fragment. ir3 keeps such an array as an addressable register range, which pins the whole range for its lifetime, and spills what does not fit. It is now a select chain. On its own (with the old probe gather below) it took Night City from 390 ms to 36 ms.
  • No arrays of structs carried between loops. The reflection probe gather copied up to four probe structs (64 floats) into a local array for the blend to walk again. The gather now returns the table index of its last candidate, and the blend re-reads the probes from the table, which costs a few cached loads and no registers. On its own (with the old UV-set array) it took Night City to 112 ms; together with the select chain, 19 ms.
  • Read what is used, when it is used. The probe volume scattered 27 halves through a float[28] before assembling nine coefficients, unpacked a second probe’s nine beside the first’s for a set crossfade, and kept eight corners’ slots, weights and validity bytes in arrays across the whole blend. giCoefficient now reads one coefficient’s halves directly, the crossfade mixes the previous set in coefficient by coefficient, and giCorner recomputes a corner where it is needed. On d1_canals_01 that took its heaviest variant from 48 registers to 32 (4 waves to 6).

When a new block goes into dev_opaque.frag, check the program on a tiler: IR3_SHADER_DEBUG=disasm and MESA_SHADER_CACHE_DISABLE=true on a headless --render, then grep "; FRAG prog" the log for instr, cat6, full and max_waves.

  1. Gate it in the shader with a new Uses* boolean on the next free bit of Features, ANDed into the block’s existing runtime gate.
  2. Add the member to SceneFeatures and to MaterialOwned or FrameOwned.
  3. Derive it in SceneFeatureKey from the same lanes the gate reads, so the bit is set whenever the gate could be true.
  4. SceneFeaturesTests fails until the shader and the enum agree.

Running the safe parts of the shader at 16-bit precision would roughly halve their register weight and double their ALU rate on Adreno. It is deliberately not done yet: GLSL precision is fixed when the SPIR-V is generated, so a client.render.halfPrecision preference needs the shader compiled twice (a mediump build beside the full one) and a per-variable argument for where 16 bits is safe. The plan is for it to default on for mobile-class devices and the lower quality presets and off on desktop High and Ultra, with the on/off difference measured and shown before any default ships.