Skip to content

Renderer Textures

Every texture a frame samples is bound through a descriptor set the renderer owns. This page is how those are allocated, retired, kept resident and filtered.

Set 1 is a per-material descriptor carrying all six texture channels as combined image samplers, in the fixed order baseColor, normal, metallicRoughness, occlusion, emissive, height (followed by the detail layer’s own). Every binding is always filled: a channel the material did not bind gets a 1×1 neutral default — opaque white (uploaded sRGB) for the four color/mask slots, a flat (128, 128, 255) tangent-space normal, and a mid-gray height sitting exactly on the march’s reference plane (both uploaded linear) — so a material that declares only a base color shades exactly as it did before PBR landed, with no permutation per channel. A bitmask on the material tells the shader which channels are real, so the normal-map path is skipped entirely rather than multiplied by a no-op. (The larger optional blocks, such as the detail layer, decals and parallax, are compiled out per draw by shader variants; the channel bits stay runtime branches.)

Descriptor pools grow; they are never a ceiling

Section titled “Descriptor pools grow; they are never a ceiling”

A VkDescriptorPool has its capacity fixed at creation, so one pool is a hard cap — and a cap is the wrong shape here twice over. It is wrong for a big map, because the right number is “however many materials the author wrote”. And it is wrong for a reload even at generous sizes, because ClientRuntime.ReloadMap deliberately meshes the entire new scene before disposing the live one — the rollback that keeps the current map fully rendered when meshing fails — so a switch’s peak is both generations’ material tables at once, not one.

So set 1 is allocated out of a DescriptorSetArena: a chain of fixed-size pool blocks (SetsPerBlock), allocating out of the first block with a free slot and creating another when none has one. Blocks carry FREE_DESCRIPTOR_SET_BIT, so a set freed mid-life returns its slot to the block that issued it — vkFreeDescriptorSets takes the pool the set came from, so every live handle is mapped back to its block.

An arena is a lifetime, and that is the part that matters:

  • The registry keeps one long-lived arena for its neutral default set and for the object cache’s sets, which outlive any one map.
  • Every RenderScene gets an arena of its own, opened with the scene and destroyed whole in Dispose (after releasing its sets, with the device idle). A reload’s two generations therefore each grow a chain of their own instead of competing for one budget.
  • Destroying a pool frees its sets implicitly, so any of that arena’s handles still aging out of the retirement queue are dropped rather than freed (MaterialSetLedger.ForgetRetiring) — freeing them individually afterwards would be a use after free.

The block size is a const, not a preference, on purpose: a pool’s capacity is fixed when the pool is created, so changing it at runtime could not affect a block that already exists, and since the arena grows on demand it changes nothing but how many vkCreateDescriptorPool calls a large map costs.

Slots still have to come back, because a client that server-hops all session loads and unloads maps over and over — an arena that only ever grows is a leak with a friendlier name. So the rule is: a material set is retired by whoever created it, on that owner’s teardown.

  • The registry allocates; it does not own. CreateMaterialSet hands back a set and tracks it (so a mip upgrade can rebuild it and a mipmaps toggle can rewrite it), but the caller is on the hook for ReleaseMaterialSet. For a map’s materials that caller is the render scene, which releases every set it asked for in Dispose — first, before anything those sets bind is destroyed.
  • Texture disposal is only a safety net. When a texture dies, every set still bound to it is retired too, because a descriptor may not outlive what it points at. That is a correctness backstop, not the accounting rule — and it cannot be, because a material that authors no texture channels binds only the registry’s neutral 1×1 defaults, which outlive every map. Nothing about such a set ever dies, so nothing would ever reclaim its slot: one leaked slot per all-default material per map load, growing the arena forever.
  • The two paths are idempotent against each other. A scene releases everything it created without caring which of its textures happened to go first, and a set a texture teardown already took is silently accepted rather than double-freed.
  • Retirement is deferred, never immediate. A handle recorded into a command buffer must outlive that buffer’s execution, so a retired set ages out over frames-in-flight + 1 render-thread pumps before vkFreeDescriptorSets gets it. The same rule the replaced image and view of a mip upgrade follow.

The bookkeeping lives in MaterialSetLedger, and the arena’s growth and block-ownership rules live in DescriptorSetArena, both deliberately split out of the registry with no device and no Vulkan calls in them, so the accounting is covered by headless tests rather than by reading the code and hoping. The regression case that matters most models the overlapping reload ordering — generation B registered in full before generation A is released — because a teardown-first test peaks at one map’s worth of sets and can never see the failure a real map switch produced.

Per-draw material state rides push constants, and the opaque layout uses exactly 128 bytes — a 64-byte model matrix in the vertex stage plus 64 bytes of material parameters in the fragment stage. 128 bytes is the Vulkan guaranteed minimum for maxPushConstantsSize, so this is the portable ceiling rather than a device-specific one, and a unit test asserts the total.

The practical consequence for future work: there is no room left for another per-draw field. Anything new has to move into a uniform buffer indexed by a handle passed through the existing bytes, not appended to the push block.

One lane is deliberately a bit field for this reason. MaterialFlags packs into the PBR vector’s w as a float, and it carries both the five texture-channel bits (each 1 << binding, so the mask and the descriptor slot can never disagree) and the material’s boolean switches — Unlit, and the alpha-mask cutout. A new per-draw boolean costs a bit rather than a lane; only a new per-draw number is out of room, which is why the cutoff itself had to displace the unlit flag out of params.y.

GPU resources stay resident across map loads

Section titled “GPU resources stay resident across map loads”

Re-uploading every mesh, table and texture a map references on every hot reload would waste seconds copying gigabytes the GPU already holds: a map pallet recompile usually changes one thing, and the rest of the map — and every dependency pallet it reads — was not touched.

So a render scene is brought up as the difference from the scene it replaces. RenderScene.Apply is the one build: a first load is the difference from an empty map, a rebuild the difference from the same map’s last build, and a switch the difference from another map. Residency is owned by the renderer, not by a scene, and every resource is identified by its content:

  • A pallet texture is the (pallet id, entry path, CRC32) triple the pallet file table already carries, plus the size cap it was loaded under. That checksum folds the entry’s texture metadata as well as its bytes, so a re-authored image and a re-cut one are both misses, and reading it is a file-table lookup rather than a read or a decode.
  • Anything the scene derives — the merged map mesh, the brush mesh, the outline index buffer, the lightmap, the sky cube, the reflection cubes and their table, the GI volume, the detail, decal, water and glass tables, the shared avatar and gizmo meshes — is keyed by a hash of exactly what it hands the device. Two builds that would make the same copy make it once.
  • A table written in place leaves the residency first. A material retune rewrites a row of the water or glass table on the device; the scene withdraws that table’s key before the write, so no later load takes it for the pristine rows the key describes. Whoever already holds it keeps it.

CompiledMapContent and MapContentDiff (in Content) say what changed, by identity rather than by position: pallet entries by path and checksum, entities by authored Guid, nodes by id, material slots by name. Every build logs one line naming the difference and what it cost — map rebuild 'test.apply.map:room.dh-map': nothing changed; uploaded 0 bytes (0 B), reused 10 resource(s) (118 KB) — which is the only line that tells an unchanged rebuild from a full one.

  • The decision is made once, in pass one. The scene build folds every material before it uploads anything, and that first pass asks the residency about each channel. A hit is bound there and never queued for decode; only a miss becomes a decode request. Skipping the decode is most of the saving — the copy is the smaller half.
  • Pass two may not ask again. The decode queue is in-order by construction, so the gating in pass two mirrors pass one’s answer exactly. Re-probing would desync it the moment two materials shared a texture: the first would upload and make it resident, the second would then skip a payload the queue had already produced, and every later channel would get the wrong pixels.
  • Eviction is a reference count, not a sweep. Each scene holds a lease (its SceneResources ledger) with one reference per resource it resolved; the scene’s disposal closes the lease and destroys only what nothing else holds. Because a switch builds the whole new scene before disposing the old one, both generations hold references across the swap, so a shared resource’s count never reaches zero and no destroy can ever land before the commit. A build that fails drops only its own references and leaves the live map fully resident.
  • client.render.textureResidency (default on) turns it off, for A/B against re-uploading everything on every reload. It is read once when a load begins and held for the whole build — a switch flipped between two slices would leave half a scene owned by the residency and half by the scene.
  • The desktop client and the offscreen renderer share it. The headless renderer the mobile shell draws through takes the same residency, so an iPad rebuild re-uploads exactly what a desktop one does.

The load summary names both halves: vram upload 0.4s (48 MB, reused 174 3.1 GB). Without the reuse half a hot reload just reads as a load with nothing to do, and the copies that did not happen are invisible.

A process decides once, at startup, what it spends on textures, and logs the decision (texture cap 2048 (automatic; 2.1 GB available to the process); decoded pixels kept up to 96 MB) so the next memory kill is readable from the client log rather than only from the OS’s report.

  • client.render.maxTextureSize (default 0 = automatic) is the largest edge a texture is loaded at. When automatic, the cap is fitted to each map before its first texture decodes: the loader costs the map’s whole texture set at every candidate cap (none, 4096, 2048, 1024, 512, 256, 128) from the pallets’ file tables alone — a mip companion’s stored size is the pixels of every level but the base, so a texture’s GPU cost is about four times it and its base edge follows — and takes the largest cap whose total fits client.render.textureMemoryShare (default half) of what the platform says the process may still take at that moment. iOS has one memory pool, kills the process that oversteps its per-process limit, and that limit does not track the device’s physical memory, so the reading is the only honest input and a device table would be wrong in the direction that looks safe. A desktop hands over no reading and gets no cap. The line the loader logs (map textures 3.70 GB at full size across 197; cap 1024 costs 0.96 GB of 2.39 GB allowed (4.90 GB available to the process)) is the number to read after a memory kill. MLTN City shows why this matters: 197 textures, 3.7 GB as RGBA on an iPad with 4.9 GB to spare — a memory tier choosing 4096 kills the process at 5.0 GB.
  • A texture over the cap is loaded from the first level of its mip chain that fits. Wherever the pallet shipped a mip companion — every dh build texture does — that happens before the PNG is decoded: the companion’s first level is half the base, so the base’s size is known from the companion alone, and the full-size pixels never exist in memory. Only a texture with no companion is decoded whole and reduced with the same box filter the runtime mip pump uses.
  • The resident-texture identity carries the cap it was loaded under, so a copy uploaded at 1024 is never reused when a later, larger cap asks for the same file.
  • client.render.decodedTextureBudgetMb (default 0 = the host’s own default: 1024 MB on a desktop, 96 MB on iOS) bounds the decoded pixels the asset cache keeps once their GPU copies exist. The accounting is process-wide across every pallet’s cache, oldest use first; an evicted image is decoded again from the pallet the next time something reads it (the editor’s material preview, the lightmap view), and the GPU copy is untouched. Without this budget, every decoded texture stays in memory for the life of the map beside its GPU copy — enough to put MLTN City’s 1 GB texture library at a 5.4 GB process footprint on an iPad.

The texture budget above is about host memory and a per-map cap. This is about the card: how much video memory the process may fill, and what a load does when it would fill it. Without it a large map allocates until the driver refuses, and one refusal in the wrong place (a shared table upload) ended the process.

  • The pool. The engine reads how much of the card it may use from VK_EXT_memory_budget (asked for wherever the device has it, never depended on) and takes client.render.vramShare (default 0.9) of the device-local heaps’ budget. Without the extension it takes the heap’s size times that share times client.render.vramHeapFallbackShare (default 0.75), since a heap size knows nothing about other processes. client.render.vramBudgetMb (default 0 = read it) overrides both and is taken as the pool as written, so client.render.vramBudgetMb 6500 rehearses an 8 GB card on a 16 GB one. On a unified-memory device the pool is also held to what the process may still take (the iPad’s os_proc_available_memory reading). The device line in the render log names which of these decided: GPU budget 14.92 GB of 15.99 GB device-local (VK_EXT_memory_budget), pool 13.43 GB at share 0.90.
  • The pool follows the budget like Source 2’s. The driver is read again every client.render.vramPollSeconds, and the pool moves toward the new reading at client.render.vramGrowMbPerSecond (default 64) upward and client.render.vramShrinkMbPerSecond (default 256) downward: it sheds four times faster than it grows, so another process taking memory is answered quickly and a budget that is only briefly larger is not chased. A change of the setting itself resizes the pool at once.
  • What counts against it is everything the device holds, plus the texture bytes of loads that have decoded and not yet uploaded, so two loads deciding at the same moment do not both spend the same room.
  • Admission. An object’s textures are decoded off the frame thread. When loading one whole would cross client.render.vramHighWater (default 0.95 of the pool), it starts at the first mip level whose chain fits, never below the resident tail (client.render.textureTailEdge, default 128 texels; a texture already that small is never reduced). The level is chosen from the entry’s header before any pixel is decoded: a PNG entry reads only its .mipchain companion and the PNG is never decoded, and a KTX 2 entry transcodes from the level it names. It goes through the same FirstLevelWithinCap rule as the global cap, as a cap on that one request. The done line for an object says how many of its textures started below their top level.
  • A map’s own textures are admitted against the pool too, as a set. A map’s whole texture list is known before the first one decodes, so the load prices it from the pallets’ file tables (the same companion-size reading the texture budget uses, over the textures the map’s materials actually name rather than the whole sibling closure) against what the pool has left under the high watermark, less what the map’s geometry and its lightmap will take (the lightmap is bound whole whatever the pool says, so its size in the storage the device chooses is set aside first; client.render.vramDecodeLightmaps makes it the decoded size). When the set fits, every texture loads whole. When it does not, the load drops the fewest top levels that make it fit, from every texture alike, and hands one level back to the small textures first, as many of them as still fit, since they cost the least to keep. Each texture is then admitted exactly as an object’s is, from its own header, so a texture that is larger than its file-table price said still cannot overspend. A texture that does not fit even at its tail is left unbound: the material shows the registry’s neutral default for that channel, one warning names it, and the load goes on. One line says what happened: map textures admitted against a 6.30 GB GPU pool (0.16 GB spent): 763 loaded, 5.81 GB at full size, 1.4 GB admitted; 2 top level(s) dropped from textures over 1448 px, 1 from the rest, 751 started below their top (26 at the 128 px tail); priced from the file tables: 737 textures, 5.81 GB whole, 0.68 GB at that drop. The pool counts the outgoing map’s textures while a switch builds the incoming one, because both are on the card until the swap, so a switch out of a large map loads the next one a little softer than a cold start would.
  • A full pool defers, it does not fail. When even the tail of a texture does not fit what the pool has left, or the device refuses an allocation while an object is uploading, the object keeps its stand-in and is deferred, not marked failed for the session. It is loaded again once the pool has the room it said it needed (client.render.vramRetrySeconds, default 1, is the least time between tries). One line says so: object 'X' deferred: needs 34 MB at its smallest, pool has 2 MB; stand-in kept. A stand-in silhouette that cannot be uploaded is drawn as its bounds instead.
  • The shared material tables survive a refused upload. A detail or decal table that grew is uploaded again whole. If the device refuses, the old table stays bound, the growth stays flagged and the upload is retried at the next flush after the retry interval. A table the upload replaces is retired through a frame-deferred queue and freed once no frame in flight can name it, instead of being held until the scene closes.
  • Reading it. The detail level of client.showFps has a vram row: what is spent against the pool, how the resident bytes divide between images and buffers, the live allocation count (against the limit where the driver has a real one) and the objects deferred. It is tinted at the high watermark and at the pool. dh render --renderInstances ends its settle line with the pool, its high watermark and the deferred count.
  • What it does not do yet. It does not rebalance when the camera moves, and it does not shrink a texture that is already resident: a texture admitted small stays small until the streaming that follows.

A UASTC KTX 2 entry (how every host reads one) is uploaded as GPU blocks rather than decoded pixels, a quarter of the video memory of RGBA8.

  • The device picks the format from its format properties. TextureBlockFormats takes BC7 on desktop and Linux and ASTC 4x4 on iOS and Android, falling back to the other and then to RGBA8, and counts a format only when both its sRGB and UNORM forms can be copied into and sampled with linear filtering at optimal tiling. The choice is logged once at device creation (KTX 2 textures upload as Bc7.). ASTC is lossless from UASTC; BC7 costs a fraction of a decibel.
  • The transcode runs where the PNG decode ran, on the decode-ahead workers, so the render thread only copies blocks. PalletAssetSource.ReadTextureForUpload hands an uploader a BlockImage, and everything that wants pixels (thumbnails, the texture viewer, the lightmap bake, the Source export) still calls ReadTexture and gets RGBA from the same entry. The two forms are cached apart.
  • client.render.textureBlocks (default on) uploads RGBA8 instead, for A/B. It is read when an uploader opens, and the resident-texture identity carries the format, so a flip never reuses a copy made in the other one. It applies on the next map load.
  • The map’s cap is fitted with the real cost. A KTX 2 entry’s base edge follows from its own size (a byte a texel across its whole chain), and it is costed at one byte a texel when the device takes blocks and four when it does not.

Textures are sampled trilinearly with anisotropic filtering — the shared sampler uses LINEAR mipmap mode across the full LOD range, and enables anisotropy (up to min(deviceLimit, 8)×) wherever the physical device advertises the samplerAnisotropy feature — so surfaces no longer shimmer or alias into noise at distance or at grazing angles (the flat map’s grid was the worst offender). Where the device lacks the feature the sampler falls back to trilinear-only. Mip chains are produced by a single shared, dependency-free CPU downsampler, DigitalHeaven.Core.Imaging.MipmapGenerator, which reduces RGBA8 pixels down to a 1×1 level (floor(log2(max(w,h))) + 1 levels). Writing the algorithm once means the build-time compiler and the runtime engine produce identical chains.

The recipe is a property of the channel, not of the image

Section titled “The recipe is a property of the channel, not of the image”

How a texel should be averaged depends entirely on what the number means, and a single filter cannot be right for a color, a normal and a roughness value at once. MipSemantic names the four recipes, and the pallet compiler picks one per image from the material channel that binds it:

SemanticChosen forWhat the filter does
ColorbaseColor, emissive on a non-cutout materialDecodes sRGB → linear, averages light, re-encodes. Alpha is averaged in place.
CutoutColorbaseColor on an alphaMode: mask materialColor, then rescales each level’s alpha so the fraction of texels clearing the material’s alphaCutoff matches the base level’s.
NormalnormalUnpacks t*2-1, averages the vectors, renormalizes, repacks — and writes the length it renormalized from into alpha.
DatametallicRoughness, occlusion, height, and anything no material bindsPlain average of all four channels in stored space — no transfer curve, no premultiplication.

Three things are worth stating explicitly, because each is a bug the pipeline had or could easily acquire:

  • Color mips must average in linear light. An sRGB texture uploads as R8G8B8A8_SRGB, so the hardware decodes every level on sample. Averaging the stored code values instead of the light they represent makes each level darker than the one above it, and a surface visibly dims as it recedes.
  • Normals must be renormalized. Two normals leaning opposite ways average to a short vector — in the limit, to nothing. Uploading that stub flattens the lighting response of every distant surface. The convention here is three-channel tangent-space RGB with +Y up (glTF / OpenGL), unpacked in dev_opaque.frag as texture(normalTex, uv).xyz * 2.0 - 1.0 with no channel negated; renormalization is sign-symmetric, so only the flat fallback for a degenerate average assumes +Z.
  • What renormalizing throws away is worth keeping. The length of the mean is the measure of how much the four parents disagreed, and it is the only record of it that survives the standing-up. Each reduced level stores it in alpha (the base level is 1 by definition, stamped on both sides so the compiler’s chain and the engine’s decoded base agree), and the fragment shader turns it into a Toksvig variance that widens the specular lobe — see Specular antialiasing. A normal map has no use for a fourth channel, so nothing is displaced. The MipContainer companion carries the recipe it was reduced with. A compiled normal map is BC5, whose two channels hold X and Y alone: there the same lengths ride beside the blocks and reach the shader as a half-size R8 texture, and Z is rebuilt from X and Y.
  • Packed data must not be premultiplied or curved. In an ORM-style map the fourth channel is a number, not an opacity. Data treats all four channels alike and applies no transfer curve, so nothing the shader reads is silently rewritten.

Every recipe reduces with a 2×2 box, whose support is exactly the destination texel. That is deliberate rather than a placeholder: a windowed kernel (bicubic, Lanczos) has negative lobes and rings, manufacturing values the source never contained — tolerable as a softness on a photograph, but on a roughness or height map simply a wrong number. A test asserts no level leaves the source’s range on any channel.

The runtime path sees only a color space, never a channel binding, so it can pick Color or Data and nothing more. That gap is closed by where mips actually come from: the compiler writes a companion for every .png/.jpg/.jpeg it packs, so an authored normal map or cutout in a pallet never reaches the runtime filter.

Generation is preferably build-time, with a hybrid, threaded runtime path as the fallback — either way a frame never stalls:

  1. Build time (pallet path). A texture a material binds ships as one UASTC KTX 2 that carries its own chain (see Block compression); this companion is for the images that stay PNG or JPEG. When such a texture is compiled into a .pallet, the compiler generates its full mip chain and stores the reduced levels (level 1 downward) in a companion MipContainer blob at the texture’s path plus .mipchain. At load the engine reads that blob, prepends the decoded base level, and uploads the whole chain in one shot — no runtime generation, no threaded swap. The base level is never duplicated on disk (it lives in the texture blob) and the companion is distinct from the runtime disk cache’s .dhmip files.

  2. Runtime (fallback). Textures without a precomputed companion — the built-in white/checker defaults, or images a game mod downloads at runtime — upload only the base level first (drawable immediately, crisp and un-mipped), while a background worker builds the full chain (reading it from the on-disk cache when present, otherwise generating and persisting it, subject to the mipmap-cache toggle above). A once-per-frame render-thread pump uploads each finished chain and swaps it into the live texture in place: because the draw list holds the same texture handle, the swap needs no resubmission, and the replaced GPU image is retired only after the frames that might still reference it have completed.

    The pump is budgeted. Each install is a fenced, blocking transfer, so a map whose textures all finish generating on the same frame would stall that frame by the length of the whole queue. client.render.mipUploadBudgetMs (default 2, range 0.1–5.5) caps how long one pump may spend; whatever does not fit stays queued for the next frame, where the only visible consequence is a texture sampling its base level for one frame longer than it otherwise would. The clock is checked after each install, never before, so the smallest legal budget still lands one chain per frame and the queue always drains. The offscreen renderer’s WaitForPendingMips deliberately bypasses the budget: a capture that installed only part of the queue would sample base level for the rest, making the image depend on worker-thread timing.

    The on-disk cache is keyed by the pixels, the dimensions and the recipe (semantic plus cutoff), so one image bound as a normal map by one material and as color by another can never share a chain. The .dhmip format version was bumped alongside the recipe work, so chains written by the old box-only algorithm are treated as misses and regenerated rather than read back.

A texture whose settings resolve to mipmaps: false (every palette unless it says otherwise) is the one exception to both paths. The compiler writes no companion, the engine uploads a single level and queues no runtime job, and its sampler key carries Mipmaps = false whatever the toggle below says, so it binds the sampler clamped to level 0. A palette is read one cell at a time, and any reduced level would be a blend of cells the palette never contained.

The Video section of the settings screen exposes a Mipmaps toggle (client.display.mipmaps, on by default) — a separate master switch for whether mip chains are sampled at all, distinct from the Mipmap cache toggle under System → Asset cache (client.cache.mipmaps), which only governs on-disk persistence of generated chains. The two are orthogonal: the cache toggle decides whether a chain is reused across runs, this one decides whether it is used for rendering.

The switch is implemented as a pure sampler swap, so it applies live — no relaunch, and no map reload or player respawn. The texture registry owns two samplers: the trilinear/anisotropic one across the full LOD range (mipmaps on) and an otherwise-identical one clamped to the base level (MaxLod = 0, anisotropy off — mipmaps off). A texture’s set-1 descriptor binds whichever the ambient MipmapPolicy selects. Chains are always generated regardless of the switch, so toggling never re-uploads an image: on a change the registry waits for device idle (a rare, user-initiated settings event) and rewrites every live texture’s descriptor set — including the fallback white texture — to point the selected sampler at the texture’s current image view. Turning mipmaps off makes distant surfaces crisper but prone to shimmer/aliasing; leaving them on is the default.

glTF materials remain unsupported by design. DigitalHeaven pallets are the material system; the renderer takes its textures from pallet materials and map slots, not from material data embedded in a GLB.