Skip to content

Offscreen Rendering

The engine can draw a map straight to disk. --render opens a map, shoots every camera it authored (an object carrying a dh.camera), and writes each one to a PNG — no window, no swapchain, no server, no simulation. It is one of the seven launch modes of the single engine executable, alongside the default windowed singleplayer, the --headless dedicated server, the --connect remote client, the --benchmark offscreen timing run, the --avatarService sidecar render loop, and --cookCollision, the collision cook dh build starts for a mesh-mode map.

DigitalHeaven.Engine.Host --render <dir> [--map <barcode|name>] [--camera <name>]
[--width <px>] [--height <px>] [--renderUi] [--renderMotion]
[--renderWindow <name>] [--spawn <asset>] [--set <key>=<value>]...
[--bakeReflections] [--renderReflectionBake] [--bakePace gentle|full]
[--renderSprings --avatar <barcode>]
FlagArgumentDefault
--renderOutput directory. Required — it is what selects this mode.—
--mapA map barcode, or a bare name if it is unambiguous.The bundled map
--palletA pallet file to open the map out of, bypassing the catalog. --map then names the map’s path inside that file. Exists for a build photographing a pallet it has not published yet.Look the map up in the catalog
--captureMapThumbnailsNone. Photographs the map through its thumbnail camera at that camera’s own thumbnailSize, and writes thumbnail.png into the run directory. --width/--height are ignored; a map with no thumbnail camera is skipped and says so.Off
--thumbnailPngA path to write the photograph to, in addition to the run directory. This is how a build lands the picture beside the map in the source tree. Needs --captureMapThumbnails.Run directory only
--view<name>=x,y,z,yaw,pitch[,fov], repeatable: an eye in world meters, yaw and pitch in degrees as the position readout prints them, and a vertical field of view (default 75). Shoots these views instead of the map’s cameras, so a picture can be taken from where a player stood without authoring a camera. Cannot be combined with --camera or --spawn.—
--cameraThe name of one camera to shoot: the name of the object carrying its dh.camera.Every camera the map authors — except under --renderUi, which takes the first
--worldSecondsWorld time the shaders see, on the same timeline as NetClient.RenderTime and wrapped by client.render.worldTimeWrap before it reaches the water’s clock — the same number a live client’s worldtime console command or its frame-stats HUD line reports, so a screenshot’s time can be matched exactly.0, which is why a plain render has always been deterministic
--worldClip<frames>@<fps>, such as 90@30. Writes every shot as a numbered run of frames (<stem>_f000.png…) instead of one still, the world clock stepped 1/fps seconds a frame from --worldSeconds and the scene composed afresh each frame, so whatever animates from the clock moves as it does in play: a loading stand-in under client.loading.standInAnimation, the water. Stitch the frames with ffmpeg. Cannot be combined with --renderUi, --renderMotion or --renderWindow.One still per shot
--bakeSet<zone>=<set>[:<fadeSeconds>], repeatable. Shows a bake zone switched to one of its sets, as a setBakeSet action fired at world second 0 would, so --worldSeconds picks the point of the crossfade drawn: --bakeSet room=off:2 --worldSeconds 1 is halfway. A set the map lacks draws its zone’s default and says so.Every zone in its default set
--probeVolume<volume>=<input>[:<value>], repeatable, applied in order. Calls a light probe volume’s input the way a map’s logic would: platformProbes=turnOff, platformProbes=setTint:1,0.3,0.3, platformProbes=setIntensity:2. A name the map lacks is warned about and ignored.Every volume lit as baked
--doorOpen<door>=<fraction>, repeatable, the fraction in [0, 1]. Draws the door that open and everything it carries with it, a dynamic light probe volume included, and the placed objects it carries (a prop door’s model), through the hierarchy a live client carries them by.Every door closed
--spinAt<mover>=<seconds>, repeatable. Draws a spinning dh.mover as if started from rest that many seconds before the shot, its spin-up included, and everything it carries turned with it.Every spin at rest
--widthOutput width in pixels.1600
--heightOutput height in pixels.900
--renderUiNone. Composites the Halcyon player UI over one camera and writes its whole transition as a sequence.Off
--renderMotionNone. Photographs camera motion blur as a labeled sequence. Cannot be combined with --renderUi.Off
--renderWindowhierarchy, inspector, scene, options, connect, connectTouch, or all. Photographs one window, cropped to that window’s own rectangle. Cannot be combined with --renderUi or --renderMotion.Off
--bakeReflectionsNone. Bakes the map’s dh.reflectionProbe cubes through this run’s renderer, writes them beside the map source (--mapSource names it) for dh build to package, and swaps them onto the loaded map before any camera is shot. Without --render it bakes and exits. Each face is rendered at its probe’s own resolution whatever the run’s size, and the cubes hold the composed scene, entities included, as the editor’s bake does.Off
--renderReflectionBakeNone. Photographs a bake: every selected camera twice, once with nothing baked (_reflBefore) and once with freshly captured cubes swapped onto the already-loaded map (_reflAfter). Writes no file.Off
--bakeLightmapsNone. Bakes the map’s lightmap on the GPU, writes it beside the map source for the next build to package, and exits. With --render it bakes first and photographs the fresh bake, provided the loaded pallet carries the same lightmap UVs.Off
--bakePacegentle or full. How hard --bakeLightmaps, --bakeLightProbes and --bakeReflections lean on the machine: gentle runs short GPU dispatches with an idle gap after each, the CPU work on half the cores and the process at below-normal priority on Windows, so a game or a headset can run beside the bake; full takes the whole machine. The same as --set client.editor.bakePace=<pace> for this run only; it never writes the profile. The bake logs its pace as it starts (bake pace: gentle (4 ms dispatches, rest 1x, 8 of 16 threads, below-normal priority)), and the phases line and a closing bake total: line report the GPU’s share of the wall clock (gpu duty). See the pace.gentle (client.editor.bakePace)
--mapSourceA .dh-map path to bake instead of the source the workspace resolves for --map. The bake lands beside that file, so a scratch copy of a map bakes without touching the workspace.The map’s own source
--spawnA spawnable asset, in any spelling the spawn command takes. Photographs the asset instead of the map.—
--spawnCamerayaw,pitch[,distance] in degrees, degrees, meters. Adds one extra view named custom, aimed and pushed back by these angles instead of the fixed four. Needs --spawn.Distance falls back to the same bounds fit the four views use
--spawnViewhead or face. head adds a view named head, framed on the avatar’s head from the front at eye height — the Head bone’s rest position if the rig maps one, else the top client.render.spawnHeadFraction of the bounds by height. face shoots the eye-look strip instead of the four views. Needs --spawn.—
--avatarA .dh-avatar barcode. Stands a pair of player pawns in front of each camera — one wearing it, one wearing nothing — so a frame photographs both halves of the pawn draw. The same flag a live client wears an avatar with. Ignored under --spawn, which photographs an asset rather than a map.—
--avatarAtx,y,z in meters. Where the --avatar pair, or --firstPerson’s one body, stands in the map (feet), instead of the world origin, so a pond, a roof or a map whose ground is nowhere near y 0 can be photographed with a pawn on it. Pair it with --view.The origin
--set<key>=<value>. Overrides one preference for this render only. Repeatable. A world.* key works too: the run stores it under its own render world, so --set world.sun.intensity=0 --set world.ambient.intensity=0 photographs a pitch-dark scene.—
--gpuTimingsNone. After the shots, times a frame from the first selected camera and writes gpu-timings.json beside them.Off
--profileNone. Records the scope profiler over every frame drawn and writes frame-profile-<timestamp>.csv and its Chrome trace beside the shots, the pair the profile verb writes. Add --gpuTimings for a warmed run of many frames.Off

Each shot is written as <mapFileStem>_<cameraName>.png — the map’s file stem, not its display name — and its path is echoed as it lands. A --spawn run uses the asset’s file stem instead. A run wearing an --avatar appends _avatars, so a frame with players in it never overwrites the empty map’s.

The mode is exclusive by construction, and the contradictions are rejected at parse time rather than half-honored:

  • --render with --headless — “different modes: one draws a map offscreen, the other runs a server.”
  • --render with --connect — --render draws a local map; a connected client takes its map from the server.
  • --map or --camera without --render — those two select what to render, so they need something to render into.
  • --worldSeconds without --render — it sets the clock an offscreen render’s shaders see; with nothing to render there is nothing for it to steer.
  • --bakeSet without --render — it switches what a render shows; a live session switches with bakeSet.switch or a setBakeSet action.
  • --probeVolume or --doorOpen without --render — they pose a map for a capture; a live session’s map logic does both.
  • --pallet without --render — same reason; and --pallet without --map, because bypassing the catalog leaves nothing to guess which map inside the file was meant.
  • --captureMapThumbnails without --render — there is nowhere to write the picture; it also needs --map.
  • --thumbnailPng without --captureMapThumbnails — it names where a photograph lands, and without the capture there is no photograph.
  • --captureMapThumbnails with --renderUi, --renderMotion, --bakeReflections, --renderWindow or --spawn — it renders at the thumbnail camera’s own thumbnailSize, which is not a size any of those get to choose.
  • --renderUi without --render — there is no offscreen frame to composite a UI over.
  • --renderMotion without --render — same reason, and there is nowhere to put a sequence.
  • --renderMotion with --renderUi — each writes its own sequence from the same camera; run them separately.
  • --renderWindow without --render — there is no offscreen frame to crop a window out of.
  • --renderWindow with --renderUi or --renderMotion — one crops to a single window, the others write a whole full-frame sequence from the same camera. Run them separately.
  • --renderReflectionBake without --render — it photographs through the offscreen renderer, so it needs one.
  • --bakeReflections without --map — the cubes are photographs of the compiled map, so a source alone is not enough.
  • --renderReflectionBake with --renderUi, --renderMotion, --bakeReflections, --renderWindow or --spawn — it writes its own before/after pair from every selected camera, and it already performs the capture --bakeReflections performs.
  • --set without --render — a live session already has a console; the flag exists only because a render has no way to type into one. --benchmark, --bakeLightmaps, --bakeLightProbes and --bakeReflections take it too, for the same reason (a bake’s client.editor.lightmapBakeMemoryMb, for instance).
  • --bakePace without a bake — it paces --bakeLightmaps, --bakeLightProbes or --bakeReflections, and nothing else. A value other than gentle or full (a number included) is refused.
  • --spawn without --render — a live client has the spawn command; this flag exists only to photograph one.
  • --spawn with --camera — --spawn frames its own views from the asset’s bounds, so there is no map camera left to name.
  • --spawnCamera or --spawnView without --spawn — both add one extra view to --spawn’s own framing, so they need an asset to frame.

A map that authors no camera cannot be rendered; the error says so and tells you to add one. So does naming a camera the map does not have — it lists the ones it does. A --spawn run is the exception: it needs the map only as a backdrop, so a camera-less map is fine.

--gpuTimings renders the first selected camera 30 frames to warm up, then 120 measured ones, and writes gpu-timings.json beside the pictures. The log prints the same numbers:

  • each GPU pass’s mean milliseconds (shadow, ao, opaque, …), read from the device’s own timestamps;
  • the frame: the fenced wall time (recording, submission and the GPU together), the CPU milliseconds spent recording the command buffer, the scene’s items, the draws and triangles every world pass recorded, the draws they culled, and the voxel meshes resident;
  • the same draws and triangles per pass group (shadows, ao, opaque, blended), and, on a map with a visibility set, the camera’s cluster, how many clusters its row sees, the placed objects judged and hidden, and the draws the camera passes skipped for it (pvs in the JSON).

On a map that places objects it also lands them cold under the live client’s pacing (a frame between pumps, the transfer under client.avatar.transferBudgetMs and its burst) and prints object streaming: the models landed, the time until the last one was installed, and the frame times while they landed. It is objectStreaming in the JSON. Compare --set client.avatar.transferBurstMs=0 against the default to see what the burst buys.

On a map with voxel volumes it then streams them cold, the way the live client does. It drops every volume and palette, stands at the map’s first spawn point at standing eye height (the camera when there is none), and runs the live client’s frame: land finished meshes on the device under client.voxel.uploadBudgetMs, compose, render. It records:

NumberWhat it is
groundMsMilliseconds until a mesh is drawn in the column under the eye
fullMsMilliseconds until everything in client.voxel.streamRadius has streamed in
loadingFrame times while it streamed: p50, p95, worst
stillThe settled view from there, timed like the camera above, with its GPU total and shadow cost
flight600 frames flying forward along the eye’s heading at 20 m/s (a sixtieth of a second each): frame times, and the mean milliseconds spent landing meshes, composing and recording
--render out/older --map io.mltn.tests.minecraft-worlds:maps/world-older.dh-map --gpuTimings

Compare two builds by running the same line on each. --set client.render.cullDraws=false --set client.voxel.regionChunks=1 reproduces the draw path from before draw culling and draw regions in the same build, and --set client.render.shaderVariants=false times every scene draw on the one whole shader instead of its variant. A capture compiles each variant before the first draw that needs it and logs it (scene variant ... compiled inline), so the frames and the timings are always the variants a live client settles on. --set client.render.pvsCulling=false draws every placed object a Source map visibility set would hide.

--spawn answers a different question from the rest of the mode: not “what does this map look like” but “what does this thing look like in the engine”. It loads a .dh-obj or .dh-avatar through the same object loader the spawn console command drives, stands it at the world origin of the chosen map, and shoots four views.

DigitalHeaven.Engine.Host --render shots --spawn nova --width 1400 --height 1400

The views are front, threeQuarter, side and back, and they are not authored anywhere — each is fitted to the object’s own bounds on both axes, so a two-meter character and a hand prop both fill the frame at the same command. Authoring a camera per asset would have put the framing in the wrong file: how big a thing is, is a property of the thing.

Every view is framed at client.render.spawnFov (default 40, vertical degrees); a narrower value is a longer lens, and the bounds fit pulls the camera back on its own to compensate, so it never crops the subject tighter.

front is the camera on the +Z side. That is where the front of an avatar out of a Unity-shaped pipeline is, and the importer does not rotate anything — so an object authored the other way round photographs its own back, and the labels are best read as “the +Z side” rather than as a claim about facing.

Unlike a live client, the load here is synchronous. A photograph has to show the same thing every time it is taken, and an asynchronous load would photograph whichever pieces a worker thread happened to have finished.

Two things to expect in the images, neither of them a rendering fault:

  • Bind pose unless the rig is posed. The object loader is a bind-pose importer, so a character stands in whatever pose its mesh was authored in — usually a T-pose. The skin frames below are the exception: they are shot from a posed skeleton. Rigs, armature links and blend shapes are applied (see the component table), so accessories ride their bones; spring chains resolve but nothing swings them in a capture, so they sit at rest.
  • Declined components. Every component or override the loader could not apply is logged by name as the object loads, so a missing accessory is a line in the log rather than a guess about the renderer.

A rigged asset gets two more frames, written once per run from the front framing: <asset>_front_bonesRest.png and <asset>_front_bonesHeadTurned.png. Both wear the editor’s bone gizmos — a marker and a name label per bone, drawn by the shipping overlay rather than by a capture-only imitation — and the second stands the rig with its standard Head bone turned 30° about its own up axis.

A rigged asset also gets a pair of skin frames from the same framing and the same two poses, with no gizmo layer: <asset>_front_skinRest.png and <asset>_front_skinHeadTurned.png. They photograph the deformation itself rather than a marker set over it. With client.render.gpuSkinning on, the turned frame shows the mesh bending at the neck; with it off, the same frame shows the frozen bind pose under a skeleton that moved — which is what makes the pair a comparison worth shooting.

The turned frame is the one carrying evidence. A marker set over a rest pose proves only that markers were drawn; a head turned away from the body proves the markers under that head moved with it and that the rigid drawables parented to it — glasses, a collar, hair — followed the bone instead of staying behind. An asset whose rig names no standard Head still gets its rest frame, with a warning, rather than nothing.

They come last, after every plain scene shot, because the UI layer they need is enabled on the way in and the bands are cleared only by a UI step — a frame shot after these ones would be photographed through a band they lit.

The four fixed views answer “what does this thing look like”; comparing an avatar against how it renders in Unity needs an arbitrary angle and a tight shot of the head, so both reuse the same bounds fit and view-naming code as the four rather than a second framing function:

DigitalHeaven.Engine.Host --render shots --spawn nova --spawnCamera 135,15
DigitalHeaven.Engine.Host --render shots --spawn nova --spawnView head

--spawnCamera 135,15 writes one extra frame, <asset>_custom.png, at yaw 135° (rear three-quarter) and pitch 15° — yaw zero is the same +Z “front” the four fixed views use, and a positive pitch looks up at the subject (so the camera sits below eye height), the same convention every other pitched camera in the engine follows. A third number is an explicit distance in meters; omitted, custom falls back to the identical bounds-fit distance the four views use.

--spawnView head writes <asset>_head.png, framed level and from the front, fitted to the avatar’s head rather than its whole body: the rig’s Head bone rest position when one is mapped, otherwise the top client.render.spawnHeadFraction (default 0.2, i.e. the top fifth) of the bounds by height. Both options can be combined with each other and with the four fixed views in the same run; neither changes the four.

--spawnView face photographs an avatar’s eyes and blinks over simulated time rather than four stills. It steps the client’s own face rig at 60 Hz for 12 seconds and offers the avatar a scripted face: ahead and to its right, looking back, for the first 4 seconds; nothing for the next 4; ahead and to its left for the last 4. The face stands inside the angle the eyes acquire a target within, derived from the rig’s own limits. The camera is a level close-up centered between the eyes, sized from the distance between them.

It writes <asset>_face.csv, every step’s EyeLookOutput (target, point, both eyes’ yaw and pitch, lid weights, the largest blink weight, and whether a saccade or a blink began). Then it shoots <asset>_face_<nn>_t<seconds>_<label>.png every 1.5 seconds (cadence), plus three frames the trace picks: atFace (settled on the face), lookAway (settled looking away from it) and blinkMidClose (a blink closing, nearest half shut). Each frame’s state is logged beside its path. The seed comes from the barcode, so two runs of one avatar shoot the same frames; --set client.anim.* retunes the brain for a run. An avatar whose rig describes no face (no eyes with limits, no eyelid shapes) is an error.

An object whose slots carry no material from a dh.renderer draws flat white. The loader records the source glTF’s material name as provenance, and the engine has no way to turn a name into a dh.material — so an untextured render is a statement about the definition, not about the renderer.

A capture of a map that places objects draws them too, with their reflection probes bound as the record names them, and a map that places none pays nothing. Each is submitted the way a live client submits a standing instance, so the picture is the one a player sees: the model where it has loaded, its compiled stand-in where it has not, and the placeholder cube where the pallet ships no stand-in. The loads run asynchronously, as they do in play, and the capture waits for them to settle before its first shot, so the image depends on the pallets and the preferences alone.

A capture is never mid-load on its own, so client.debug.holdObjectLoads holds models back on purpose. It takes a fraction: each barcode has a fixed place in [0, 1), and a model whose place is below the fraction never loads, leaving its stand-in up. The same objects are held on every run and every machine, and raising the fraction only adds to them.

DigitalHeaven.Engine.Host --render shots --map gm_construct --set client.debug.holdObjectLoads=0.5
DigitalHeaven.Engine.Host --render boxes --map gm_construct --set client.debug.holdObjectLoads=0.5 --set client.loading.standIns=false

The second line is the same moment with stand-ins switched off: every held object is the placeholder cube, as every object was before stand-ins existed. The log names what it settled to, models and stand-ins counted once per object and placeholder cubes once per placement (instances: 16 placed object(s) settled in 0.4 s, holding 0.5 of loads back; 9 model(s) drawn, 7 standing in, 0 drawn as the placeholder cube).

A render resolves the map’s look exactly as the map author authored it. The mode builds a virgin preference store — freshly constructed, with no profile ever loaded into it — and no world context, so every world.* value falls through to what the map’s lighting and render blocks seeded, and every viewer-owned comfort setting takes its engine default.

That is the point, not a limitation. No saved console edit, no tuning session, and no profile on the machine that happens to be running the render can leak into the image. The same map renders the same way on a developer’s machine, on a fresh clone, and on a build agent that has never had the game open.

--set is the one deliberate way in, and it is deliberately explicit: every override is spelled out on the command line, so the image is still reproducible from the command that made it. It is applied to the virgin store before anything reads from it, and an unknown key or an unparsable value is an error rather than a silent no-op — which is what makes it usable for A/B comparisons of a single effect.

DigitalHeaven.Engine.Host --render out/ --set client.render.aoHalfResolution=false --set client.render.aoDenoise=0

Logic entities draw in rest pose: a door closed, a button unpressed, and a light lit if and only if it was authored startOn. Nothing is simulated, so there is no tick count, no elapsed time and no interaction state to agree on.

Every shot this mode takes composes them the same way, through one shared composer — the plain camera renders, the physics and selection timelines, the reflection bake, the probe cook and the map thumbnail. That is not decoration: the thumbnail once composed nothing, so it shipped a picture of the map’s geometry with none of the buttons, doors, lamps, plates or crates standing in it, while the editor’s camera card previewed the live scene and promised all of them. A visual entity that cannot be composed at all — its authored position or size carries fewer than three numbers — is named in the log rather than quietly left out, because a hole in a picture is indistinguishable from a rendering fault.

A dh.camera authored "projection": "orthographic" renders a parallel projection, framed by orthoSize — a half-height in meters — instead of by fov. Nothing converges, so a straight-down shot is a survey of the map rather than a photograph of it: two equal walls at opposite ends of the frame measure the same number of pixels, and the image can be read off with a ruler.

DigitalHeaven.Engine.Host --render out/ --camera topOrtho

A perspective frustum is infinite and needs no far plane. A parallel one is a finite box, so an orthographic camera that does not author far would frame nothing. The render derives it from the map bounds — the furthest geometry along the view direction, plus a margin — and logs the number it settled on, so a shot from 140 m up reaches the ground without the author computing anything.

Two effects have no parallel counterpart and are switched off for the shot, each naming itself in the log:

Off under orthographicWhy
Camera motion blurIts velocity field is reconstructed from depth through the near-plane algebra of a reversed-Z infinite projection, which a finite parallel box does not satisfy. --renderMotion refuses an orthographic camera outright rather than writing frames the filter never touched.
Panini wideningA correction for the stretch a wide pinhole puts in the corners. A parallel projection has no such stretch, so the widening distance is zeroed and every derivation returns its input unchanged.

Everything else — the sky, the cascaded shadows, ambient occlusion, water, the outline pass — is fitted for the parallel case rather than left reading a field of view that is no longer there.

By default a render contains no UI at all — it is a picture of the world. --renderUi composites the Halcyon layer over the tonemapped frame in exactly the position the live client puts it: after the resolve, color-only, loading rather than clearing what is underneath.

The developer tooling takes part too. The debug HUD, the gizmos and the DigitalHeaven overlay are plain draw commands rather than an immediate-mode context that only exists next to a window, so the offscreen path can composite them like anything else. There is one hand-off per band — PaintTooling for the gizmos, the readouts and the frame-stats card that sit under the player’s screens, PaintHud for the speedometer, spliced right alongside Tooling and under the screens too, PaintCrosshair for the reticle above both readouts and still under the screens, and PaintOverlay for the overlay’s own chrome over the screens — and each takes a callback rather than handing out its list, because the ordering is the contract: after the step, which clears the list, and before the shot, which records it. Which band a surface goes into is read off the surface’s own Band constant (PerformanceOverlay, EffectsReadoutOverlay, SpeedometerOverlay, CrosshairOverlay), so a capture cannot photograph a stack the live client does not draw.

A UI that animates cannot be represented by one still, so --renderUi writes a sequence instead: the panel mounts hidden, is shown, and the transition is sampled at both ends and at every quarter between, then once well after it settles, then the same again on the way out. Each frame is named <mapFileStem>_<cameraName>_<phase><milliseconds>ms.png:

harbor_skyline_enter0000ms.png just mounted — the recipe starts at zero opacity, so this frame is the world alone
harbor_skyline_enter0037ms.png partway in: offset by part of the travel, partly opaque
harbor_skyline_enter0075ms.png the arrive point — placed and fully opaque, still visibly soft
harbor_skyline_enter0112ms.png nearly resolved; only a trace of blur left
harbor_skyline_enter0150ms.png the transition's end — sharp, and no layer at all
harbor_skyline_settled1000ms.png a second later, to prove nothing keeps moving
harbor_skyline_exit0037ms.png …and back out, sampled the same way

The pinned camera walks a run of sequences back to back, so one pass photographs every surface the engine draws: the demo panel (enter/settled/exit above), the settings screen (settingsBrowse, settingsSearch, settingsSwitches, and settingsMenu — a dropdown held open, which is the one frame that needs a synthesized press, since whether a menu is open is the control’s own state and there is no preference to write instead), the pause menu (pauseMenuEnter…, pauseMenuSettled, pauseMenuSettings — settings as it is actually reached, from the menu — and pauseMenuExit…), the main menu (mainMenu…), which is the same screen with the no-session link set, and the console in each of its three placements (consoleSheet, consoleCard, consoleDock). Every sequence labels its frames separately so none can overwrite another’s images.

Two of them photograph the stack contract rather than a surface. The window run (windows…) drags the options window, swaps which of the two windows is in front, and then spends one Escape: windowsEscapeClearedBoth is the frame after it, bare, because leaving the pause layer takes the console with it, and windowsRepausedRestored is the re-pause, which brings back only the window that belongs to the layer. The in-game console run (inGameConsoleOpen, inGameConsoleClosed, pausedConsoleDimmed, menuAfterConsole) photographs the headline claim of the console contract: inGameConsoleOpen is the console with no menu and no dim behind it — the world is still running — and pausedConsoleDimmed is the same console over the pause layer, so the two images differ in exactly the thing being claimed. A still frame cannot show a clock, so what the pair actually proves is that the surface is reachable without the layer; the behavior itself is pinned in ClientUiStateTests and ClientUiStackTests.

One more is the text specimen run (textSubpixel, textGrayscale, textFade…), which A/Bs the two antialiasing paths on frames that differ in nothing else.

Another is the chat run, and like the in-game console pair its evidence is the relation between two images. chatFeed is the passive feed as it reads during play — no plate, no field, no keyboard — and chatPromptOpen is the same feed with the prompt over it. Every line visible in the first is at exactly the same pixels in the second: the closed and open states are one tree, the input row is mounted and reserving its height even while chat is closed, and the surface is anchored by its bottom edge, so opening adds a field, a plate and more history without moving anything the player was already reading. Then chatClose… samples the close across the transition, because the claim there is that the prompt retraces its entrance instead of blinking out. A render has no session, so the lines are seeded by the command — a join, a leave, a server line and several player lines, all stamped at the capture’s single instant so the fade can never make which lines are legible depend on how many frames the run took to get there.

Another is the appearance run (appearanceSeedDefault, appearanceSeedAlternate), which searches the settings screen for the Interface page’s groups and then photographs them twice with the clock held still, changing only client.ui.themeColor. Two frames that differ in nothing but the seed are the only honest evidence that the tinted surfaces are derived from it: any card, chip or track that fails to move between them is a hardcoded color, and the image says so at a glance in a way no unit test does.

The last is the developer run, and it is built to be read as a diff. It walks client.debug.hud one level at a time with the overlay down — hudStats (the frame-stats card and the position readout), hudGizmos (+ the axis, entity and scene-pivot gizmos and the sound cues), hudSignal (+ the rush figures) — so consecutive images differ by exactly one level and each one says what that level draws. Three images taken at the top level would say nothing. Both corners then go through every background at the compact, summary and netcode levels, at the desktop scale and at twice it (where the padding and the margins grow with the text): hudStatsBare, hudSummaryBare and hudSignalBare on nothing, hudStatsCard, hudSummaryCard and hudSignalCard on the opaque card, then the same with a Scale200 suffix, and hudStatsScale200, hudSummaryScale200 and hudSignalScale200 on the shipped plate (hudSummary itself is the summary card at the desktop scale on the shipped plate). The frame-stats card in every one of them is drawn over an authored frame history rather than a measured one, seeded afresh before each frame, so its frame-time graph is the same graph in every run. Then it raises the DigitalHeaven overlay, steps past its fade, and shoots it settled: overlayBar with nothing open, one frame per window archetype, four frames with the pointer parked on something (overlayBarHover, overlayBarPress, overlayControlHover, overlayControlPress — the only pictures there are of the chrome’s hover and pressed paints), and overlayWindows with all three windows cascaded. The run ends with the overlay put away, so nothing concatenated after it inherits a surface it did not ask for.

Beside it is the HUD-layering run, which holds the frame-stats card still at one level and moves the surfaces around it: hudOverPauseMenu (the block dimmed under the menu’s scrim, along with the gizmos), hudUnderConsole (the same menu with the console raised over both), and hudOverMainMenu (the same block under the main menu, which takes no scrim and no vignette). What they photograph is the block in Tooling. Three images with one constant and one variable are what makes the layering claim visible; a capture is a still, so it photographs the two ends of the session ramp rather than any point along it. Like the developer run it ends driven back out, leaving nothing raised.

The speedometer run is the player-facing counterpart, and makes the same claim as the HUD-layering run, since Hud lands at Tooling’s seam. Four zero-length steps photograph the readout at a stand, at the engine’s own declared walk and sprint speeds and at a fall — speedometer0 through speedometer3, differing in the digits and nothing else, which is how the fixed-width number field is visible at all — then speedometerOverMenu raises the pause menu over it, dimming it under the scrim, and speedometerUnderConsole raises the console over it too. Both are needed because a boundary photographed only on its true side is not a picture of a boundary — though the console frame is a weaker picture than it sounds: at the default sheet geometry the console does not reach the bottom band, so the two share a screen rather than one hiding the other. A capture has no pawn, so the one number the card exists to print is requested as data; it has no interaction state either, so what those two frames show is the layering under the gate a live client would apply.

The crosshair run exists to give the Crosshair band a hand-off, so a capture actually contains its pixels. crosshairInPlay is the reticle during play with client.debug.hud at its top level, so the scene-pivot gizmo is projecting into the middle of the frame and the mark has to be on top of it — the half the band exists for. crosshairUnderMenu and crosshairUnderOverlay request the reticle with a player screen up, which is the gate held open on purpose, because what needs photographing there is the stack doing the covering rather than ClientUiState.CrosshairVisible declining to draw anything. A live client never reaches that state, and that is the point — a capture has no interaction state, so the reticle is a step field rather than a predicate the harness evaluates. The gate itself is untouched.

The touch controls run photographs the on-screen action cluster — the one set of surfaces no frame in this command had ever shown, because they are mounted only inside a session at the touch metrics and this command has neither a session nor a touchscreen. Both are stated as data: a step carries Touchscreen for the metrics and TouchGameplay for the session, and the cluster’s two conditions arrive the same way, since the live buttons ask a callback about a world the capture does not have. Three frames, and the evidence is in what is missing from each. touchControlsResting is a plain session — the joystick, Jump and the menu corner and nothing else. touchControls adds the two conditional buttons at once: Use, because the session is looking at something it could act on, and Noclip, because the server said this client may fly. touchControlsNoclipEngaged differs from it in exactly one thing, the confirmed noclip state being on, so a diff of that pair is a picture of the engaged style and of nothing else. The slot to Noclip’s right stays empty in all three; it is held for a sprint control whose placement is not settled.

The console’s frames are driven, not staged. The half-typed command line is fed through the model’s own controllers, so the completion panel is ranked by the real ranker and windowed by the real cursor. The sheet sequence adds three extra frames the other two placements do not need, because they photograph the shared console body rather than a placement: consoleSheetCycled (one Tab into the candidate list), consoleSheetFiltered (the header filter narrowing the log with a category chip muted), and consoleSheetSelected (a multi-line text selection over the last few lines of the scrollback).

consoleHelp and consoleHelp2x are last in the run and photograph the console after help and help client.render.vram have been submitted: the capture runs a submitted line against a console of its own whose only world is the capture’s preference store, and prints its answer into the scrollback. They are last because those lines stay in the ring and every console frame before them was shot without them.

The UI clock is stepped by fixed amounts supplied by the harness, never by a wall clock, and the pointer is delivered as unavailable so nothing can hover. Two surfaces read a clock the harness cannot step, and both are pinned rather than stepped: the caret’s blink is held solid with CaretBlinkSeconds: 0, because a blink would make the caret’s presence in a PNG a coin flip, and the overlay’s toast arrival ramp is held at zero with client.overlay.toastFadeIn, because a Notification stamps DateTime.UtcNow and is asked its own age — which made overlayBar differ between two runs by however far into a 0.3 s fade the machine had got. Pinned, every toast is photographed settled, which is the state those frames were always described as showing. One frame is the exception and says so: the open-dropdown shot presses the anchor’s own rectangle before it is photographed, because whether a menu is open belongs to the control rather than to any model, and it puts the pointer back out of reach immediately after — so the menu appears open with nothing hovered. The client.ui.* preferences resolve against the same virgin store the world settings do, so the transition’s duration, travel and blur are the engine’s declared defaults — the demo panel itself is requested by the flag rather than by client.ui.demo, which a virgin store always reports as off.

Choosing the ground a frame is judged against

Section titled “Choosing the ground a frame is judged against”

Everything above composites the interface over whatever the loaded map happened to put behind it. That is fine while the surface under test is opaque and useless the moment it is not. A scrim is an alpha: what a picture proves about it depends entirely on the value underneath, and every frame in the run above sits over the same dim interior, which is the one place a dark scrim has no work to do.

A capture step can therefore name the camera it is shot from. UiCaptureStep.Camera is a map camera’s name, resolved against the map’s authored set at the moment the frame is taken; a step that names none is shot from the pinned camera like everything else. Only the shot moves — the developer overlays are built once against the pinned camera, which the frames using this are content with because none of them photographs a world-anchored surface. They photograph a screen over a wall.

An unknown name stops the run. It is not a fallback to the pinned camera, deliberately: the whole value of naming a room is that the frame’s backdrop is the one it asked for, so substituting a different one would write exactly the image the frame exists to rule out, under the name of the one it wanted. The message names the step, the camera it asked for, and every camera the map does have — because the mistake is nearly always a typo or a lane pointed at the wrong map.

The rooms those frames need ship with the engine. render-eval is a purpose-built map rather than a corner of an existing one, because a room that is also somebody’s level changes when the level does, and a reference ground that drifts is worse than none.

Four sealed boxes stand in a row on the flat grid plain, and a row of material panels stands out in the sun beside them. Each has an authored camera named after it:

CameraThe room
whiteFlat white walls, floor and ceiling. The end where a dark scrim has to do all of its work.
grayThe same room at mid-gray — the value most interfaces are actually read over.
blackFlat black. Where a scrim has nothing to do and can only make an already-dark room unreadable.
bloomThe black room holding a panel pushed far past display white, so a highlight blooming through a surface shows up as a wash.
materialsA row of panels out in the sun: the dev grids, the flat orange, the noise sheet, and one albedo/normal/occlusion material whose look depends on the whole shading chain.
waterGrazingThree rimmed pools in a row — ocean, lake, pool presets, nearest first — seen across their surfaces at 5° down, where Fresnel hands the far water to the horizon sky.
waterAboveThe ocean pool at 78° down: three descending steps and the floor, where the water absorption gradient reads against the step grid.
waterPlanStraight down over all three pools, for comparing the presets on one floor.
texFilterThree cubes wearing one pure black-and-white tile under point, bilinear and trilinear sampling, each with a strip running away from the camera. The cube face reads magnification; the strip reads minification and the mip step.
parallaxTwo identical brick slabs with white cubes sunk into them. The left slab marches its height field but writes the triangle’s depth, so its cubes are cut off along a straight line; the right one writes the marched depth, so the bricks in front bite into the same cubes.

The room materials are unlit on purpose. A shaded wall would put the lighting rig into the evidence along with the surface being judged, and the point of these grounds is to be one known value. The bloom panel is the exception and has to be lit, since an unlit material outputs tint times base and ignores emission outright.

The boxes have no doors. A person walks in with noclip; a camera never needed one. The map’s two spawn points stand outside, by the material row and in front of the white box.

white is first in author order, which matters twice: it is what --renderUi pins to, and it is the anchor the render-evaluation lane is gated on. RenderCommand composes those frames only when the loaded map authors a camera by that name — otherwise every ordinary render, including the two that verify-render-determinism.bat shoots on the bundled debug map, would hit the hard error above and a determinism check that cannot start is not a passing one. The gate is on the camera name rather than the map’s barcode, so a fork of the map or a second one carrying the same rooms gets the frames too.

The lane writes four frames: renderEvalPauseOverWhite, renderEvalPauseOverBlack, renderEvalPauseOverBloom and renderEvalOptionsOverWhite.

--renderUi is a regression sheet: a hundred-odd full frames, shot to be diffed. --renderWindow answers the other question, the one you ask while a widget is still being designed — what does this window look like right now — and it answers it with one tight picture per window instead of a frame you have to hunt through.

DigitalHeaven.Engine.Host --render .tmp/window-shots --renderWindow all

That writes window_hierarchy.png, window_inspector.png, window_scene.png, window_options.png, window_connect.png and window_connectTouch.png, each cropped to that window’s own rectangle plus a uniform 16-pixel margin, on a plain backdrop with no world, no toolbar and no other window in frame. A single name — hierarchy, inspector, scene, options, connect, connectTouch — writes just that one. A name that is not one of those is refused at parse time and the message lists the ones that would have worked.

Not every shot is an editor window. connect raises the client’s main menu and opens the Connect window over it, against a fixed sample of starred and recent servers (UiCaptureConnectModel) so the picture never depends on which servers this machine happens to have joined. connectTouch is the same window at touch metrics: every step of that shot states its own Touchscreen flag rather than leaving one standing, because --renderWindow all drives every shot in one process and a metric left set would be inherited by whatever is shot next. Each shot still opens with the same teardown pair every window shot opens with, and the labeled frame at the end is the one that writes the file.

The crop rectangle comes from the layout the frame itself used: the window’s Element.Rect, looked up by its reconciliation key through the same search the capture’s pointer aims with. Nothing scans pixels for an edge, so a window that changes size changes the picture’s size with it, and a window whose chrome fails to draw produces a warning rather than a plausible crop of the wrong thing. UiTree.UiScale is a layout scale rather than a transform, so that rectangle is already in the framebuffer’s own space and no conversion stands between the two.

The backdrop is one opaque fill painted into the Tooling band, which splices in over the world and under every screen — so the loaded map is simply not in the picture, without the command having to touch a clear color, choose a map, or care about the lighting.

That holds while the world is the frame. A docked game view makes the world a picture a screen paints, and the Tooling band then rises over the screens so its readouts are not covered by the pane — which puts a viewport-wide fill over the editor and every window in it. So the fill is withheld whenever a pane holds the world, and it is withheld rather than shrunk to the pane, because the pane and the editor’s floating windows are one tree with no band between them: a pane-sized fill would black out whatever window is standing over the pane. The world stays confined to the pane’s own rectangle, and it is pinned exactly as tightly as the rest of the frame — so a window whose crop overlaps the pane shows the world there instead of flat color, and two runs of the same command still agree byte for byte.

Two renders of the same window are byte-identical, on the same terms as everything else on this page: fixed hierarchy and inspector fixtures, a stepped UI clock, a pointer that cannot hover, and a caret held solid.

A crop that comes out one flat color is refused outright rather than written. A wash is exactly as deterministic as the window it was meant to show, so every diff over it passes forever — the same blind spot an empty game view opens, and the same answer: state the condition instead of trusting the diff.

Camera motion blur is the first effect whose output depends on the previous frame, and that breaks a still outright: every ordinary capture rewinds the renderer’s frame-to-frame history first (see Determinism below), so the filter is handed a camera with no predecessor, correctly calls that a discontinuity, and leaves the frame sharp. --render with --set client.render.motionBlur=true therefore produces exactly the same image as without it. That is not a bug in either half; it is why --renderMotion exists.

--renderMotion walks a fixed timeline of phases. Each rewinds once, drives one unphotographed frame to establish a previous camera, and then shoots the frame with the motion between the two in force. Each image is named <mapFileStem>_<cameraName>_<phase>.png:

PhaseWhat it isWhat it should look like
stillStatic camera, filter onPerfectly sharp — and byte-identical to stillBlurOff
walk4 m/sSharp: under the onset speed
onsetExactly the onset speedStill sharp: the ramp is zero at its start
sprint18 m/sBlurred
sprintZeroAmount / sprintZeroShutterThe same sprint with the strength and the shutter each switched offBoth byte-identical to sprintBlurOff
fall24 m/s straight downThe strongest blur the effect produces
flickA mouse flick with the feet plantedPerfectly sharp, however large the screen velocity is
flickMovingThe same flick while sprintingBlurred, with rotation contributing its capped share
teleportA 40 m jump in one frame, at full gainSharp — the discontinuity cancels it
hitchA 200 ms frame while sprintingSharp — past the hitch gate
lowFpsThe same sprint at 40 fpsVisibly weaker than sprint, and the same streak length

Every phase above that claims sharpness also writes a <phase>BlurOff twin — stillBlurOff, walkBlurOff, onsetBlurOff, sprintBlurOff, fallBlurOff, flickBlurOff, teleportBlurOff, hitchBlurOff — at the identical camera pose, yaw, gain and frame duration, varying only the master toggle. That pairing is the point: a walk, a flick, a teleport and a hitch all move or turn the camera, so none of them can be the same bytes as the static frame, and comparing any of them to stillBlurOff would be a check that can only ever fail — and whose failure would say nothing about the blur.

The whole timeline is a pure function of the onset and saturation speeds, so it is reproducible run to run and a --set on either moves the phases with it. Everything the timeline does not override still comes from the preference store, so --set client.render.motionBlurShutter=4 A/Bs the entire sequence at the maximum exposure.

Aim it at a camera looking at detailed geometry with real depth complexity. A flat wall would hide the exact artifact the sequence exists to catch — Jimenez’s own note on the technique is that it “works great for constant or low-frequency backgrounds, but it creates artifacts on detailed ones”, which is why the resolve this engine uses is not the one from those slides.

--renderPhysics photographs a solver rather than a scene. Where every other capture on this page shoots one instant, this one loads a map, mints every physics body its instance placements carry, and steps the world forward at a fixed 60 Hz tick, stopping at five ticks to shoot every authored camera at each of them:

DigitalHeaven.Engine.Host --render .tmp/physics --renderPhysics --map core:maps/physics-eval
TickWhat it catches
0The authored pose: both columns still in the air, before gravity has been applied once
30Half a second in: free fall, the first crates reaching the ground
90The collapse — where a stack is loudest and a solver change shows first
240Settling: the pile is in its final shape and the last crates are still nudging
600Ten seconds. Everything must be asleep; a crate still moving here is the bug

Images are named <mapFileStem>_<cameraName>_t<tick>.png, zero-padded so a directory listing is in tick order.

The frames are the outer loop and the cameras the inner one, because a solver only runs forward: all of a tick’s cameras are shot before the world is stepped again, so the whole sheet is one simulation rather than one per camera.

Alongside the images the run writes physics-eval.transcript.txt — for every captured tick, every prop by name, its position rounded to a millimeter, and whether it is awake or asleep:

# physics transcript of core:maps/physics-eval.dh-map at 60 Hz
tick 600
zone1Box01 -12.781 0.350 -0.800 asleep
zone1Box02 -11.980 0.350 -0.800 asleep

It is the more useful half of the output. A picture says two runs looked alike; the transcript says which crate moved and when, and it is readable with no GPU in the machine — which is also why the round trip is covered by ordinary unit tests while the images are left to the batch script.

A map whose autoCollision is "mesh" is refused rather than photographed: the probe installs boxes and hulls, not cooked triangle meshes, so its crates would fall through a floor that was never there.

The room this needs ships with the engine, for the same reason render-eval does: a yard that is also somebody’s level changes when the level does.

Its centerpiece is two identical columns of twenty-four crates hanging in the air a stone’s throw apart over the flat plain — the lowest four already four meters up, every layer above them turned, nudged sideways and separated by a gap — so both columns drop, scatter and settle the moment the map loads. The blue pile is client-simulated — the server never spawns it, never gives it a body and never sends a byte about it — and the orange pile is the ordinary server-driven one whose poses arrive over the wire. If the two piles land differently, the difference is the netcode, not the solver.

CameraThe station
overviewBoth piles at once, from above the spawn. Also the map’s thumbnail.
zone1The blue pile, client-simulated.
zone2The orange pile, server-driven. The same twenty-four crates, the same offsets and turns, twenty-four meters over.
slopesTwo ramps straddling the default slope limit — one at 30°, one at 60° — each with a blue and an orange pair dropped straight onto its face, seen from the downhill side so the runout is in frame.
logicThe sliding door with its button and its pressure plate: a solid the server drives, standing beside solids it does not.
waterA rimmed pool with a blue and an orange crate hanging above it.

The map also cannot yet carry a spring-bone rig. A .dh-map places brushes, colliders and logic entities; there is no entity type that places a dh.object, so an avatar-shaped station has nowhere to hang. When one exists, it belongs here.

The third bundled evaluation map is for the lightmapper, and unlike the other two it is deliberately unfinished. render-eval is unbaked and its rooms are unlit on purpose; a bake over all of it would be big and would say nothing readable, so the lighting gets a yard of its own: a 40 m matte floor carrying an open-fronted room with a baked lamp inside it, a canopy on four pillars for the sun to shadow under, a cylinder for a smooth terminator, and three cubes at different sizes — every solid its own named node in bake-eval.glb, wearing one lit, fully rough white. lighting.lightmap is on with bakeSun: true, so the shadows on the ground are the atlas rather than the cascades; two realtime lights, a reflection probe over the room and a light probe volume over the yard are the rest of it.

CameraThe view
anchorEye level from the spawn, the whole yard in one frame. First in author order, so it is what --renderUi pins to.
overviewFrom above and behind the spawn. Also the map’s thumbnail.

The stations beyond these are meant to be set up by hand in the level editor, because the map exists to exercise the editor’s lighting workflow as much as the baker; Engine/design-notes/bake-eval-howto.md walks that. What the map cannot show is the same list the baker cannot express: brushes are not baked, the atlas carries no bounce and no direction, and a "mode": "baked" light reaches a moving object only through the probe volume.

DigitalHeaven.Engine.Host --render .tmp/bake --map core:maps/bake-eval
DigitalHeaven.Engine.Host --render .tmp/bake-atlas --map core:maps/bake-eval --debugView lightmap
DigitalHeaven.Engine.Host --render .tmp/bake-uv --map core:maps/bake-eval --debugView lightmapTexels

--renderSprings is the spring harness: it puts one avatar’s dh.springBone chains through a fixed list of scripted motions and writes what they did, as pictures and as numbers. It is how a chain’s numbers get tuned by eye without a live session, and how a tuned chain is checked for regressions.

DigitalHeaven.Engine.Host --render <dir> --renderSprings --avatar <barcode>
[--springMotion rest,yaw90,yaw180,yaw360,yaw720,whip,nod,jump,snap180,walk,run,sprint,kick,sit,crouch,prone]
[--springOverride <file.json>]
[--springSweep "<chain>:<field>=v1,v2,..."]
[--springFrameHz 144]
[--set <key>=<value>]...
FlagArgumentDefault
--renderSpringsNone. Needs --render and --avatar. Cannot be combined with --renderUi, --renderMotion, --renderPhysics, --firstPerson, --spawn or --ladder.Off
--springMotionComma-separated motion tokens, run in the order given.Every motion
--springOverrideAn object layer laid over the avatar in memory: dh.springBone components on the parts in its children (or on a record’s parts, under objects), dh.springCollider components on its root. A field left out inherits. No pallet rebuild.None
--springSweepOne chain, one field, several values; the whole run happens once per value.None
--springFrameHzThe display rate frames are stepped at, 10 to 1000.144
--width / --heightOne tile’s size.480 x 640

The avatar is drawn by the same code a live pawn is drawn by: PlayerPawnVisual stepping a SpringBoneRig with the transform the draw uses, at the fitted size. Nothing about the solver is special-cased for the harness. It watches the solver through a read-only probe (ISpringBoneProbe) that sees every tick and every drawn frame, so a run with the harness attached draws the same bytes as one without. client.anim.springTick, client.anim.springMaxSubsteps, client.anim.springSweep and client.anim.springCollisionPasses are read from the run’s store, so --set reaches them.

Every motion but rest starts with 3 seconds of standing still, so a chain starts settled rather than still sagging from its bind pose. Every motion ends with a hold, which is what settling is measured over. The leg and head motions are bone deltas on the rig’s humanoid roles, about the avatar’s own side-to-side axis. The engine has no gait of its own, so these are harness poses, not the game’s.

MotionScriptWhat it shows
restStand for 3 s from spawnThe first-tick pop and the gravity sag
yaw90 … yaw720Turn in place at that many degrees per second for 1 s, hold 2 sA chain swung out by turning
whipWalk at 4 m/s for 1 s, stop, hold 2 s, walk 1 s more, hold 2 sStop-start overshoot
nodThe neck and the head each pitch ±15° at 2 Hz for 2 sEars and whiskers on a moving head
jumpA 0.5 m parabola and a hard landing, hold 2 sVertical whip and gravity
snap180180° in one frame, hold 2 sThe one-frame discontinuity
walk1.4 m/s for 3 s; thighs ±25° at 1 Hz in antiphase, knees 0 to 60°A tail over a gait
run4 m/s for 3 s; thighs ±45° at 1.4 Hz, knees 0 to 100°The same, faster
sprint5.6 m/s for 2 s; Minecraft’s 1.4 · cos(13.3 t) swing, knee rigidThe fastest real limb
kickOne thigh swings 60° back and returns in 0.2 sA limb faster than any gait
sitOver 0.5 s thighs fold 90° forward and knees 90°, lowered until the thigh joints meet the floor; hold 3 sResting contact with the legs
crouchThighs 60°, knees 110°, lowered so the feet stay down; hold 2 sThe backs of the thighs rising into a tail
proneOver 0.5 s the whole body pitches 90° forward about the hips onto its face, lowered until the hips are 0.12 m off the floor; hold 4 sA tail hanging off a back that faces up, as on a ragdoll lying on its belly

The numbers are authored in SpringCaptureTimeline, never read from a preference.

Everything lands in <dir>/springs/<avatar-stem>/:

FileContents
<motion>.pngThe strip: six moments left to right, three rows each. The moments are the motion’s start, middle and end, 0.25 s and 0.75 s into the hold, and the end of the run. A motion with no duration (rest, snap180) uses 0, 0.1, 0.25, 0.5 and 1 s after it instead.
<motion>_f<frame>_lit.pngRow one: the avatar lit, as a player sees it
<motion>_f<frame>_xray.pngRow two: the x-ray. The body is hidden. Each chain is drawn as it is drawn this frame, in its own color, over its rest pose ghosted in light gray, one dot per slot. An authored endpoint is a hollow dot.
<motion>_f<frame>_head.pngRow three: the head close up, from three-quarters in front of the body’s own facing, with the chains drawn over it. This is where ears and whiskers are big enough to read.
springs.csvOne row per motion and chain
springs-ticks.csvOne row per solver tick, motion and chain: motion,chain,tick,seconds,tipX,tipY,tipZ, for plotting
summary.txtThe chains, the override, and one verdict line per finding

springs.csv has these columns:

ColumnMeaning
motion, chainThe motion token, and the chain’s name: its target path, qualified by the import that brought it (see below)
bonesReal bones in the chain, the anchor included and an endpoint not
endpointThe authored endpointPosition, or empty
settleSecondsTime after the motion ends until the tip stays within max(1 mm, 2% of the chain’s length) of where it finishes, for at least 0.25 s. never when it is still moving at the end of the run.
overshootDegreesThe first swing past the final direction, measured about the anchor
ringCountHow many times the tip swings through the final direction before it settles
maxTipDisplacementM, maxTipAngleDegreesThe farthest the drawn tip stood from where the animation alone would put it, from the motion’s start to the end of the run, in meters and as an angle about the anchor

A chain’s tip is its deepest slot, the endpoint when one is authored. Settling, overshoot and ringing are read off the solver’s tick states. Tip displacement is read off drawn frames.

summary.txt prints WARN for a chain of two bones with no endpointPosition (“only its base bone swings”) and FAIL for an unresolved chain, such as an override naming a path the avatar does not have, for two chains printing one name, which no override could tell apart, and for a model or object that did not import, since every chain on it then reads as a missing target. The run exits non-zero on any FAIL.

An override file is an object layer, laid over the avatar as one layer of overrides, the same layer a placed instance or a staged prefab carries: components on the root, children overriding the avatar’s parts by path, objects overriding the records it adds by id. Comments and trailing commas are allowed, as in any definition file.

mayu-tune.json
{
"children": [
{ "path": "Armature/Hips/Tail Root",
"components": [ { "$type": "dh.springBone", "pull": 0.35, "gravity": 0.2 } ] },
{ "path": "Armature/Hips/Spine/Chest/Neck/Head/Hair Root/Whiskers L",
"components": [ { "$type": "dh.springBone", "pull": 0.5, "endpointPosition": [0, 0.04, 0] } ] }
]
}
  • A chain is tuned by a dh.springBone on the part that roots it. Like every override, it states only what is being tuned: a field it leaves out inherits the avatar’s number. A part the avatar has no chain on gains one.
  • A chain an added record declares is tuned through that record. See below.
  • A dh.springCollider on the root is laid over the collider with the same id, field by field.
  • Any other component type, any other member, or a field the component does not read (a typo such as "pul"), is refused by name.

summary.txt echoes the layer, so a run always says what it was tuned with.

An accessory that carries its own chains, such as whiskers or a collar bell, declares them in its own .dh-obj, and the avatar adds it as a record in objects. Those chains are named <record>#<path>: the record’s id, then #, then the path of the chain’s root part in the accessory.

mltn Taidum adds its whiskers as the record whiskers and its collar as collar, so summary.txt prints:

chain 3 collar#Armature/CollarRoot/Bell_Anchor: 3 bone(s), endpoint
chain 4 whiskers#WhiskersArmature/Head/EyeBrow_L: 3 bone(s)

and an override reaches them through the record’s own children:

{
"objects": [
{ "id": "whiskers", "children": [
{ "path": "WhiskersArmature/Head/EyeBrow_L", "components": [
{ "$type": "dh.springBone", "pull": 0.8, "damping": 0.3, "limitType": "angle", "maxAngleX": 20 } ] } ] },
{ "id": "collar", "children": [
{ "path": "Armature/CollarRoot/Bell_Anchor", "components": [
{ "$type": "dh.springBone", "pull": 0.5, "damping": 0.12, "gravity": 0.07, "limitType": "polar", "maxAngleX": 60,
"maxAngleZ": 30, "limitRotation": [-13, 0, 0], "endpointPosition": [0, 0.0239, 0] } ] } ] }
]
}
  • The entry is laid over the accessory’s own chain, exactly as an entry for one of the avatar’s own chains is. Neither the avatar’s children nor the path the node ends up at after the armature link reaches it.
  • The # is what keeps the name unique: an avatar and an accessory may both declare a chain on Armature/Hips/Tail, and they print as Armature/Hips/Tail and acc#Armature/Hips/Tail.
  • A chain inside a record nested in a record is printed, but a layer cannot reach it.

Copy the chain’s numbers from the accessory’s .dh-obj when in doubt about what it holds; a one-value sweep (--springSweep "EyeBrow_L:pull=0.5001") prints the whole layer it swept in summary.txt.

--springSweep "<chain>:<field>=v1,v2,..." runs everything once per value, each run in its own folder, sweep-<field>=<value>/. The chain is its whole printed name, or just the last segment (after the last / or #) when only one chain ends that way. A vector value is space-separated numbers. The sweep is laid over the override file when both are given.

Worked example: a whisker’s endpointPosition

Section titled “Worked example: a whisker’s endpointPosition”

A whisker is a two-bone chain: an anchor and one tip bone. The anchor aims at the tip, and the tip has nothing to aim at, so without an endpoint only the anchor turns and the whisker swings as one stiff piece from its root. endpointPosition hangs a virtual segment past the tip, in the tip bone’s local space. A bone authored in Blender points down its own +Y, so [0, L, 0] extends the whisker by L meters along itself.

  1. Look first. Run the whiskers’ motions on the untouched avatar:

    DigitalHeaven.Engine.Host --render .tmp/springs --renderSprings
    --avatar tech.azuki.avatars.mayu:mayu-tora.dh-avatar --springMotion yaw180,nod

    summary.txt lists WARN .../Whiskers L: 2 bones and no endpointPosition; only its base bone swings.

  2. Sweep the length. The whisker mesh runs a few centimeters past its tip bone, so try a few:

    DigitalHeaven.Engine.Host --render .tmp/springs --renderSprings
    --avatar tech.azuki.avatars.mayu:mayu-tora.dh-avatar --springMotion yaw180,nod
    --springSweep "Whiskers L:endpointPosition=0 0.02 0,0 0.04 0,0 0.06 0"
  3. Read the head row. In each sweep-endpointPosition=.../yaw180.png, the hollow dot is the endpoint. The right length puts it on the end of the whisker mesh in the first tile, and keeps the drawn whisker bending along its length rather than pivoting at the root in the middle tiles. If the dot points off the whisker, the tip bone is not +Y-along-itself: try the axis the dot should have gone along.

  4. Read the numbers. Compare the whisker’s rows across the folders’ springs.csv. On the draft chains (pull 0.5, no endpoint yet) yaw180 gave:

    endpointPositionsettleSecondsmaxTipAngleDegrees
    [0, 0.02, 0]0.83121.5
    [0, 0.04, 0]0.77109.4
  5. Tune the feel. Put the chosen endpoint in an override file with the whisker’s other numbers, then sweep pull or damping over it the same way:

    --springOverride whiskers.json --springSweep "Whiskers L:pull=0.4,0.5,0.6"
  6. Copy the winner into source. Write the same numbers into the chain’s dh.springBone in the .dh-avatar, mirror them onto Whiskers R, rebuild the pallet, and run once more with no override. The strip has to match the one you picked, and summary.txt must have no WARN for the whiskers.

The run is a pure function of the avatar, the motions, the override and --springFrameHz: frames step by that fixed rate and nothing reads a clock. Two runs write byte-identical strips, raw frames and CSVs.

--renderAnim --avatar <barcode> photographs the engine’s built-in animation: one pawn walked, run, jumped and flown through the live pawn visual, strip by strip, with a CSV of the phase, the blend weights, the foot heights and the foot skate beside each strip. It needs --avatar and cannot be combined with any other capture mode. Pass --set client.render.msaa=1 when two runs must be byte-identical.

Add --animVideo to render every frame of the locomotion clips instead, at --animFps (default 60), narrowed with --animClip; Engine/scripts/encode-anim-video.ts turns the frame folders into MP4s.

Two runs of the same scene produce byte-identical PNGs. That is the whole reason the feature exists — render diffing and regression images only mean anything if a pixel that moved moved because the content changed.

Getting there means pinning every input that would otherwise drift between runs:

PinnedWhy it would otherwise drift
The view is built directly as a render view, never through ClientCameraClientCamera carries the speed-FOV spring, the roll spring and a sub-tick interpolation alpha — all of which are functions of when you looked, not where from
The frame-in-flight slot is reset before every captureThe windowed path rotates through its in-flight slots, so the same shot would land in a different slot depending on how many frames preceded it
Every frame-to-frame history is dropped: the temporal accumulation and the motion filter’s previous cameraBoth would otherwise make a shot depend on whatever frame happened to precede it. For motion blur this is also what makes a still unblurrable — see above
The scene extent is re-baselined rather than grownThe renderer’s scene target only ever grows in normal play; without the reset, a wide earlier shot in the same run would leave a later narrow one rendering into a larger target
Background mip generation is drained before any captureMip chains are built on a worker thread; a capture that raced it would sample different mips depending on thread timing
Logic props are walked in author orderLight ranking breaks ties by submission order, so an unordered walk would reorder the light set
Panini is offWith the warp off the scene extent is exactly the output extent, always — and the shot is plain rectilinear, which is what a reference image should be
The preference store is virginSee above

--set client.render.rebuildDevice=N tears down every Vulkan object the run has stood up and builds them again N times, after the map is loaded and after every setting is applied, but before the first shot. Everything resident is re-transferred from the CPU source it came from — the scene from the pallets it is still holding open, the glyph atlas and icon sheet from their texels, the object cache’s shapes from its store — and the renderer’s live settings are carried across as one snapshot, so the pictures are the same pictures.

That is exactly the claim it exists to check: a run with rebuildDevice set must produce byte-identical PNGs to the same run without it. Each pass logs its wall time, the device’s resident allocation total and the process’s private bytes, so a long run shows drift as numbers.

Terminal window
dh render out/plain --map core:maps/surface-eval
dh render out/rebuilt --map core:maps/surface-eval --set client.render.rebuildDevice=25

In the live client the same rebuild is the gpu.rebuild console command, which takes the same repeat count.

--avatarService is a different shape of offscreen mode: not one shot and exit, but a long-lived render loop with no map, no swapchain and no window, driven by line-oriented commands on stdin and replying on stdout, for a host process that owns its own world and only wants a DigitalHeaven avatar drawn into it. It exists for the Minecraft mod, whose Blaze3D renderer has no compute pass or storage buffers to run DigitalHeaven’s own skinning shaders in, so the engine renders the avatar out of process instead and the host composites the result over its own frame.

DigitalHeaven.Engine.Host --avatarService --frameFile <path> --frameBytes <n>
DigitalHeaven.Engine.Host --avatarService --frameSection <name> --frameBytes <n>
FlagArgumentDefault
--avatarServiceNone. Required — it is what selects this mode.—
--frameFilePath to a file both processes memory-map. The engine writes rendered frames into it; nothing else touches it. The Java mods put it on a tmpfs on Linux, so it never reaches the disk.—
--frameSectionInstead of --frameFile: the name of shared memory the caller created. On Windows a page-file-backed section, opened with MemoryMappedFile.OpenExisting; on macOS a POSIX shared-memory object, opened with shm_open and mmap, whose name the service unlinks as soon as it is open. Its pages stay in memory.—
--frameBytesThe mapping’s size in bytes, sized by the caller for the largest viewport it will request (64 + width * height * 8 + 32: a header, then RGBA8 color, then float32 depth, then the hand item frames).—

A mapped regular file is written back to disk for nothing (on Linux continuously while its frame changes, on Windows when the mapping closes), which is why the callers avoid one where they can. Engine/tools/FrameWritebackBench drives the service’s own AvatarFrameFile writer at 60 fps, with no game, to measure what each backing writes to disk. The numbers and the rules for choosing are in the contract note’s “Frame backing” section.

Stdout carries only protocol replies — the process’s own Console.Out is redirected to stderr before anything else can log a line into it, so a caller parsing every byte of stdout never sees a stray sentence from the engine’s own startup banner or logging. The host exits when stdin closes.

The commands are list (every avatar the workspace can spawn), load <barcode>, frame <width> <height> <camera...> <pawn...> (renders one frame into the mapped file, replying with a sequence number once the bytes are there, in the caller’s own coordinate and angle conventions so it does no conversion), and quit. The full protocol — every field of the frame command, the mapped file’s exact byte layout, and the compositing contract the caller is expected to hold up on its side — is design-notes/minecraft-avatar-service-contract.md at the repository root; this mode is built against that file and it changes first.

A caller that darkens the frame itself, as Project Zomboid does, adds glow1 to a frame. The service then renders the same pose a second time with every light off and writes that glow plane after the frame, so the caller can dim the lit part and add back what the materials light themselves (unlit surfaces, emission) undimmed. The file has to be sized for it (64 + width * height * 12 + 32). Only frames that ask pay for the second pass; AvatarGlowPlaneTests holds it to the engine’s own dark render.

A caller whose own time is stopped adds pause1 to a frame or view. The avatar’s springs, eye look and blink then hold their last simulated state while the frame still renders, and the first frame without it is a zero-length step, so a resume never catches up the paused time. AvatarSimulationClock is the one place a frame’s step length and hold are decided; AvatarPauseTests holds it.

A host composites the frame by its alpha and depth, so blended parts (whiskers, glass, hair cards) need both of their own. HeadlessRenderer.ApplyCoverageCapture(true, cutoff) turns on the coverage capture mode:

  • The scene alpha is coverage. The scene clears to (0, 0, 0, 0) and draws no sky. Opaque surfaces write alpha 1, and the blended pipeline (SrcAlpha/OneMinusSrcAlpha color, One/OneMinusSrcAlpha alpha) accumulates color premultiplied by a coverage of 1 − Π(1 − a).
  • The resolve writes straight alpha. It tonemaps C / A and writes A to alpha, and it writes zero where A is zero. The mode rides the spare y lane of the resolve’s cut-fade push constant, so the block does not grow. Every other effect in the resolve (vignette, chromatic aberration, sharpen, the cut fade, dither) touches RGB only. Bloom adds RGB where the coverage may be zero, which is one reason the service turns it off.
  • Blended parts write depth. After the blended color, the blended draws are recorded again with depth write on, compare GreaterOrEqual (reversed-Z) and every color channel masked. They are drawn as alpha masks at client.render.coverageDepthCutoff (default 0.05), so a nearly clear texel writes no depth. CaptureDepth then returns the nearest surface at least that opaque, so a whisker in front of the body reads the whisker’s depth.

Glass and water compose their pixel from a copy of the scene and write alpha 1, so they are not coverage-correct under the mode. Off (the default), every frame is byte-identical to one rendered before the mode existed. PixelSwizzle.ApplyCoverage keeps the resolve’s alpha where depth was written, so an ordinary capture stays opaque and a coverage capture keeps each part’s own coverage. CoverageCaptureTests holds the renderer to this, and AvatarTranslucentCoverageTests holds the service frame to it.

AvatarServiceCommand and AvatarService (Engine/DigitalHeaven.Engine.Host/AvatarService/) are the implementation, tested in AvatarServiceTests (DigitalHeaven.Engine.Tests). The service renders through the live client’s own pawn path (PlayerPawnVisual), so an avatar worn here is fitted, scaled and posed exactly as one worn in the engine’s own client — it is not a second, parallel avatar renderer. As with --render, a Release build is required for interactive use: the renderer draws at roughly 45 ms/frame in Debug versus roughly 6 ms/frame in Release at 1080p.

--previewService <dir> is the avatar service’s sibling for an authoring app: a long-lived render loop with no map, no swapchain and no window that answers one request line with one picture. DigitalHeaven Studio starts it, because Studio owns no device and references no engine project.

DigitalHeaven.Engine.Host --previewService <dir>

The option takes no other. Stdout carries the protocol alone (every log line goes to stderr), the process prints ready once it can serve and exits when stdin closes or a quit line arrives.

LineFields, separated by tabs
Requestrender, id, kind (avatar, object or material), edge in pixels (16 to 1024), yaw, elevation, zoom, barcode
Reply, drawnok, id, file name, width, height
Reply, refusederror, id, one line saying why

The id is the caller’s name for the picture and its file name: the service writes <id>.png into <dir> (to a temporary name and then moved, so a reader never sees half a file), and a request whose file is already there is answered without drawing, which makes an id named by a content fingerprint a cache. Core.Previews (PreviewRequest, PreviewReply, PreviewProtocol.Fingerprint) is the one spelling of all of it for both ends.

  • An avatar or an object is loaded through the avatar service’s own renderer configuration (clear background, coverage alpha), framed from its bounds by PreviewFraming (the camera orbits the middle of the bounds at the distance the bounding sphere needs, from the +Z side at yaw zero) and drawn at rest. client.previewService.fieldOfViewDegrees and client.previewService.margin tune the framing.
  • A material is brought up on demand in an empty scene that stands on the core pallet (RenderScene.CreateEmpty over a pallet, HeadlessRenderer.LoadMaterial) and photographed by the material preview’s own rig on the body its kind is read by, as the inspector’s thumbnails are. Pallets the material names are opened through the workspace’s catalog. Yaw, elevation and zoom do not apply to it.
  • A pallet that changed since the last request (any built pallet’s path, size or write time) makes the service stand the scene up again, so nothing is photographed as it was before a rebuild.
  • A request it cannot draw is answered with error and the loop goes on.

PreviewServiceTests covers the command line, the framing and real renders; PreviewProtocolTests in Core.Tests covers the lines and the fingerprint.

The offscreen path is not a separate renderer. The swapchain and the offscreen target sit behind one output seam, so shadow cascades, the scene pass, bloom and the HDR resolve are the same passes running in the same order — only the last hop differs, presenting a frame versus reading one back. A headless graphics device simply requests no surface or swapchain support. The readback format matches the swapchain’s, so no gamma math happens on the way out, and the PNG is written by a small hand-rolled encoder rather than an image library.

See CONTRIBUTING.md → Offscreen Rendering for the full architecture — the output seam, the device setup and the encoder.