Audio
Audio is a mixer with no device of its own: something outside it pumps it, which is what lets the same code run a desktop output, an iOS session and an offline test.
DigitalHeaven.Engine.Audio is a hand-rolled audio stack built on two vendored native libraries:
- miniaudio owns the output device and a node-graph mixer running at a fixed 48 kHz.
- Steam Audio (
phonon) provides per-source binaural HRTF 3D spatialization, wired in as a DSP node. Spatialization parameters are handed to the audio thread through an atomic swap, so the game thread never blocks the mixer.
The stack ships the device and mixer, 2D sound playback, and one 3D spatialized source.
The engine is device-less; an output pumps it
Section titled “The engine is device-less; an output pumps it”The ma_engine is always initialized with noDevice. It holds the node graph, every voice, every gain and the Steam Audio wiring, and it never talks to hardware. Sound comes out because a separate ma_device is attached to it and calls ma_engine_read_pcm_frames from its data callback.
That split is the whole reason a device change is survivable. Replacing an output means tearing down a ma_device and initializing another one; the engine underneath is not touched, so nothing audible is reset — playing voices keep their position, master/SFX/UI gains, per-voice volume and pitch, 3D source positions, the listener transform and the HRTF state all carry straight across. Sounds do not restart, and nothing re-loads.
Three shapes exist, chosen by AudioOutputMode:
| Mode | What pumps the mixer |
|---|---|
Default | A supervised real device that follows the system default. The engine uses this. |
Silent | miniaudio’s null backend — headless, no hardware, real-time. |
Manual | Nothing; the caller pumps ReadInterleaved for deterministic offline tests. |
Following the default output device
Section titled “Following the default output device”The engine plays on whatever the OS says the default output is, and moves when that moves — headphones plugged in, a USB interface connected or removed, the output switched in the OS mixer. No relaunch, and no silence that needs one. There is no device picker: following the default is the only behavior.
Two signals feed one decision, because neither alone is enough:
- A poll of the default device (
client.audio.device.pollInterval, 0.5 s) is what notices the default moving to a different device while ours is still perfectly healthy. No backend raises a notification for that — a reroute notification fires when your device is taken away, not when a different one becomes preferred — so it is polled withma_context_get_device_info(…, NULL, …), whose returned id blob is the identity that gets compared. - A device notification callback (
ma_device_notification_type_stopped) catches a device that is invalidated or removed, immediately, without waiting out a poll interval.
miniaudio’s own WASAPI automatic stream routing is disabled (wasapi.noAutoStreamRouting). It half-solves the problem — it reroutes on a default change but discards the result, and when the last device disappears it stops the device and never restarts it, which is exactly the “goes silent until you relaunch” failure — and two authorities racing to reopen one device is worse than one authority that always wins. Backends that reroute a live device internally (CoreAudio) are still absorbed: the output reports its identity by reading the device live, so a device that was rerouted underneath us already matches the new default and no redundant switch happens.
Both signals arm one small state machine (AudioOutputSupervisor), which runs on its own thread and never on the frame thread — opening a WASAPI endpoint takes tens of milliseconds and would otherwise hitch the render. Its rules:
- Debounce. A change must settle for
client.audio.device.debounce(0.25 s) before it is acted on, and each further change restarts that timer. A burst — an unplug and replug, or a dock that enumerates in stages — therefore costs exactly one reinitialization, never one per notification. Opening and closing only ever happen on that one thread, so overlapping reinitializations are impossible by construction. - A failure never costs sound that is already playing. If the new default will not open, playback stays on the device it is already on and the attempt is retried on a doubling wait (
client.audio.device.retryInterval→client.audio.device.maxRetryInterval). Only when there is nothing playing does it fall back to silence. - No device at all is a normal state. With no playback device — including at startup — the mixer runs on the null backend: still real-time, so voices finish and loops loop and nothing piles up to blare later, just inaudible. When a device appears the poll picks it up and hardware resumes on its own. The engine never crashes, never blocks and never gives up.
- No click. Every output fades in over
client.audio.device.fade(8 ms), and a planned teardown fades out first, so neither end of a switch snaps. An unplanned loss cannot be faded — the hardware is already gone — but nothing stale is emitted either: miniaudio pre-silences the output buffer and the callback never writes past what the engine produced.
Each real transition writes one line on the audio log channel — output: <device>, output follows the new default: <device>, no playback device — running silent, could not open <device> — …. Nothing is written for a poll that found no change, so this is a line per transition, not per frame. There is deliberately no toast: following the default is meant to be invisible, and a user who just unplugged their headphones does not need to be told.
Portability: the mechanism is backend-neutral. The default-device query and the notification callback are core miniaudio, so macOS (CoreAudio) and Linux (ALSA/PulseAudio) follow the default by the same path — with CoreAudio additionally doing some of the work itself, which the live identity read absorbs. Only wasapi.noAutoStreamRouting is Windows-specific, and it is ignored elsewhere. miniaudio is prebuilt in the repo for win-x64 and linux-x64; macOS needs the recipe in native/build-miniaudio.md, which also builds the dh_output_* shims.
The reconnection logic is testable headlessly: the device layer sits behind IAudioOutputBackend / IAudioOutput (the same split as IPacingWaiter under DeadlineWait), so debounce, coalescing, backoff, the silent fallback and the resume are all exercised against a fake with no hardware.
Sound file formats
Section titled “Sound file formats”Every sound is decoded by miniaudio, whichever entry point loads it: LoadSound, CreateSource and PlayOneShot2D all hand a file path to the same ma_decoder behind the resource manager. Each sound is decoded in full when it loads (MA_SOUND_FLAG_DECODE). Nothing streams yet, music included. The decoder chooses by content, not by name. The extension only decides which decoder is tried first, and each decoder checks the file’s magic bytes before accepting it, so a mislabeled file still decodes as what it actually is.
| Format | Decoder | Notes |
|---|---|---|
| WAV (RIFF, RIFX, RF64, Wave64) | dr_wav | |
| AIFF | dr_wav | Big-endian 8/16/24/32-bit PCM at any rate, including the 80-bit extended rate field. |
| AIFF-C | dr_wav | NONE, sowt, raw , fl32/FL32, fl64/FL64, alaw/ALAW and ulaw/ULAW. Refused: ima4, twos, in24, in32, MAC3, MAC6, GSM and any other type. |
| Ogg Vorbis | stb_vorbis 1.22 | Compiled into the miniaudio binary (native/stb_vorbis.c). |
| FLAC, Ogg FLAC | dr_flac | |
| MP3 | dr_mp3 |
Ogg Opus, Speex and anything unrecognized are refused. Nothing fails silently. A file miniaudio will not open raises an AudioException that names the path and the format its magic bytes declare (AudioFileKind.Sniff), for example it is AIFF-C 'ima4', which the engine cannot decode. The cues that tolerate a missing sound (AudioOneShots, FallWind, SurfaceScrape) log that message once per file as a warning on the audio channel through AudioLoadFailures, rather than dropping it.
AudioFileProbe.Scan / Decode read a file through that same decoder at its own rate and channel count, block by block, and report the kind, rate, channels, declared and decoded frame counts and the peak. The format tests use it, and so did the check that all 198 Ogg and AIFF files in ULTRAKILL’s pallet decode to the end.
Audio on or off
Section titled “Audio on or off”client.audio.enabled (bool, default true) decides whether the client opens an audio device at all. It is read once, by the shared ClientAudioComposition both hosts compose through, so a change applies on the next launch rather than live — the point of switching audio off is that the mixing thread never exists, which a live toggle could not promise. False and the desktop host’s audio system stays null (every consumer already tolerates that and degrades to silence, and the play console command says so) while the iOS host never constructs its AVAudioSession adapter. One line is written on the audio channel at startup either way, naming the setting, so a log states its own audio configuration. It exists so a device with no cable can be flipped from the in-game console and relaunched; see the iOS lifecycle and diagnostics section of the engine overview.
Volume controls
Section titled “Volume controls”One-shots route through one of a small set of volume buses (categories). The mixer’s Master gain sits over the whole mix; under it, each category applies its own multiplier to the fire-and-forget one-shots routed to it. The Audio section of the settings screen exposes one live slider per bus, all stored as linear 0..1 fractions and shown as percentages:
- Master volume (
client.audio.master) scales the whole mix via the mixer’s master gain. - SFX volume (
client.audio.sfx) attenuates the SFX bus — gameplay one-shots (footsteps, landings, fall pain, the spawn gasp) and the fall-wind loop — on top of the master volume. - UI / Menu volume (
client.audio.ui) attenuates the UI bus — interface/menu one-shots (the noclip toggle blip; future menu sounds) — on top of the master volume.
The bus is chosen per call: PlayOneShot2D(path, volume, pitch, AudioCategory category = AudioCategory.Sfx) scales the caller’s volume by that category’s multiplier (SfxVolume for Sfx, UiVolume for Ui) inside the audio system. The category-less overloads default to SFX, so gameplay callers are unchanged. None of the buses affect loaded Sounds or 3D sources.
The fall-wind loop (FallWind) is an SFX sound but plays through a loaded looping Sound rather than PlayOneShot2D, so it applies the SFX multiplier itself: it reads IAudioSystem.SfxVolume and scales the loop’s live gain by it every frame, so moving the SFX slider mid-fall quiets the wind at once.
The noclip blip is the one UI sound. It is routed through the UI bus at a fixed 0.5 base gain (ClientHost.NoclipVolume) so the debug toggle sits under gameplay — the base drop stacks with the sliders: effective = NoclipVolume × UiVolume × Master.
All buses apply live — the host pushes the client.audio.* preferences onto the audio system every frame, so a slider or a loaded profile takes effect with no restart.
Output device preferences
Section titled “Output device preferences”The device-following timings are pushed the same way and are retunable from the console with no relaunch. They are not on the settings screen; the defaults are meant to be right.
| Preference | Default | What it does |
|---|---|---|
client.audio.device.pollInterval | 0.5 | Seconds between checks of the system default output device. Bounds how long a switch takes to begin. |
client.audio.device.debounce | 0.25 | Seconds a detected change must settle before it is acted on; restarted by each further change. |
client.audio.device.retryInterval | 1 | First wait before a failed open is retried; doubles per consecutive failure. |
client.audio.device.maxRetryInterval | 8 | Ceiling the doubling retry wait is clamped to. |
client.audio.device.fade | 0.008 | Gain ramp at each end of a switch, so a device that starts mid-waveform does not click. |
Gameplay sound files (footsteps, impacts) ship as loose files beside the core pallet at content/core/sounds/ — they are not yet bundled into the .pallet blob. SoundLibrary resolves them from AppContext.BaseDirectory/content/core/sounds at runtime; on a no-content build they resolve to nothing and audio degrades to silence.
One movement-feel layer, every host
Section titled “One movement-feel layer, every host”Everything the local player hears and feels while moving — footsteps, the jump scuff, the landing thud and its pain grunt (MovementAudio), the fall/flight wind (FallWind), the surface scrape (SurfaceScrape) and the camera punch that rides the same landing edge (FallImpactRoll) — is one implementation, MovementFeel, in the shared client-frame assembly DigitalHeaven.Engine.Client.Frame — the feel is one part of the frame both hosts drive, not a layer beside it. The desktop host calls it from ClientHost.DriveMovementAudioTick; the iOS app calls it from NetworkSession.DriveMovementFeelTick. Neither owns a line of the logic, so the iPad’s feel is the desktop’s feel by construction rather than by discipline (see CONTRIBUTING.md § “Hosts compose shared logic; adapters stay thin”).
It is a separate assembly for one reason: it needs both the prediction state in DigitalHeaven.Engine.Client and an IAudioSystem from DigitalHeaven.Engine.Audio, and Client deliberately does not reference Audio (Audio references nothing at all). A layer needing both has nowhere else to live.
MovementFeel reads its live tuning straight from the client PreferenceStore on every tick — client.audio.*, client.feel.* and client.camera.fallImpact.* — so a console or settings edit retunes the feel with no relaunch and no host carries a preference-push block of its own. It also re-checks the host’s audio system by reference each tick: a host that disposes and recreates its mixer (as the iOS adapter does across every interruption and background/foreground cycle) transparently gets fresh looping voices on the new device instead of writing to voices the old mixer already took down.
MovementAudio’s footstep, landing and wall-hit samples all play through the one shared one-shot player, AudioOneShots.PlayAt (the same helper HitDressing uses for hitscan impact sounds) — never a second PlayOneShot2D call site of its own. The local player’s own sounds are positionless by design (the listener is its head), so they call PlayAt with the listener and the sound both at Vector3.Zero, a co-located point at full gain. Another player’s footstep is placed: it plays from where that pawn’s foot came down and falls off with its distance from the listener.
Footsteps are the feet landing
Section titled “Footsteps are the feet landing”A footstep is never inferred from distance. Every pawn’s stride clock (GaitClock, see Animation) plants each foot as its leg phase wraps, and the legs are drawn landing on those same plants. ClientEntityPoses collects the frame’s plants as PawnFootfalls, and each host hands them to MovementFeel.PlayFootfalls once a frame, after the local body is submitted. Each footfall is exactly one footstep, and a frame with no footfall is silent, so the step heard and the step seen cannot drift apart.
- Every pawn has footsteps. The clock ticks for a humanoid, an avatar with no legs to animate, the hull box and the stand-in shown while an avatar loads, at the same cadence for the same speed. It ticks for your own pawn with the body hidden in first person.
- Other players cost nothing on the wire. Each client runs its own clock for every pawn from the motion it already draws them with (the drawn position and the replicated
Grounded,NoclipandFlyingflags). There are no per-step messages. - No speed threshold. The clock does not advance while standing, airborne, in noclip, flying or frozen, so it plants nothing then. A pawn that stops finishes its step, which is a soft step, and then falls silent. Volume still scales with the walking speed, from a soft scuff at a creep to full gain at a sprint.
- The touchdown is a step. A local foot planted in the frame a landing already sounded stays silent, so a landing never plays twice.
- Surface kind. Your own steps use the ground kind under the last predicted tick. Other players’ steps use the default set for now; their ground is not resolved on this client yet.
Movement audio timing
Section titled “Movement audio timing”The feel layer’s landing, jump, wall-hit, wind and scrape are driven from the predicted motion, and stepped once per simulation tick inside the client’s fixed-tick loop — never once per rendered frame. Footsteps are the exception: they play once per frame on the stride clock’s plants, the clock the legs are drawn on. Movement edges can live and die inside a single 60 Hz tick, so sampling them at frame rate would alias them away non-deterministically.
The landing thud keys off an explicit CharacterFlags.Landed edge the mover raises on the tick of fresh ground contact, not the persistent Grounded flag. This matters for a platform bounce: a landing immediately consumed by a buffered jump (Space tapped within the jump buffer of touchdown) never latches Grounded — the mover clears it in the same tick it applies the jump impulse — but it still thuds. Landed is transient and never networked (SnapshotCodec.ToReplicaFlags maps only the persistent stance flags); the local predictor recomputes it during reconciliation replay.
A landing hard enough to hurt layers a fall-pain grunt on top of the impact thud. Whether it hurts depends on the landing value, impactSpeed * groundHardness, compared against the live client.feel.landingHurtValue preference (default 9.9, clamped 0.5–25) — impact speed in m/s, hardness the ground material’s own 0..1 value from its surface preset. Stone’s hardness of 0.9 means it hurts at a descent speed around 11 m/s, while grass (hardness 0.15) would need an unreachable ~66 m/s to ever cross it — soft ground genuinely cushions a fall instead of every surface sharing one number. The same landing value drives the fall-impact view punch’s hard bonus (see Camera & View Effects), so the grunt and the extra snap fire together off the same math, declared once as FeelPreferences.DefaultLandingHurtValue in the shared preferences assembly and read live by both effects. MovementFeel reads the preference per tick, so a console edit retunes the “this hurt” line without a restart.
Landings are judged by the speed they cost
Section titled “Landings are judged by the speed they cost”A bunny hop that keeps its speed through the touchdown is not a crash, so the heavy impact is judged by what the landing took from the player, not by how fast they fell. LandingLoss (in Engine.Client, beside FallDescent, driven by the same predicted Landed flag on desktop and mobile) reads the speed on the last tick before touchdown and again client.feel.landingLossWindowTicks ticks after it, counting all three axes. The loss is the difference, never negative, so a hop that holds or gains horizontal speed loses about nothing.
On the touchdown tick the light footstep plays at once (when the fall cleared the audible floor). When the window closes, a loss of at least client.feel.landingHeavyLoss layers the heavy impact on top, with volume and pitch scaled from the loss. At 60 ticks a second the default window delays the heavy cue by 50 ms, which reads as the body arriving just after the foot; the light cue is never delayed, so the touchdown still lands on the frame it happens. Only the local player’s landings are voiced; other players are heard by their footsteps alone.
| Preference | Default | Range | Meaning |
|---|---|---|---|
client.feel.landingLossWindowTicks | 3 | 0-30 | Ticks after touchdown the speed is re-read; 0 judges the landing on its own tick. |
client.feel.landingHeavyLoss | 2.0 | 0-25 | Speed (m/s) a landing must cost before the heavy impact plays. |
client.feel.fallHurtKeepsMomentum | false | on/off | Experimental. A landing hurts less by the share of speed kept through it. Off leaves the landing value exactly impactSpeed * groundHardness. |
client.feel.fallHurtMomentumRelief | 0.5 | 0-1 | With the toggle on, the most a fully kept landing reduces the landing value by: value * (1 - keptShare * relief). |
The engine has no health or fall damage yet; the landing value that decides the pain grunt is the stand-in, and the toggle scales that value. The fall-impact view punch is unchanged.
Ground and wall material identity
Section titled “Ground and wall material identity”Both the landing thud and the footstep loop need to know what the player just touched, not just that they touched it. CharacterMover.TryGroundSurface (Engine/DigitalHeaven.Engine/Movement/CharacterMover.cs:545) and CharacterMover.TryWallSurface (same file, line 899) resolve a SurfaceSlotInfo (SurfaceKind + Hardness) from a short physics probe — straight down for the ground, along the slide’s blocked normal for a wall — the same per-triangle material-slot lookup (MaterialSlotNames → the material’s surface preset) that SampleSurfaceFriction already used for movement grip, so the identity and the friction are read from one resolution path rather than two. A miss (no ground/wall, or geometry with no material slot) reports SurfaceSlotInfo.Default rather than throwing, so a fall onto un-slotted geometry still plays something.
Footstep, landing and wall-hit sounds all pick their sample by kind. SurfaceSoundSets.ResolveFootstepSet(SurfaceKind) (Engine/DigitalHeaven.Engine.Client/SurfaceSoundSets.cs) maps a kind to its footsteps/{name}/{name}-N variant folder, falling back to the always-shipped stone set — and logging the fallback exactly once per kind per process on the audio channel — when a kind has no dedicated folder of its own. MovementAudio takes this resolver as a constructor dependency (defaulting to SurfaceSoundSets.ResolveFootstepSet) and reuses it for every sound kind (footstep, landing thud, wall-hit clang), rather than keeping a separate set per kind — one resolution, one fallback log, three playback sites. Four kinds ship dedicated folders: stone (5 variants, the fallback set and every kind’s safety net), dirt (4), grass (3) and metal (2), all under Engine/content/core/sounds/footsteps/. Grass and metal are audibly distinct from stone and from each other, so the surface identity is actually heard, not just modeled.
A wall hit raises the identical landing-shaped event, keyed by the wall’s own kind. MovementAudio.UpdateWallHit mirrors Update’s landing path exactly — same impactSpeed * hardness landing-value math, same hard-impact tier, same fall-pain-grunt layering — but resolves its footstep/impact sample from the wall’s SurfaceSlotInfo, independent of whatever the ground underfoot resolved. MovementFeel.Tick (Engine/DigitalHeaven.Engine.Client.Frame/MovementFeel.cs) calls it when the mover reports a hard-enough lateral impact (wallHit, wallImpactSpeed, wallSurface parameters), and keeps the louder, rarer landing sound as the tick’s reported Sound if both fire in the same tick — a wall hit alone still plays and is reported on its own.
Noclip flight has its own wind curve. While the local pawn is flying, FallWind swaps the target gain to a second, direction-agnostic curve driven by the total velocity magnitude rather than descent speed — so climbing, strafing and diving at the same speed all rush the same amount — and tops out at a lower ceiling so free flight never roars like a real long fall. Both curves feed the same smoothed envelope, so toggling noclip mid-flight eases between them instead of popping. Three live preferences tune it:
| Preference | Default | Meaning |
|---|---|---|
client.audio.noclipWind.startSpeed | 10 m/s | Total speed below which flight is silent (just under the default world.noclipSpeed cruise of 12, so drifting is near-silent). |
client.audio.noclipWind.fullSpeed | 30 m/s | Total speed at which the flight wind reaches its ceiling (full-throttle sprint flight: 12 × 2.5). |
client.audio.noclipWind.maxGain | 0.35 | Ceiling gain, well below the falling wind’s own 0.62. |
The falling curve has the matching three, so the same knobs exist on both sides:
| Preference | Default | Meaning |
|---|---|---|
client.audio.fallWind.startSpeed | 12 m/s | Airborne speed below which falling is silent. This is the dial for how eager the wind is: because the signal is a magnitude, a sprint-jump leaves the ground at sqrt(10.16² + 5.08²) = 11.36 m/s, and 12 is what keeps takeoff quiet. Deliberately equal to client.audio.surfaceScrape.startSpeed, but not to client.camera.speedFov.startSpeed — the view opens earlier (at 9) by design, because a sprint is meant to be seen and not heard. |
client.audio.fallWind.fullSpeed | 24 m/s | Airborne speed at which the fall wind reaches its ceiling. |
client.audio.fallWind.maxGain | 0.62 | Ceiling gain for a genuine fall. |
Sliding across a surface scrapes. The grounded twin of the fall wind, SurfaceScrape, is a second looping voice driven off the same SpeedRush signal — which on a grounded tick is the horizontal ground speed. It is silent through ordinary locomotion (the default walk is 5.08 m/s and a sprint 10.16 m/s) and opens up only once something other than legs is moving the player: a slick slope, a preserved bunny hop, a launcher. Airborne and noclip ticks target silence, so leaving the ground is a clean handover to the fall wind rather than two loops layered. Both the gain and the pitch ride the one ramp, so a faster slide is louder and brighter — that is what carries a single sample across the whole speed range.
| Preference | Default | Meaning |
|---|---|---|
client.audio.surfaceScrape.startSpeed | 12 m/s | Ground speed below which sliding is silent. The dial for what counts as “sliding” rather than “running” — above the 10.16 m/s sprint by design, and deliberately equal to client.audio.fallWind.startSpeed so the grounded rush and the airborne one open at one number. |
client.audio.surfaceScrape.fullSpeed | 18 m/s | Ground speed at which the scrape reaches its ceiling gain and top pitch. |
client.audio.surfaceScrape.maxGain | 0.45 | Ceiling gain — under the fall wind’s 0.62, because the scrape is a texture beneath the footsteps and impacts. |
client.audio.surfaceScrape.minPitch | 0.85 | Pitch multiplier at the deadzone edge: a low, heavy drag. |
client.audio.surfaceScrape.maxPitch | 1.25 | Pitch multiplier at the full-gain speed: brighter and more abrasive. |
One sample, every surface. The mover already resolves the grip of the dh.material underfoot (surfaceprop) each tick, so a per-surface scrape is the natural eventual hook, but it is not wired: the repo ships no per-surface scrape assets to select between, and the mover keeps surfaceFriction as a tick-local, so surfacing it to the client would mean widening the replicated character-motion struct for a cosmetic sound. When scrape assets per surface exist, the selection belongs beside the footstep-variant selection in MovementAudio and the friction plumbing can be paid for once, for both. The scrape sample itself is synthesized, not imported — Engine/scripts/make-surface-scrape.ps1 generates the seamless loop deterministically (see Engine/content/core/sounds/PROVENANCE.md), so unlike the HL2 placeholders it carries no third-party licensing.
The client reads the noclip state locally from the predictor (ClientPredictor.PredictedNoclip) — the effect is purely cosmetic and adds no wire state.
The wind does not compute its own speed. Both curves are fed the shared SpeedRush signal — the single scalar the host reduces each tick’s predicted motion to (the full velocity magnitude whenever the pawn is off the ground, whether falling or flying; the horizontal ground speed while grounded) — and both are that signal through the shared SpeedRush.Ramp deadzone/clamp, scaled by their own ceiling. The speed field of view is handed the very same float on the very same tick, so the rush you hear and the rush you see cannot drift apart. See Camera & View Effects.
The client.debugSoundCues convar toggles the SoundCueOverlay, a rolling top-right panel listing recent audio cues — kind, computed volume and pitch, age, and whether each played or was suppressed (a touchdown too soft to clear the audible floor is logged as a suppressed landing). Its landing readout shows each measured descent against the live roll / pain / max thresholds, making it the tuning companion for client.camera.* and client.feel.landingHurtValue. It is a developer tool for diagnosing audio-timing issues, off by default.
Native libraries
Section titled “Native libraries”The native binaries are vendored and stored via git-LFS (contributors need git-LFS installed to check them out). A per-RID resolver (NativeAudioResolver, installed with NativeLibrary.SetDllImportResolver) maps the platform-neutral miniaudio / phonon P/Invoke names to the right file at runtime, so no DllImport hard-codes a filename. miniaudio is built here for win-x64 and linux-x64 (build-win.ps1, build-posix.sh; flags in Engine/DigitalHeaven.Engine.Audio/native/build-miniaudio.md). Each binary carries a .built-from stamp of the three source files it was compiled from, and a test fails when one no longer matches the tree, so a change to miniaudio_impl.c cannot ship with an older binary. The iOS and Android builds apply the same check to their archives.
The iOS host statically links ios-arm64 miniaudio and Steam Audio archives plus Steam Audio’s pinned PFFFT/MySOFA/zlib companions; the shared resolver returns the main executable’s handle, with symbol retention derived from the production imports. Two AOT lessons are baked into the bindings: the resolver is asserted explicitly at the first managed entry point of each library (AudioDevice.Create, SteamAudioContext) because Mono’s full-AOT iOS runtime can dispatch a raw DllImport without running its declaring type’s static constructor, and the Steam Audio simdLevel ceiling is NEON on arm64 (phonon.h aliases it to the SSE2 slot; passing AVX2 there names an instruction set the CPU does not have).
Live iOS audio (validated on an iPad Pro M4). IosAudioSession is the platform adapter: it owns the AVAudioSession (Playback category, MixWithOthers), creates the shared AudioSystem.CreateForDevice(frameSize: 512) once the session is active, feeds the listener from the rendered eye every frame, applies the three volume preferences and drains route transitions to the log, and suspends/recreates the system across interruption, route change and background/foreground. Gameplay audio comes from the same MovementFeel layer the desktop client uses (see above) — spawn cue, footsteps, landing, fall wind, surface scrape — through the same IAudioSystem; the adapter holds no audio policy of its own. The core sounds ship inside the app bundle at content/core/sounds/, where SoundLibrary resolves them exactly as on desktop. The separate DH_PROBE_MODE=audio-offline DSP smoke (AudioSystem.CreateOffline, manual PCM pumps) remains the device-free regression path and never opens a playback device.
Licensing: miniaudio is public-domain / MIT-0, and Steam Audio is Apache-2.0. See the third-party notices under Engine/DigitalHeaven.Engine.Audio/.
Publishing a distributable build
Section titled “Publishing a distributable build”A plain dotnet build ships every RID’s natives side by side (native/<rid>/ from Audio, runtimes/<rid>/ from Physics, plus the NuGet-restored runtimes/<rid>/ for Silk.NET and friends) — convenient for development, but a shipped build only needs one platform’s worth. Engine/scripts/publish.bat produces that trimmed build:
Engine\scripts\publish.bat # win-x64 (default)Engine\scripts\publish.bat linux-x64Engine\scripts\publish.bat osx-arm64It runs a self-contained dotnet publish of DigitalHeaven.Engine.Host and drops the result at builds\engine\<rid>\ (git-ignored). Relative to a plain dotnet build -c Release, it strips PDBs, XML doc files, and every RID’s natives but the one requested — the last of those via -p:DhPublishRid, a plain MSBuild property Audio’s and Physics’s native-copy items condition on. It is not $(RuntimeIdentifier): the .NET SDK strips that property before handing global properties down a ProjectReference, so it never reaches a referenced project no matter how the top-level project tries to forward it; an ordinary property name has no such special-casing. The script prints the published folder’s total size when it finishes.
content/, Halcyon.Graphics.ShaderCompiler.exe, and the compiled core pallet still ship — those come from the Host’s own publish-time copy targets, independent of the RID trimming above.
Ambisonics, occlusion/reverb, GPU acceleration (TrueAudio Next), streaming, and a formal dh.sound asset type are planned but not yet implemented.