Visemes
2026-10-02. Built from three research reports (standards, DH’s avatars today, and BONELAB’s voice; every claim there carries its source). The proposal is kept as written; the decisions at the end are mltn’s and binding.
What exists today
Section titled “What exists today”- The avatar format already has visemes. A rig maps the 15 standard viseme keys (the Oculus set) to blendshapes (
RigDefinition.cs:43-49,Viseme.cs).- The compiler warns when a mapped shape is missing.
- The Unity runtime keeps mapped shapes on import.
- The VRChat export writes a full viseme lip-sync setup.
- Nothing moves a mouth at runtime on any host.
- The engine, Unity hosts, BONELAB and the GMod export all drive eyes and eyelids only.
dh.speakSound.mouth(loudness of a reaction clip drives the mouth) is declared but implemented nowhere.
- mltn’s avatars:
- Mayu: all 15 visemes as
vrc.v_*, mapped in its rig. It also hasMouth Open/Jaw Open, MMD あいうえお and the ARKit/Unified sets. The rig has itsJawbone commented out (intentional?). - Taidum: all 15 visemes under bare names (
aa,PP, …sil, matching its VRChat v4 descriptor), plusJaw_Openand a mappedJawbone. Its rig declares no visemes and its avatar names no target mesh, so hosts that import only referenced blendshapes (BONELAB, MegaBonk) probably drop the shapes. The fix is a pallet-source edit, so it goes through the pallet session.
- Mayu: all 15 visemes as
How other platforms do it
Section titled “How other platforms do it”- Oculus 15 is the shared set. It matches MPEG-4 index for index. VRChat, ChilloutVR, Resonite (plus laughter) and Ready Player Me use it, and
vrc.v_<oculus>names are auto-detected by VRChat, CVR and Resonite. - Lossy sets: VRM 0.x/1.0, VTube Studio and MMD あいうえお are vowels only; the visible loss is the lips never closing on p/b/m.
- ARKit is a set of separate mouth and jaw movements, not visemes; each viseme needs a recipe.
- Microsoft/Azure’s 22 IDs don’t map one-to-one onto Oculus.
- VRChat still uses Oculus Lipsync (OVRLipSync) per today’s docs. Each client analyzes the voice it hears locally; nothing is synced over the network. Animators get only the winning viseme index (0–14).
- Smoothing: OVRLipSync’s smoothing runs 1–100, with samples defaulting to 70; ChilloutVR exposes the same range.
- Latency and window size: unpublished; it has to be measured.
- VRChat’s beta face tracking (2026-09-24) adds multi-mesh visemes and “Viseme Dampen” against face tracking.
- VRM 1.0’s
overrideMouthlets an expression (e.g. happy) block or soften lip sync, which is worth borrowing.
BONELAB: where the voice comes from
Section titled “BONELAB: where the voice comes from”- The stock game has no mouth drive.
- LabFusion (v1.14.2):
- Capture: Unity
Microphone; G.711 A-law plus Deflate, not Opus. - Playback: a
VoiceSourceper player, attached toHeadSFX.mouthSrc. JawFlapper: tilts the humanoid Jaw bone with loudness (up to 40°) for the local and remote players; no blendshapes.- Gating: everything is gated on
NetworkInfo.HasServer. Plain single player with Fusion loaded gives nothing; a lobby you host alone should work. There’s no “hear self”. - What we can read: public pieces we can reach by reflection, as the mod already does with Fusion: local loudness (
VoiceInfo.VoiceAmplitude), each remote player’s amplitude, and the raw samples (Harmony hooks plusVoiceConverter.Decode). - Hazard: if our mod holds the mic when a Fusion server starts, Fusion’s voice goes silent. We must release the mic whenever a server runs.
- Capture: Unity
- Our own mic in single player:
MicrophoneandAudioClip.GetDatasurvive in the IL2CPP build (Fusion uses them). The other options are out: Steam voice (stripped binding), NAudio (Windows-only) and Meta voice (needs platform init). - Turning sound into visemes:
- OVRLipSync is ruled out: discontinued, a native DLL, and an unclear license.
- uLipSync: MIT, with an MFCC core of about 600 lines. It can be ported to plain C# (no Burst or Jobs) as a shared Core helper usable by every host, at an estimated well under 0.1 ms per voice.
- Minimum route: loudness drives one mouth-open shape.
Proposed architecture
Section titled “Proposed architecture”- Format (pick one):
- A: the Oculus 15 in camelCase as DH’s set. Fallback chain: full set → a fixed map of missing visemes onto the vowels present → one mouth-open shape driven by loudness → the jaw bone driven by loudness.
- B: a required core (
sil,pp, five vowels), with the rest optional. - C: ARKit-like mouth movements, with visemes as recipes over them.
- A shared Core mouth brain, modeled on the eye brain. Input per frame: 15 weights, loudness and laughter. Smoothing and attack/release are preferences. Output: weights for the avatar’s resolved shapes or jaw. It combines with authored weights by an explicit rule (today the engine takes the max, so a driven mouth couldn’t close one an author left open).
- Analysis in Core: port uLipSync’s MFCC core to managed C#. Loudness-only is the first stage and the fallback.
- A voice-source adapter per host.
- BONELAB: with a Fusion server running, read Fusion’s audio and never open the mic; without a server, open Unity
Microphonebehind a preference and release it when a server starts. - Analysis: done locally on every client for every voice it hears, as VRChat does, so nothing new crosses the network.
- BONELAB: with a Fusion server running, read Fusion’s audio and never open the mic; without a server, open Unity
- The same brain drives
dh.speakSound.mouthfrom the clip’s samples. - Housekeeping:
- fold the VRChat exporter’s own viseme loop into
NativeAvatarSlots.VisemeShapes; - call
Viseme.IsValidin the compiler; - add a platform-support table to
rig.mdx.
- fold the VRChat exporter’s own viseme loop into
Decisions (mltn, 2026-10-02)
Section titled “Decisions (mltn, 2026-10-02)”- Format: the Oculus 15 are the reference set, and every viseme is optional. Whatever is missing is substituted as well as the avatar allows:
- first, synthesized by blending two or three of the shapes present at chosen weights (recipes);
- else, the nearest shape substitutes;
- with no mouth shapes at all, the Jaw bone opens to an authored offset. The jaw is configured and driven in code the same way eye look is (bone rotation limits).
- Analysis: both. Volume-only and uLipSync’s shape analysis are two modes of one driver. Volume ships first, and the port is added on top.
- Fusion’s jaw:
- If the avatar has DH visemes, the DH driver runs and Fusion’s jaw flap is suppressed.
- If not, and the audio is Fusion’s voice in a server, DH runs nothing and lets Fusion flap the jaw. DH makes sure the DH avatar’s jaw is wired so Fusion can use it natively.
- Settings disable DH visemes for remote players, or entirely (own avatar too).
- Single-player mic: off until enabled in the DigitalHeaven menu. It is released whenever a Fusion server runs.
- Mayu’s Jaw: it was commented out because a mapped jaw breaks some hosts (VRChat; MegaBonk, where Taidum’s mouth gapes permanently). Decision: map it, and let each integration decide jaw use with a sensible, overridable default: off where it breaks, and off when the avatar has visemes. The pallet edit goes through the pallet session.
- MMD あいうえお: they count as a fallback source of vowel shapes in the substitution chain.
Standing principle (mltn): one standardized pipeline, not competing implementations. Sources (own mic, Fusion voices, dh.speakSound clips) feed one analysis stage (volume or uLipSync), which feeds one mouth brain, which feeds the avatar’s resolved shapes, recipes or jaw.
dh.speakSoundplays through the avatar’s voice channel and drives the mouth exactly like speech: real visemes when the avatar has them, volume on the jaw when not. This works with or without Fusion.- The only case DH leaves to Fusion: Fusion’s own voice, in a server, on an avatar without visemes.
Stage 1 as built
Section titled “Stage 1 as built”- Core contract (
DigitalHeaven.Core.Animation):MouthFrame(the 15 weights inViseme.Allorder, loudness, laughter, and whether it carries visemes at all) andIMouthAnalyzer(Processover interleaved samples,Current,Reset).VolumeMouthAnalyzeris the loudness-only analyzer; the uLipSync port is the second. - Resolution (
MouthRig): per viseme direct, recipe (MouthRecipes, MMD vowels as sources), substitute or missing, then the open fallback: a mouth-open shape, else the Jaw bone toward the rig’s newjawRotationLimits.open. See the rig page. - Mouth brain (
MouthBrain): visemes smoothed on OVRLipSync’s 1..100 scale, loudness through a noise floor and gain ontoaaor the open fallback, every number a preference (MouthSettingFields). It owns the mouth channels while the voice is heard instead of combining by max. - Clips:
MouthClipFeedfeeds adh.speakSoundclip’s played samples to the same analyzer, so a host lane only adds a source. - Housekeeping:
Viseme.IsValidruns in the compiler; the VRChat exporter usesNativeAvatarSlots.VisemeShapes; a referenced-only Unity import keeps every shape the resolution could use.