Skip to main content

Lipsync (Addon)

The lipsync addon parses incoming audio streams (from a microphone or a remote peer) and maps frequency formants to real-time mouth shapes on a 3D avatar face.

It runs locally with a heuristic viseme mapper, requiring zero machine learning runtimes or model downloads.


How It Works​

  1. LipsyncMouth: A Script component that processes a MediaStream chunk-by-chunk using Web Audio FFT, extracting vowel-formant features.
  2. StylizedFace: A canvas-based decal texture displaying blinking eyes and a mouth that responds to viseme weight changes.
  3. VisemeWeights: Six numeric weights from 0 to 1: jawOpen, aa, oo, oh, ee, and consonant. Use the exported ZERO_VISEME for the rest pose.

For a manually driven face, supply all six weights. ZERO_VISEME is exported from the main package:

import * as xb from 'xrblocks';

const face = new xb.StylizedFace();
face.setVisemes({...xb.ZERO_VISEME, jawOpen: 0.5, aa: 0.7});
face.setVisemes(xb.ZERO_VISEME);

Single User Setup (Microphone Test)​

To drive a mouth shape using the user's local microphone feed:

import * as xb from 'xrblocks';
import {LipsyncMouth} from 'xrblocks/addons/lipsync/index.js';

class MicTestScript extends xb.Script {
async init() {
// 1. Request microphone access
const stream = await navigator.mediaDevices.getUserMedia({audio: true});

// 2. Create a StylizedFace decal and attach to a pivot mesh (e.g. head sphere)
this.face = new xb.StylizedFace({showEyes: true});
this.add(this.face);

// 3. Create the mouth driver linking stream and face
this.mouthDriver = new LipsyncMouth(stream, {
target: this.face,
});

// Add driver to the scene so it ticks/updates
this.add(this.mouthDriver);
}
}

Netblocks Integration (Multiplayer Avatars)​

You can pair lipsync with the Netblocks multiplayer addon so that users see other peers' mouth shapes move when they speak over WebRTC voice channels.

The following script joins a local two-tab room in init(), installs the voice listeners, and then enables the microphone. When a remote audio track arrives, it creates a LipsyncMouth targeting that peer's avatar face. onSession() is not an xb.Script lifecycle hook; the demo's NetSample superclass calls that method explicitly.

import * as xb from 'xrblocks';
import {
enableNet,
BroadcastChannelTransport,
} from 'xrblocks/addons/netblocks/src/index.js';
import {LipsyncMouth} from 'xrblocks/addons/lipsync/index.js';

class MultiplayerLipsync extends xb.Script {
constructor() {
super();
this.drivers = new Map(); // Keep track of active drivers by peer ID
this.unsubscribe = [];
}

async init() {
this.sharedCtx = xb.core.sound.listener.context;
this.net = enableNet();
const session = await this.net.joinRoom('lipsync-room', {
transport: new BroadcastChannelTransport(),
});

// Triggers when a peer adds a voice stream
const offTrack = session.voice.onTrack((peerId, stream) => {
const remoteUser = session.users.get(peerId);
if (!remoteUser) return;

this.removeDriver(peerId);

// Create new mouth driver pointing to the remote user's avatar face
const driver = new LipsyncMouth(stream, {
target: remoteUser.avatar.face,
audioContext: this.sharedCtx,
});

remoteUser.avatar.add(driver);
this.drivers.set(peerId, driver);
});

const offTrackRemoved = session.voice.onTrackRemoved((peerId) =>
this.removeDriver(peerId)
);
this.unsubscribe.push(offTrack, offTrackRemoved);

await session.voice.enable(session.transport.remotePeerIds);
}

removeDriver(peerId) {
const driver = this.drivers.get(peerId);
if (driver) {
driver.dispose();
driver.removeFromParent();
this.drivers.delete(peerId);
}
}

dispose() {
for (const unsubscribe of this.unsubscribe) unsubscribe();
for (const peerId of this.drivers.keys()) this.removeDriver(peerId);
this.net?.leaveRoom();
}
}

xb.add(new MultiplayerLipsync());
await xb.init();

Resource Disposal Guidelines​

To prevent audio pipeline memory leaks, manage lifecycle hooks correctly:

  • Mouth Driver: Disposing the driver (mouthDriver.dispose()) disconnects Web Audio nodes, but does not close your shared AudioContext or stop the underlying MediaStream audio tracks.
  • MediaStream: Stop tracks acquired by your own getUserMedia() call when they are no longer needed. Let Netblocks manage its voice streams through net.leaveRoom(); mouth drivers should not stop those shared tracks.
  • Target Face: Disposing the driver resets the target face back to its rest/silent pose. The face mesh itself is not deleted and must be disposed separately.

For complete reference implementation details, see src/addons/lipsync/README.md or the code samples under samples/avatar_lab/lipsync_puppet/ and demos/netblocks/lipsync/.