Skip to main content
Version: Next

Render WebGPU effects into the camera

This guide shows how to render each frame of your published camera yourself with WebGPU. Whatever you draw is what other peers receive.

The rendering code is plain WebGPU on both platforms. The integration differs:

  • In React Native, return a WebGPU-rendered track from a camera track middleware. createCameraFrameProcessorSession from @fishjam-cloud/video-effects calls your worklet for every frame of the Fishjam camera, with the camera as a GPU texture and an output texture to draw into. It handles the output track, GPU synchronization with the video encoder, timestamps and frame lifetimes. Peers receive the result as your regular camera track.
  • On the web, render into a <canvas>, capture it with captureStream, and publish the stream as a custom source.

If you only need background blur or a background image, use the ready-made background effects instead of writing shaders.

Prerequisites​

  • The packages (@fishjam-cloud/video-effects 0.1.5 or newer, react-native-webgpu 0.10.1 or newer) and the Babel and minimum OS configuration from Background effects. The segmentation model and the Metro and expo-asset setup are not needed.
  • typegpu in your app, if you write your shaders in TypeGPU as below
npm install typegpu

Publish the camera through your own shaders​

First time?

The WebGPU effects tutorial builds this pipeline step by step in a working app, from a passthrough render pass to a watermark overlay and a color effect.

The example below publishes the camera in grayscale. It is a camera track middleware: it builds the pipeline, starts a frame processor on the raw camera track, and returns the processor's track in its place. To apply your own effect, replace the fragment stage.

note

The shaders are written in TypeGPU (TGSL): typed TypeScript functions compiled to WGSL by unplugin-typegpu. TypeGPU is not required; you can hand-write WGSL and prepend the bindings' bindingDeclarations yourself.

import tgpu from "typegpu"; import * as d from "typegpu/data"; import { dot } from "typegpu/std"; import type { TrackMiddleware } from "@fishjam-cloud/react-native-client"; import { createCameraFrameProcessorSession, createCameraShaderBindings, getCameraWebGpuDevice, getOutputSurfaceFormat, type CameraFrameKernel, } from "@fishjam-cloud/video-effects/fishjam-react-native"; // Full-screen triangle; uv spans the visible area. const vertexMain = tgpu.vertexFn({ in: { vertexIndex: d.builtin.vertexIndex }, out: { position: d.builtin.position, uv: d.location(0, d.vec2f) }, })((input) => { const positions = [d.vec2f(-1, -1), d.vec2f(3, -1), d.vec2f(-1, 3)]; const p = positions[input.vertexIndex]; return { position: d.vec4f(p.x, p.y, 0, 1), uv: d.vec2f((p.x + 1) * 0.5, 1 - (p.y + 1) * 0.5), }; }); function createGrayscalePipeline(device: GPUDevice) { const cameraBindings = createCameraShaderBindings(device, { cameraPixelLayout: "rgb", }); const fragmentMain = tgpu.fragmentFn({ in: { uv: d.location(0, d.vec2f) }, out: d.vec4f, })((input) => { const color = cameraBindings.sampleCamera(input.uv); const gray = dot(color.xyz, d.vec3f(0.299, 0.587, 0.114)); return d.vec4f(gray, gray, gray, 1); }); // TypeGPU cannot emit the camera's external-texture binding itself, so // prepend cameraBindings.bindingDeclarations to the resolved WGSL. const module = device.createShaderModule({ code: cameraBindings.bindingDeclarations + tgpu.resolve([vertexMain, fragmentMain]), }); const pipeline = device.createRenderPipeline({ layout: device.createPipelineLayout({ bindGroupLayouts: [cameraBindings.bindGroupLayout], }), vertex: { module, entryPoint: "vertexMain" }, fragment: { module, entryPoint: "fragmentMain", targets: [{ format: getOutputSurfaceFormat() }], }, }); return { cameraBindings, pipeline }; } export const grayscale: TrackMiddleware = async (track) => { const device = await getCameraWebGpuDevice(); const { cameraBindings, pipeline } = createGrayscalePipeline(device); const frameKernel: CameraFrameKernel = (frame, render) => { "worklet"; render(({ commandEncoder, outputView, cameraBindGroup }) => { const pass = commandEncoder.beginRenderPass({ colorAttachments: [ { view: outputView, loadOp: "clear", storeOp: "store" }, ], }); pass.setPipeline(pipeline); pass.setBindGroup(0, cameraBindGroup!); pass.draw(3); pass.end(); }); }; const session = await createCameraFrameProcessorSession({ track, device, width: 720, height: 1280, cameraShaderBindings: cameraBindings, frameKernel, }); return { track: session.track, onClear: () => void session.dispose() }; };

Switch it on like any camera middleware:

import { useCamera } from "@fishjam-cloud/react-native-client"; const { setCameraTrackMiddleware } = useCamera(); await setCameraTrackMiddleware(grayscale); // null restores the plain camera

How the example works:

  • getCameraWebGpuDevice() returns the app-wide GPUDevice, requested with the features the camera import needs. Build your pipelines, textures and bind groups on this device.
  • createCameraShaderBindings(device, { cameraPixelLayout: "rgb" }) gives your shaders sampleCamera(uv), which returns upright RGB. The Fishjam camera arrives as RGB on both platforms, so the layout is always "rgb".
  • TypeGPU cannot emit the camera's texture_external binding, so cameraBindings.bindingDeclarations is prepended to the resolved WGSL. The fragment stage targets getOutputSurfaceFormat() (rgba8unorm on Android, bgra8unorm on iOS).
  • Passing cameraShaderBindings to the session makes the render context carry a ready-made cameraBindGroup, rebuilt every frame because the camera's external texture expires with each frame.
  • frameKernel is a worklet. It runs on the camera frame thread for every frame, encodes one render pass, and the session submits it. The pipeline it uses is built on the JS thread and copied into the worklet.
  • The middleware returns session.track, and onClear disposes the session when the middleware is replaced, cleared, or the camera stops.

Rules inside the frame kernel​

The kernel receives frame, with timestampNanoseconds, isFrontCamera, width, height and rotationDegrees, and render. The function you pass to render(...) receives a WebGpuFrameRenderContext with the device, queue, commandEncoder, the live cameraTexture (a GPUExternalTexture, already upright), the output surface (outputTexture, outputView, outputWidth, outputHeight), and the upright camera size (cameraWidth, cameraHeight).

Rules your worklet must follow
  • Always draw into the provided outputView. Calling outputTexture.createView() per frame leaks native wrappers on the frame runtime, because GPUTextureView has no release API.
  • Call render(...) at most once per frame. Skipping it drops the frame; nothing is published for it.
  • Don't call queue.submit() yourself. The session submits your passes and synchronizes with the video encoder.
  • Finish GPU uploads before you create the session. Helpers like queue.copyExternalImageToTexture submit work internally. Running them from the JS thread while frames flow races the session's submissions and can crash the app, so upload textures in the middleware before createCameraFrameProcessorSession.
  • Camera bind groups cannot be cached across frames, because the external texture changes every frame. Use cameraBindGroup, or call createCameraBindGroup inside the worklet each frame.
  • Capture only what a worklet can copy. The kernel's closure is copied to the camera frame runtime when the session starts: GPU objects, numbers, strings and other worklets work. React state and TypeGPU root objects don't, and later changes on the JS thread are not seen by the kernel.

Going further​

  • Overlays: a frame may contain more than one render pass. Encode additional passes into the same outputView (with loadOp: "load") after the camera pass to draw watermarks or other content on top.
  • Aspect ratio: the grayscale example stretches the camera to 720×1280. createCameraPassthroughPipeline and encodeCameraPassthrough draw the camera cropped to fill the output, given a crop from computeAspectFillCrop. To sample a cropped camera from your own shaders, createCameraTextureResolver and resolveCameraTexture render it into an owned rgba8unorm texture first, at the cost of one extra render pass per frame.
  • Reading frames without drawing: to run on-device inference on camera frames without changing the published video, see Process camera frames in a worklet.
  • Output size: width and height set the published resolution. The camera is imported at its own resolution and your passes decide how it maps onto the output.

The full toolkit is documented in the Video Effects API reference.

Platform notes​

  • The Fishjam camera arrives as RGB on both platforms. On Android the camera tap converts the camera texture before handing it over; on iOS the camera buffer is imported directly.
  • getOutputSurfaceFormat() returns the published surface format: rgba8unorm on Android, bgra8unorm on iOS. Use it for your fragment targets instead of hard-coding a format.
  • The published frames are not mirrored. The local self-view mirrors the front camera when you render it with mirror={true}.