Skip to content

Mock microphone input

Use case: your app records the user’s voice — speech-to-text, a voice memo, a language-learning app grading pronunciation, an audio level meter. You want an automated test that speaks into the microphone.

▶ Runnable sample: mock-microphone-input.spec.ts / live demo — how to run: examples/README

Playwright has no API for microphone input, so every solution has to replace the microphone somewhere in the stack:

  • Chromium’s launch flag (--use-file-for-fake-audio-capture) — built in and dependency-free, but Chromium-only, WAV-only, and fixed for the whole browser session.
  • OS-level virtual audio devices — cross-browser, but miserable to set up in CI.
  • Patching getUserMedia in the page — cross-browser and controllable at runtime; this is what the library below does.

Chromium can replace the microphone with a WAV file at launch — no extra dependencies:

playwright.config.ts
export default defineConfig({
projects: [
{
name: 'chromium',
use: {
launchOptions: {
args: [
'--use-fake-ui-for-media-stream', // auto-accept the permission prompt
'--use-fake-device-for-media-stream', // provide fake capture devices
'--use-file-for-fake-audio-capture=tests/fixtures/hello-world.wav',
],
},
},
},
],
});

The file must be WAV; convert anything with ffmpeg:

Terminal window
ffmpeg -i hello-world.mp3 tests/fixtures/hello-world.wav

By default the file loops forever. Append %noloop to play it once and then stay silent:

--use-file-for-fake-audio-capture=tests/fixtures/hello-world.wav%noloop

This is a good fit when one utterance per browser session is enough and Chromium is your only target. The limits are real, though: the file is fixed at launch — you can’t say one thing on the first screen and another thing later — and Firefox/WebKit get nothing.

One more sharp edge: on macOS the fake capture device is unreliable — getUserMedia() can stall indefinitely (with or without --no-sandbox, headed or headless; see e.g. crbug.com/1032604). Treat this approach as Linux/CI-only and use the library below for local runs on a Mac.

▶ Runnable sample: mock-microphone-input-flag.spec.ts (skipped on macOS for the reason above)

Cross-browser and runtime control: playwright-audio-mocking

Section titled “Cross-browser and runtime control: playwright-audio-mocking”

playwright-audio-mocking replaces navigator.mediaDevices.getUserMedia before the page loads and backs it with a real Web Audio MediaStream. Anything downstream — MediaRecorder, AnalyserNode, WebRTC, speech APIs — works unmodified, on Chromium, Firefox and WebKit.

Terminal window
npm install -D playwright-audio-mocking
import { test, expect } from '@playwright/test';
import { mockMicrophone } from 'playwright-audio-mocking';
test('transcribes speech', async ({ page }) => {
const mic = await mockMicrophone(page); // install BEFORE page.goto()
await page.goto('/voice-memo');
await page.getByRole('button', { name: 'Record' }).click();
await mic.play('tests/fixtures/hello-world.wav'); // stream into the mic
await mic.waitForEnd();
await page.getByRole('button', { name: 'Stop' }).click();
await expect(page.getByTestId('transcript')).toContainText('hello world');
});

Unlike the launch flag, playback is fully scriptable mid-test:

await mic.play('fixtures/greeting.mp3'); // any format the browser decodes
await mic.play('fixtures/question.mp3'); // switch files mid-stream
await mic.play('fixtures/hold-music.mp3', { loop: true });
await mic.pause(); // silence, position kept
await mic.resume();
await mic.stop(); // silence, position reset
const { playing, position, duration } = await mic.status();

Pass bytes instead of a path — for example straight from a TTS API, so each test can say something unique:

await mic.play(Buffer.from(await synthesizeSpeech('add milk to my shopping list')));
  • Call mockMicrophone(page) before page.goto() — it’s injected as an init script.
  • When using the library on Chromium, launch with its RECOMMENDED_CHROMIUM_ARGS (--autoplay-policy=no-user-gesture-required) so Web Audio can start without a user gesture.
  • The mock replaces the audio track only; getUserMedia({ audio, video }) passes the video constraint through to the real implementation by default.