Skip to main content

How it works

Audio files are automatically transcribed to text using speech recognition, then the transcript is analyzed by all enabled text-based policies. This means any policy that works on text (toxicity, hate, PII, wordlists, guidelines, etc.) also works on audio with zero additional configuration.

Supported audio formats

Any format FFmpeg can decode is supported. All audio is internally converted to 16 kHz mono WAV before transcription.

Limits

Transcription quality

You can configure transcription quality per channel in the dashboard under Content > Audio > Transcription quality.

Usage and billing

Each audio moderation request costs 2 units: 1 for transcription + 1 for policy analysis.