Emergency server help: get in touch

Call Recording Transcription for Asterisk, FreePBX and Vicidial

Upload an Asterisk, FreePBX or Vicidial call recording (.gsm, WAV49, .wav, µ-law, A-law) and get a timestamped transcript with caller and agent separated. TXT, SRT and JSON downloads.

Last updated
October 6, 2026

Upload a call recording and get back a written transcript with timestamps, split into caller and agent when the recording is stereo. It reads the files Asterisk, FreePBX, Issabel and Vicidial write to disk, including raw .gsm, GSM-in-WAV (WAV49), G.711 µ-law and A-law and signed-linear .sln, so you do not have to convert anything first. Download the result as plain text, SRT subtitles or JSON.

Short answer: Choose the recording, leave the language on “Detect automatically”, tick the consent box and press Transcribe. A 3-minute call usually comes back in well under a minute. Transcription runs on srvScripts’ own servers with the open-source Whisper model; recordings are not sent to any third-party AI service.

Drop one recording here or choose it.
Asterisk, FreePBX, Issabel and Vicidial formats as they are on disk (.gsm, GSM-in-WAV, .wav, .ulaw, .alaw, .sln) plus MP3, M4A, OGG and FLAC.

Before your first upload on this visit we ask you to confirm privacy and consent. Privacy policy

How it works

  1. You choose a recording. The page checks the length in your browser so you know straight away whether it fits your plan.
  2. The file is uploaded over HTTPS directly to the transcription server. WordPress only hands out a short-lived upload ticket; it never stores the audio.
  3. The server converts the audio to 16 kHz mono (or keeps both channels for stereo calls), skips long silences such as hold music gaps, and runs the speech model.
  4. You see the transcript line by line with timestamps, and can copy it or download TXT, SRT or JSON.
  5. The recording is deleted from the transcription server as soon as the transcript is ready.

Supported recording formats

FormatWhere you find itNotes
.gsm (raw GSM 06.10)Asterisk / FreePBX default for older systems, VicidialHeaderless 8 kHz; read directly, no conversion needed
.WAV / WAV49 (GSM in WAV)Asterisk format=wav49, many Vicidial installsAlways mono, so caller and agent cannot be separated
.wav (PCM, µ-law, A-law)FreePBX default, call centre exportsStereo files can be split into caller and agent
.ulaw, .alaw, .sln, .sln16Asterisk MixMonitor raw formatsHeaderless; the extension tells the server how to read it
MP3, M4A, OGG/Opus, FLACCloud PBXs, 3CX exports, converted filesAny sample rate

If a file does not play anywhere, convert it first with the call recording converter (it runs in your browser) to check it is a real recording.

Limits

Without an accountFree accountPro
Recordings2 a day3 a day60 hours a month
Length per fileUp to 3 minutesUp to 30 minutesUp to 5 hours
ModelStandardStandardStandard or Accurate
Caller / agent split (stereo)YesYesYes
DownloadsTXT, SRT, JSONTXT, SRT, JSONTXT, SRT, JSON

How accurate is it on phone audio?

We measured both models on a 2 vCPU lab server without a GPU, using 232 seconds of recorded speech from six speakers (LibriSpeech, CC BY 4.0). We then converted the same audio to the two formats phone systems use most, 8 kHz G.711 µ-law and 8 kHz GSM, and transcribed it again. The figures are word error rate: the share of words that were wrong, missing or added. Lower is better.

ModelClean audio8 kHz µ-law (phone line)8 kHz GSM (Asterisk .gsm)Speed on 2 vCPU
Standard (Whisper base)4.8%5.3%7.3%about 6× faster than real time
Accurate (Whisper small)2.7%2.2%3.8%about 2× faster than real time

In practice that means roughly one word in 25 is wrong on a GSM recording with the accurate model. Real calls are harder than our test audio: background noise, people talking over each other and strong accents raise the error rate. Names, e-mail addresses and account numbers are the words most often misheard, so check those against the recording.

In a stereo test call (one speaker per channel, 8 kHz µ-law) with caller / agent split on, the accurate model scored 0.0% errors on the caller side and 2.1% on the agent side.

Getting caller and agent on separate channels

The transcript can only name “Caller” and “Agent” if they are on different channels of a stereo recording. Asterisk’s MixMonitor mixes both directions into one mono file by default. Its r() and t() options also write the received and transmitted audio to separate files, which you can join into one stereo WAV with sox:

; extensions.conf - record both legs separately as well as the mixed file
exten => _X.,n,MixMonitor(${UNIQUEID}.wav,r(${UNIQUEID}-in.wav)t(${UNIQUEID}-out.wav))

# afterwards: -in = what this channel sent (the caller, when MixMonitor runs on the
# caller's channel), -out = what it heard (the agent). Left = caller, right = agent.
sox -M 1700000000.123-in.wav 1700000000.123-out.wav 1700000000.123-stereo.wav

Upload the -stereo.wav file and tick “Stereo recording”. GSM-in-WAV (WAV49) is always mono, so record in wav if you want the split.

Privacy and consent

  • Recordings are processed on servers run by srvScripts. They are not sent to OpenAI, Google or any other AI provider.
  • The recording is deleted as soon as the transcript is ready. The transcript is kept for 24 hours so you can download it again, then deleted.
  • We do not use your recordings or transcripts to train models.
  • Call recordings usually contain personal data. Make sure you are allowed to process them: many countries require callers to be told the call is recorded, and some require every party’s consent.

Frequently asked questions

Which languages are supported?

English, German, Spanish, French, Italian, Dutch, Portuguese, Arabic, Urdu and Hindi can be picked from the list, and “Detect automatically” recognises around 90 more. Accuracy is best for English and the major European languages.

Can it tell who is speaking on a mono recording?

Not yet. On a stereo recording it labels the left channel as Caller and the right channel as Agent. On a mono recording all lines are unlabelled. Record calls in stereo if you need the split.

Why is my recording rejected as too long?

Each plan has a maximum length per file: 3 minutes without an account, 30 minutes with a free account and 5 hours on Pro. The server checks the real length after upload, so renaming or re-encoding the file does not change it.

Is there an API for transcribing many recordings automatically?

An API for PBX servers is planned, so recordings can be sent for transcription when the call ends. Until then, Pro members can upload files one at a time here.

Can I run the same thing on my own server?

Yes. The model is the open-source Whisper (MIT licence) running on faster-whisper. On a 2 vCPU server without a GPU the standard model transcribes about six times faster than real time and needs under 1 GB of RAM.

Free website test

Is your website set up right?

Check SSL, security headers, redirects, robots.txt, sitemap, llms.txt and security.txt in one test. It takes about 30 seconds.