Emergency server help: get in touch

Transcribe Asterisk and FreePBX Call Recordings on Your Own Server (faster-whisper, Tested)

A Python script that transcribes Asterisk, FreePBX and Vicidial recordings (.gsm, WAV49, µ-law) on your own CPU server with faster-whisper, splits stereo calls into caller and agent, and writes TXT, SRT and JSON.

Published 8 min read

Short answer: Install faster-whisper and ffmpeg on the PBX or a separate server and run the script below on a recording. It reads Asterisk’s formats as they are on disk (raw .gsm, GSM-in-WAV, µ-law, A-law, .sln), splits stereo calls into caller and agent, and writes TXT, SRT and JSON. No GPU and no cloud service needed: on 4 CPU cores the standard model transcribes about three times faster than real time and the more accurate one about twice as fast.

We ran everything below on our lab server (AlmaLinux 9.8, 4 vCPU, no GPU) on 6 October 2026 with public test audio: the JFK sample from whisper.cpp (MIT) and a two-speaker call built from LibriSpeech (CC BY 4.0), converted to the formats a PBX writes. If you would rather not run it yourself, our call transcription tool uses the same engine.

Rather not run it yourself? Our call recording transcription tool does all of this for you: upload a GSM, WAV or MP3 file and get a timestamped transcript, with caller and agent split on stereo recordings. It runs on our own server, deletes the audio straight away and is free for short files; Pro covers 5-hour files and 60 hours a month.

What you need

  • Python 3.9 or later and ffmpeg/ffprobe (from your distribution, RPM Fusion on AlmaLinux, or a static build).
  • RAM: about 0.75 GB for the base model and 1.4 GB for small in our measurements.
  • Disk: the models download once (about 150 MB for base, 500 MB for small).
  • Run it as an unprivileged user that can read the recordings, not as root, when you put it into production.
python3 -m venv /opt/transcribe
/opt/transcribe/bin/pip install faster-whisper
dnf install ffmpeg        # or: apt install ffmpeg

The script

Save it as transcribe-call.py. It decodes the recording to 16 kHz mono with ffmpeg (each channel separately with --split), runs Whisper with voice activity detection so long silences and hold gaps are skipped, and writes the three output files next to the recording.

#!/usr/bin/env python3
"""transcribe-call.py - transcribe a PBX call recording on your own server (faster-whisper, CPU).

Reads what Asterisk/FreePBX/Vicidial write: .wav (PCM, u-law, A-law, GSM-in-WAV), raw .gsm/.ulaw/.alaw/.sln, MP3.
Stereo recordings can be split into caller (left) and agent (right). Writes TXT, SRT and JSON next to the input.

  python3 transcribe-call.py CALL.gsm [--model base|small] [--split] [--lang en] [--threads N]
"""
import argparse, json, os, subprocess, sys, tempfile, time

RAW = {'.gsm': ['-f', 'gsm', '-ar', '8000'], '.ulaw': ['-f', 'mulaw', '-ar', '8000'], '.ul': ['-f', 'mulaw', '-ar', '8000'],
       '.alaw': ['-f', 'alaw', '-ar', '8000'], '.al': ['-f', 'alaw', '-ar', '8000'],
       '.sln': ['-f', 's16le', '-ar', '8000', '-ac', '1'], '.sln16': ['-f', 's16le', '-ar', '16000', '-ac', '1']}

def decode(src, dst, channel=None):
    """Any input -> 16 kHz mono 16-bit WAV (one channel of a stereo file when channel is 0 or 1)."""
    cmd = ['ffmpeg', '-v', 'error', '-y'] + RAW.get(os.path.splitext(src)[1].lower(), []) + ['-i', src]
    if channel is not None:
        cmd += ['-af', 'pan=mono|c0=c%d' % channel]
    subprocess.run(cmd + ['-ac', '1', '-ar', '16000', '-c:a', 'pcm_s16le', dst], check=True)

def channels(src):
    if os.path.splitext(src)[1].lower() in RAW:
        return 1
    out = subprocess.run(['ffprobe', '-v', 'error', '-select_streams', 'a:0', '-show_entries', 'stream=channels',
                          '-of', 'csv=p=0', src], capture_output=True, text=True).stdout.strip()
    return int(out or 1)

def ts(t, sep=','):
    ms = int(round(t * 1000)); h, ms = divmod(ms, 3600000); m, ms = divmod(ms, 60000); s, ms = divmod(ms, 1000)
    return '%02d:%02d:%02d%s%03d' % (h, m, s, sep, ms)

def main():
    ap = argparse.ArgumentParser(description=__doc__.split('\n')[0])
    ap.add_argument('file'); ap.add_argument('--model', default='base'); ap.add_argument('--lang', default=None)
    ap.add_argument('--split', action='store_true', help='stereo: left = Caller, right = Agent')
    ap.add_argument('--threads', type=int, default=os.cpu_count() or 2)
    a = ap.parse_args()
    from faster_whisper import WhisperModel          # pip install faster-whisper
    t0 = time.time()
    model = WhisperModel(a.model, device='cpu', compute_type='int8', cpu_threads=a.threads)
    legs = [(None, '')]
    if a.split:
        if channels(a.file) < 2:
            sys.exit('--split needs a stereo recording; this file is mono (GSM-in-WAV is always mono)')
        legs = [(0, 'Caller'), (1, 'Agent')]
    segs, info = [], None
    with tempfile.TemporaryDirectory() as tmp:
        for ch, who in legs:
            wav = os.path.join(tmp, 'leg%s.wav' % ch)
            decode(a.file, wav, ch)
            it, info = model.transcribe(wav, language=a.lang, beam_size=5, vad_filter=True)
            segs += [{'start': round(s.start, 2), 'end': round(s.end, 2), 'speaker': who, 'text': s.text.strip()} for s in it]
    segs.sort(key=lambda s: s['start'])
    took = time.time() - t0
    base = os.path.splitext(a.file)[0]
    with open(base + '.txt', 'w') as f:
        f.writelines('[%s] %s%s\n' % (ts(s['start'], '.')[3:8], s['speaker'] + ': ' if s['speaker'] else '', s['text']) for s in segs)
    with open(base + '.srt', 'w') as f:
        for i, s in enumerate(segs, 1):
            f.write('%d\n%s --> %s\n%s%s\n\n' % (i, ts(s['start']), ts(s['end']), '[%s] ' % s['speaker'] if s['speaker'] else '', s['text']))
    with open(base + '.json', 'w') as f:
        json.dump({'file': os.path.basename(a.file), 'model': a.model, 'language': info.language,
                   'language_probability': round(info.language_probability, 3), 'duration': round(info.duration, 2),
                   'processing_time': round(took, 1), 'segments': segs}, f, indent=2)
    print(open(base + '.txt').read(), end='')
    print('-- %s: %.1f s of audio, language %s (%.0f%%), model %s, %.1f s on %d CPU threads (%.2fx real time) -> %s.txt/.srt/.json'
          % (os.path.basename(a.file), info.duration, info.language, info.language_probability * 100, a.model, took, a.threads,
             info.duration / took if took else 0, os.path.basename(base)))

if __name__ == '__main__':
    main()

Transcribe a recording

/opt/transcribe/bin/python transcribe-call.py /var/spool/asterisk/monitor/2026/10/06/call.gsm --model base
/opt/transcribe/bin/python transcribe-call.py call-stereo.wav --model small --split

The first run downloads the model. Our results, on a raw Asterisk .gsm file and a 59-second stereo µ-law call:

Terminal: A raw .gsm file with the base model (3.4x real time) and a stereo µ-law call split into caller and agent with the small model (1.9x real time). AlmaLinux 9.8, 4 vCPU, 6 Oct 2026.
A raw .gsm file with the base model (3.4x real time) and a stereo µ-law call split into caller and agent with the small model (1.9x real time). AlmaLinux 9.8, 4 vCPU, 6 Oct 2026. IP addresses masked.

Note the one mistake on the GSM file: “my fellow American” instead of “Americans”. GSM at 13 kbit/s loses detail, and the small base model is the first to suffer. Use --model small for anything you will rely on.

Terminal: The SRT file with caller and agent labels, and the check that refuses --split on a mono GSM-in-WAV file. 6 Oct 2026.
The SRT file with caller and agent labels, and the check that refuses –split on a mono GSM-in-WAV file. 6 Oct 2026. IP addresses masked.

Accuracy and speed

We measured word error rate (lower is better) on 232 seconds of LibriSpeech speech from six speakers, as clean audio and converted to the two most common PBX formats, on a 2 vCPU server:

ModelClean8 kHz µ-law8 kHz GSMSpeed on 2 vCPURAM
base4.8%5.3%7.3%about 6x real time750 MB
small2.7%2.2%3.8%about 2x real time1.3 GB

Real calls score worse than test audio: background noise, cross-talk and accents all add errors, and names, e-mail addresses and numbers are the words most often wrong. The medium and large models are more accurate but slower than real time on a small CPU server; use a GPU for those.

Getting caller and agent on separate channels

--split needs a stereo recording with one party per channel. Asterisk’s MixMonitor writes mono by default; its r() and t() options also save each direction to its own file, and sox -M joins them into one stereo file. GSM-in-WAV (wav49) is always mono, so record in wav if you want the split. Details and the dialplan line are in our transcription tool page.

Run it automatically after each call

The simplest way is a cron job that picks up finished recordings, the same pattern as our recordings-to-MP3 script: process files older than a few minutes that do not have a .txt yet. Run one transcription at a time, or the CPU is shared and every job gets slower.

# /etc/cron.d/transcribe-calls  (as the asterisk user, one file per run)
*/5 * * * * asterisk f=$(find /var/spool/asterisk/monitor -name "*.wav" -mmin +5 -mmin -1440 | while read r; do [ -e "${r%.*}.txt" ] || { echo "$r"; break; }; done); [ -n "$f" ] && flock -n /tmp/transcribe.lock /opt/transcribe/bin/python /opt/transcribe/transcribe-call.py "$f" --model small >/dev/null 2>&1

Privacy

Transcripts are as sensitive as the recordings. Keep them with the same permissions, delete them on the same schedule, and check that callers are told about recording and transcription where the law requires it. Running the model on your own server means the audio never leaves it; Whisper and faster-whisper are MIT-licensed and work offline once the model is downloaded.

See also: Asterisk Call Recording with MixMonitor: Formats, Stereo, Retention

Frequently asked questions

Does this need a GPU?

No. The base and small models run on CPU with int8 quantisation. A GPU makes the large models practical and is 20 to 50 times faster.

Which languages work?

Whisper recognises about 99 languages and detects the language itself; pass –lang to skip detection. Accuracy is best for English and the major European languages.

Can it tell speakers apart on a mono recording?

Not this script. It labels caller and agent only on stereo recordings with one party per channel. Speaker diarisation on mono audio needs an extra model such as pyannote.

Free website test

Is your website set up right?

Check SSL, security headers, redirects, robots.txt, sitemap, llms.txt and security.txt in one test. It takes about 30 seconds.