Short answer: AudioSocket is a small TCP protocol built into Asterisk (18 and later) that streams a call’s audio to your own server and plays back whatever audio that server returns. You call it from the dialplan with AudioSocket(${UUID()},bot.example.com:9092) or dial it as a channel with Dial(AudioSocket/bot.example.com:9092/${UUID()}/c(slin16)). Your bot service receives 16-bit signed linear PCM in 20 ms frames, runs speech-to-text, an LLM and text-to-speech, and writes audio frames back on the same socket.
We ran these commands on our lab server (Debian 12, FreePBX 17.0.33, Asterisk 22.11) on 6 October 2026, with a minimal echo server on 127.0.0.1 standing in for the AI service. The protocol details below come from the official AudioSocket documentation and the Asterisk 22 source, and the frame sizes from that test.
Table of Contents
What AudioSocket is (and is not)
AudioSocket gives you raw audio over a plain TCP connection: no SIP, no RTP, no WebSocket framing. That makes it the shortest path from a phone call to code you control. It is a good fit when:
- you want a voice bot, IVR or real-time transcription on calls that already reach Asterisk;
- your bot runs on the same host or private network as the PBX;
- you would rather write a TCP server than a SIP user agent.
It is not a full voice-AI platform. Asterisk only moves audio and DTMF. Turn-taking, barge-in, silence detection, speech recognition and speech synthesis are all your service’s job. There is no encryption on the socket, so keep it on localhost, a private network or a VPN.
Asterisk 22.11 on our lab shipped three modules for it, all loaded by default on FreePBX 17:
Module Description Use Count Status Support Level
app_audiosocket.so AudioSocket Application 0 Running extended
chan_audiosocket.so AudioSocket Channel 0 Running extended
res_audiosocket.so AudioSocket support 2 Running extended
Check yours with asterisk -rx "module show like audiosocket". “extended” support level means community-maintained rather than core.
The protocol in one table
Every message has a 3-byte header: 1 byte type, then a 2-byte payload length (unsigned, big-endian), then the payload. These are the message types in the official documentation and in res_audiosocket.h:
| Type | Meaning | Payload |
|---|---|---|
0x00 | Hang up / terminate | None (length 0) |
0x01 | Call UUID (sent first by Asterisk) | 16 bytes, binary UUID |
0x03 | DTMF digit | 1 ASCII byte |
0x10 | Audio, 8 kHz | 16-bit signed linear, mono, little-endian |
0x11–0x18 | Audio at 12, 16, 24, 32, 44.1, 48, 96 and 192 kHz | Same sample format |
0xff | Error | Optional error code |
Both directions use the same framing. Your server sends audio back as type 0x10 (or the rate matching the call) and can end the call by sending 0x00 with length 0.
Option 1: the AudioSocket() dialplan application
The application form is the simplest. Its syntax in Asterisk 22 is AudioSocket(uuid,service): the UUID must be a standard UUID string, and the service is host:port (IPv6 in brackets, for example [::1]:9092). It does not answer the call itself.
This is the context we used on lab3. On FreePBX, put it in extensions_custom.conf and point a Custom Destination or a misc application at it:
[srvs-test-voip]
exten => as-app,1,Answer()
same => n,AudioSocket(${UUID()},127.0.0.1:9092)
same => n,Hangup()
In this mode, the help text says Asterisk sends 16-bit, 8 kHz mono PCM. Our test confirmed it: every audio message was type 0x10 with a 320-byte payload, which is 160 samples x 2 bytes, or 20 ms of audio.
Option 2: the AudioSocket channel and higher sample rates
The channel driver lets you Dial() the bot like any other endpoint, which keeps normal features such as dial timeouts, DIALSTATUS and bridging. The dial string, from chan_audiosocket.c in Asterisk 22, is:
AudioSocket/host:port/uuid[/c(codec)]
The c() option picks the format sent to your server. Wideband audio improves speech recognition, so slin16 is a good default for AI bots:
exten => as-dial,1,Answer()
same => n,Dial(AudioSocket/127.0.0.1:9092/${UUID()}/c(slin16))
same => n,Hangup()
On lab3 this produced type 0x12 messages (16 kHz) with 640-byte payloads, again 20 ms each. The Asterisk log shows the channel being created and bridged:
app_dial.c: Called AudioSocket/127.0.0.1:9092/33a30935-f9da-4f04-b85a-e2767a5a4ceb/c(slin16)
app_dial.c: AudioSocket/127.0.0.1:9092-33a30935-f9da-4f04-b85a-e2767a5a4ceb answered Local/as-dial@srvs-test-voip-00000009;2
bridge_channel.c: Channel AudioSocket/127.0.0.1:9092-33a30935-f9da-4f04-b85a-e2767a5a4ceb joined 'simple_bridge' basic-bridge <b534e699-...>
The channel driver also sets AUDIOSOCKET_UUID and AUDIOSOCKET_SERVICE channel variables, which helps when you log or correlate calls.
A minimal test server
Before wiring in an AI stack, prove the plumbing with an echo server. This is the script we ran on lab3. It logs each message type, echoes audio back so the caller hears themselves, and hangs up after a few seconds:
#!/usr/bin/env python3
"""Minimal AudioSocket test server: logs message kinds, echoes audio back, hangs up after N seconds."""
import socket, struct, sys, time, uuid
HOST, PORT, SECONDS = '127.0.0.1', 9092, float(sys.argv[1]) if len(sys.argv) > 1 else 5.0
KINDS = {0x00: 'hangup', 0x01: 'uuid', 0x03: 'dtmf', 0xff: 'error'}
KINDS.update({k: 'audio' for k in range(0x10, 0x19)})
def read_exact(conn, n):
buf = b''
while len(buf) < n:
chunk = conn.recv(n - len(buf))
if not chunk:
return None
buf += chunk
return buf
srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
srv.bind((HOST, PORT)); srv.listen(1); srv.settimeout(30)
conn, _ = srv.accept()
start, counts, sizes = time.time(), {}, set()
while True:
hdr = read_exact(conn, 3)
if hdr is None:
print('socket closed by Asterisk'); break
kind, length = hdr[0], struct.unpack('>H', hdr[1:3])[0]
payload = read_exact(conn, length) if length else b''
counts[hex(kind)] = counts.get(hex(kind), 0) + 1
if kind == 0x01:
print('uuid', uuid.UUID(bytes=payload))
elif KINDS.get(kind) == 'audio':
sizes.add(length)
conn.sendall(hdr + payload) # echo the caller's audio back
elif kind == 0x00:
print('hangup from Asterisk'); break
if time.time() - start > SECONDS:
conn.sendall(b'\x00\x00\x00') # kind 0x00, length 0 = hang up
print('sent hangup'); break
print('message counts', counts, 'audio payload sizes', sorted(sizes))
conn.close(); srv.close()
Start it, then send a test call into the context with a Local channel that plays a prompt:
python3 as_echo.py 3 &
asterisk -rx "channel originate Local/as-app@srvs-test-voip application Playback demo-congrats"
Output from our two runs (application form, then channel form with c(slin16)):
uuid 8b66172c-b1d3-4514-ace5-9fea1bbd3b7d
sent hangup
message counts {'0x1': 1, '0x10': 152} audio payload sizes [320]
uuid 33a30935-f9da-4f04-b85a-e2767a5a4ceb
sent hangup
message counts {'0x1': 1, '0x12': 152} audio payload sizes [640]
The UUID message arrives first, then a steady stream of 20 ms audio frames (152 in about 3 seconds). That is the contract your AI service has to handle.
Architecture of an AI voice agent
A production bot replaces the echo with a pipeline. A common shape:
- Asterisk answers the call and connects it to your service with AudioSocket (one TCP connection per call, keyed by the UUID).
- Your service buffers incoming PCM and runs voice activity detection to decide when the caller has finished speaking.
- Speech-to-text (a streaming cloud API or a local model such as Whisper) turns speech into text.
- The LLM decides the reply and any actions (look up an order, transfer, end call).
- Text-to-speech produces audio, which your service resamples to the call’s rate and sends back in 20 ms frames.
- Barge-in: when new caller speech arrives while you are talking, stop sending audio.
Some speech-to-speech APIs accept audio in and return audio out, which collapses steps 3-5. Either way, AudioSocket only sees PCM frames. Check each AI provider’s own documentation for the audio formats it accepts; many expect 16 kHz or 24 kHz PCM, which is where c(slin16) or c(slin24) helps.
To hand the call to a human, end the AudioSocket session and let the dialplan continue to a queue or extension. In the Asterisk 22 source (app_audiosocket.c), only a hangup message (0x00) or the server closing the socket counts as a normal exit; the application then returns and the dialplan carries on with the next priority. Any other failure returns an error and the call is hung up. A common pattern is to set a channel variable through AMI or ARI before closing the socket, then branch on it after AudioSocket().
The same source sets a hard 2-second inactivity timer: if neither the caller nor your server produces a frame for 2000 ms, Asterisk logs Reached timeout after 2000 ms of no activity on AudioSocket connection and drops the call. Keep sending audio (silence frames are fine) while your AI service is thinking.
Latency budget and tuning
Callers notice delays above roughly a second. The budget is spent on: end-of-speech detection, speech-to-text, LLM time to first token, text-to-speech time to first audio, and the network. AudioSocket itself adds very little (20 ms framing and one TCP hop). Practical rules:
- Run the bot close to the PBX, ideally on the same host or LAN, and keep cloud AI calls in the same region.
- Stream everything: send partial transcripts to the LLM and start sending TTS audio before the full sentence is synthesised.
- Pace outgoing audio at real time (one 20 ms frame every 20 ms). Dumping a long reply at once makes barge-in hard to handle.
- Send silence frames while waiting on the LLM, so the 2-second inactivity timeout never fires.
- Use wideband (
slin16) for better recognition; resample in your service, not in Asterisk dialplan. - Log the UUID and timestamps for each stage so you can see where the latency goes.
Security and common problems
- Never expose the AudioSocket port to the internet. The protocol has no authentication or encryption. Bind your server to 127.0.0.1 or a private address and firewall it.
- Call drops immediately: nothing is listening on host:port, or a firewall blocks it. Check
/var/log/asterisk/fullaround the AudioSocket line. - Call drops during a long pause:
Reached timeout after 2000 ms of no activitymeans your server stopped sending and the caller was silent. Send silence frames. - “Failed to parse UUID”: the first argument must be a real UUID string; use
${UUID()}. - Silence on the call: your server is not sending audio back, or sends the wrong type byte or a length that does not match the payload.
- Chipmunk or slow audio: sample-rate mismatch. Send audio back at the same rate type you receive.
- No channel type registered for AudioSocket:
chan_audiosocket.sois not loaded; checkmodule show like audiosocket.
Official documentation: Asterisk: AudioSocket protocol · Asterisk: AudioSocket() application · Asterisk source: chan_audiosocket.c
Related: Install Asterisk 22 LTS on Debian 13: Step-by-Step Guide · Asterisk Recordings to MP3: Bulk Convert FreePBX Call Recordings · VoIP Codec Bandwidth: G.711 vs G.729 vs Opus, Per-Call Figures · VoIP Call Quality: Jitter, Packet Loss and MOS Explained With 3 Targets
See also: Asterisk WebRTC Setup: WSS Transport, Certificates, webrtc=yes · Asterisk Versions and EOL Dates: LTS Support Table and Upgrade Plan
Frequently asked questions
Which Asterisk versions support AudioSocket?
The AudioSocket application has been in Asterisk since 18.0.0. Asterisk 20 and 22 LTS include it, and FreePBX 17 with Asterisk 22 loads the modules by default.
What audio format does AudioSocket send?
Signed linear 16-bit mono PCM, little-endian, in 20 ms frames. The application sends 8 kHz; the channel driver can send other rates with the c() option, for example c(slin16) for 16 kHz.
Is AudioSocket encrypted?
No. It is plain TCP without authentication, so keep it on localhost, a private network or a VPN.
How does the bot end its part of the call?
Send a message with type 0x00 and length 0, or close the socket. The AudioSocket() application returns normally and the dialplan continues, so you can transfer to a queue or hang up there.
Should I use AudioSocket or ARI external media?
AudioSocket is simpler: one TCP stream per call and no RTP handling. ARI gives more call control. Many bots use AudioSocket for audio and AMI or ARI for transfers.