
An audio codec is the format and method used to encode voice audio into digital data, and to decode it back into sound. In voice AI, the codec and sample rate a system uses directly affect audio quality, bandwidth usage, and how accurately downstream speech recognition and voice generation perform. The two codecs that come up constantly in telephony and voice AI are PCM and mu-law, along with the sample rate each is typically paired with.
PCM is an uncompressed way of representing audio: it takes a sound wave and measures its amplitude at regular intervals, storing each measurement as a number. Linear PCM at 16-bit depth is the standard used by most speech recognition and text-to-speech models, since it preserves the full resolution of the audio signal without the quality loss that comes from compression. The trade-off is size: uncompressed PCM audio takes up meaningfully more bandwidth and storage than a compressed format carrying the same duration of speech.
Mu-law and A-law are companding algorithms defined under the ITU-T G.711 standard, and they’re the default audio codecs used across most of the world’s telephone networks, mu-law in North America and Japan, A-law almost everywhere else. Instead of storing full-resolution amplitude values like linear PCM, companding compresses the dynamic range non-linearly, allocating more precision to quieter sounds and less to louder ones, which mirrors how human hearing perceives loudness. This lets G.711 represent voice-quality audio in 8 bits per sample instead of 16, roughly halving the data size compared to linear PCM at the same sample rate, at the cost of some audio fidelity.
Sample rate is how many times per second the audio is measured, expressed in Hertz (Hz). Traditional telephony runs at 8kHz, sometimes called narrowband, which is sufficient to capture the frequency range of the human voice needed for intelligibility (roughly 300Hz to 3,400Hz) but noticeably strips out the fuller tone a caller hears in person. Voice AI systems, particularly speech recognition and text-to-speech models, often perform meaningfully better at 16kHz or higher, sometimes called wideband, since that extra frequency range carries information that improves transcription accuracy and makes synthesized speech sound more natural.
A live phone call typically arrives at a voice AI system as 8kHz mu-law or A-law audio, matching standard telephony infrastructure, while many speech recognition and generation models are trained on 16kHz or higher linear PCM audio. Bridging that gap requires transcoding, converting from one codec and sample rate to another, and resampling, converting from one sample rate to another. Done well, this conversion is nearly transparent to call quality; done poorly, it introduces audible artifacts or subtly degrades transcription accuracy, which is one of the less visible reasons two voice AI systems can sound noticeably different even when using similar underlying models.
What’s the difference between mu-law and A-law?
They’re both G.711 companding algorithms with the same underlying idea but different formulas and regional adoption: mu-law is standard in North America and Japan, while A-law is used across most of Europe and the rest of the world. A call crossing between regions using different standards typically needs to be transcoded.
Why don’t voice AI systems just always use the highest possible sample rate?
Higher sample rates mean more data per second of audio, which increases bandwidth use and processing load. Since standard telephony audio arrives at 8kHz regardless, there’s often limited benefit to processing at a much higher rate unless the original audio genuinely carries that additional detail.
Does codec choice affect voice AI latency?
Indirectly, yes. Heavier transcoding and resampling steps add small amounts of processing time, and using a codec that closely matches what the AI pipeline expects natively reduces the number of conversion steps needed before the audio can be processed.

Transform customer engagement with our AI Voice Assistant. More than a bot, it’s your conversational partner, fluent in Hindi, English, and Hinglish. Available 24/7, it learns continuously for meaningful, personalised interactions.

Utilizing Gen AI and Natural Language Processing (NLP) capabilities, the House of AI transforms customer conversations into engaging, human-like experiences. It goes deep into understanding context, sentiment, and intent, enabling dynamic, personalized responses that boost engagement and loyalty.

At Exotel, we understand the frustration of support engineers, service managers, IT personnel, sales representatives and customers when placed on hold. ExoInsights provides users with just the right and relevant answer, tailored to their specific queries. It simplifies access to accurate information, making the decision-making process more efficient and user-friendly.

Automate cross-channel conversation quality analysis against your SOPs and KPIs to maintain top-tier service quality and agent efficiency, effortlessly.