ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

Project Webpage
[ code]

Abstract

Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms. Our approach combines large-scale WavLM distillation with a redesigned transformer-based architecture, a scalar spherical quantizer, and a latency-aware streaming decoder. Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate. Despite its 842M parameters, ZipCodec achieves real-time single-stream inference on a consumer-grade CPU.

Please note:

Speech Resynthesis

Sample A

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample B

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample C

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample D

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample E

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample F

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample G

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample H

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample I

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Sample J

Reference      
EnCodec (1.50 kbps)    
AudioDec (1.60 kbps)    
HILCodec (1.50 kbps)    
Mimi (0.83 kbps)    
PAST (1.00 kbps)    
FocalCodec-Stream (0.80 kbps)    
ZipCodec (0.80 kbps)    
FocalCodec (0.65 kbps)    



Voice Conversion

Input      
Reference      
FocalCodec-Stream (0.80 kbps)      
ZipCodec (0.80 kbps)      
FocalCodec (0.65 kbps)