- Zig 99.5%
- C 0.4%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0196Afvr1xSduWMdFeb1AVd3 |
||
| .forgejo/workflows | ||
| conformance | ||
| LICENSES | ||
| src | ||
| tests | ||
| tools | ||
| .gitignore | ||
| build.zig | ||
| build.zig.zon | ||
| flake.lock | ||
| flake.nix | ||
| README.md | ||
| REUSE.toml | ||
zig-opus
An Opus decoder and encoder, written in Zig against RFC 6716 as RFC 8251 updates it. It is a library: it takes packets as a transport delivers them and gives back PCM, and takes PCM and gives back packets, and has no files, no threads and no container of its own.
The API documentation is generated from the doc comments, where most of the explanation lives, and is published at https://jeff.jcollie.page/zig-opus/.
Where it is
Opus is two codecs behind one range coder. SILK is a linear-prediction codec for speech. CELT is a transform codec for music and for anything that needs low delay. A hybrid mode runs both at once. The first byte of a packet says which of them it uses.
It decodes every Opus packet, sample for sample as libopus's decoder does:
- The range decoder of Section 4.1, which every symbol goes through, and a range encoder to test it with.
- Packets: the table of contents and the four ways frames are packed into a packet, with every requirement of Section 3.4 checked.
- SILK (Section 4.2), the linear prediction codec for speech: mono and stereo, at 8, 12 and 16 kHz, resampled to the rate asked for, with its own concealment of lost frames, comfort noise, and the redundant copies of the last frame that forward error correction uses. SILK decodes in fixed point, even in libopus's floating point build, and so does this, with the same arithmetic.
- CELT (Section 4.3), the transform codec for music, with packet loss concealment: the last pitch period repeated through a linear predictor, fading as the signal was, and after 100 ms, or in hybrid mode, noise at the bands' last energies.
- The Opus layer:
Decodertakes a packet as a transport delivers it and gives back float samples, interleaved, at 48, 24, 16, 12 or 8 kHz. SILK, hybrid and CELT packets, the transitions between them, cross-faded through the redundant CELT frames the encoder adds or through concealment, lost packets, anddecodeFecfor a lost packet's copy in the next one.
zig build reference checks each of these against libopus: the tables,
every fixed-point helper, the filters and resampler, whole SILK and CELT
frames from random bytes, and streams libopus's encoder made in every
mode, switching between them, with packets lost and recovered, down to
the range decoder's final state.
It conforms: the twelve RFC 8251 test vectors, decoded in mono and in
stereo at every output rate as libopus's opus_demo decodes them, each
packet ending in the encoder's range coder state, and each output within
opus_compare's measure of the reference decoder's. conformance/ is
that test, a package of its own so that the vectors, 75 MB of them, are
fetched only by whoever runs it:
cd conformance
nix develop .. -c zig build run -- 48000 24000 16000 12000 8000
It encodes too, byte for byte as libopus's floating point encoder does from
the same input and settings: Encoder takes 16-bit or float PCM,
interleaved, at 48, 24, 16, 12 or 8 kHz, and gives back packets.
- CELT: pre-emphasis and the pitch pre-filter, transient detection, the forward MDCT, energy quantization, the analyses that pick time-frequency resolution, spreading, dynamic allocation, the allocation trim and stereo coding, the variable bitrate, and the band coder, which it shares with the decoder.
- SILK: voice activity detection, the pitch estimator, the noise shaping and prediction analyses in floating point, the parameter quantizers and noise shaping quantizers in fixed point, the loop that keeps a frame within its bits, stereo mid/side coding, in-band redundancy, DTX, and the resamplers from the API rate.
- The Opus layer: the tonality analysis and its small recurrent network, which tell speech from music and hear the input's bandwidth; the choice of mode, bandwidth and channels from those and the bitrate; redundant CELT frames where the mode changes; frames of up to 120 ms put together from shorter ones; VBR, constrained VBR and CBR with its padding; forward error correction and DTX.
zig build reference checks every layer against libopus's own
functions, and whole packets from both encoders byte for byte, through
rising and falling bitrates, every frame size, speech, music and silence.
It is the encoder in libopus's floating point build, so the same float
arithmetic in the same order; where libopus calls the C library's exp,
log or sqrt, so does this, and the two agree as long as the C
library's and Zig's do.
libopus computes some kernels with SSE, SSE2, SSE4.1 or AVX2: inner
products, cross-correlations, the pitch post-filter and the encoder's
pulse search, and its networks' layers and activations. Each adds in its
own order, so libopus's output depends on the CPU it runs on. Built for
x86-64, libopus presumes SSE and SSE2 and picks SSE4.1 or AVX2 at run
time, so it has three levels there, and Arch names them: .sse2,
.sse4_1 and .avx2, with .generic for libopus built for any other
CPU. Encoder, Decoder and the networks detect the level as libopus
does, and take another with setArch. Each level is computed as libopus
computes it, with Zig's vectors and fused multiply-adds where libopus
fuses them, so the results are the same on any x86-64 CPU, with SIMD
where it has it. The exception is the approximate reciprocal
instructions that the pulse search and the networks' activations use:
their results differ between Intel's CPUs and AMD's, so they come from
the CPU itself, and the x86 levels are for x86-64 only.
zig build reference builds libopus as its configure does for x86-64,
with the run-time choice made by the tests. It
compares every level kernel by kernel, network layer by layer, and in
whole streams: encoding, decoding, deep concealment and DRED.
Beyond one or two channels: MultistreamEncoder and MultistreamDecoder
code up to 255 channels as several streams in one packet, laid out as
RFC 7845's channel mapping families say (mono and stereo, the Vorbis
orders up to 7.1, ambisonics, or no layout at all), and
ProjectionEncoder and ProjectionDecoder do RFC 8486's family 3,
ambisonics mixed through a matrix into streams and demixed after. For
surround, the encoder works out how far each channel's bands sit below
the masking of the rest of the mix, and each stream spends fewer bits
where it is masked; the LFE is coded narrowband and stereo pairs in CELT.
These take an allocator, for their streams. zig build reference
compares all of them with libopus, byte for byte and sample for sample.
It also has the neural parts libopus 1.5 added, which are optional there and here:
- DRED, Deep Audio Redundancy: up to about a second of the audio
before each packet, coded at a very low bitrate by the encoder half of a
rate-distortion-optimized variational autoencoder and carried in the
packet's padding as an Opus extension. After a burst of losses, the next
packet that arrives holds enough to rebuild what was lost.
Encoderadds it once given adnn.dred.Encoderand a duration (setDred,setDredDuration), setting part of the bitrate aside for it as the expected packet loss says.dnn.dred.Decoderfinds and decodes it in a packet, andDecoder.decodeDredturns it back into audio. A decoder that does not know DRED skips the extension. - Deep packet loss concealment: LPCNet's features of the audio
before a loss, a network that predicts the features that come next (or
DRED's, when there are some), and the FARGAN vocoder to speak them.
Decoder.setDeepPlcgives a decoder adnn.Plc; from complexity 5, set withsetComplexity, it conceals every loss, and below that only where DRED has recovered something. - OSCE, Opus Speech Coding Enhancement: a network that sharpens what
SILK decodes at 16 kHz in 20 ms frames, working from features of the
frame SILK already has (its predictors, pitch, gains and bitrate) to
steer adaptive comb, convolution and shaping filters. LACE is the
smaller, NoLACE the larger.
Decoder.setOscegives a decoder adnn.osce.Osce; it uses LACE from complexity 6 and NoLACE from 7.
They need the networks' weights, about 4.9 MB, which libopus keeps as C
source. -Ddred builds them in from libopus's release, where they are
then dnn.weights.blob; without it the library has the code and no
weights. Other weights in libopus's blob format can be parsed with
dnn.nnet.Weights.parse.
All three are checked against libopus built with ENABLE_DEEP_PLC,
ENABLE_DRED and ENABLE_OSCE. That changes libopus's classic
concealment as well: CELT conceals with pitch for longer, and SILK does
not fade in after a loss at 16 kHz. A decoder given a dnn.Plc or a
dnn.osce.Osce behaves as that build does, and one without behaves as
the ordinary build. The comparison covers libopus's portable C and each
of its x86-64 levels: every layer, LPCNet's features, the concealment,
DRED's latents and features, OSCE's features and adaptive filters, and
whole packets and decoded audio, sample for sample, through concealment,
forward error correction, DRED and enhancement in every mode.
The first user is Pipit, which is sent Opus by Music Assistant over Sendspin as raw packets, 20 ms of 48 kHz stereo each.
Building it
It builds with Zig 0.17.0, which the Nix devshell provides. The last release
for Zig 0.16 is v0.1.0, and the zig-0.16 branch is where any fix for it
would go.
git clone https://git.jcollie.dev/jeff/zig-opus.git
cd zig-opus
nix develop -c zig build test --summary all
nix develop -c zig build reference --summary all # compared with libopus
nix develop -c zig build test -Ddred # with the networks' weights built in
nix develop -c zig build docs-serve # the API documentation, at http://localhost:8000
zig build reference builds libopus from its release, a lazy dependency
fetched only for this, and compares each layer here with the libopus
function it corresponds to: the same input to both, and the same answer
expected back. A decoder is only right if it makes the arithmetic decisions
the encoder did, and libopus is the encoder.
Fuzzing
A packet comes off a network, so whatever a peer sends must not crash the
receiver. tests/fuzz.zig holds properties, not examples, for the packet
parser and the repacketizer, the extension parser, the decoder over whole
streams with losses and FEC, the reader of DRED's latents, the encoder at
any settings and any samples (NaNs included), and the range coder:
whatever arrives, the call returns, stays inside its buffers, and what it
gives back holds together.
Their seeds are real: packets in every mode, bandwidth, frame size and
channel count, and streams of them, encoded at build time by
tools/fuzz_corpus.zig with this library's own encoder. zig build test
runs the properties on those seeds. zig build fuzz --fuzz hands them to
Zig's own coverage-guided fuzzer, and zig build fuzz-run mutates them in
the loop in tools/fuzz.zig, which runs every target for a fixed time from
a seed that repeats a run exactly:
nix develop -c zig build fuzz --fuzz # until interrupted
nix develop -c zig build fuzz --fuzz=1M -Dfuzz-filter=decoding # a bounded run of the decoders
nix develop -c zig build fuzz-run # a minute of each
nix develop -c zig build fuzz-run -- --seconds 600 --target decoding
nix develop -c zig build fuzz-run -- --input fuzz-findings/x.bin --target packets
A failing input, a panic included, is written to fuzz-findings/ with the
command that runs it again.
Where this lives
-
Forgejo, at https://git.jcollie.dev/jeff/zig-opus.
-
Tangled, at https://tangled.org/jcollie.dev/zig-opus.
-
Radicle, as
rad:z3JXrVBxL8K8HiRNhxksdtPBbQDV8. A Radicle repository is only findable by its ID, so that string is the whole address:rad clone rad:z3JXrVBxL8K8HiRNhxksdtPBbQDV8
License
MIT, following the REUSE specification; reuse lint checks it.
The codec is a port of libopus, the reference implementation, whose files
are under the two-clause BSD license (CELT and the Opus layer) and the
three-clause one (SILK). A file here that follows
one of them keeps that file's copyright lines and license beside this
project's: MIT AND BSD-2-Clause, for instance, with the libopus authors
named in its SPDX-FileCopyrightText.
References cited
Kept in the zig-opus Zotero collection.
- Valin, JM., K. Vos, and T. Terriberry. Definition of the Opus Audio Codec. RFC 6716. Internet Engineering Task Force (IETF), September 2012. https://www.rfc-editor.org/info/rfc6716. The decoder is the normative part; this follows its Section 3 for packets and Section 4.1 for the range decoder, and the reference implementation it carries for the arithmetic those have to reproduce exactly.
- Valin, JM., and K. Vos. Updates to the Opus Audio Codec. RFC 8251. Internet Engineering Task Force (IETF), October 2017. https://www.rfc-editor.org/info/rfc8251. Corrections to the decoder, and the test vectors conformance is measured against.
- Terriberry, T., R. Lee, and R. Giles. Ogg Encapsulation for the Opus Audio Codec. RFC 7845. Internet Engineering Task Force (IETF), April 2016. https://www.rfc-editor.org/info/rfc7845. Pre-skip, the 48 kHz timebase every duration in Opus is counted in, and the channel mapping families the multistream encoder and decoder follow.
- Skoglund, J., and M. Graczyk. Ambisonics in an Ogg Opus Container. RFC 8486. Internet Engineering Task Force (IETF), October 2018. https://www.rfc-editor.org/info/rfc8486. Channel mapping families 2 and 3: ambisonics, and its projection through a demixing matrix.
- Valin, JM., and J. Buethe. Deep Audio Redundancy (DRED) Extension for the Opus Codec. Internet-Draft draft-ietf-mlcodec-opus-dred-07. Internet Engineering Task Force (IETF), August 2026. Work in progress. https://datatracker.ietf.org/doc/draft-ietf-mlcodec-opus-dred/. What DRED carries and how it is decoded.
- Terriberry, T., and JM. Valin. Extension Formatting for the Opus Codec. Internet-Draft draft-ietf-mlcodec-opus-extension-06. Internet Engineering Task Force (IETF), July 2026. Work in progress. https://datatracker.ietf.org/doc/draft-ietf-mlcodec-opus-extension/. How extensions such as DRED sit in a packet's padding.
- Büthe, J., JM. Valin, and A. Mustafa. "LACE: A Light-Weight, Causal Model for Enhancing Coded Speech Through Adaptive Convolutions." In 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 1–5. IEEE, October 2023. https://doi.org/10.1109/WASPAA58266.2023.10248150. LACE: the feature network and the adaptive comb and convolution filters it steers.
- Büthe, J., A. Mustafa, JM. Valin, K. Helwani, and M. M. Goodwin. "NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping." In ICASSP 2024 – 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 476–480. IEEE, April 2024. https://doi.org/10.1109/ICASSP48485.2024.10448332. NoLACE: LACE with the adaptive temporal shaping added.
- Buethe, J., and JM. Valin. Integration of Speech Codec Enhancement Algorithms into the Opus Codec. Internet-Draft draft-ietf-mlcodec-opus-speech-coding-enhancement-04. Internet Engineering Task Force (IETF), July 2026. Work in progress. https://datatracker.ietf.org/doc/draft-ietf-mlcodec-opus-speech-coding-enhancement/. What an enhancement such as LACE or NoLACE must do to sit in an Opus decoder.
- Xiph.Org Foundation and contributors. libopus, version 1.5.2.
https://opus-codec.org/. The reference implementation: what this is a
port of, file by file, and what
zig build referencecompares it with.