Converting between UTF-8 and the legacy character encodings, as the WHATWG Encoding Standard defines them.
  • Zig 99.5%
  • Python 0.4%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Jeffrey C. Ollie e074fe74f1
All checks were successful
test / test (push) Successful in 14m30s
test / docs (push) Successful in 6m28s
Spell Wansung syllables without a code in eight bytes; version 0.3.1
A syllable outside KS X 1001's 2,350 is now written as KS X 1001 spells
it: the Hangul filler A4 D4, then the initial, vowel and final from row
4, with the filler again for no final. Any whole eight-byte sequence
decodes, including one for a syllable that has a code of its own. A4 D4
not followed by a syllable is the filler, U+3164, and what came after
it is decoded afresh rather than knocked out of step.

Against CPython's euc_kr the encoders now agree on every scalar value,
and a new `syllables` crosscheck mode finds the two decoders agree on
all 830,584 eight-byte sequences about which are syllables and which
syllable each is. They still differ over a filler that begins nothing,
where CPython reports an error for the A4 and resumes mid-pair.

`Bytes` and the `ByteQueue` restore buffer grow from four to eight, so
the encoder returns the sequence whole and a failed one can hand back
its six bytes. The jamo tables and syllable arithmetic move to
`codecs/hangul.zig`, shared with Johab.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S2VYemLKKMU3LHzL5tyqUS
2026-10-03 23:18:29 -05:00
.forgejo/workflows Move to Zig 0.17.0 2026-10-03 01:20:10 -05:00
data Add EBCDIC, as a type of its own rather than an Encoding 2026-09-15 19:34:53 -05:00
LICENSES Add EBCDIC, as a type of its own rather than an Encoding 2026-09-15 19:34:53 -05:00
src Spell Wansung syllables without a code in eight bytes; version 0.3.1 2026-10-03 23:18:29 -05:00
tests Spell Wansung syllables without a code in eight bytes; version 0.3.1 2026-10-03 23:18:29 -05:00
tools Spell Wansung syllables without a code in eight bytes; version 0.3.1 2026-10-03 23:18:29 -05:00
.gitignore Remove the Nix package 2026-10-03 01:22:15 -05:00
build.zig Move to Zig 0.17.0 2026-10-03 01:20:10 -05:00
build.zig.zon Spell Wansung syllables without a code in eight bytes; version 0.3.1 2026-10-03 23:18:29 -05:00
flake.lock Move to Zig 0.17.0 2026-10-03 01:20:10 -05:00
flake.nix Remove the Nix package 2026-10-03 01:22:15 -05:00
README.md Spell Wansung syllables without a code in eight bytes; version 0.3.1 2026-10-03 23:18:29 -05:00
REUSE.toml Add EBCDIC, as a type of its own rather than an Encoding 2026-09-15 19:34:53 -05:00

zig-charset

Converting between UTF-8 and the legacy character encodings — windows-1252, Shift_JIS, Big5, gb18030 and the rest — as the WHATWG Encoding Standard defines them.

The API documentation is generated from the doc comments and published at https://jeff.jcollie.page/zig-charset/; zig build docs-serve reads the same thing locally.

const charset = @import("charset");

// A label off a Content-Type header, if it names an encoding at all.
const encoding = charset.Encoding.fromLabel("Shift-JIS") orelse return error.UnknownCharset;

const utf8 = try charset.decodeAlloc(gpa, encoding, bytes, .{});
defer gpa.free(utf8);

What is in it

The forty encodings the standard defines, and no others:

UTF-8 UTF-8
Single-byte IBM866, ISO-8859-2 through -16, ISO-8859-8-I, KOI8-R, KOI8-U, macintosh, windows-874, windows-1250 through -1258, x-mac-cyrillic
Chinese GBK, gb18030, Big5
Japanese EUC-JP, ISO-2022-JP, Shift_JIS
Korean EUC-KR
Other replacement, UTF-16BE, UTF-16LE, x-user-defined

Plus EBCDIC, which the standard does not define, as a separate type: cp037 (US/Canada), cp500 (International), cp875 (Greek) and cp1026 (Turkish). See EBCDIC below for why it is kept apart.

And two Korean encodings the standard leaves out, as another separate type: Wansung, which is KS X 1001 in EUC form without code page 949's additions, and Johab, KS X 1001 Annex 3 (code page 1361). See Korean outside the standard.

That the list is closed is the standard's doing and is deliberate. A label the server and the client resolve differently is an attack rather than a feature, which is also why the labels for ISO-2022-KR, ISO-2022-CN and HZ-GB-2312 all name replacement, an encoding whose entire output is one U+FFFD.

Several of these are not what their name says. iso-8859-1, latin1, ascii and us-ascii are all labels for windows-1252, because content labelled that way is windows-1252 often enough that decoding it as Latin-1 gets the wrong answer more often than the right one. EUC-KR is really Windows code page 949, Big5 includes the Hong Kong Supplementary Character Set, and Shift_JIS is code page 932. The standard documents each of these, and this library follows it rather than the registered definitions; see differences from other implementations.

Using it

zig fetch --save git+https://git.jcollie.dev/jeff/zig-charset.git
// build.zig
const charset = b.dependency("charset", .{});
exe.root_module.addImport("charset", charset.module("charset"));

It needs Zig 0.17.0 and has no dependencies of its own. The last release for Zig 0.16.0 is the v0.1.0 tag, and fixes for it go on the zig-0.16 branch.

One-shot

const utf8 = try charset.decodeAlloc(gpa, .shift_jis, bytes, .{});
const bytes = try charset.encodeAlloc(gpa, .windows_1252, "café", .{});

charset.decodeBomAlloc is the same but lets a leading byte order mark override the encoding you pass, which is what HTML and XML both want; the mark itself is not part of the result.

Streaming

Text off a socket does not arrive on character boundaries. Decoder and Encoder keep the state a character split across two chunks needs — a lead byte, a half-read escape sequence, a leading surrogate — so that it decodes as itself rather than as two replacement characters.

var decoder: charset.Decoder = .init(.euc_jp, .{});
while (try readChunk()) |chunk| try decoder.decode(chunk, writer);
try decoder.finish(writer);

finish is not optional. A sequence left half-read at the end of the text is an error, and nothing else can tell that apart from a sequence whose remaining bytes have not arrived yet. For the encoder it is what returns ISO-2022-JP to ASCII mode before the text ends.

When the bytes do not mean anything

Decoding does not fail by default: a byte sequence that means nothing becomes U+FFFD, which is what leaves a document with one bad byte in it still a document. Pass .{ .error_mode = .fatal } when the input is a checked format rather than prose, and you get error.InvalidByteSequence instead.

Encoding fails by default, with error.UnmappableCodePoint and the offending scalar value in Encoder.unmappable — most of these encodings can write only a few thousand characters, and losing the rest silently is worse than stopping. .{ .error_mode = .html } is the other choice the standard defines: it writes &#20013; and carries on, which HTML form submission requires and which cannot be told apart from those characters typed literally. Use UTF-8 and the question does not arise.

Three encodings have no encoder at all — replacement, UTF-16BE and UTF-16LE — so Encoder.init returns error.NoEncoder for them. Encoding.outputEncoding maps those three to UTF-8 and is what URL parsing and form submission are meant to use; Encoder.initOutput applies it for you and cannot fail.

EBCDIC

EBCDIC is not in the Encoding Standard, and it is not a member of Encoding. It has a type of its own:

const utf8 = try charset.decodeEbcdicAlloc(gpa, .cp037, bytes, .{});
const bytes = try charset.encodeEbcdicAlloc(gpa, .cp037, "HELLO", .{});

// Or streaming, through the same Decoder and Encoder as everything else.
var decoder: charset.Decoder = .initEbcdic(.cp500, .{});

Keeping it out of Encoding is the point rather than an inconvenience. Encoding.fromLabel refuses everything outside the standard, and that is what stops a Content-Type header choosing an encoding nobody vetted; there is no fromLabel for EBCDIC at all, so it can only be reached by naming a code page in code. Everything below the type is shared — the same decoder and encoder, the same error modes, the same streaming behaviour, the same table generator.

Two things about EBCDIC catch people out. It shares nothing with ASCII, not even the letters: A–I, J–R and S–Z sit in three runs with gaps between them, a hole punched by the card codes it grew out of, so code that arranges bytes by arithmetic gets it wrong. And every one of the 256 positions is assigned in these four pages, so decoding cannot fail — Microsoft's tables spell an unassigned position as U+001A SUBSTITUTE rather than leaving it out.

Adding a code page is a mapping file in data/ and a line in the generator, provided the mapping comes from somewhere citable.

Korean outside the standard

The standard's EUC-KR is Windows code page 949: KS X 1001 plus the Unified Hangul Code, which fits the 8,822 syllables KS X 1001 left out into byte ranges EUC leaves free. That is what the web sends. The two older ways of writing Korean have a type of their own, charset.Korean, kept out of Encoding for the same reason EBCDIC is:

const utf8 = try charset.decodeKoreanAlloc(gpa, .johab, bytes, .{});
const bytes = try charset.encodeKoreanAlloc(gpa, .wansung, "한글", .{});

var decoder: charset.Decoder = .initKorean(.johab, .{});
  • Wansung (완성형, "precomposed") is KS X 1001 in EUC form and nothing else: both bytes A1 to FE. It has codes for 2,350 Hangul syllables. A syllable outside them, such as 똠, is written as eight bytes, the way KS X 1001 defines: the Hangul filler A4 D4, then the initial consonant, vowel and final consonant from row 4, with the filler again for no final. So 똠 is A4 D4 A4 A8 A4 C7 A4 B1. A code page 949 pair decodes as an error.
  • Johab (조합형, "combining") spells Hangul out rather than looking it up: a one bit, then five bits each for the initial consonant, the vowel and the final consonant. So all 11,172 syllables have a code. KS X 1001's symbols and Hanja are moved into the remaining byte ranges by arithmetic.

Neither needs a table of its own. KS X 1001 is the part of the standard's euc-kr index whose bytes are both A1 to FE, including the euro and registered signs the 1998 revision added, and both encodings read it from there.

Johab's jamo deserve a note. Unicode's Compatibility Jamo has one code point per letter, but Johab has a code per position: ㄱ can be written as an initial or as a final. This follows Unicode.org's JOHAB.TXT, which gives each letter one code. A consonant that can begin a syllable is an initial, and the eleven clusters that cannot (ㄳ, ㄵ, ㄶ, ㄺ to ㅀ, ㅄ) are finals. The other spellings decode as errors, so whatever decodes also encodes back to the same bytes.

Differences from other implementations

These are the standard's departures from the registered definitions, not this library's. Each is there because deployed content needed it.

  • windows-125x and windows-874 fill the positions Microsoft never assigned with the C1 control of the same number, rather than rejecting them. There are a handful in each of the windows-125x pages and two dozen in windows-874.
  • KOI8-U takes ten positions from KOI8-R where RFC 2319 takes eight. The eight are the Ukrainian letters є і ї ґ and their capitals; the extra two are AE and BE, which the standard reads as ў and Ў — KOI8-RU's Belarusian short U — and RFC 2319 leaves as the box-drawing characters ╝ and ╬. CPython's koi8_u follows the RFC, so those two positions are the whole of the difference.
  • Shift_JIS decodes pointers 8836 to 10715 into the private use area (the Windows end-user defined characters) and refuses to encode them again. Its encoder skips pointers 8272 to 8835, so U+2170 comes out as FA 40 rather than the EE EF that some implementations produce.
  • Big5 decodes the Hong Kong extensions but will not write them, and takes the last pointer for six code points where everything else takes the first. Four pointers decode to two code points each, a letter plus a combining mark.
  • GBK writes the euro sign as 80 and will not use the four-byte form; gb18030 does the opposite. gb18030's encoder keeps eighteen private use characters where GB18030-2005 had them, so those do not round trip, and U+E5E5 cannot be encoded at all.
  • Shift_JIS, EUC-JP and ISO-2022-JP write U+00A5 as 5C and U+203E as 7E, which is why a path written on a Japanese system comes back with a yen sign where the backslash was.
  • EUC-JP reads JIS X 0212 and will not write it.

The rest are not the standard's doing, because these encodings are not in it.

  • Johab follows JOHAB.TXT except at 5C, which is the backslash rather than the won sign (the file's own notes say the backslash "might be a better idea", and Windows and CPython both use it), and at D9 E6 and D9 E7, the euro and registered signs, which the file predates. Against CPython the encoders agree on every scalar value. CPython's decoder also accepts the sixteen spellings of a jamo as a final that JOHAB.TXT leaves out (84 42 for ㄱ, and so on), and reads 84 41, all three fields filled, as U+3000. Both are errors here.
  • Wansung agrees with CPython's euc_kr on every scalar value, and on which eight-byte sequences are syllables and which syllable each one is. The two part company over A4 D4 that does not begin a syllable. Here it is U+3164 HANGUL FILLER, which is what KS X 1001 has at that position, and whatever followed it is decoded afresh. CPython reports an error for the A4 alone and carries on from the D4, which puts every pair after it out of step: A4 D4 A4 A3 A4 BF A4 D4, a sequence that fails because ㄳ cannot begin a syllable, is filler, ㄳ, ㅏ, filler here and <EFBFBD>渡$엘<EFBFBD> there.
  • CP875 encodes U+001A to 3F where CPython gives FD. Microsoft's table spells six unassigned positions as SUBSTITUTE alongside the real one, and 3F is the position every other EBCDIC page keeps SUBSTITUTE in — including, in CPython, cp037, cp500 and cp1026.

The html encoder error mode also departs from the letter of the standard for EBCDIC alone. The standard says to emit the bytes 0x26 0x23 … 0x3B, which works because every encoding it defines spells ASCII as ASCII; EBCDIC does not, so what is written is those characters through the encoder. For the standard's own encodings the two are the same bytes.

tools/crosscheck.py walks every input of every encoding against CPython's codecs module and prints what differs; the list above is what it finds, and anything else appearing there would be a bug.

It also prints what it skipped, which is utf-16le, utf-16be and replacement asked to encode. The standard makes those three decode-only — Encoding.canEncode is this library's answer about which — so there is nothing to compare, and the run says so rather than failing.

The tables

src/tables/ is generated from the index files the standard publishes — and, for EBCDIC, from Microsoft's mapping tables as Unicode.org publishes them — which are committed in data/ so that a build needs no network and so that a change upstream shows up in review as a diff to its input rather than as an unexplained change to a mapping.

nix develop -c zig build gen-tables    # then read `git diff`

The workflow regenerates and checks for a dirty tree, so an index that was updated without the tables being remade fails the build.

Testing

nix develop -c zig build test            # unit and conformance tests
nix develop -c python3 tools/crosscheck.py   # against CPython's codecs
nix develop -c zig build fuzz-run        # our own fuzzing loop
nix develop -c zig build fuzz --fuzz     # Zig's fuzzer, until interrupted

The conformance tests are spot checks taken from the standard's worked examples and from arithmetic done by hand against the algorithms — a test whose expected value was read out of the same table the code reads proves nothing.

The fuzz targets check three properties: decoding cannot crash and cannot produce invalid UTF-8; where a chunk boundary falls does not change the answer; and text that has already been through an encoding once comes out of it unchanged the second time. zig build fuzz --fuzz hands them to Zig's fuzzer, with coverage feedback, until it is interrupted; zig build fuzz-run is a loop of our own over a corpus that stops after a fixed number of inputs, which is what CI runs. See tools/fuzz.zig.

Where this lives

git clone https://git.jcollie.dev/jeff/zig-charset.git
rad clone rad:z2SapCVyHHcNssHjcrBvpA2XXuQnW

A Radicle repository is findable only by its ID, so that string is the one thing a reader needs in order to seed or clone it.

References cited

The first is the whole specification this implements: the encodings and their labels, the decoder and encoder algorithm for each, and the index files that data/ is a copy of. The second is where the four EBCDIC tables come from.

The third was consulted and deliberately not used. It does not say which code page it is, it states that the translation is not bidirectional, and it collapses everything unmappable to 0x1A, so nothing built on it could round trip.

JOHAB.TXT is what Johab is checked against; see differences from other implementations for the three positions where this library departs from it. Windows' own table for code page 1361 was consulted and not followed. It decodes a standalone jamo to the conjoining Hangul Jamo block rather than the Compatibility Jamo that KS X 1001 itself uses, and a few of its entries disagree with the arithmetic Annex 3 defines. RFC 1557 is the EUC-KR that Wansung is: KS X 1001 in EUC form, with nothing added.

All of them are in the zig-charset Zotero collection.

Licensing

MIT, and the project follows REUSE — nix develop -c reuse lint checks it.

The index files under data/ are the WHATWG's, under the CC-BY 4.0 that covers the Encoding Standard. The tables generated from them are MIT AND BSD-3-Clause, the standard's own terms for portions of it incorporated into source code. The EBCDIC mapping files and the table generated from them carry the Unicode licence.