- Zig 99.5%
- Python 0.4%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
A syllable outside KS X 1001's 2,350 is now written as KS X 1001 spells it: the Hangul filler A4 D4, then the initial, vowel and final from row 4, with the filler again for no final. Any whole eight-byte sequence decodes, including one for a syllable that has a code of its own. A4 D4 not followed by a syllable is the filler, U+3164, and what came after it is decoded afresh rather than knocked out of step. Against CPython's euc_kr the encoders now agree on every scalar value, and a new `syllables` crosscheck mode finds the two decoders agree on all 830,584 eight-byte sequences about which are syllables and which syllable each is. They still differ over a filler that begins nothing, where CPython reports an error for the A4 and resumes mid-pair. `Bytes` and the `ByteQueue` restore buffer grow from four to eight, so the encoder returns the sequence whole and a failed one can hand back its six bytes. The jamo tables and syllable arithmetic move to `codecs/hangul.zig`, shared with Johab. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S2VYemLKKMU3LHzL5tyqUS |
||
| .forgejo/workflows | ||
| data | ||
| LICENSES | ||
| src | ||
| tests | ||
| tools | ||
| .gitignore | ||
| build.zig | ||
| build.zig.zon | ||
| flake.lock | ||
| flake.nix | ||
| README.md | ||
| REUSE.toml | ||
zig-charset
Converting between UTF-8 and the legacy character encodings — windows-1252, Shift_JIS, Big5, gb18030 and the rest — as the WHATWG Encoding Standard defines them.
The API documentation is generated from the doc comments and published at
https://jeff.jcollie.page/zig-charset/; zig build docs-serve reads the
same thing locally.
const charset = @import("charset");
// A label off a Content-Type header, if it names an encoding at all.
const encoding = charset.Encoding.fromLabel("Shift-JIS") orelse return error.UnknownCharset;
const utf8 = try charset.decodeAlloc(gpa, encoding, bytes, .{});
defer gpa.free(utf8);
What is in it
The forty encodings the standard defines, and no others:
| UTF-8 | UTF-8 |
| Single-byte | IBM866, ISO-8859-2 through -16, ISO-8859-8-I, KOI8-R, KOI8-U, macintosh, windows-874, windows-1250 through -1258, x-mac-cyrillic |
| Chinese | GBK, gb18030, Big5 |
| Japanese | EUC-JP, ISO-2022-JP, Shift_JIS |
| Korean | EUC-KR |
| Other | replacement, UTF-16BE, UTF-16LE, x-user-defined |
Plus EBCDIC, which the standard does not define, as a separate type:
cp037 (US/Canada), cp500 (International), cp875 (Greek) and cp1026
(Turkish). See EBCDIC below for why it is kept apart.
And two Korean encodings the standard leaves out, as another separate type: Wansung, which is KS X 1001 in EUC form without code page 949's additions, and Johab, KS X 1001 Annex 3 (code page 1361). See Korean outside the standard.
That the list is closed is the standard's doing and is deliberate. A label the
server and the client resolve differently is an attack rather than a feature,
which is also why the labels for ISO-2022-KR, ISO-2022-CN and HZ-GB-2312 all
name replacement, an encoding whose entire output is one U+FFFD.
Several of these are not what their name says. iso-8859-1, latin1,
ascii and us-ascii are all labels for windows-1252, because content
labelled that way is windows-1252 often enough that decoding it as Latin-1
gets the wrong answer more often than the right one. EUC-KR is really
Windows code page 949, Big5 includes the Hong Kong Supplementary Character
Set, and Shift_JIS is code page 932. The standard documents each of these,
and this library follows it rather than the registered definitions; see
differences from other implementations.
Using it
zig fetch --save git+https://git.jcollie.dev/jeff/zig-charset.git
// build.zig
const charset = b.dependency("charset", .{});
exe.root_module.addImport("charset", charset.module("charset"));
It needs Zig 0.17.0 and has no dependencies of its own. The last release for
Zig 0.16.0 is the v0.1.0 tag, and fixes for it go on the zig-0.16 branch.
One-shot
const utf8 = try charset.decodeAlloc(gpa, .shift_jis, bytes, .{});
const bytes = try charset.encodeAlloc(gpa, .windows_1252, "café", .{});
charset.decodeBomAlloc is the same but lets a leading byte order mark
override the encoding you pass, which is what HTML and XML both want; the mark
itself is not part of the result.
Streaming
Text off a socket does not arrive on character boundaries. Decoder and
Encoder keep the state a character split across two chunks needs — a lead
byte, a half-read escape sequence, a leading surrogate — so that it decodes as
itself rather than as two replacement characters.
var decoder: charset.Decoder = .init(.euc_jp, .{});
while (try readChunk()) |chunk| try decoder.decode(chunk, writer);
try decoder.finish(writer);
finish is not optional. A sequence left half-read at the end of the text is
an error, and nothing else can tell that apart from a sequence whose remaining
bytes have not arrived yet. For the encoder it is what returns ISO-2022-JP to
ASCII mode before the text ends.
When the bytes do not mean anything
Decoding does not fail by default: a byte sequence that means nothing becomes
U+FFFD, which is what leaves a document with one bad byte in it still a
document. Pass .{ .error_mode = .fatal } when the input is a checked format
rather than prose, and you get error.InvalidByteSequence instead.
Encoding fails by default, with error.UnmappableCodePoint and the offending
scalar value in Encoder.unmappable — most of these encodings can write only
a few thousand characters, and losing the rest silently is worse than
stopping. .{ .error_mode = .html } is the other choice the standard defines:
it writes 中 and carries on, which HTML form submission requires and
which cannot be told apart from those characters typed literally. Use UTF-8
and the question does not arise.
Three encodings have no encoder at all — replacement, UTF-16BE and
UTF-16LE — so Encoder.init returns error.NoEncoder for them.
Encoding.outputEncoding maps those three to UTF-8 and is what URL parsing
and form submission are meant to use; Encoder.initOutput applies it for you
and cannot fail.
EBCDIC
EBCDIC is not in the Encoding Standard, and it is not a member of Encoding.
It has a type of its own:
const utf8 = try charset.decodeEbcdicAlloc(gpa, .cp037, bytes, .{});
const bytes = try charset.encodeEbcdicAlloc(gpa, .cp037, "HELLO", .{});
// Or streaming, through the same Decoder and Encoder as everything else.
var decoder: charset.Decoder = .initEbcdic(.cp500, .{});
Keeping it out of Encoding is the point rather than an inconvenience.
Encoding.fromLabel refuses everything outside the standard, and that is what
stops a Content-Type header choosing an encoding nobody vetted; there is no
fromLabel for EBCDIC at all, so it can only be reached by naming a code page
in code. Everything below the type is shared — the same decoder and encoder,
the same error modes, the same streaming behaviour, the same table generator.
Two things about EBCDIC catch people out. It shares nothing with ASCII, not
even the letters: A–I, J–R and S–Z sit in three runs with gaps
between them, a hole punched by the card codes it grew out of, so code that
arranges bytes by arithmetic gets it wrong. And every one of the 256 positions
is assigned in these four pages, so decoding cannot fail — Microsoft's tables
spell an unassigned position as U+001A SUBSTITUTE rather than leaving it out.
Adding a code page is a mapping file in data/ and a line in the generator,
provided the mapping comes from somewhere citable.
Korean outside the standard
The standard's EUC-KR is Windows code page 949: KS X 1001 plus the Unified
Hangul Code, which fits the 8,822 syllables KS X 1001 left out into byte
ranges EUC leaves free. That is what the web sends. The two older ways of
writing Korean have a type of their own, charset.Korean, kept out of
Encoding for the same reason EBCDIC is:
const utf8 = try charset.decodeKoreanAlloc(gpa, .johab, bytes, .{});
const bytes = try charset.encodeKoreanAlloc(gpa, .wansung, "한글", .{});
var decoder: charset.Decoder = .initKorean(.johab, .{});
- Wansung (완성형, "precomposed") is KS X 1001 in EUC form and nothing
else: both bytes
A1toFE. It has codes for 2,350 Hangul syllables. A syllable outside them, such as 똠, is written as eight bytes, the way KS X 1001 defines: the Hangul fillerA4 D4, then the initial consonant, vowel and final consonant from row 4, with the filler again for no final. So 똠 isA4 D4 A4 A8 A4 C7 A4 B1. A code page 949 pair decodes as an error. - Johab (조합형, "combining") spells Hangul out rather than looking it up: a one bit, then five bits each for the initial consonant, the vowel and the final consonant. So all 11,172 syllables have a code. KS X 1001's symbols and Hanja are moved into the remaining byte ranges by arithmetic.
Neither needs a table of its own. KS X 1001 is the part of the standard's
euc-kr index whose bytes are both A1 to FE, including the euro and
registered signs the 1998 revision added, and both encodings read it from
there.
Johab's jamo deserve a note. Unicode's Compatibility Jamo has one code point
per letter, but Johab has a code per position: ㄱ can be written as an
initial or as a final. This follows Unicode.org's JOHAB.TXT, which gives
each letter one code. A consonant that can begin a syllable is an initial,
and the eleven clusters that cannot (ㄳ, ㄵ, ㄶ, ㄺ to ㅀ, ㅄ) are finals. The
other spellings decode as errors, so whatever decodes also encodes back to
the same bytes.
Differences from other implementations
These are the standard's departures from the registered definitions, not this library's. Each is there because deployed content needed it.
- windows-125x and windows-874 fill the positions Microsoft never assigned with the C1 control of the same number, rather than rejecting them. There are a handful in each of the windows-125x pages and two dozen in windows-874.
- KOI8-U takes ten positions from KOI8-R where RFC 2319 takes
eight. The eight are the Ukrainian letters є і ї ґ and their capitals; the
extra two are
AEandBE, which the standard reads as ў and Ў — KOI8-RU's Belarusian short U — and RFC 2319 leaves as the box-drawing characters ╝ and ╬. CPython'skoi8_ufollows the RFC, so those two positions are the whole of the difference. - Shift_JIS decodes pointers 8836 to 10715 into the private use area (the
Windows end-user defined characters) and refuses to encode them again. Its
encoder skips pointers 8272 to 8835, so U+2170 comes out as
FA 40rather than theEE EFthat some implementations produce. - Big5 decodes the Hong Kong extensions but will not write them, and takes the last pointer for six code points where everything else takes the first. Four pointers decode to two code points each, a letter plus a combining mark.
- GBK writes the euro sign as
80and will not use the four-byte form; gb18030 does the opposite. gb18030's encoder keeps eighteen private use characters where GB18030-2005 had them, so those do not round trip, and U+E5E5 cannot be encoded at all. - Shift_JIS, EUC-JP and ISO-2022-JP write U+00A5 as
5Cand U+203E as7E, which is why a path written on a Japanese system comes back with a yen sign where the backslash was. - EUC-JP reads JIS X 0212 and will not write it.
The rest are not the standard's doing, because these encodings are not in it.
- Johab follows
JOHAB.TXTexcept at5C, which is the backslash rather than the won sign (the file's own notes say the backslash "might be a better idea", and Windows and CPython both use it), and atD9 E6andD9 E7, the euro and registered signs, which the file predates. Against CPython the encoders agree on every scalar value. CPython's decoder also accepts the sixteen spellings of a jamo as a final thatJOHAB.TXTleaves out (84 42for ㄱ, and so on), and reads84 41, all three fields filled, as U+3000. Both are errors here. - Wansung agrees with CPython's
euc_kron every scalar value, and on which eight-byte sequences are syllables and which syllable each one is. The two part company overA4 D4that does not begin a syllable. Here it is U+3164 HANGUL FILLER, which is what KS X 1001 has at that position, and whatever followed it is decoded afresh. CPython reports an error for theA4alone and carries on from theD4, which puts every pair after it out of step:A4 D4 A4 A3 A4 BF A4 D4, a sequence that fails because ㄳ cannot begin a syllable, is filler, ㄳ, ㅏ, filler here and<EFBFBD>渡$엘<EFBFBD>there. - CP875 encodes U+001A to
3Fwhere CPython givesFD. Microsoft's table spells six unassigned positions as SUBSTITUTE alongside the real one, and3Fis the position every other EBCDIC page keeps SUBSTITUTE in — including, in CPython,cp037,cp500andcp1026.
The html encoder error mode also departs from the letter of the standard for
EBCDIC alone. The standard says to emit the bytes 0x26 0x23 … 0x3B, which
works because every encoding it defines spells ASCII as ASCII; EBCDIC does
not, so what is written is those characters through the encoder. For the
standard's own encodings the two are the same bytes.
tools/crosscheck.py walks every input of every encoding against CPython's
codecs module and prints what differs; the list above is what it finds, and
anything else appearing there would be a bug.
It also prints what it skipped, which is utf-16le, utf-16be and
replacement asked to encode. The standard makes those three decode-only —
Encoding.canEncode is this library's answer about which — so there is
nothing to compare, and the run says so rather than failing.
The tables
src/tables/ is generated from the index files the standard publishes — and,
for EBCDIC, from Microsoft's mapping tables as Unicode.org publishes them —
which are committed in data/ so that a build needs no network and so that a
change upstream shows up in review as a diff to its input rather than as an
unexplained change to a mapping.
nix develop -c zig build gen-tables # then read `git diff`
The workflow regenerates and checks for a dirty tree, so an index that was updated without the tables being remade fails the build.
Testing
nix develop -c zig build test # unit and conformance tests
nix develop -c python3 tools/crosscheck.py # against CPython's codecs
nix develop -c zig build fuzz-run # our own fuzzing loop
nix develop -c zig build fuzz --fuzz # Zig's fuzzer, until interrupted
The conformance tests are spot checks taken from the standard's worked examples and from arithmetic done by hand against the algorithms — a test whose expected value was read out of the same table the code reads proves nothing.
The fuzz targets check three properties: decoding cannot crash and cannot
produce invalid UTF-8; where a chunk boundary falls does not change the
answer; and text that has already been through an encoding once comes out of
it unchanged the second time. zig build fuzz --fuzz hands them to Zig's
fuzzer, with coverage feedback, until it is interrupted; zig build fuzz-run
is a loop of our own over a corpus that stops after a fixed number of inputs,
which is what CI runs. See tools/fuzz.zig.
Where this lives
- Forgejo: https://git.jcollie.dev/jeff/zig-charset
- Tangled: https://tangled.org/jcollie.dev/zig-charset
- Radicle:
rad:z2SapCVyHHcNssHjcrBvpA2XXuQnW
git clone https://git.jcollie.dev/jeff/zig-charset.git
rad clone rad:z2SapCVyHHcNssHjcrBvpA2XXuQnW
A Radicle repository is findable only by its ID, so that string is the one thing a reader needs in order to seed or clone it.
References cited
- van Kesteren, A. (2026). Encoding Standard (Living Standard, 21 May 2026). WHATWG. https://encoding.spec.whatwg.org/
- Unicode Consortium. EBCDIC code page mapping tables
(
MAPPINGS/VENDORS/MICSFT/EBCDIC). https://www.unicode.org/Public/MAPPINGS/VENDORS/MICSFT/EBCDIC/ - IBM. EBCDIC to ASCII conversion table (InfoSphere Information Server 11.3.0). https://www.ibm.com/docs/en/iis/11.3.0?topic=tables-ebcdic-ascii
- Shin, J. (2011). Johab to Unicode table (
JOHAB.TXT, version 1.1). Unicode Consortium. https://www.unicode.org/Public/MAPPINGS/OBSOLETE/EASTASIA/KSC/JOHAB.TXT - Microsoft. Windows best fit table for code page 1361
(
bestfit1361.txt). Unicode Consortium. https://www.unicode.org/Public/MAPPINGS/VENDORS/MICSFT/WindowsBestFit/bestfit1361.txt - Choi, U., Chon, K., & Park, H. (1993). Korean Character Encoding for Internet Messages (RFC 1557). RFC Editor. https://www.rfc-editor.org/info/rfc1557
The first is the whole specification this implements: the encodings and their
labels, the decoder and encoder algorithm for each, and the index files that
data/ is a copy of. The second is where the four EBCDIC tables come from.
The third was consulted and deliberately not used. It does not say which code
page it is, it states that the translation is not bidirectional, and it
collapses everything unmappable to 0x1A, so nothing built on it could round
trip.
JOHAB.TXT is what Johab is checked against; see differences from other
implementations for the three
positions where this library departs from it. Windows' own table for code page
1361 was consulted and not followed. It decodes a standalone jamo to the
conjoining Hangul Jamo block rather than the Compatibility Jamo that KS X 1001
itself uses, and a few of its entries disagree with the arithmetic Annex 3
defines. RFC 1557 is the EUC-KR that Wansung is: KS X 1001 in EUC form, with
nothing added.
All of them are in the zig-charset Zotero collection.
Licensing
MIT, and the project follows REUSE —
nix develop -c reuse lint checks it.
The index files under data/ are the WHATWG's, under the CC-BY 4.0 that
covers the Encoding Standard. The tables generated from them are
MIT AND BSD-3-Clause, the standard's own terms for portions of it
incorporated into source code. The EBCDIC mapping files and the table
generated from them carry the Unicode licence.