- Zig 93.7%
- Python 4.3%
- Nix 2%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
All three packages have dropped the zig_ prefix from their names, and the dependency keys follow: font, font_config and font_renderer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UkN5DesmFtC4GPiqTrmkn8 |
||
| .forgejo/workflows | ||
| LICENSES | ||
| src | ||
| tests | ||
| tools | ||
| .gitignore | ||
| build.zig | ||
| build.zig.zon | ||
| build.zig.zon.nix | ||
| flake.lock | ||
| flake.nix | ||
| package.nix | ||
| README.md | ||
| REUSE.toml | ||
zig-font-shaper
A Zig 0.17 library that shapes text: given a run of text and a font decoded
by zig-font, it picks the glyphs the font draws the text with and works out
where each one goes, applying the font's GSUB and GPOS features —
ligatures, contextual forms, kerning, mark attachment, cursive joining —
the way HarfBuzz does.
It is a port of HarfBuzz's OpenType shaper, kept close enough to the original
that it gives the same glyphs, clusters and positions, to the unit. That is
checked against HarfBuzz's own shaping tests, which run as part of the test
suite: of the 5,639 cases imported from HarfBuzz 13.2.1, 5,633 pass, 3 fail
(see HarfBuzz's tests), and 3 are skipped because they ask
for things outside the OpenType shaper: HarfBuzz's fallback shaper, and
synthetic slant and emboldening.
The API reference is generated from the doc comments, and is published from
the default branch to https://jeff.jcollie.page/zig-font-shaper/. To read
it locally, run zig build docs-serve, which serves it at
http://127.0.0.1:8000/; use
-Ddocs-port=N for a different port. The pages have to be served rather than
opened from disk, because the viewer fetches its data at runtime and a
browser refuses to do that from a file:// page.
Quick start
const std = @import("std");
const font = @import("font");
const shaper = @import("font_shaper");
pub fn example(gpa: std.mem.Allocator, bytes: []const u8) !void {
var f = try font.Font.parse(gpa, bytes, .{});
defer f.deinit();
// The face positions in font units unless told otherwise.
var face: shaper.Face = .init(gpa, &f, &.{});
var buffer: shaper.Buffer = .init(gpa);
defer buffer.deinit();
try buffer.addUtf8("office", 0, null);
// Script and direction from the text; set them to override.
buffer.guessSegmentProperties();
try shaper.shapeBuffer(gpa, &face, &buffer, &.{.init("liga", 1)});
for (buffer.info.items, buffer.pos.items) |info, pos| {
// info.codepoint is now the glyph ID, info.cluster the byte offset
// of the text it came from.
std.debug.print("glyph {d} from {d}: advance {d}, offset {d},{d}\n", .{
info.codepoint, info.cluster, pos.x_advance, pos.x_offset, pos.y_offset,
});
}
}
To shape many runs alike, make a shaper.Plan once with
Plan.init(gpa, &face, .{ .direction = .ltr, .script = .latin, ... }) and
call shaper.shape.shape(&plan, &buffer) for each; shapeBuffer makes one
plan per call.
A variable font is shaped at a point in its design space by giving the face
normalized coordinates, which shaper.variations.normalize works out from
axis values the way HarfBuzz does:
var coords: [8]font.F2Dot14 = @splat(.{ .raw = 0 });
const axes = if (f.fvar) |fvar| fvar.axes.len else 0;
shaper.variations.normalize(&f, &.{.{ .tag = std.mem.readInt(u32, "wght", .big), .value = 700 }}, null, coords[0..axes]);
var face: shaper.Face = .init(gpa, &f, coords[0..axes]);
To use it from another project, run
zig fetch --save git+https://git.jcollie.dev/jeff/zig-font-shaper.git and
import the module in build.zig:
const shaper = b.dependency("font_shaper", .{ .target = target, .optimize = optimize });
exe.root_module.addImport("font_shaper", shaper.module("font_shaper"));
With zig-font-renderer
The font_shaper_text module joins the shaper to zig-font-renderer and
zig-font-config, replacing the renderer's unshaped fallback_layout:
const text = @import("font_shaper_text");
var list = try loader.resolveName("sans-serif");
defer list.deinit();
var layout = try text.layout(gpa, &list, 24, "Hello, world", .{});
defer layout.deinit(gpa);
for (layout.runs) |run| try renderer.drawRun(gpa, &surface, run, x, y, .{});
The text is split into items wherever the face that has its characters, or
its script, changes, and each item is shaped on its own with the text around
it as context. layout.clusters says, for each glyph, which byte of the text
it came from. Items are laid out in the order they come in the text: text
that mixes right-to-left and left-to-right scripts on one line needs the
Unicode bidirectional algorithm to order its items first, and that is not
done here.
On the command line
zig build run -- FONT-FILE TEXT shapes a line of text and prints the
glyphs in the format hb-shape --no-glyph-names prints them, with a subset
of hb-shape's options (--features, --variations, --direction,
--script, --language, --font-size, --unicodes and the output ones),
so the two can be compared directly. The tool is zig-shape in the package.
What it shapes
The pipeline is HarfBuzz's, stage by stage: Unicode properties and
grapheme clusters, dotted circles before stray marks, mirroring for
right-to-left text, normalization against what the font can draw, the
feature map with its stages and pauses, GSUB, advances from hmtx and
HVAR (or gvar's phantom points), GPOS or else the legacy kern table,
mark zeroing and fallback mark positioning, default ignorables, and the
glyph flags that say where text may be broken without reshaping.
Every lookup type of GSUB and GPOS is applied, contextual and
chained ones included, with lookup flags, mark filtering sets, feature
variations, and Device and VariationIndex adjustments in variable fonts.
Features can be set over ranges of clusters, with HarfBuzz's feature string
syntax (kern=0, +liga[3:5], aalt=2).
Script and language select the font's script and language system with HarfBuzz's tables of OpenType script and language tags, BCP 47 tags included.
Shapers: all of HarfBuzz's OpenType shapers. The default shaper covers
Latin, Greek, Cyrillic, CJK and the other scripts without a shaper of their
own. The Arabic shaper, for Arabic and Syriac, does their joining forms, mark
reordering, stch stretching, and shaping from the Unicode presentation
forms when a font has no GSUB. The Indic shaper, for Devanagari, Bengali,
Gurmukhi, Gujarati, Oriya, Tamil, Telugu, Kannada and Malayalam, finds
syllables and base consonants and reorders reph and pre-base forms, for old-
and new-specification fonts. The Khmer and Myanmar shapers reorder their
syllables likewise. The Universal Shaping Engine takes Tibetan, Sinhala,
Javanese, Balinese, Mongolian and the many other complex scripts, and Indic
fonts with dev3-style tags. The Thai shaper splits SARA AM, for Lao too,
and positions marks with the private-use glyphs of fonts without Thai
GSUB; the Hebrew shaper composes the presentation forms older fonts need;
the Hangul shaper composes or decomposes syllables to suit the font. The
Zawgyi shaper leaves Zawgyi-encoded Myanmar to the font.
Not done: Apple's morx, kerx and trak tables; the kern table's
state-machine subtables; hinting-dependent positioning (anchor points and
Device tables at a pixel size), since the shaper works in font units
scaled, not at a ppem; and glyph extents for bitmap and COLR glyphs, which
only fallback mark positioning uses.
Unicode data comes from uucode, at Unicode 18; HarfBuzz 13.2.1 is at Unicode 17, which can only matter for characters new in 18.
How it works
The source follows HarfBuzz's layout so that a reader of one can find their
way in the other: Buffer.zig is hb_buffer_t, Map.zig is hb_ot_map_t,
apply.zig is the matching machinery of hb-ot-layout-gsubgpos.hh,
gsub.zig and gpos.zig the lookup subtables, shape.zig is
hb-ot-shape.cc, and so on. Functions keep HarfBuzz's names in their doc
comments.
Agreeing to the unit takes reproducing more than the algorithms. Values
are scaled the way HarfBuzz scales them, in 16.16 fixed point and rounding
each value rather than the sum; roundf is HarfBuzz's, which rounds halves
up rather than away from zero; axis values are normalized through floating
point as HarfBuzz does; the cmap subtable is chosen as HarfBuzz chooses it,
symbol and Arabic PUA remappings included; and where a font's data is
ambiguous — duplicate glyphs in a coverage table the Arabic fallback builds —
HarfBuzz's binary search is reproduced to land on the same entry.
Memory. A Buffer owns its glyphs, and a Plan owns what it worked out;
both take an allocator and free with deinit. A Face borrows its font and
coordinates, and uses its allocator only for scratch space. Shaping does no
I/O.
Hostile fonts are what the fuzzer is for: it edits fonts and shapes with them, looking for panics, leaks and hangs. The buffer limits its length and the number of operations it allows as HarfBuzz does, and nested lookups are limited in depth; a font that hits a limit gets whatever shaping was done up to it.
Building and testing
The toolchain comes from Nix, as does hb-shape from the same HarfBuzz
release the tests were imported from.
$ nix develop
$ zig build test --summary all # unit tests, HarfBuzz's tests, the renderer glue
$ zig build hbtest -- -v arabic # HarfBuzz's tests in detail, filtered by name
$ zig build test --fuzz=1M # a bounded fuzzing run
$ zig build coverage # kcov report in zig-out/coverage
$ zig build docs # API reference in zig-out/docs
$ zig build run -- font.ttf 'text' # shape a line, as hb-shape prints it
$ nix build .#zig-font-shaper # the package, tests included, in the sandbox
After changing a dependency in build.zig.zon, regenerate the Nix expression
for it:
$ nix develop -c zon2nix --17 --nix=build.zig.zon.nix build.zig.zon
HarfBuzz's tests
tests/hb holds HarfBuzz's shaping tests — its own, the Adobe Annotated
OpenType Specification tests, and the Unicode text rendering tests — with
the fonts they use. tools/import_hb_tests.py copies them out of the
HarfBuzz source that nixpkgs builds hb-shape from, and runs every case
through that hb-shape twice: once to confirm it agrees with the expected
output recorded in HarfBuzz, and once with --no-glyph-names, whose output
is what is kept, so the tests compare glyph IDs and need no glyph names. The
handful of cases that release does not reproduce with its own font functions
are left out, with the reason, in tests/hb/dropped.txt. The AAT and
platform-shaper suites are not imported.
$ nix develop -c python3 tools/import_hb_tests.py
tests/hb/known-failures.txt lists the cases expected to fail: a .dfont
collection, which zig-font does not read, and two that print the extents of
bitmap and COLR glyphs. The test fails if any other case fails, or if a
listed one passes; zig build hbtest -- --update rewrites the list.
Generated tables
Tables are translated from HarfBuzz's generated sources by scripts in
tools: the BCP 47 to OpenType language tags (gen_ot_tag_table.py), the
Arabic presentation forms (gen_arabic_table.py), the legacy Arabic PUA
encodings (gen_arabic_pua.py), the character categories of the Indic,
Khmer and Myanmar shapers and of the Universal Shaping Engine
(gen_indic_table.py, which compiles HarfBuzz's own tables into a small C++
program to read them out, and so wants a compiler), and the vowel
sequences that get a dotted circle (gen_vowel_constraints.py). Each takes
the HarfBuzz source as its argument; run zig fmt on what it writes.
The syllable scanners HarfBuzz compiles with Ragel are translated by
gen_ragel_machine.py, which takes one hb-ot-shaper-*-machine.hh and a
name. Ragel's tables carry over unchanged; its driver loop is written once,
in src/shapers/ragel.zig, and each machine's actions become data for it. The table of canonical
compositions is generated at build time from uucode by
tools/gen_compose.zig.
Where this lives
git clone https://git.jcollie.dev/jeff/zig-font-shaper.git
It is on Radicle as
rad:z44vbkcWYHEAZUUDiGA7Gc8vniqfb. A Radicle repository can only be found
by its identifier, so that is all a peer needs to fetch it:
rad clone rad:z44vbkcWYHEAZUUDiGA7Gc8vniqfb
License
The code is MIT-licensed, apart from what is ported or translated from
HarfBuzz, which also carries HarfBuzz's own "Old MIT" license
(MIT-Modern-Variant) and copyright lines, file by file. HarfBuzz's test
data keeps its licenses: HarfBuzz's own under MIT-Modern-Variant, the
Adobe and Unicode suites under Apache-2.0. The project follows the
REUSE specification; reuse lint checks it.
References cited
These are kept in the zig-font-shaper Zotero collection.
- Esfahbod, Behdad, and the HarfBuzz contributors. HarfBuzz, version 13.2.1, text shaping library. https://github.com/harfbuzz/harfbuzz
- Hadley, Josh. Unicode Text Segmentation. Unicode Standard Annex #29, revision 49, Unicode 18.0.0, 1 September 2026. https://www.unicode.org/reports/tr29/tr29-49.html
- Karoonboonyanan, Theppitak. Thai OpenType Shaping. https://linux.thai.net/~thep/th-otf/shaping.html
- Microsoft Corporation. Creating and Supporting OpenType Fonts for the Universal Shaping Engine. https://learn.microsoft.com/en-us/typography/script-development/use
- Microsoft Corporation. Developing OpenType Fonts for Arabic Script. https://learn.microsoft.com/en-us/typography/script-development/arabic
- Microsoft Corporation. Developing OpenType Fonts for Devanagari Script. https://learn.microsoft.com/en-us/typography/script-development/devanagari
- Microsoft Corporation. Developing OpenType Fonts for Khmer Script. https://learn.microsoft.com/en-us/typography/script-development/khmer
- Microsoft Corporation. Developing OpenType Fonts for Myanmar Script. https://learn.microsoft.com/en-us/typography/script-development/myanmar
- Microsoft Corporation. Developing OpenType Fonts for Syriac Script. https://learn.microsoft.com/en-us/typography/script-development/syriac
- Microsoft Corporation. OpenType Language System Tags. https://learn.microsoft.com/en-us/typography/opentype/spec/languagetags
- Microsoft Corporation. OpenType Specification, version 1.9.1, May 2024. https://learn.microsoft.com/en-us/typography/opentype/spec/
- Microsoft Corporation. Script Tags. https://learn.microsoft.com/en-us/typography/opentype/spec/scripttags
- Phillips, A. and M. Davis. Tags for Identifying Languages. BCP 47, RFC 5646, September 2009. https://www.rfc-editor.org/info/rfc5646
- Pournader, Roozbeh, Bob Hallissy and Lorna Evans. Unicode Arabic Mark Rendering. Unicode Standard Annex #53, revision 13, Unicode 18.0.0, 2 September 2026. https://www.unicode.org/reports/tr53/tr53-13.html
- The Unicode Consortium. "Middle East-I", chapter 9 of The Unicode Standard, Version 18.0, 2026. https://www.unicode.org/versions/Unicode18.0.0/
- The Unicode Consortium. "South and Central Asia-I", chapter 12 of The Unicode Standard, Version 18.0, 2026. https://www.unicode.org/versions/Unicode18.0.0/
- The Unicode Consortium. "Southeast Asia", chapter 16 of The Unicode Standard, Version 18.0, 2026. https://www.unicode.org/versions/Unicode18.0.0/
- The Unicode Consortium. "East Asia", chapter 18 of The Unicode Standard, Version 18.0, 2026. https://www.unicode.org/versions/Unicode18.0.0/
- Whistler, Ken. Unicode Normalization Forms. Unicode Standard Annex #15, revision 57, Unicode 17.0.0, 30 July 2025. https://www.unicode.org/reports/tr15/tr15-57.html