// dictforge
DICTFORGE
Validation-driven dictionary training that beats zstd's own trainer by 3–32%.
// in plain terms
Compression tools like zstd can use a dictionary, a stored set of common patterns, to shrink small files further. The standard tool for building these dictionaries leaves gains unclaimed. We build them by search and measurement instead: 3–32% smaller output, in a format every stock zstd already understands.
// 01
Defect in the reference trainer
zstd's dictionary trainer never searches over dictionary size, targets a hard-coded compression level regardless of the caller's, and past a threshold its objective silently prefers a build that failed to fill the requested size. A request for a larger dictionary can return a smaller, worse one, with no warning at any verbosity level.
// 02
Method
We treat dictionary construction as search: generate candidates across sizes and trainer families, judge each by measured compression on a held-out split, and never return a candidate worse than the incumbent. The output is a 100% standard zstd dictionary, with zero runtime dependency and drop-in use in any deployment.
// 03
Gains by corpus
SMALLER OUTPUT VS zstd --train — LEVEL 19, HELD-OUT DATA
+32.1% — weblogs
+28.2% — gharchive
+28.2% — csvrows
+16.6% — apijson
+11.5% — github_users
at level 3: +3.1% to +14.3% — most of that from dictionary size; the method's own margin shows at level 19
The paper states this rather than leaving it implicit: at level 3 most of the gain is dictionary size itself, and the method's own margin shows at level 19.
| corpus | level 3 | level 19 |
|---|---|---|
| weblogs | +14.3% | +32.1% |
| gharchive | +9.4% | +28.2% |
| csvrows | +6.7% | +28.2% |
| apijson | +6.1% | +16.6% |
| github_users | +3.1% | +11.5% |
// 04
Root cause
The defect is pinned to exact source lines in facebook/zstd: each candidate's score is seeded with its own emitted size, so past a threshold the optimizer deterministically prefers the most degenerate candidate, and the diagnostics that would reveal it are dead code. Four patches are drafted against upstream.
// the thread
Status
The method is the standard machine-learning practice of measuring on held-out data, applied to a setting where it had not been used. Remaining before the DCC 2027 deadline: fix the one contamination issue gating the headline numbers, and send the four zstd patches, which are still drafts.
// papers
this page is the TL;DR — the papers are the full story. drafts, provided as-is.