// dictforge

DICTFORGE

SystemsSUBMITTED
+0%
over zstd's own trainer

Validation-driven dictionary training that beats zstd's own trainer by 3–32%.

// in plain terms

Compression tools like zstd can use a dictionary, a stored set of common patterns, to shrink small files further. The standard tool for building these dictionaries leaves gains unclaimed. We build them by search and measurement instead: 3–32% smaller output, in a format every stock zstd already understands.

// 01

Defect in the reference trainer

zstd's dictionary trainer never searches over dictionary size, targets a hard-coded compression level regardless of the caller's, and past a threshold its objective silently prefers a build that failed to fill the requested size. A request for a larger dictionary can return a smaller, worse one, with no warning at any verbosity level.

// 02

Method

We treat dictionary construction as search: generate candidates across sizes and trainer families, judge each by measured compression on a held-out split, and never return a candidate worse than the incumbent. The output is a 100% standard zstd dictionary, with zero runtime dependency and drop-in use in any deployment.

// 03

Gains by corpus

SMALLER OUTPUT VS zstd --train — LEVEL 19, HELD-OUT DATA

+32.1% — weblogs

+28.2% — gharchive

+28.2% — csvrows

+16.6% — apijson

+11.5% — github_users

at level 3: +3.1% to +14.3% — most of that from dictionary size; the method's own margin shows at level 19

The paper states this rather than leaving it implicit: at level 3 most of the gain is dictionary size itself, and the method's own margin shows at level 19.

corpuslevel 3level 19
weblogs+14.3%+32.1%
gharchive+9.4%+28.2%
csvrows+6.7%+28.2%
apijson+6.1%+16.6%
github_users+3.1%+11.5%

// 04

Root cause

The defect is pinned to exact source lines in facebook/zstd: each candidate's score is seeded with its own emitted size, so past a threshold the optimizer deterministically prefers the most degenerate candidate, and the diagnostics that would reveal it are dead code. Four patches are drafted against upstream.

// the thread

Status

The method is the standard machine-learning practice of measuring on held-out data, applied to a setting where it had not been used. Remaining before the DCC 2027 deadline: fix the one contamination issue gating the headline numbers, and send the four zstd patches, which are still drafts.

// papers

this page is the TL;DR — the papers are the full story. drafts, provided as-is.

questions about this work → contact@mericanii.com