// q27b
QWEN 3.8 27B ON AN RTX 3060
Qwen 3.8 27B at 38 t/s and Qwen 3.6 35B-A3B at 56 t/s on a used RTX 3060 12GB, quality held to within one HumanEval task.
// in plain terms
The RTX 3060 12GB is the most common graphics card on the Steam Hardware Survey and sells for about $250 used. Two open Qwen coding models, quantized to about 3 bits per weight, run on it at 1.7× and 2.5× the stock llama.cpp rate with no measurable quality loss. The configuration is published as an open-source launcher, with the full record of what was measured and what failed.
// 01
Installed base
The Steam Hardware Survey for August 2026 places 49.2% of its panel at 12 GB of video memory or more; the RTX 3060 12GB is the single most common card at about $250 used. The current 12 GB card, the RTX 5070, lists at $549 and fits no larger a model. For this tier the binding constraint is memory, not compute.
// 02
Models and quality
Compression to 3.2 bits per weight costs one HumanEval task against the uncompressed original. On multi-file editing the MoE model scores higher (p=0.039); the dense model produces well-formed edits that are more often wrong.
| model | file | HumanEval-164 | editing, 34 tasks |
|---|---|---|---|
| Qwen3.6-35B-A3B (MoE) | 16.0 GiB | 153 / 164 | 9 first try, 16 overall |
| Qwen3.8-27B (dense) | 10.2 GiB | 152 / 164 | 4 first try, 8 overall |
| 27B uncompressed original | 27.1 GiB | 153 / 164 | 6 first try, 10 overall |
// 03
Throughput
Two machines: a headless Proxmox VM with the card passed through, and a Windows 10 desktop with the same card driving the monitor through WSL2. Windows holds 0.3 to 1 GB of the card, which costs one context step. Electricity at the card's 170 W cap: 15 to 22 cents per million tokens.
| MoE, Linux | dense, Linux | MoE, WSL2 | dense, WSL2 | |
|---|---|---|---|---|
| stock decode | 22.2 t/s | 22.5 t/s | 18.7 t/s | 21.9 t/s |
| configured decode | 55.9 t/s | 38.4 t/s | 41.9 t/s | 34.6 t/s |
| editing | ~188 t/s | 113–246 t/s | not run | not run |
| time to finish, same task | 4.24× sooner | 3.82× sooner | 4.49× sooner | 3.74× sooner |
| context served | 16,384 | 12,288 | 12,288 | 8,192 |
// 04
Configuration
- +Thinking mode disabled. Same pass rate, 2.8× less time per task.
- +Built-in MTP draft head at depth 2. +72% on the dense model, quality-neutral by paired HumanEval.
- +N-gram matcher chained before the draft head. 2.8 to 6.1× on editing, byte-identical output, no effect on generation.
- +Placement-aware quantization. Each expert tensor's format chosen by the processor that executes it: +6.9% on the MoE model. A control with more bits and no CPU fast path ran 5% slower, isolating the kernel path as the mechanism.
// 05
Negative results
- +22 approaches failed and are documented with their cost: expert deferral, expert caching, kernel rewrites, third-party drafters, pruning, alternative KV formats, and two retractions.
- +HumanEval is blind to a class of damage. A published expert-deferral technique left HumanEval unchanged (p=1.0) while first-attempt editing solves fell from 13 of 34 to 3 of 34.
- +Every gain that shipped came from a flag, a file already on the machine, or a small patch. Approaches that required new machinery failed on 12 GB.
// the thread
Implication
Consumer hardware already in circulation carries more capacity than its age suggests, and that capacity grows without the hardware changing: every gain measured here came from the model file or the runtime, not the card. The installed base is not a fixed ceiling; it is a floor the software keeps raising.
// papers
this page is the TL;DR — the papers are the full story. drafts, provided as-is.