// q27b

QWEN 3.8 27B ON AN RTX 3060

InferenceSHIPPED
0.0
tokens/s on a 12 GB card, stock 22

Qwen 3.8 27B at 38 t/s and Qwen 3.6 35B-A3B at 56 t/s on a used RTX 3060 12GB, quality held to within one HumanEval task.

// in plain terms

The RTX 3060 12GB is the most common graphics card on the Steam Hardware Survey and sells for about $250 used. Two open Qwen coding models, quantized to about 3 bits per weight, run on it at 1.7× and 2.5× the stock llama.cpp rate with no measurable quality loss. The configuration is published as an open-source launcher, with the full record of what was measured and what failed.

// 01

Installed base

The Steam Hardware Survey for August 2026 places 49.2% of its panel at 12 GB of video memory or more; the RTX 3060 12GB is the single most common card at about $250 used. The current 12 GB card, the RTX 5070, lists at $549 and fits no larger a model. For this tier the binding constraint is memory, not compute.

// 02

Models and quality

Compression to 3.2 bits per weight costs one HumanEval task against the uncompressed original. On multi-file editing the MoE model scores higher (p=0.039); the dense model produces well-formed edits that are more often wrong.

modelfileHumanEval-164editing, 34 tasks
Qwen3.6-35B-A3B (MoE)16.0 GiB153 / 1649 first try, 16 overall
Qwen3.8-27B (dense)10.2 GiB152 / 1644 first try, 8 overall
27B uncompressed original27.1 GiB153 / 1646 first try, 10 overall

// 03

Throughput

Two machines: a headless Proxmox VM with the card passed through, and a Windows 10 desktop with the same card driving the monitor through WSL2. Windows holds 0.3 to 1 GB of the card, which costs one context step. Electricity at the card's 170 W cap: 15 to 22 cents per million tokens.

MoE, Linuxdense, LinuxMoE, WSL2dense, WSL2
stock decode22.2 t/s22.5 t/s18.7 t/s21.9 t/s
configured decode55.9 t/s38.4 t/s41.9 t/s34.6 t/s
editing~188 t/s113–246 t/snot runnot run
time to finish, same task4.24× sooner3.82× sooner4.49× sooner3.74× sooner
context served16,38412,28812,2888,192

// 04

Configuration

  • +Thinking mode disabled. Same pass rate, 2.8× less time per task.
  • +Built-in MTP draft head at depth 2. +72% on the dense model, quality-neutral by paired HumanEval.
  • +N-gram matcher chained before the draft head. 2.8 to 6.1× on editing, byte-identical output, no effect on generation.
  • +Placement-aware quantization. Each expert tensor's format chosen by the processor that executes it: +6.9% on the MoE model. A control with more bits and no CPU fast path ran 5% slower, isolating the kernel path as the mechanism.

// 05

Negative results

  • +22 approaches failed and are documented with their cost: expert deferral, expert caching, kernel rewrites, third-party drafters, pruning, alternative KV formats, and two retractions.
  • +HumanEval is blind to a class of damage. A published expert-deferral technique left HumanEval unchanged (p=1.0) while first-attempt editing solves fell from 13 of 34 to 3 of 34.
  • +Every gain that shipped came from a flag, a file already on the machine, or a small patch. Approaches that required new machinery failed on 12 GB.

// the thread

Implication

Consumer hardware already in circulation carries more capacity than its age suggests, and that capacity grows without the hardware changing: every gain measured here came from the model file or the runtime, not the card. The installed base is not a fixed ceiling; it is a floor the software keeps raising.

// papers

this page is the TL;DR — the papers are the full story. drafts, provided as-is.

questions about this work → contact@mericanii.com