Rumination 2026-07-17, not a purchase order. Front-door README + NOTES N-1..N-4 + BUILDS.md. [ezra]
4.5 KiB
local-horse — the two priced builds (as of 2026-07-17)
⚠️ Prices dated 2026-07-17, ±15% on used parts and worse on RAM/GPU (see NOTES N-3). Re-price before buying. Budget target: ~$20k. Constraint: no Mac. Both shapes serve an OpenAI-compatible endpoint on the tailnet (llama.cpp / vLLM, headless Debian, Dockge like everything else).
Build 1 — The Workhorse (one big card + cheap capacity)
Quiet, <1kW under load, standard 120V outlet, racks next to the Mill. The pick (NOTES N-4).
| Part | Spec | Est. price (2026-07-17) |
|---|---|---|
| GPU | RTX Pro 6000 Blackwell Max-Q 96GB (300W blower, 1.8TB/s) | $9,500–13,000 |
| CPU | EPYC 7763 used (64c Milan, 8-ch DDR4; full 8 CCDs = full bandwidth) | ~$1,200 |
| Board | Supermicro H12SSL-i | ~$650 |
| RAM | 8× 64GB DDR4-3200 RDIMM refurb (512GB, ~205GB/s) | ~$2,000–2,800 |
| Storage | 4TB NVMe | ~$450 |
| Chassis/PSU | 4U + 1300W + cooler + fans | ~$900 |
| Total | ~$15k–19k |
What it runs (MoE trick: hot path/attention in VRAM, cold experts in system RAM):
| Model | Fit | Speed (est.) | Capability, plainly |
|---|---|---|---|
| GPT-OSS-120B (5B active) | Entirely in VRAM @ 4-bit | 100+ tok/s | Reliable daily hand — summaries, scripts, small fixes. ~Sonnet-3.5 tier. |
| Qwen3-Coder 480B-A35B @ 4-bit (~270GB) | Hybrid VRAM+RAM | ~15–25 tok/s | The prize. Real agentic coding on a known repo. Sonnet-4-era. |
| GLM 5.2 quantized | Hybrid (footprint unverified) | similar | Current open SWE-Bench champ; same tier or a notch up. |
| Kimi K2.6 (1T) @ Q3 | Barely (608GB total, tight) | single digits | Party trick, not a daily driver. |
| K3 (2.8T) | No | — | Cloud only. Its distills should land in this envelope. |
Build 2 — The Feral Cats (8× used RTX 3090, 192GB VRAM)
Everything model-resident → prefill screams. Also screams literally; needs a 240V circuit (~3kW load), ~150W combined idle 24/7, eight points of used-silicon failure.
| Part | Spec | Est. price (2026-07-17) |
|---|---|---|
| GPUs | 8× used RTX 3090 24GB (~$850 ea; GDDR6X dodged the DRAM spike) | ~$6,800 |
| CPU | EPYC 7402 used (Rome — just needs the PCIe lanes) | ~$300 |
| Board | ASRock Rack ROMED8-2T (7× PCIe x16) | ~$700 |
| RAM | 256GB DDR4 refurb | ~$1,000 |
| Risers/frame | Bifurcation risers + open frame or 4U conversion | ~$500 |
| Power | 2× 1600W PSU + sync board, 240V | ~$650 |
| Storage/misc | NVMe, fans, zip ties, regret | ~$550 |
| Total | ~$10.5k–12k |
| Model | Fit | Speed (est.) | Capability, plainly |
|---|---|---|---|
| Qwen3-235B-A22B @ 4-bit (~130GB) | Fully in VRAM — the rig's sweet spot | fast decode AND prefill | Strong all-rounder one tier below Coder-480B. Best agentic feel per dollar here. |
| GPT-OSS-120B | Fully in VRAM, headroom | blazing | Same as workhorse, faster. |
| 70B-class dense (Llama/Qwen) | Trivially | blazing | Helper tier. |
| Qwen3-Coder 480B / GLM 5.2 | Doesn't fit at useful quant | offload defeats the rig's point | The one thing the workhorse does that this can't. |
| K2.6 / K3 | No | — | — |
The trade in one sentence
Workhorse reaches the best open coding models slowly-but-surely and behaves like an appliance; the cats run one tier down blisteringly fast for half the money, paid for in a 240V run, fan roar, and jank. For agentic loops (99% context re-reads — NOTES N-1) all-VRAM prefill genuinely matters, but model ceiling matters more: a smarter slow model wastes fewer turns than a fast one going in circles.
Why one big card and not 8 (the physics, kept for re-derivation)
- Power/noise: 8×350W = 2.8kW load — over a 15A/120V circuit's ceiling; plus ~150W idle, forever.
- Interconnect tax: split models cross PCIe (~32GB/s) every token vs 1.8TB/s on-card; consumer cards have no P2P/NVLink (disabled on 4090+). Multi-GPU wins at batch throughput (many users), not single-user latency — and the counter has one user.
- MoE changed the buy: ~35B active params need one fast GPU + cheap capacity, not 8 GPUs of compute.
Sources (as read 2026-07-17)
- Open-model landscape: digitalapplied.com (models↔hardware matching) · llmconfigurator.com (local coding report) · modal.com (SWE-bench open models) · mindstudio.ai (agentic open-source 2026)
- Hardware/pricing: videocardz.com (RTX Pro 6000 list $13,250) · thundercompute.com (build pricing) · corewavelabs.com + pcserverandparts.com (DRAM crisis) · tomshardware.com RAM price index · memory.net
- K3: venturebeat.com · cnbc.com · marktechpost.com · simonwillison.net · platform.kimi.ai quickstart