Fit
Tell it the machine. Every rung on the shelf is badged against that budget, live.
1-click • secure • instant
Shelf
Not a factory you operate — files that exist right now, sizes read from the published repos.
Request a build
Any model, any lane, built with the Pollard method on a cook's hardware. You get a quote; your card is charged only when the build passes the verification gate and is delivered.
pollard-verify must pass and the standard card ships with it. Fail the gate, no charge.pollard-taskeval report included. Ask in the notes.Cooks
Every cook ships the same method, the same gate, the same card. Their Hugging Face attribution is on every build they deliver.
Get Pollard
Python 3.9+. The install script also builds the llama.cpp runtime so the whole chain — measure, allocate, export, verify — runs end to end. Studio is the desktop app over all of it.
Apple Silicon or Intel. Metal runtime built by the installer; MLX lane native.
CUDA on RTX builds every lane; Pollard's own K2 tokenizer fix ships in the runtime patch set.
CUDA, Vulkan or CPU. Same script; set POLLARD_GPU to force a backend.
Device scan, the shelf, every lane, chat with what you built. Same look as this page.
Docs
The measured allocation — protect attention, router and embeddings, crush the expert body, spend bits where the routing profile says — is the same in every lane. Only the format changes.
Know what the machine can run before you download anything, then build for it.
The win over uniform quants is the calibration step: measure, then allocate on the measurement. Start from f16/bf16.
pollard-calcfits resident · streaming-viable · too bigpollard-smoothprecondition outliers (low-bit lanes)pollard-probeper-tensor sensitivity, any boxpollard-automapthe measured mix, dense or MoEpollard-verifyreconstruction gate — no pass, no shippollard-cardthe standard model card, every repo matchesPick by where the model will actually run.
GGUFllama.cpp · Ollama · LM Studio — trellis mix to ~1-bit; the smallest buildvLLM · SGLangGPTQ 4/8-bit dynamic mix (Marlin) for tensor-parallel cluster servingGPTQINT3/INT4 full-Hessian error-feedback for torch / HFMLXApple Silicon, mixed 4/8-bit at Metal speedEXL3exllamav3 — Pollard beats EXL3 on its own allocator: 8.670 vs 8.699 @4bpwMX · NVFP4Blackwell FP4 tensor cores via compressed-tensors; W4A16 on any vLLM GPUpollard-sensitivityeach tensor's true KL costpollard-expertsmeasured expert usage (MoE)pollard-pruneREAP-style expert pruningpollard-rotateincoherence rotation (QuIP# / QuaRot)pollard-packwafer-scale capacity plannerpollard-klKL-to-f16 — the judging metricpollard-evaltop-1 agreement + KL, with chartpollard-taskevaltask suites; point it at your own evalspollard-serve-evalA/B on the served stack (vLLM / SGLang)pollard-doctordiagnose · predict · repair any model, any lanepollard-runmeasured expert placement, RAM-streamingpollard-noderun on every box so Studio totals a cluster's RAMggml-rpcpool machines' memory over the networkpollard-onboardaudit a new architecture, emit a PR-ready contributionskills/pollardagent skill — routes any model down the right lane