POCKET-Darwin-180B-GGUF

A 180B frontier model from a laptop to DGX Spark. A 4-bit GGUF build of Darwin-180B-RSI-R3 (111.3 GB, 4 files) for llama.cpp. It runs on a gaming laptop with an 8 GB GPU and 32 GB RAM, a 128 GB mini PC, a CPU-only server, a single datacenter GPU, or NVIDIA DGX Spark.

  • Quality kept: MMLU-Pro (2,000 questions, paired) 87.65% vs 87.65% for the BF16 original, difference +0.00 points, 95% CI [βˆ’0.95, +1.00].
  • Runs on a laptop: 4.17 tokens/s on an RTX 5060 Laptop GPU (8 GB) with 32 GB RAM.
  • Runs without a GPU: 18–21 tokens/s on one server CPU socket, 78.8 GB peak RAM.
  • Runs on a mini PC: any mini PC with 128 GB RAM holds the whole model in memory, no GPU needed.
  • Built-in confidence: ZTC AUC 0.758 reading its own state; the 20% of questions it is most sure about are 97.5% correct. Probe file and script included in ztc/.
  • Shorter answers: on the same 2,000 questions, answers average 3,694 tokens vs 4,322 for the same-format parent build (βˆ’14.5%).

From an 8-GPU server to a laptop

Darwin-180B-RSI original (BF16) POCKET-Darwin-180B (this model)
Size 360 GB (131 files) 111.3 GB (4 files)
Hardware 4–8Γ— NVIDIA B200, or 8Γ— NVIDIA H100 (80 GB) a laptop with an 8 GB GPU and 32 GB RAM Β· a 128 GB mini PC Β· a CPU-only server Β· one DGX Spark
Hardware cost (industry estimate) an 8Γ— H100 server: roughly USD 350,000+ a gaming laptop: roughly USD 1,500
MMLU-Pro 87.65% 87.65% (identical)

At BF16 the original weighs 360 GB, more than four 80 GB H100s can hold, so in practice it is served on an 8-GPU H100 server; we ran it on 4Γ— B200, and the original card recommends 8Γ— B200. POCKET brings that class of model to a personal computer.

Darwin + POCKET in one picture

  • Darwin (model-level RSI): built on Qwen3.8-Flash-Next (180B MoE, 512 routed experts). The model solves verifiable problems, keeps only its own checked-correct solutions and trains on them, no human-written answers. Only attention paths and shared experts are trained; the experts, router, per-layer n-gram embeddings and MTP head stay intact.
  • POCKET (on-device): a graft onto the widely used UD-Q4_K_XL layout, exactly the 300 R3-changed tensors re-encoded in the layout's own format, every other byte identical; about 3B active parameters per token; experts read on demand from SSD via memory mapping.

What this is

Darwin-180B-RSI is Qwen3.8-Flash-Next improved by model-level recursive self-improvement (RSI): the model solves verifiable problems, keeps only its own solutions that check out, and trains on them, no human-written solutions. R3 is the second round on top of R1.

Only 300 tensors changed between the parent and R3 (attention paths and shared experts; all 512 routed experts, the router, the per-layer n-gram embeddings and the MTP head are unchanged). So this build takes the widely used Unsloth UD-Q4_K_XL layout of the parent and replaces exactly those 300 tensors with R3's weights, encoded in the same format the layout uses for them (Q8_0). Every other byte is identical to the parent build.

Verified before release:

  • all 300 target tensors found, same type and size, read back equal to R3, 300/300
  • conversion verified against the parent build: 570 untouched tensors byte-identical
  • loads and runs in llama.cpp b11048

Quality, before vs after quantization

MMLU-Pro, 2,000 questions stratified by subject (fixed by question-id hash, never chosen by looking at results), 0-shot CoT, temperature 1.0, top-p 0.95, top-k 20, min-p 0, seed 7, up to 131K tokens. Paired bootstrap 95% CI on the same questions.

Build Accuracy Paired difference
Darwin-180B-RSI-R3 BF16 (vLLM) 87.65% -
POCKET-Darwin-180B-GGUF (this, llama.cpp) 87.65% vs R3 BF16: +0.00 [βˆ’0.95, +1.00]

Zero truncated answers. 4-bit quantization costs nothing measurable.

The RSI gain carries over

Model-level RSI lifted R3 over R1 on held-out SuperGPQA by +1.03 points (95% CI [+0.05, +2.00], statistically significant): 1,000 questions never used in training or selection. Because this build keeps every R3-changed tensor at 8-bit and matches R3 BF16 exactly on MMLU-Pro, it carries that gain to a single machine. See the R3 card.

Same budget, more answers: R0 vs R3 inside the same 4-bit build

This build is the parent's public UD-Q4_K_XL file with only the 300 RSI-changed tensors swapped in, so comparing it with the parent 4-bit build isolates the effect of RSI: everything else is byte-identical.

Held-out SuperGPQA, 1,000 questions never used in training or selection, 4 samples each, 16K generation budget, same settings for both:

Build Mean of 4 Single sample Majority of 4
Parent Qwen3.8-Flash-Next UD-Q4_K_XL (R0) 59.10% 58.70% 63.60%
POCKET-Darwin-180B (R3, this) 61.55% 62.20% 65.90%

Paired difference (mean of 4): +2.45 points, 95% CI [+1.54, +3.36]. Majority of 4 is over answered samples; ties go to the answer that appears first in sample order (seeds 21 to 24).

POCKET reasons more concisely, so more of its answers finish within the same budget (675 vs 798 of 4,000 samples reached the 16K cap) and it uses fewer tokens: 13% less per question (geometric mean of per-question ratios) and 10% less in total (6,716 to 6,051 tokens on average). The largest gains are on the longest questions, where the parent often runs out of budget before answering.

Reference, the original Darwin-180B-RSI (BF16) benchmark results

POCKET carries the weights of the Darwin-180B-RSI line. For reference, these are the official results of the original BF16 model (Darwin-180B-RSI), #1 on seven Hugging Face official leaderboards.

Benchmark Score Setting Leaderboard
GPQA Diamond (198) 94.44 majority vote over up to 16 samples Β· 131,072-token thinking budget #1
MMLU-Pro (12,032) 88.12 single sample Β· 131,072-token thinking budget #1
AIME 2026 (30) 100.0 majority vote over 16 samples (mean accuracy 98.75) Β· 131,072-token thinking budget #1
HMMT Feb 2026 (33) 100.0 majority vote over 16 samples (mean accuracy 96.59) Β· 131,072-token thinking budget #1
MMMU-Pro (vision, 1,730) 79.48 majority vote over 3 samples Β· 131,072-token thinking budget #1
LEXam (law, MCQ 4-choice, 1,655) 68.94 majority vote over 4 samples (single sample 60.54 Β· mean 61.42) Β· 32,768-token thinking budget #1
LEXam-hard (law, open-ended, 518) 45.72 single sample Β· 32,768-token thinking budget (60 truncated answers regenerated at 120K) Β· judged by DeepSeek-R1-0528 per the official eval.yaml #1

Head-to-head with frontier models (published scores):

Model AIME 2026 GPQA Diamond MMLU-Pro MMMU-Pro HMMT Feb 2026 LEXam LEXam-hard
🧬 Darwin-180B-RSI (ours Β· πŸ‡°πŸ‡·) 100 πŸ₯‡ 94.44 πŸ₯‡ 88.12 πŸ₯‡ 79.48 πŸ₯‡ 100 πŸ₯‡ 68.94 πŸ₯‡ 45.72 πŸ₯‡
Inkling (Thinking Machines) - - - - - - 40.82
Kimi-K3 (Moonshot AI) - 93.5 - - - - 29.54
Kimi-K2.6 (Moonshot AI) 96.4 90.5 - 79.4 92.7 - 36.18
DeepSeek-V4-Pro (DeepSeek) - 90.1 87.5 - - - 38.93
Qwen3.5-397B-A17B (Alibaba) 93.33 88.4 87.8 - 87.88 - -
MiniMax-M2.1 (MiniMax) - 80.81 88 - - - -
GLM-5 (Zhipu AI) 95.83 86 86 - 86.36 - -
Intern-S2-Preview (Shanghai AI Lab) - - 88 76.88 87.31 - -
Step-3.5-Flash (StepFun) 96.67 83.5 84.4 - 86.36 - -
DeepSeek-R1 (DeepSeek) - - - - - 52.41 -
Qwen3-235B-A22B-Thinking-2507 (Alibaba) - - - - - 48.19 -

Settings for the reference results: BF16, vLLM, temperature 1.0 Β· top-p 0.95 Β· top-k 20, up to 131,072 thinking tokens; voting as listed per benchmark (see the original card).

ZTC: zero-token confidence

ZTC reads a "will this answer be right?" signal from the model's own hidden state at the last prompt token, before it writes a single token. One forward pass over the prompt, no answer tokens, no second model. Use it to send uncertain questions to a second pass, more samples, or a human.

Included in this repository: ztc/ with the probe file (32 KB), a usage script for llama.cpp, and a one-line llama.cpp patch for hidden-state output.

Self-readout on MMLU-Pro (2,000 questions; POCKET answers, ZTC reads POCKET's own state; 5-fold cross-validated):

ZTC AUC 0.758
Accuracy of the 20% most confident questions 97.5%
Accuracy over all questions 87.6%
Wrong answers caught by reviewing only the 20% least confident 46% of all wrong answers

Cross-model readout (same 400 questions with answers written by another model, same readout procedure):

Build ZTC AUC
Qwen3.8-Flash-Next BF16 (parent original) 0.715
Qwen3.8-Flash-Next 4-bit (same layout) 0.722
POCKET-Darwin-180B (this model) 0.742

Where it runs

Per token the model activates only about 3B parameters (10 of 512 experts). llama.cpp memory-maps the file, so the experts a token needs can be read straight from a fast SSD, the whole 111 GB does not have to fit in memory. More memory simply means more of the model stays resident and it runs faster.

Measured

Device Specs Speed (generation) How it runs
πŸ’» Gaming laptop NVIDIA GeForce RTX 5060 Laptop (8 GB VRAM) Β· 32 GB RAM Β· NVMe SSD 4.17 tok/s llama.cpp with the model memory-mapped from the SSD, experts are read on demand
πŸ–₯️ CPU-only server AMD EPYC 9365, 1 socket Β· 16 threads, no GPU 18.4–21.0 tok/s (prompt 54–68 tok/s) whole model in RAM, peak 78.8 GB, load 85 s
🟩 NVIDIA DGX Spark GB10 · 128 GB unified memory 39.25 tok/s (single stream, reference build of identical size and layout) fully resident on one Spark, no swap, no throttling
⚑ NVIDIA B200 1 GPU (183 GB) runs 8–16 parallel sessions of 131K context 110–140 GB VRAM

Device classes and recommended specs

Class Example devices Memory Disk Mode
Entry, laptop / desktop RTX 4060/5060 laptops, RTX 3060–5090 desktops 32 GB RAM + 8 GB+ VRAM NVMe SSD, 120 GB free experts streamed from SSD; speed follows SSD and RAM bandwidth
🧊 Mini PC mini PCs with 128 GB RAM (DDR5 SO-DIMM or LPDDR5X), e.g. AMD Ryzen AI Max+ 395 mini PCs, Intel Core Ultra / AMD Ryzen mini PCs 128 GB RAM NVMe SSD, 120 GB free whole model in RAM, CPU only, or with the integrated GPU
Workstation desktop with 96–128 GB RAM + any NVIDIA GPU 96–128 GB RAM NVMe SSD experts in RAM (--cpu-moe), rest on GPU
Unified-memory AI PC NVIDIA DGX Spark Β· Apple M-series 128 GB+ Β· AMD Ryzen AI Max+ 395 (128 GB) 128 GB NVMe SSD fully resident
CPU server AMD EPYC / Intel Xeon, 96 GB+ RAM 96 GB+ any whole model in RAM (--load-mode none)
Datacenter GPU NVIDIA B200 / H200 (one card), 2Γ— RTX PRO 6000 141 GB+ VRAM any fully on GPU, many parallel users

Peak RAM in the CPU run is below the file size because llama.cpp reads the per-layer n-gram embeddings from the file on demand.

Run

Requires llama.cpp b11048 or newer (architecture qwen4exp).

# DGX Spark / single large GPU
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 -fa on -c 131072 --jinja
# Laptop / desktop with 8 GB+ VRAM and 32 GB+ RAM (experts streamed from SSD)
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 999 --cpu-moe -fa on -c 8192 --jinja
# CPU only
llama-server -m POCKET-Darwin-180B-UD-Q4_K_XL-00001-of-00004.gguf -ngl 0 -t 16 -c 8192 --jinja --load-mode none

Recommended sampling: temperature 1.0, top-p 0.95, top-k 20. It is a reasoning model, give it room (β‰₯ 2,048 output tokens); the reasoning is returned in reasoning_content.

Text model (the vision encoder is left out to keep the build compact).

Coming next

  • DGX Spark throughput figures for this build.

Credits and license

Downloads last month
7,423
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FINAL-Bench/POCKET-Darwin-180B-GGUF

Quantized
(1)
this model

Spaces using FINAL-Bench/POCKET-Darwin-180B-GGUF 2

Collections including FINAL-Bench/POCKET-Darwin-180B-GGUF

Papers for FINAL-Bench/POCKET-Darwin-180B-GGUF

Article mentioning FINAL-Bench/POCKET-Darwin-180B-GGUF