Qwengram-2B

Qwengram-2B logo

Frozen Qwen3.5-2B plus an R=1 reader at decoder IDX2/IDX8 and linear750 dynamic arbitration, using external Qwen3.8-Flash-Next PLE memory.

Canonical endpoint: REAL-15M + linear750 (15,000,064 reader tokens; 749,568 calibration tokens). The matched milestone study found lower full-validation, five-domain mean and LAMBADA NLL at 15M than at 10M, with paired 95% confidence intervals below zero. Four individual domains also improve; code and benchmark accuracy differences remain inconclusive.

The canonical checkpoint reduces frozen full-validation perplexity by 3.700% (14.141361 stock to 13.618098). The release now selects 15M for its stronger overall language-model results, superseding the original conservative rule that retained 10M. All model files below contain the 15M reader and its matching linear750 arbiter; there is no separate 10M model variant in the current release. See the four-arm comparison and paired intervals.

GGUF files

QwenGram-2B-BF16.gguf, QwenGram-2B-Q8_0.gguf, QwenGram-2B-Q6_K.gguf and QwenGram-2B-Q4_K_M.gguf contain the backbone and 11 FP32 reader/arbiter tensors. The reader and arbiter stay FP32 in every precision. The PLE is a required external file, not embedded in these GGUFs. reader.safetensors and arbiter.pt contain the same canonical 15M checkpoint separately. SHA256.json records all artifact hashes.

Required PLE sidecar

Download Ivan Fioravanti's Q4_1 PLE GGUF. Credit for this PLE conversion belongs to Ivan. Its SHA256 is 66db3ab390f4dd5063ecc89cc180f4713898577682347001bf64ab8e328527a1. The approximately 32 GB file is mapped on the host; only selected rows are dequantized for each token. It is not loaded as a 32 GB GPU allocation.

Build and run

Use the Qwengram llama.cpp fork, commit 068fcb42662453bec15298ba6bd59f552190a468, which includes 2B support:

git clone https://github.com/Ninnix/llama.cpp-qwengram.git
cd llama.cpp-qwengram
git checkout 068fcb42662453bec15298ba6bd59f552190a468
cmake -S . -B build-qwengram-cpu -DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_EXAMPLES=ON
cmake --build build-qwengram-cpu -j --target llama-completion
export QWENGRAM_PLE=/path/to/Qwen3.8-Flash-Next-PLE-Q4_1.gguf
build-qwengram-cpu/bin/llama-completion -m /path/to/QwenGram-2B-Q8_0.gguf -p 'The capital of France is' -n 16 -no-cnv -ngl 0

For Vulkan, build with -DGGML_VULKAN=ON. On the tested AMD BC-250, BF16, Q8_0, Q6_K, Q4_K_M with full Vulkan offload (-ngl 99) matched the corresponding CPU eight-token greedy continuation for The capital of France is. These short checks do not establish broad GPU parity.

The fork supports both the original 0.8B reader and the 2048-wide 2B reader. It changes reader dimensions and the corresponding inverse-square-root scale; injection placement, hashing, PLE lookup and arbitration semantics stay the same. Stock upstream llama.cpp does not execute this custom reader. MTP and embedding-only inputs are unsupported for this Qwengram runtime.

Frozen evaluation

Canonical REAL-15M + linear750, using the original FP8 PLE and frozen Kaggle evaluation suite.

Metric Frozen stock Canonical Qwengram-2B
Full-validation NLL 2.649104 2.611400
Full-validation perplexity 14.141361 13.618098
General NLL 2.841681 2.789140
Code NLL 1.361171 1.352904
Math NLL 1.317180 1.303235
Scientific NLL 2.073600 2.046080
Multilingual NLL 3.309704 3.267818
Five-domain mean NLL 2.180667 2.151835
LAMBADA-1000 NLL 1.825668 1.786991
LAMBADA-1000 accuracy 54.1% 54.6%
HellaSwag-1000 accuracy 45.6% 46.8%

These frozen study metrics are not quantized GGUF measurements. Quantized GGUF retention is evaluated separately below.

GGUF runtime retention

The matched CPU test scores 8,128 tokens from the first 64 consecutive 256-token WikiText-2 raw test chunks, scoring the last 127 tokens per chunk. All runs use eight threads and context/batch/microbatch 256, with no warmup. Reader gain is NLL(stock) - NLL(Qwengram); retention divides each quantized gain by the BF16 gain. Paired 95% intervals use 10,000 resamples of 16 consecutive four-chunk blocks, seed 1234. The external PLE is Ivan Fioravanti's Q4_1 sidecar.

Precision Stock NLL Qwengram NLL Reader gain [95% CI] Gain retention [95% CI] Perplexity reduction vs stock
BF16 2.538120 2.486529 0.051591 [0.037129, 0.066689] 100% 5.03%
Q8_0 2.539465 2.487978 0.051487 [0.037551, 0.066078] 99.8% [98.3%, 101.8%] 5.02%
Q6_K 2.551537 2.497033 0.054504 [0.039340, 0.070659] 105.6% [98.1%, 112.6%] 5.30%
Q4_K_M 2.577519 2.525079 0.052440 [0.037291, 0.067937] 101.6% [95.1%, 108.2%] 5.11%

These measurements use the canonical 15M reader. All 335 backbone tensors and nine tokenizer fields match the stock controls; all 11 reader and arbiter tensors remain bit-exact FP32. See tensor verification, generation checks and the matched runtime report for hashes, commands, logs and per-chunk scores. The quantized-sidecar runtime test and the FP8 PLE frozen study above are separate benchmarks.

The target model revision is 15852e8c16360a2fea060d615a32b45270f8a8fc. Full provenance is in qwengram-2b.json, the evaluation artifacts, and runtime/. This is an experimental text-generation release; vision has not been validated.

Downloads last month
982
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ninnix96/Qwengram-2B

Finetuned
Qwen/Qwen3.5-2B
Quantized
(235)
this model

Collection including Ninnix96/Qwengram-2B