Instructions to use Ninnix96/Qwengram-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Ninnix96/Qwengram-2B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ninnix96/Qwengram-2B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Ninnix96/Qwengram-2B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ninnix96/Qwengram-2B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Ninnix96/Qwengram-2B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Ninnix96/Qwengram-2B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Ninnix96/Qwengram-2B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Ninnix96/Qwengram-2B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Ninnix96/Qwengram-2B:Q4_K_M
Use Docker
docker model run hf.co/Ninnix96/Qwengram-2B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Ninnix96/Qwengram-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ninnix96/Qwengram-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ninnix96/Qwengram-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Ninnix96/Qwengram-2B:Q4_K_M
- Ollama
How to use Ninnix96/Qwengram-2B with Ollama:
ollama run hf.co/Ninnix96/Qwengram-2B:Q4_K_M
- Unsloth Desktop
- Pi
How to use Ninnix96/Qwengram-2B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-2B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Ninnix96/Qwengram-2B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Ninnix96/Qwengram-2B with Docker Model Runner:
docker model run hf.co/Ninnix96/Qwengram-2B:Q4_K_M
- Lemonade
How to use Ninnix96/Qwengram-2B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Ninnix96/Qwengram-2B:Q4_K_M
Run and chat with the model
lemonade run user.Qwengram-2B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Ninnix96/Qwengram-2B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-2B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Ninnix96/Qwengram-2B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Ninnix96/Qwengram-2B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-2B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Ninnix96/Qwengram-2B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwengram-2B
Frozen Qwen3.5-2B plus an R=1 reader at decoder IDX2/IDX8 and linear750 dynamic arbitration, using external Qwen3.8-Flash-Next PLE memory.
Canonical endpoint: REAL-15M + linear750 (15,000,064 reader tokens; 749,568 calibration tokens). The matched milestone study found lower full-validation, five-domain mean and LAMBADA NLL at 15M than at 10M, with paired 95% confidence intervals below zero. Four individual domains also improve; code and benchmark accuracy differences remain inconclusive.
The canonical checkpoint reduces frozen full-validation perplexity by 3.700% (14.141361 stock to 13.618098). The release now selects 15M for its stronger overall language-model results, superseding the original conservative rule that retained 10M. All model files below contain the 15M reader and its matching linear750 arbiter; there is no separate 10M model variant in the current release. See the four-arm comparison and paired intervals.
GGUF files
QwenGram-2B-BF16.gguf, QwenGram-2B-Q8_0.gguf, QwenGram-2B-Q6_K.gguf and QwenGram-2B-Q4_K_M.gguf
contain the backbone and 11 FP32 reader/arbiter tensors. The reader and arbiter
stay FP32 in every precision. The PLE is a required external file, not embedded
in these GGUFs. reader.safetensors and arbiter.pt contain the same canonical
15M checkpoint separately. SHA256.json records all artifact hashes.
Required PLE sidecar
Download Ivan Fioravanti's Q4_1 PLE GGUF.
Credit for this PLE conversion belongs to Ivan. Its SHA256 is
66db3ab390f4dd5063ecc89cc180f4713898577682347001bf64ab8e328527a1.
The approximately 32 GB file is mapped on the host; only selected rows are
dequantized for each token. It is not loaded as a 32 GB GPU allocation.
Build and run
Use the Qwengram llama.cpp fork, commit
068fcb42662453bec15298ba6bd59f552190a468, which includes 2B support:
git clone https://github.com/Ninnix/llama.cpp-qwengram.git
cd llama.cpp-qwengram
git checkout 068fcb42662453bec15298ba6bd59f552190a468
cmake -S . -B build-qwengram-cpu -DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_EXAMPLES=ON
cmake --build build-qwengram-cpu -j --target llama-completion
export QWENGRAM_PLE=/path/to/Qwen3.8-Flash-Next-PLE-Q4_1.gguf
build-qwengram-cpu/bin/llama-completion -m /path/to/QwenGram-2B-Q8_0.gguf -p 'The capital of France is' -n 16 -no-cnv -ngl 0
For Vulkan, build with -DGGML_VULKAN=ON. On the tested AMD BC-250, BF16, Q8_0, Q6_K, Q4_K_M with full Vulkan offload (-ngl 99) matched the corresponding CPU eight-token greedy continuation for The capital of France is. These short checks do not establish broad GPU parity.
The fork supports both the original 0.8B reader and the 2048-wide 2B reader. It changes reader dimensions and the corresponding inverse-square-root scale; injection placement, hashing, PLE lookup and arbitration semantics stay the same. Stock upstream llama.cpp does not execute this custom reader. MTP and embedding-only inputs are unsupported for this Qwengram runtime.
Frozen evaluation
Canonical REAL-15M + linear750, using the original FP8 PLE and frozen Kaggle evaluation suite.
| Metric | Frozen stock | Canonical Qwengram-2B |
|---|---|---|
| Full-validation NLL | 2.649104 | 2.611400 |
| Full-validation perplexity | 14.141361 | 13.618098 |
| General NLL | 2.841681 | 2.789140 |
| Code NLL | 1.361171 | 1.352904 |
| Math NLL | 1.317180 | 1.303235 |
| Scientific NLL | 2.073600 | 2.046080 |
| Multilingual NLL | 3.309704 | 3.267818 |
| Five-domain mean NLL | 2.180667 | 2.151835 |
| LAMBADA-1000 NLL | 1.825668 | 1.786991 |
| LAMBADA-1000 accuracy | 54.1% | 54.6% |
| HellaSwag-1000 accuracy | 45.6% | 46.8% |
These frozen study metrics are not quantized GGUF measurements. Quantized GGUF retention is evaluated separately below.
GGUF runtime retention
The matched CPU test scores 8,128 tokens from the first 64 consecutive 256-token WikiText-2 raw test chunks, scoring the last 127 tokens per chunk. All runs use eight threads and context/batch/microbatch 256, with no warmup. Reader gain is NLL(stock) - NLL(Qwengram); retention divides each quantized gain by the BF16 gain. Paired 95% intervals use 10,000 resamples of 16 consecutive four-chunk blocks, seed 1234. The external PLE is Ivan Fioravanti's Q4_1 sidecar.
| Precision | Stock NLL | Qwengram NLL | Reader gain [95% CI] | Gain retention [95% CI] | Perplexity reduction vs stock |
|---|---|---|---|---|---|
| BF16 | 2.538120 | 2.486529 | 0.051591 [0.037129, 0.066689] | 100% | 5.03% |
| Q8_0 | 2.539465 | 2.487978 | 0.051487 [0.037551, 0.066078] | 99.8% [98.3%, 101.8%] | 5.02% |
| Q6_K | 2.551537 | 2.497033 | 0.054504 [0.039340, 0.070659] | 105.6% [98.1%, 112.6%] | 5.30% |
| Q4_K_M | 2.577519 | 2.525079 | 0.052440 [0.037291, 0.067937] | 101.6% [95.1%, 108.2%] | 5.11% |
These measurements use the canonical 15M reader. All 335 backbone tensors and nine tokenizer fields match the stock controls; all 11 reader and arbiter tensors remain bit-exact FP32. See tensor verification, generation checks and the matched runtime report for hashes, commands, logs and per-chunk scores. The quantized-sidecar runtime test and the FP8 PLE frozen study above are separate benchmarks.
The target model revision is 15852e8c16360a2fea060d615a32b45270f8a8fc.
Full provenance is in qwengram-2b.json, the evaluation artifacts, and runtime/.
This is an experimental text-generation release; vision has not been validated.
- Downloads last month
- 982
4-bit
6-bit
8-bit
16-bit