Instructions to use m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Apertus v1.5 8B text — MLX MXFP4
Base model |
Apertus |
mlx-lm |
oMLX |
The family
Format: MLX | Weights: MXFP4, 4.28 GB | License: Apache 2.0
Updated on 2026-09-19 with a better build. Same format, same file size, same loaders: the weights are now calibrated (GPTQ on the MXFP4 grid, a Hadamard rotation, a chosen exponent per group) instead of rounded to the nearest value. Perplexity went from 11.71 to 10.85 (+12.5% to +4.3% against the 8-bit build), and the KL against the 8-bit fell by three quarters in all four languages measured. If you downloaded this repository before that date, download it again.
This repository holds the text branch of Apertus 1.5 8B, converted to MLX and quantized to MXFP4. It runs on Apple silicon through mlx-lm.
Apertus 1.5 is the fully open model of the Swiss AI Initiative, built at EPFL, ETH Zurich and the Swiss National Supercomputing Centre on open data.
This conversion was made independently of the Apertus release.
Model summary
| Base model | swiss-ai/Apertus-v1.5-8B |
| Parameters | 8.05B (8,053,338,240) |
| Architecture | ApertusForCausalLM, 32 layers, hidden size 4096, 32 attention heads, 8 key-value heads |
| Activation | xIELU |
| Vocabulary | 131072 text tokens |
| Context length | 262144 |
| Quantization | MXFP4, 4.250 bits per weight, calibrated |
| Size on disk | 4.28 GB |
| Format | MLX safetensors |
| Modality | text in, text out |
| License | Apache 2.0, with the Apertus 1.5 acceptable use policy |
* The model size the sidebar of this page reports is smaller than 8.05B. That figure counts the elements of the stored tensors, and MLX packs quantized weights into 32-bit containers. The parameter count of the model is the one in the table.
The family
Six builds of the same text branch, measured on the same corpus with the same script. Perplexity is measured on the test split of Salesforce/wikitext, configuration wikitext-2-raw-v1, over 200 non-overlapping windows of 512 tokens, which is 102,400 scored tokens.
| Build | Bits per weight | Size | Perplexity | Against the 8-bit | Use it when |
|---|---|---|---|---|---|
| 8-bit | 8.500 | 8.54 GB | 10.4088 | reference | quality first, and the reference every other build here is measured against. |
| MXFP8 | 8.250 | 8.31 GB | 10.4748 | +0.6% | you want the floating-point format. In perplexity it sits next to the 6-bit, which is smaller. |
| 6-bit | 6.500 | 6.54 GB | 10.4618 | +0.5% | two GB less than the 8-bit for half a percent of perplexity. |
| 5-bit | 5.500 | 5.54 GB | 10.5261 | +1.1% | three GB less than the 8-bit, and still within one and a half percent of it. |
| 4-bit DWQ | 4.500 | 4.53 GB | 10.9691 | +5.4% | the smallest build that stays close: distillation buys back most of what plain 4-bit gives away, at the same file size. |
| MXFP4 *(this repository)* | 4.250 | 4.28 GB | 10.8522 | +4.3% | the smallest build here, and now calibrated: closer to the 8-bit than the 4-bit DWQ, a quarter of a GB lighter. |
* The 8-bit build is the reference because a bfloat16 run does not fit the 24 GB machine these were made on: its 15 GB of weights page to disk. The step from 8 bits to 6 costs half a percent, which makes a large gap between bfloat16 and 8 bits unlikely. That is an inference, not a measurement.
* Perplexity compares quantizations of one model on one corpus. It says nothing about how this model compares to a different model, and nothing about how well it follows instructions.
What was left out
These variants were measured on the same corpus and are not published. The table is here so that nobody repeats the work.
| Variant | Bits per weight | Size | Perplexity | Why it is not published |
|---|---|---|---|---|
| MXFP4, rounded to nearest (published here until 2026-09-19) | 4.250 | 4.28 GB | 11.71 | replaced by the calibrated MXFP4, same size |
| mixed 4/6, affine | 5.000 | 5.03 GB | 11.24 | bigger than the 4-bit DWQ build, and weaker |
| 4-bit, affine, no distillation | 4.500 | 4.53 GB | 11.59 | same size as the 4-bit DWQ build, and weaker |
| 3-bit, group size 32 | 4.000 | 4.03 GB | 66.06 | unusable |
| 3-bit DWQ, group size 64 | 3.500 | 3.52 GB | 50.91 | unusable |
| mixed 3/6, affine | 4.250 | 4.28 GB | 140.93 | unusable |
| 3-bit, group size 64 | 3.500 | 3.52 GB | 169.79 | unusable |
Three bits break this model. At a group size of 64 it answers fluently and wrongly, completing "The capital of Switzerland is" with "not a good idea"; at 32 it answers "Bern" and still reaches a perplexity of 66. Distillation recovers a large share of the loss, from 169.79 to 50.91, and the result is still five times the reference.
The two four-bit rows are here for a different reason: they work, but each is beaten by a published build that is the same size or smaller. Nothing between 4.25 and 4.5 bits per weight was worth publishing besides the two that are.
Need images or audio? The whole model, this same decoder plus the vision and audio towers in float32, is a separate family, converted to MLX for mlx-vlm. Start at Apertus v1.5 8B MLX.
This build
MXFP4 stores each weight as a 4-bit float and shares one 8-bit power-of-two exponent across every group of 32. The format is the one mlx_lm.convert --q-mode mxfp4 writes, byte for byte, and any loader that reads MXFP4 reads this build. What changed on 2026-09-19 is which values sit inside that format.
The first version of this repository rounded every weight to the nearest representable value, with the exponent MLX picks on its own: the smallest that never clips the largest weight of the group, which on many groups throws away one of the eight levels to save a single weight. This build is calibrated instead, in three steps that each move the error somewhere cheaper:
- the exponent of every group is chosen, between the one MLX would pick and the one below it, by the error it causes on the layer's output rather than on the weights;
- GPTQ: every rounded column pushes its error onto the columns not yet rounded, in proportion to how the layer's inputs correlate on 128 windows of calibration text, and each layer is calibrated on the inputs produced by the layers already quantized;
- a Hadamard rotation of the residual stream, applied to the weights and never at run time: the output of the model is unchanged (verified at 1e-6 in float32), the architecture is unchanged, and the weights inside each group of 32 become easier to round. The RMSNorm gains are folded into the projections that read them, which is why the norm weights of this build are all 1.
Rotation alone makes things worse (perplexity 13.36 with plain rounding); it pays only together with GPTQ. Nothing of the evaluation text was used for calibration.
Its perplexity is 10.8522, +4.3% against the 8-bit build, on the protocol described above.
Beyond perplexity, on text held out from calibration, against the 8-bit build:
| previous MXFP4 | this build | |
|---|---|---|
| KL per token, English and code | 0.208 | 0.050 |
| KL per token, Italian | 0.231 | 0.065 |
| KL per token, German | 0.220 | 0.063 |
| KL per token, French | 0.215 | 0.063 |
| same top-1 token as the 8-bit | 80.1% | 89.1% |
| multiple choice, mean of 7 tasks (8-bit: 68.8) | 63.6 | 67.2 |
| same answer as the 8-bit on those tasks | 83.2% | 91.5% |
The multiple-choice tasks are ARC-Challenge, HellaSwag, Winogrande and Global-MMLU in English, Italian, German and French, 3,467 questions scored by log-likelihood.
# GPTQ on the MXFP4 grid, Hadamard rotation of the residual stream,
# act-order with static groups; 128 calibration windows of 512 tokens
python scripts/mxfp4_gptq.py --metodo gptq --esponente gptq \
--campioni 128 --ruota --ordine --uscita <folder>
Built on 2026-09-19 on macOS 27.0 with mlx-lm 0.31.3. The MXFP4 bytes are written by the conversion's own packer, identical to mx.quantize when given the same exponents, and every tensor was read back and checked. The perplexity above was measured with mlx 0.31.2 and again with mlx 0.32.2: the same to the fourth decimal.
Run it
pip install mlx-lm
mlx_lm.generate --model m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4 \
--prompt "Name the capital of Switzerland and say one sentence about it." \
--max-tokens 200
From Python:
from mlx_lm import load, generate
model, tokenizer = load("m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4")
messages = [{"role": "user", "content": "Name the capital of Switzerland."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
print(generate(model, tokenizer, prompt=prompt, max_tokens=200))
Behind an OpenAI-compatible endpoint:
mlx_lm.server --model m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4
In oMLX, search for m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4 and download
it. Deliberation is read natively: the model opens it with <think> and
closes it with </think>, which is what oMLX looks for.
What the text branch is
The upstream release is multimodal. It reads images and audio as well as text,
and its architecture is Apertus1p5ForConditionalGeneration, which neither
mlx-lm 0.31.3 nor mlx-vlm 0.6.17 implements.
This repository holds the text decoder on its own, declared as
ApertusForCausalLM so that mlx-lm loads it. The token embedding goes from
266752 rows to 131072. The rows that go are the image and audio codebooks,
which start at index 131272 and 262344 in the upstream vocabulary. The
upstream output_vocab_size is already 131072, so the language modelling head
is untouched and no text token is lost.
The model reads and writes text. It does not accept images or audio. The whole model, towers included, is a different family: Apertus v1.5 8B MLX.
Limitations
The quantization inherits every limitation of the upstream model, which its model card describes. No output filter ships with these weights.
Quality here is measured by perplexity on one English corpus. Apertus 1.5 is multilingual, and the effect of quantization on languages other than English is not measured in this collection. Instruction following, reasoning and tool use are also unmeasured.
License and acceptable use
The weights stay under the Apache 2.0 license of the upstream release. Use is also subject to the Apertus 1.5 acceptable use policy and privacy policy:
For removal of personal or copyrighted data, write to the Swiss AI Initiative at llm-privacy-requests@swiss-ai.org or llm-copyright-requests@swiss-ai.org.
Credits
The model is the work of the Swiss AI Initiative. This repository adds the MLX conversion, the quantization and the measurements above.
@misc{ApertusV15,
author = {{Swiss AI Initiative}},
title = {Apertus v1.5},
year = {2026},
howpublished = {\url{https://huggingface.co/swiss-ai/Apertus-v1.5-8B}},
note = {EPFL, ETH Zurich, and the Swiss National Supercomputing Centre}
}
- Downloads last month
- 455
4-bit
Model tree for m1rkocasu/Apertus-v1.5-8B-text-MLX-mxfp4
Base model
swiss-ai/Apertus-v1.5-8B