AlphaZero Chess
Play the model ยท Source and documentation
The current production model is checkpoint 1026 from the final chess run. Its evaluated and deployed artifact is
production/final-generation-1026/model.int8.onnx. Older production and compressed models remain available under
their descriptive directories for reproducibility.
Stable aliases
| Alias | Contents |
|---|---|
latest.int8.onnx |
Exact bytes of the evaluated and deployed generation-1026 INT8 QAT ONNX artifact. |
latest.pt |
Unfused generation-1026 PyTorch training weights. |
latest.jit.pt |
Fused floating-point TorchScript export from the same generation-1026 weights. |
The TorchScript file is a compatibility export, not the quantized artifact used for the reported evaluation or live deployment. Consumers that need reproducibility should pin a full Hugging Face commit and use the descriptive path.
Production model
| Checkpoint | Generation 1026 |
| Network | 14 residual layers, width 160, scaled post-activation blocks |
| Residual context | Global pooling every second block |
| Policy head | Chess from-to attention, key size 128 |
| Input/output | 52x8x8 input, 1,880 chess actions, W/D/L value |
| Trainable parameters | 6,315,378 |
| Production precision | INT8 QAT through TensorRT |
| Training hardware | 8x RTX 4070 SUPER |
| Reported compute window | 2.5 days |
| Reported compute cost | USD 43.20 |
The final recipe used SGD with Nesterov momentum 0.9, a learning-rate schedule from 0.1 to a 0.01 floor, gradient clipping at 1.0, replay ratio 4, a 20-million-position replay capacity, policy loss weight 1.0, and 800 self-play visits. Self-play used the TensorRT INT8 QAT path.
Strength by search budget
Each reported rung contains 100 opening-paired games against Stockfish 13. The headline for each budget uses the opponent whose score is closest to 50%. Elo values use Marco Meloni's SSDF-calibrated Stockfish 13 node-strength curve and are ladder ratings, not FIDE, online-platform, or current full-strength Stockfish ratings.
| Model search budget | Opponent | W/D/L | Approximate calibrated Elo |
|---|---|---|---|
| Policy only | Stockfish 13, 1,000 nodes | 32/24/44 | 1,658 [1,608โ1,710] |
| 100 searches | Stockfish 13, 10,000 nodes | 39/18/43 | 2,456 [2,400โ2,512] |
| 1,000 searches | Stockfish 13, 50,000 nodes | 21/48/31 | 2,925 [2,875โ2,977] |
| 10,000 searches | Stockfish 13, 100,000 nodes | 30/44/26 | 3,114 [3,065โ3,163] |
| 100,000 searches | Stockfish 13, 200,000 nodes | 25/56/19 | 3,251 [3,206โ3,297] |
The 100- and 1,000-search evaluations used one parallel search, the 10,000-search evaluation used four, and the 100,000-search evaluation used 16. Bracketed intervals are 95% paired-match estimates conditional on the fixed Stockfish anchors. Calibration and rating-list uncertainty are additional.
Artifacts and provenance
| File | Purpose | SHA-256 |
|---|---|---|
production/final-generation-1026/model.int8.onnx |
Evaluated and deployed INT8 QAT model | d634abacae3c874eac6ded89f6af861eb81b509da638b5ad710587b1a08be658 |
production/final-generation-1026/model.pt |
Raw PyTorch checkpoint weights | c92a363b041a18d0ef93b852ac1c6d58716ae9a22b4e62d543de297c4ec5f904 |
production/final-generation-1026/model.jit.pt |
Floating-point TorchScript compatibility export | 68c30ba27a961bfc024d6a8fc0b72a47b74e2b67df1cc011379f8f3dcb67edb2 |
The ONNX digest exactly matches the artifact used by the final evaluation matrix. The interactive deployment pins that digest and builds hardware-specific TensorRT engines into persistent Modal storage.
Follow-up attention model (archived, not deployed)
After the paper, one more self-play run trained a 10-layer, 192-wide attention network built like Lc0's T1 (with
smolgen) using AdamW on 8x RTX 4080 SUPER. It is stronger than the production model but is not used by the play
site or the latest aliases, which remain generation 1026.
| Network | 10 post-norm encoder layers, embedding 192, 6 heads, squared-ReLU feed-forward 768 |
| Attention bias | Smolgen (32/192/192), one template bank shared by all layers |
| Policy head | Chess from-to attention, key size 192 |
| Input/output | 52x8x8 input, 1,880 chess actions, W/D/L value |
| Parameters | 11,908,643 playing network; 12,022,145 with the auxiliary training heads |
| Precision | TensorRT float16 |
| Training | ~56 h of run time, about USD 89 at USD 1.60/h |
| Model search budget | Opponent | W/D/L | Approximate calibrated Elo |
|---|---|---|---|
| 100,000 searches (8 parallel) | Stockfish 13, 200,000 nodes | 39/52/9 | 3,338 [3,293โ3,381] |
Same protocol and calibration as the table above, except eight-way rather than sixteen-way parallelism and float16 rather than INT8 serving. Its 64-search training ladder settled near 2,480, against about 2,360 for the CNN lineage.
| File | Purpose | SHA-256 |
|---|---|---|
production/attention-10x192-generation-896/model.fp16.onnx |
Evaluated float16 ONNX (fixed batch 320) | 5c26b42bc7adf33ef20b6d5e6507d8bf01e023b7b2fbbfd53caa0fcc57e6ab91 |
production/attention-10x192-generation-896/model.pt |
Raw PyTorch weights of the evaluated checkpoint | bcaa612c1adea1e29620b1768be493c3c42a891584c4b87d9f924c6ab0e10a41 |
production/attention-10x192-generation-950/model.pt |
Raw PyTorch weights of the final checkpoint (not match-evaluated) | 4d676d4bd4889c04429a16eda8f073d5b35324617158696d72c7eee750873c5d |
Each folder also holds metadata.json and the training checkpoint manifest; generation 896 includes its resolved
experiment configuration. model.pt lists the shared smolgen template bank once per layer, so its tensors sum to
19.1M entries. Full results and caveats are in the
attention run record;
training curves and logs are in the
run dataset.
Archived models
production/v93-generation-782/is the earlier INT8 deployment candidate.production/v34-generation-1465/is the previous 14x160 TorchScript production model.compressed/v34-replay-8x56-seed-20260827/is the 474,069-parameter replay-distilled v34 student.model_4_day_run.ptandmodel_4_day_run.jit.ptare legacy root-level artifacts retained for compatibility.
Loading
The production ONNX has fixed deployment dimensions. Real play requires the repository's 52-plane chess encoder, action mapping, legal-action mask, and search engine; the interactive site and UCI adapter provide complete integrations.
from huggingface_hub import hf_hub_download
path = hf_hub_download(
repo_id='BertilBraun/alphazero-chess',
filename='production/final-generation-1026/model.int8.onnx',
revision='PIN_A_FULL_HUGGING_FACE_COMMIT',
)
For floating-point PyTorch inference, load production/final-generation-1026/model.jit.pt with torch.jit.load.
The .pt file is intended for reconstructing the training network, not direct inference. Optimizer and QAT state are
intentionally omitted.
The project code, documentation, and model artifacts are available under the MIT License. Third-party materials retain their own terms.