AlphaZero Chess

Play the model ยท Source and documentation

The current production model is checkpoint 1026 from the final chess run. Its evaluated and deployed artifact is production/final-generation-1026/model.int8.onnx. Older production and compressed models remain available under their descriptive directories for reproducibility.

Stable aliases

Alias Contents
latest.int8.onnx Exact bytes of the evaluated and deployed generation-1026 INT8 QAT ONNX artifact.
latest.pt Unfused generation-1026 PyTorch training weights.
latest.jit.pt Fused floating-point TorchScript export from the same generation-1026 weights.

The TorchScript file is a compatibility export, not the quantized artifact used for the reported evaluation or live deployment. Consumers that need reproducibility should pin a full Hugging Face commit and use the descriptive path.

Production model

Checkpoint Generation 1026
Network 14 residual layers, width 160, scaled post-activation blocks
Residual context Global pooling every second block
Policy head Chess from-to attention, key size 128
Input/output 52x8x8 input, 1,880 chess actions, W/D/L value
Trainable parameters 6,315,378
Production precision INT8 QAT through TensorRT
Training hardware 8x RTX 4070 SUPER
Reported compute window 2.5 days
Reported compute cost USD 43.20

The final recipe used SGD with Nesterov momentum 0.9, a learning-rate schedule from 0.1 to a 0.01 floor, gradient clipping at 1.0, replay ratio 4, a 20-million-position replay capacity, policy loss weight 1.0, and 800 self-play visits. Self-play used the TensorRT INT8 QAT path.

Strength by search budget

Each reported rung contains 100 opening-paired games against Stockfish 13. The headline for each budget uses the opponent whose score is closest to 50%. Elo values use Marco Meloni's SSDF-calibrated Stockfish 13 node-strength curve and are ladder ratings, not FIDE, online-platform, or current full-strength Stockfish ratings.

Model search budget Opponent W/D/L Approximate calibrated Elo
Policy only Stockfish 13, 1,000 nodes 32/24/44 1,658 [1,608โ€“1,710]
100 searches Stockfish 13, 10,000 nodes 39/18/43 2,456 [2,400โ€“2,512]
1,000 searches Stockfish 13, 50,000 nodes 21/48/31 2,925 [2,875โ€“2,977]
10,000 searches Stockfish 13, 100,000 nodes 30/44/26 3,114 [3,065โ€“3,163]
100,000 searches Stockfish 13, 200,000 nodes 25/56/19 3,251 [3,206โ€“3,297]

The 100- and 1,000-search evaluations used one parallel search, the 10,000-search evaluation used four, and the 100,000-search evaluation used 16. Bracketed intervals are 95% paired-match estimates conditional on the fixed Stockfish anchors. Calibration and rating-list uncertainty are additional.

Artifacts and provenance

File Purpose SHA-256
production/final-generation-1026/model.int8.onnx Evaluated and deployed INT8 QAT model d634abacae3c874eac6ded89f6af861eb81b509da638b5ad710587b1a08be658
production/final-generation-1026/model.pt Raw PyTorch checkpoint weights c92a363b041a18d0ef93b852ac1c6d58716ae9a22b4e62d543de297c4ec5f904
production/final-generation-1026/model.jit.pt Floating-point TorchScript compatibility export 68c30ba27a961bfc024d6a8fc0b72a47b74e2b67df1cc011379f8f3dcb67edb2

The ONNX digest exactly matches the artifact used by the final evaluation matrix. The interactive deployment pins that digest and builds hardware-specific TensorRT engines into persistent Modal storage.

Follow-up attention model (archived, not deployed)

After the paper, one more self-play run trained a 10-layer, 192-wide attention network built like Lc0's T1 (with smolgen) using AdamW on 8x RTX 4080 SUPER. It is stronger than the production model but is not used by the play site or the latest aliases, which remain generation 1026.

Network 10 post-norm encoder layers, embedding 192, 6 heads, squared-ReLU feed-forward 768
Attention bias Smolgen (32/192/192), one template bank shared by all layers
Policy head Chess from-to attention, key size 192
Input/output 52x8x8 input, 1,880 chess actions, W/D/L value
Parameters 11,908,643 playing network; 12,022,145 with the auxiliary training heads
Precision TensorRT float16
Training ~56 h of run time, about USD 89 at USD 1.60/h
Model search budget Opponent W/D/L Approximate calibrated Elo
100,000 searches (8 parallel) Stockfish 13, 200,000 nodes 39/52/9 3,338 [3,293โ€“3,381]

Same protocol and calibration as the table above, except eight-way rather than sixteen-way parallelism and float16 rather than INT8 serving. Its 64-search training ladder settled near 2,480, against about 2,360 for the CNN lineage.

File Purpose SHA-256
production/attention-10x192-generation-896/model.fp16.onnx Evaluated float16 ONNX (fixed batch 320) 5c26b42bc7adf33ef20b6d5e6507d8bf01e023b7b2fbbfd53caa0fcc57e6ab91
production/attention-10x192-generation-896/model.pt Raw PyTorch weights of the evaluated checkpoint bcaa612c1adea1e29620b1768be493c3c42a891584c4b87d9f924c6ab0e10a41
production/attention-10x192-generation-950/model.pt Raw PyTorch weights of the final checkpoint (not match-evaluated) 4d676d4bd4889c04429a16eda8f073d5b35324617158696d72c7eee750873c5d

Each folder also holds metadata.json and the training checkpoint manifest; generation 896 includes its resolved experiment configuration. model.pt lists the shared smolgen template bank once per layer, so its tensors sum to 19.1M entries. Full results and caveats are in the attention run record; training curves and logs are in the run dataset.

Archived models

  • production/v93-generation-782/ is the earlier INT8 deployment candidate.
  • production/v34-generation-1465/ is the previous 14x160 TorchScript production model.
  • compressed/v34-replay-8x56-seed-20260827/ is the 474,069-parameter replay-distilled v34 student.
  • model_4_day_run.pt and model_4_day_run.jit.pt are legacy root-level artifacts retained for compatibility.

Loading

The production ONNX has fixed deployment dimensions. Real play requires the repository's 52-plane chess encoder, action mapping, legal-action mask, and search engine; the interactive site and UCI adapter provide complete integrations.

from huggingface_hub import hf_hub_download

path = hf_hub_download(
    repo_id='BertilBraun/alphazero-chess',
    filename='production/final-generation-1026/model.int8.onnx',
    revision='PIN_A_FULL_HUGGING_FACE_COMMIT',
)

For floating-point PyTorch inference, load production/final-generation-1026/model.jit.pt with torch.jit.load. The .pt file is intended for reconstructing the training network, not direct inference. Optimizer and QAT state are intentionally omitted.

The project code, documentation, and model artifacts are available under the MIT License. Third-party materials retain their own terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support