Brokefish, 24-hour network
A network trained by Brokefish, a chess
engine that learns only from games it plays against itself and trains on a single
consumer GPU. This network is the final checkpoint of a 24-hour run on an RTX 4060
Laptop (run name t24h-adamw-int8).
Training used no human games, no games or evaluations from other engines, no pretrained weights and no opening books. The engine knows the rules of chess and nothing about which positions are good.
What the experiments have found so far, with their caveats: findings.md.
Results
Against the published AlphaGateau model (the final checkpoint of its 500-iteration run, trained for 13.7 days on eight RTX A5000s), over 200 games:
| W | D | L | score | 95 % interval | Elo difference |
|---|---|---|---|---|---|
| 13 | 141 | 46 | 0.4175 | [0.351, 0.487] | −58 |
Both engines searched 128 simulations per move with Gumbel MuZero over their top 16
moves. The 100 openings were 8 uniformly random plies, each played once with each
colour, and games reaching 300 plies were scored as draws. On our side the network
ran in fp16. The games are in the GitHub repository under logs/h2h-t24hfull.pgn.
This network has no rating on any published Elo scale.
Model
A transformer whose 32 input tokens are the 32 pieces of the board, each keeping its slot for the whole game: 8 layers, width 256, 8 heads, feed-forward width 1024, 6,383,360 parameters, about 400 MFLOPs per evaluation.
Outputs, unmasked:
- policy: 32 × 64 logits, one per (piece, destination square). Illegal moves must be masked by the caller.
- promotion: 32 × 4 logits, one set per piece.
- value: one scalar in [−1, 1], from the side to move's point of view.
Training
| hardware | 1 × RTX 4060 Laptop, 8 GB |
| wall clock | 24 h (20.6 h self-play, 3.1 h gradient steps) |
| optimiser steps | 10,303, batch 4096 |
| self-play games | 430,122 |
| positions generated | 51.8 M |
| search during self-play | Gumbel MuZero, 128 simulations, top 16 moves |
| optimiser | AdamW, lr 1e-3 cosine to 5e-5, weight decay 0.01 |
| inference during self-play | int8 CUDA kernel |
training.json holds the full training configuration.
Files
model.safetensors: the weights, fp32, as thestate_dictofBrokefishNet.brokefish-t24h.pt: the same weights with the training configuration, in the format the Brokefish scripts load directly.training.json: the training configuration and run statistics.
Usage
The network needs the Brokefish code for its input encoding, its CUDA kernels and the search.
git clone https://github.com/TheoBoyer/brokefish && cd brokefish
uv sync
uvx --from huggingface_hub hf download Theob/Brokefish brokefish-t24h.pt --local-dir checkpoints
Loading it in Python:
from brokefish.eval.layer0 import load_net_state
from brokefish.nn.model import net_for_state
state = load_net_state("checkpoints/brokefish-t24h.pt", device="cpu")
net = net_for_state(state)
net.load_state_dict(state)
Playing a 200-game match between two checkpoints:
scripts/h2h-bf.sh checkpoints/brokefish-t24h.pt <other-checkpoint> <tag>
Limitations
- In the AlphaGateau match this network lost material without compensation about 2.2 times as often as its opponent.
- The kernels target NVIDIA Ada GPUs (sm89) and were only tested on an RTX 4060 Laptop. A plain PyTorch implementation exists, but it is much slower.