Brokefish, 24-hour network

A network trained by Brokefish, a chess engine that learns only from games it plays against itself and trains on a single consumer GPU. This network is the final checkpoint of a 24-hour run on an RTX 4060 Laptop (run name t24h-adamw-int8).

Training used no human games, no games or evaluations from other engines, no pretrained weights and no opening books. The engine knows the rules of chess and nothing about which positions are good.

What the experiments have found so far, with their caveats: findings.md.

Results

Against the published AlphaGateau model (the final checkpoint of its 500-iteration run, trained for 13.7 days on eight RTX A5000s), over 200 games:

W D L score 95 % interval Elo difference
13 141 46 0.4175 [0.351, 0.487] −58

Both engines searched 128 simulations per move with Gumbel MuZero over their top 16 moves. The 100 openings were 8 uniformly random plies, each played once with each colour, and games reaching 300 plies were scored as draws. On our side the network ran in fp16. The games are in the GitHub repository under logs/h2h-t24hfull.pgn.

This network has no rating on any published Elo scale.

Model

A transformer whose 32 input tokens are the 32 pieces of the board, each keeping its slot for the whole game: 8 layers, width 256, 8 heads, feed-forward width 1024, 6,383,360 parameters, about 400 MFLOPs per evaluation.

Outputs, unmasked:

  • policy: 32 × 64 logits, one per (piece, destination square). Illegal moves must be masked by the caller.
  • promotion: 32 × 4 logits, one set per piece.
  • value: one scalar in [−1, 1], from the side to move's point of view.

Training

hardware 1 × RTX 4060 Laptop, 8 GB
wall clock 24 h (20.6 h self-play, 3.1 h gradient steps)
optimiser steps 10,303, batch 4096
self-play games 430,122
positions generated 51.8 M
search during self-play Gumbel MuZero, 128 simulations, top 16 moves
optimiser AdamW, lr 1e-3 cosine to 5e-5, weight decay 0.01
inference during self-play int8 CUDA kernel

training.json holds the full training configuration.

Files

  • model.safetensors: the weights, fp32, as the state_dict of BrokefishNet.
  • brokefish-t24h.pt: the same weights with the training configuration, in the format the Brokefish scripts load directly.
  • training.json: the training configuration and run statistics.

Usage

The network needs the Brokefish code for its input encoding, its CUDA kernels and the search.

git clone https://github.com/TheoBoyer/brokefish && cd brokefish
uv sync
uvx --from huggingface_hub hf download Theob/Brokefish brokefish-t24h.pt --local-dir checkpoints

Loading it in Python:

from brokefish.eval.layer0 import load_net_state
from brokefish.nn.model import net_for_state

state = load_net_state("checkpoints/brokefish-t24h.pt", device="cpu")
net = net_for_state(state)
net.load_state_dict(state)

Playing a 200-game match between two checkpoints:

scripts/h2h-bf.sh checkpoints/brokefish-t24h.pt <other-checkpoint> <tag>

Limitations

  • In the AlphaGateau match this network lost material without compensation about 2.2 times as often as its opponent.
  • The kernels target NVIDIA Ada GPUs (sm89) and were only tested on an RTX 4060 Laptop. A plain PyTorch implementation exists, but it is much slower.
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
6.38M params
Tensor type
F32
·
Video Preview
loading

Paper for Theob/Brokefish