Open-Dcoder-0.5B-mixture-mdm-step2000 (CDLM-0.5B)

CDLM-0.5B, the 0.5B corrective diffusion language model of Corrective Diffusion Language Models (NeurIPS 2026): fredzzp/open-dcoder-0.5B continued for 2,000 steps with the CDLM objective (absorbing corruption plus uniform replacement of 10% of the still-visible tokens, with a cross-entropy term on the replaced positions, weight 0.1). The weights are identical to Shuibai12138/CDLM-0.5B. Use this repository id with the paper's evaluation pipeline, because its name contains open-dcoder. The matched MDLM baseline is Shuibai12138/Open-Dcoder-0.5B-baseline-mdm-step2000.

Training

Initialisation fredzzp/open-dcoder-0.5B
Objective mixture_prob=0.1, noise_token_wt=0.1, clean_token_wt=0.0
Data nvidia/Nemotron-Pretraining-SFT-v1, subset Nemotron-SFT-Code, revision af7991c59eeb (gated; NVIDIA Data Agreement for Model Training)
Steps 2,000 (step 2,000 of a 20,345,053-step schedule; lr still in warmup)
Optimiser AdamW, peak lr 3e-4, cosine, 20,345 warmup steps (lr at step 2,000 = 2.95e-5), weight decay 0.01, grad clip 1.0, bf16
Batch global 12 sequences x 4,096 packed tokens (micro 3 x 4 GPUs)
Frozen lm_head, embed_tokens
Seed 42
Provenance step 2,000 of run Open-Dcoder-0.5B-continue-mixture-mdm-20251115-115315; identified by replaying the published weights (step-2001 loss bitwise equal to the run's log)
Code zhangshuibai/CDLM, tag v1.0.1-corrective-training: ARM=cdlm bash training/scripts/train_0.5b.sh; training/configs/cdlm/ holds the run's resolved configuration

Usage

This is a masked diffusion language model with bidirectional attention. Loading it with AutoModelForCausalLM gives a causal Qwen2 model and wrong outputs. Use the evaluation pipeline of the code repository, which loads models whose name contains open-dcoder with the diffusion Qwen2 implementation (bidirectional attention, shifted logits); see its README for the environment and commands, and pass --model_name Shuibai12138/Open-Dcoder-0.5B-mixture-mdm-step2000.

Implementation note on the gradient

The 0.5B training code computes the per-token cross-entropy terms with LigerFusedLinearCrossEntropyLoss(reduction="none") from liger-kernel 0.5.8. Its backward pass scales every token's gradient by the upstream gradient of the first token, so the per-token weights of the objective (the 1/|S| and 1/t factors and the noise-term weight) are not applied in the update: each micro-batch receives the first token's weight times the unweighted sum of the per-token gradients, a micro-batch whose first target token is unsupervised receives no gradient, and when the first target token is a replaced one the noise term adds gradient on every valid position, clean tokens included. Logged losses are correct. This model was trained with that code; see training/README.md ("Effective gradient") in the code repository.

Licence

MIT. Base model: fredzzp/open-dcoder-0.5B, licensed under the Apache License 2.0; this model is a derivative (continued training) of it, and the base model's licence and notices apply to the parts derived from it.

Citation

@inproceedings{zhang2026corrective,
  title     = {Corrective Diffusion Language Models},
  author    = {Zhang, Shuibai and Peng, Fred Zhangzhi and Zhang, Yiheng and Pan, Jin and Chrysos, Grigorios G.},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}
Downloads last month
758
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shuibai12138/Open-Dcoder-0.5B-mixture-mdm-step2000

Finetuned
(5)
this model

Dataset used to train Shuibai12138/Open-Dcoder-0.5B-mixture-mdm-step2000

Collection including Shuibai12138/Open-Dcoder-0.5B-mixture-mdm-step2000