CDLM-0.5B

CDLM-0.5B is the 0.5B corrective diffusion language model of Corrective Diffusion Language Models (NeurIPS 2026): fredzzp/open-dcoder-0.5B continued for 2,000 steps with the CDLM objective, which supervises visible corrupted tokens so that token confidence becomes error-aware. Paper: arXiv:2512.15596. Code: zhangshuibai/CDLM.

The same weights are published as Shuibai12138/Open-Dcoder-0.5B-mixture-mdm-step2000, whose card lists the full training configuration and provenance.

Usage

This is a masked diffusion language model with bidirectional attention. Loading it with AutoModelForCausalLM (as an earlier version of this card showed) gives a causal Qwen2 model and wrong outputs. Use the pipeline of the code repository. The pipeline loads the diffusion Qwen2 implementation only when the model name contains open-dcoder, which this repository id does not: pass --model_name Shuibai12138/Open-Dcoder-0.5B-mixture-mdm-step2000 (identical weights), or download this repository to a local directory whose name contains open-dcoder.

Implementation note on the gradient

The 0.5B training code computes the per-token cross-entropy terms with LigerFusedLinearCrossEntropyLoss(reduction="none") from liger-kernel 0.5.8. Its backward pass scales every token's gradient by the upstream gradient of the first token, so the per-token weights of the objective (the 1/|S| and 1/t factors and the noise-term weight) are not applied in the update: each micro-batch receives the first token's weight times the unweighted sum of the per-token gradients, a micro-batch whose first target token is unsupervised receives no gradient, and when the first target token is a replaced one the noise term adds gradient on every valid position, clean tokens included. Logged losses are correct. This model was trained with that code; see training/README.md ("Effective gradient") in the code repository.

Licence

MIT (see the licence field above). Base model: fredzzp/open-dcoder-0.5B, licensed under the Apache License 2.0; this model is a derivative (continued training) of it, and the base model's licence and notices apply to the parts derived from it.

Citation

@inproceedings{zhang2026corrective,
  title     = {Corrective Diffusion Language Models},
  author    = {Zhang, Shuibai and Peng, Fred Zhangzhi and Zhang, Yiheng and Pan, Jin and Chrysos, Grigorios G.},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}
Downloads last month
562
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
Input a message to start chatting with Shuibai12138/CDLM-0.5B.

Model tree for Shuibai12138/CDLM-0.5B

Finetuned
(5)
this model
Quantizations
1 model

Dataset used to train Shuibai12138/CDLM-0.5B

Collection including Shuibai12138/CDLM-0.5B

Paper for Shuibai12138/CDLM-0.5B