model card: link to the write-up
Browse files
README.md
CHANGED
|
@@ -22,6 +22,7 @@ A 210M-parameter text-to-image diffusion transformer trained from scratch on one
|
|
| 22 |
4.2M images at 256² area with aspect-ratio buckets). Frozen FLUX.2 autoencoder (32-channel latent) and frozen
|
| 23 |
flan-t5-base text encoder; the transformer, recipe, data pipeline and evaluation are original.
|
| 24 |
|
|
|
|
| 25 |
- 🎨 Try it in the browser: https://huggingface.co/spaces/ivanmikhnenkov/tinydit
|
| 26 |
- 💻 Code, figures and every design decision with sources: https://github.com/ivanmikhnenkov/tinydit (tag `v1-pretrain`)
|
| 27 |
- 🌐 Author: https://ivanmikhnenkov.com
|
|
|
|
| 22 |
4.2M images at 256² area with aspect-ratio buckets). Frozen FLUX.2 autoencoder (32-channel latent) and frozen
|
| 23 |
flan-t5-base text encoder; the transformer, recipe, data pipeline and evaluation are original.
|
| 24 |
|
| 25 |
+
- 📝 Write-up: https://huggingface.co/blog/ivanmikhnenkov/tinydit-text-to-image-from-scratch-one-gpu
|
| 26 |
- 🎨 Try it in the browser: https://huggingface.co/spaces/ivanmikhnenkov/tinydit
|
| 27 |
- 💻 Code, figures and every design decision with sources: https://github.com/ivanmikhnenkov/tinydit (tag `v1-pretrain`)
|
| 28 |
- 🌐 Author: https://ivanmikhnenkov.com
|