Paper 2602.14041
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- Published
- Feb 2026
- Research lab
- Independent
- Citations
- 5
- GitHub
- 481 stars
01 In brief
Summary
BitDance is a scalable autoregressive image generation model that predicts binary visual tokens instead of codebook indices, scaling the vocabulary to 2^256 states.
It introduces a binary diffusion head to sample from this large discrete space, and a next-patch diffusion method for parallel multi-token prediction.
On ImageNet 256×256, BitDance achieves an FID of 1.24, the best among AR models, and outperforms a 1.4B-parameter parallel AR model with only 260M parameters, achieving an 8.7× speedup.
For text-to-image generation, a 14B model trained on large-scale multimodal tokens achieves state-of-the-art scores among AR models (GenEval 0.86, DPG-Bench 88.28, OneIG-EN 0.532, OneIG-ZH 0.512) and generates 1024×1024 images over 30× faster than prior AR models.
The tokenizer uses group-wise LFQ to scale entropy, and the diffusion head models joint bit distributions, avoiding the independence assumption of bit-wise classification.
Next-patch diffusion uses block-wise causal attention to model intra-patch dependencies, improving parallel generation quality.
Ablations confirm the benefits of binary tokens over continuous VAEs, the diffusion head over classification heads, and the patch-wise order and block masks.
02 From the paper
Abstract
We present BitDance, a scalable autoregressive (AR) image generator that predicts binary visual tokens instead of codebook indices. With high-entropy binary latents, BitDance lets each token represent up to $2^{256}$ states, yielding a compact yet highly expressive discrete representation. Sampling from such a huge token space is difficult with standard classification. To resolve this, BitDance uses a binary diffusion head: instead of predicting an index with softmax, it employs continuous-space diffusion to generate the binary tokens. Furthermore, we propose next-patch diffusion, a new decoding method that predicts multiple tokens in parallel with high accuracy, greatly speeding up inference. On ImageNet 256x256, BitDance achieves an FID of 1.24, the best among AR models. With next-patch diffusion, BitDance beats state-of-the-art parallel AR models that use 1.4B parameters, while using 5.4x fewer parameters (260M) and achieving 8.7x speedup. For text-to-image generation, BitDance trains on large-scale multimodal tokens and generates high-resolution, photorealistic images efficiently, showing strong performance and favorable scaling. When generating 1024x1024 images, BitDance achieves a speedup of over 30x compared to prior AR models. We release code and models to facilitate further research on AR foundation models. Code and models are available at: https://github.com/shallowdream204/BitDance.