Paper 2509.20427
Seedream 4.0: Toward Next-generation Multimodal Image Generation
- Published
- Sep 2025
- Research lab
- Independent
- Citations
- 226
- GitHub
- Not linked
01 In brief
Summary
Seedream 4.0 is a multimodal image generation system by ByteDance Seed that unifies text-to-image (T2I) synthesis, image editing, and multi-image composition in a single framework.
It uses an efficient diffusion transformer (DiT) with a high-compression VAE, reducing image tokens and enabling native 1K-4K resolution generation.
The model is pretrained on billions of text-image pairs, with a redesigned data pipeline for knowledge-centric concepts like formulas and instructional content.
Post-training integrates a VLM (based on Seed1.5-VL) for prompt engineering, and joint training of T2I and editing tasks via continuing training, SFT, and RLHF.
Inference acceleration combines adversarial distillation, distribution matching, quantization, and speculative decoding, achieving up to 1.4 seconds for a 2K image.
Evaluations show Seedream 4.0 ranks first on the Artificial Analysis Arena for both T2I and image editing (as of 09/18/2025).
It supports precise editing, multi-image reference and output, visual signal control, in-context reasoning, advanced text rendering, and adaptive aspect ratios up to 4K.
A scaled version, Seedream 4.5, outperforms 4.0 in all tasks.
Both models are accessible on Volcano Engine.
02 From the paper
Abstract
We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image composition within a single framework. We develop a highly efficient diffusion transformer with a powerful VAE which also can reduce the number of image tokens considerably. This allows for efficient training of our model, and enables it to fast generate native high-resolution images (e.g., 1K-4K). Seedream 4.0 is pretrained on billions of text-image pairs spanning diverse taxonomies and knowledge-centric concepts. Comprehensive data collection across hundreds of vertical scenarios, coupled with optimized strategies, ensures stable and large-scale training, with strong generalization. By incorporating a carefully fine-tuned VLM model, we perform multi-modal post-training for training both T2I and image editing tasks jointly. For inference acceleration, we integrate adversarial distillation, distribution matching, and quantization, as well as speculative decoding. It achieves an inference time of up to 1.8 seconds for generating a 2K image (without a LLM/VLM as PE model). Comprehensive evaluations reveal that Seedream 4.0 can achieve state-of-the-art results on both T2I and multimodal image editing. In particular, it demonstrates exceptional multimodal capabilities in complex tasks, including precise image editing and in-context reasoning, and also allows for multi-image reference, and can generate multiple output images. This extends traditional T2I systems into an more interactive and multidimensional creative tool, pushing the boundary of generative AI for both creativity and professional applications. We further scale our model and data as Seedream 4.5. Seedream 4.0 and Seedream 4.5 are accessible on Volcano Engine https://www.volcengine.com/experience/ark?launch=seedream.