The year/Independent research

Paper 2604.14148

Seedance 2.0: Advancing Video Generation for World Complexity

Published
Apr 2026
Research lab
Independent
Citations
78
GitHub
Not linked

01 In brief

Summary

Seedance 2.0, released by ByteDance in early February 2026, is a native multimodal audio-video generation model that supports text, image, audio, and video inputs.

It generates 4-15 second clips at 480p/720p, with a Fast version for low latency.

The model excels in real-world complexity, multimodal reference and editing, high-fidelity binaural audio, and productivity applications.

Evaluations on SeedVideoBench 2.0 show Seedance 2.0 ranks first across all dimensions in T2V, I2V, and R2V tasks, outperforming competitors like Kling 3.0, Sora 2 Pro, and Veo 3.1.

On Arena.AI, it ranks #1 on both T2V and I2V leaderboards with Elo scores of 1450 and 1449.

Key improvements include motion stability, instruction following, audio expressiveness, and audio-visual sync.

The model supports 20 of 22 multimodal task types, including exclusive features like visual effects reference and video continuation.

It is accessible on Doubao, Jimeng, and Volcano Engine under model ID doubao-seedance-2-0-260128.

Remaining limitations include minor deformation artifacts, motion plausibility in edge cases, and lip-sync errors in multi-speaker scenes.

The model is designed for professional content production, reducing costs and production cycles across advertising, film, game animation, and commentary videos.

02 From the paper

Abstract

Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and large-scale architecture for multi-modal audio-video joint generation. This allows it to support four input modalities: text, image, audio, and video, by integrating one of the most comprehensive suites of multi-modal content reference and editing capabilities available in the industry to date. It delivers substantial, well-rounded improvements across all key sub-dimensions of video and audio generation. In both expert evaluations and public user tests, the model has demonstrated performance on par with the leading levels in the field. Seedance 2.0 supports direct generation of audio-video content with durations ranging from 4 to 15 seconds, with native output resolutions of 480p and 720p. For multi-modal inputs as reference, its current open platform supports up to 3 video clips, 9 images, and 3 audio clips. In addition, we provide Seedance 2.0 Fast version, an accelerated variant of Seedance 2.0 designed to boost generation speed for low-latency scenarios. Seedance 2.0 has delivered significant improvements to its foundational generation capabilities and multi-modal generation performance, bringing an enhanced creative experience for end users.