The year/Topics/Video generation and world models

Research collection

Video generation and world models

Generating or simulating video: long-form and interactive video generation, controllable video synthesis, and generative video world models.

Papers
106
Research labs
3
Official code
82

150 of 106 papers in this collection

01

Independent research

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

VideoCoCo is an agentic dual-engine framework for physically consistent text-to-video generation. It uses executable Blender code as a process-level chain of thought. A coding agent synthesizes a Blender program from a text prompt, which is executed in a sandbox to produce a deterministic, low-fidelity spatiotemporal draft. A generative video engine then…

Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, et al.
Published
Jul 2026
Citations
0
Code
Not linked
02

Independent research

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

The paper DistillAlign revisits autoregressive video distillation from a distributional perspective. It argues that existing multi-stage pipelines, which separate initialization (e.g., ODE or consistency distillation) from DMD refinement, often have misaligned target distributions. Since DMD is mode-seeking, a good initialization must match the mode…

Jiaxing Li, Kai Zou, Cindy Zhou, Kaichen Huang, et al.
Published
Jul 2026
Citations
0
Code
97 stars
03

Independent research

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

ABot-World-0 is an action-conditioned video world model for real-time, long-horizon closed-loop interaction, deployable on a single NVIDIA RTX 5090 GPU. It uses raw keyboard inputs as a unified control interface for both scene roaming and third-person character control, with reference-character memory for identity consistency. The model is trained on…

Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, et al.
Published
Jul 2026
Citations
0
Code
1.8K stars
04

Independent research

Generative World Renderer at the Speed of Play

AlayaRenderer-Flash is a real-time generative world renderer that accelerates the offline AlayaRenderer from 0.56 FPS to 31.54 FPS, enabling interactive, prompt-controllable gameplay. It reformulates the original renderer into a few-step autoregressive streaming model with three key improvements: autoregressive generation over unbounded G-buffer streams,…

Guixu Lin, Zheng-Hui Huang, Siqi Yang, Ming-Hsuan Yang, et al.
Published
Jul 2026
Citations
0
Code
53 stars
05

Independent research

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

HOMIE is a framework for human-object centric video personalization (HOCVP) that unifies inter-subject (distinct subjects) and intra-subject (multiple references of the same subject) personalization. It addresses limitations of existing methods by integrating Multimodal Large Language Models (MLLMs) while preserving the text encoder, avoiding costly…

Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, et al.
Published
Jul 2026
Citations
0
Code
166 stars
06

Independent research

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

AlayaWorld is an interactive long-horizon video world model that generates 24-fps video at 540p and 720p, built on a 15B video diffusion transformer. It generates short latent chunks autoregressively under camera trajectories and switchable text prompts, using a bounded visual context that combines a persistent sink frame, compressed temporal history,…

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, et al.
Published
Jul 2026
Citations
0
Code
834 stars
07

Independent research

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

LingBot-Video is a DiT-based video pretraining paradigm for embodied intelligence, introduced as the first large-scale open-source Mixture-of-Experts (MoE) video foundation model. It addresses domain mismatch in video generation by using a sparse MoE framework for better capacity-efficiency trade-off, a data profiling engine that augments internet videos…

Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, et al.
Published
Jul 2026
Citations
3
Code
908 stars
08

Independent research

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld is a full-stack, open-source framework for building interactive generative worlds, fine-tuned from LTX-2.3. It addresses four key challenges: control, consistency, stability, and runtime. For control, it combines a 3D cache rendered along the camera trajectory with AdaLN-style camera modulation, and supports prompt-driven actions via a…

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, et al.
Published
Jul 2026
Citations
1
Code
Not linked
09

Independent research

Vidu S1: A Real-Time Interactive Video Generation Model

Vidu S1 is a real-time interactive video generation model that enables users to control digital characters via voice instructions during generation, supporting infinite-length video without blurring or drift. Built with TurboDiffusion and TurboServe, it outputs 540p video at up to 42 FPS on consumer GPUs. The model uses a three-stage training pipeline:…

Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, et al.
Published
Jul 2026
Citations
2
Code
240 stars
10

Independent research

Orca: The World is in Your Mind

Orca, developed by the Beijing Academy of Artificial Intelligence, is a general world foundation model that learns a unified world latent space from multimodal signals (vision and language) using Next-State-Prediction modeling. It employs two complementary learning paradigms: unconscious learning captures dense natural state transitions from continuous…

Yihao Wang, Yuheng Ji, Mingyu Cao, Yanqing Shen, et al.
Published
Jun 2026
Citations
0
Code
906 stars
11

Independent research

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

LiveEdit is a novel streaming video editing framework that performs causal, chunk-by-chunk editing with high fidelity and ultra-low latency. It addresses two core issues: attention distribution shift and spatial-temporal token redundancy. The method uses a three-stage distillation pipeline: Stage 1 tunes a bidirectional DiT for editing, Stage 2 transitions…

Xinyu Wang, Chongbo Zhao, Fangneng Zhan, Yue Ma
Published
Jun 2026
Citations
2
Code
147 stars
12

Independent research

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

DomainShuttle is a novel framework for open-domain subject-driven text-to-video (S2V) generation, addressing both in-domain (high subject fidelity) and cross-domain (flexible adaptation of subject-irrelevant features) scenarios. It introduces three key components: Domain-MoT, which decouples video and reference features and uses domain-aware AdaLN for…

Nan Chen, Yiyang Cai, Rongchang Xie, Junwen Pan, et al.
Published
Jun 2026
Citations
1
Code
165 stars
13

Independent research

Looped World Models

The paper introduces Looped World Models (LoopWM), the first looped transformer architecture for world modeling, addressing the tension between deep computation for faithful long-horizon simulation and the high cost and error accumulation of deep models. LoopWM iteratively refines latent environment states through a parameter-shared transformer block with…

Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang, Jinrui Zeng, et al.
Published
Jun 2026
Citations
1
Code
Not linked
14

Independent research

DreamX-World 1.0: A General-Purpose Interactive World Model

DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation, supporting camera navigation, revisits, and promptable events across photorealistic, game-style, and stylized domains. It uses a data engine combining Unreal Engine rendering, gameplay recordings, and real-world videos. The model…

DreamX Team, Yancheng Bai, Rui Chen, Xiangxiang Chu, et al.
Published
Jun 2026
Citations
5
Code
748 stars
15

Independent research

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

OmniDirector is a framework for cloning camera motion from reference videos to animate source images, supporting multi-shot sequences without requiring cross-paired training data. It introduces a 'camera grid' representation, which renders camera parameters as a grid motion video within an empty 3D scene, decoupling camera motion from content and enabling…

Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, et al.
Published
Jun 2026
Citations
5
Code
79 stars
16

Z.ai / GLM

SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

SCAIL-2 is an end-to-end framework for controlled character animation that bypasses intermediate representations like pose skeletons or masked backgrounds, which cause information loss. It directly concatenates driving videos to the sequence, allowing the model to capture all visual information. To address the lack of end-to-end data, the authors unify…

Wenhao Yan, Fengjia Guo, Zhuoyi Yang, Jie Tang
Published
Jun 2026
Citations
1
Code
1.1K stars
17

Independent research

Latent Spatial Memory for Video World Models

This paper introduces latent spatial memory, a persistent 3D cache for video world models that stores scene information directly in the diffusion latent space, avoiding the pixel-space round trip of RGB point-cloud memories. The authors propose Mirage, a framework that constructs the memory by lifting latent tokens into 3D via depth-guided back-projection…

Weijie Wang, Haoyu Zhao, Yifan Yang, Feng Chen, et al.
Published
Jun 2026
Citations
2
Code
297 stars
18

NVIDIA

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA introduces Cosmos 3, a family of omnimodal world models that jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. It subsumes vision-language models, video generators, world simulators, and world-action models into a single framework, supporting flexible input-output…

NVIDIA, :, Aditi, Niket Agarwal, et al.
Published
Jun 2026
Citations
36
Code
Not linked
19

arXiv.org

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

minWM is a full-stack open-source framework for converting bidirectional text-to-video (T2V) or text-and-image-to-video (TI2V) diffusion foundation models into camera-controllable, few-step autoregressive (AR) world models for real-time interaction. The pipeline has two phases: first, fine-tuning the bidirectional model with camera control via PRoPE…

Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou, et al.
Published
May 2026
Citations
6
Code
766 stars
20

NVIDIA

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

The paper introduces Gamma-World, a generative multi-agent world model for interactive simulation that scales beyond two players. It addresses limitations of prior work like Solaris, which uses dense attention and learned per-slot identities, by proposing two key innovations. First, Simplex Rotary Agent Encoding extends 3D RoPE with an agent axis,…

Fangfu Liu, Kai He, Tianchang Shen, Tianshi Cao, et al.
Published
May 2026
Citations
2
Code
Not linked
21

arXiv.org

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation

WBENCH is a comprehensive multi-turn benchmark for evaluating interactive video world models across five dimensions: video quality, setting adherence, interaction adherence, consistency, and physics compliance. It contains 289 test cases and 1,058 interaction turns, covering diverse scenes, styles, subjects, and both first- and third-person perspectives.…

Kaining Ying, Hengrui Hu, Siyu Ren, Jiamu Li, et al.
Published
May 2026
Citations
7
Code
175 stars
22

arXiv.org

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

EvalVerse is a comprehensive evaluation framework for professional cinematic video generation that addresses the gap between basic prompt-following and true cinematic quality. It introduces a pipeline-aware taxonomy mirroring the filmmaking workflow (pre-production, production, post-production) with 3 stages, 7 aspects, 18 dimensions, 45 sub-dimensions,…

Songlin Yang, Haobin Zhong, Ruilin Zhang, Xiaotong Zhao, et al.
Published
May 2026
Citations
2
Code
Not linked
23

NVIDIA

LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

LongLive-2.0 is an NVFP4-based parallel infrastructure for long video generation, co-designing training and inference. For training, it introduces Balanced SP, a sequence-parallel autoregressive (AR) training method that pairs clean-history and noisy-target temporal chunks on each GPU, enabling a natural teacher-forcing mask and SP-aware chunked VAE…

Yukang Chen, Luozhou Wang, Wei Huang, Shuai Yang, et al.
Published
May 2026
Citations
11
Code
2.5K stars
24

arXiv.org

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

The paper introduces MIGA, a training-free method for infinite-frame long video generation that builds on frame-level autoregressive frameworks like FIFO-Diffusion. MIGA addresses two key limitations: the training-inference gap and long-term consistency. It proposes a Two-Stage Training-Inference Alignment (TTA) mechanism that reduces the noise span of…

X. Feng, J. Zhu, M. Wu, C. Chen, et al.
Published
May 2026
Citations
1
Code
Not linked
25

arXiv.org

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

FashionChameleon is a real-time and interactive framework for human-garment video customization, enabling users to switch garments during generation while preserving motion coherence. It uses three key techniques: a Teacher Model with In-Context Learning trained on single-garment data to implicitly handle garment switching; Streaming Distillation with…

Quanjian Song, Yefeng Shen, Mengting Chen, Hao Sun, et al.
Published
May 2026
Citations
3
Code
256 stars
26

NVIDIA

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

SANA-WM is a 2.6B-parameter open-source world model for generating one-minute, 720p videos with precise 6-DoF camera control. It uses a hybrid linear diffusion transformer combining frame-wise Gated DeltaNet (GDN) and softmax attention for efficient long-context modeling, a dual-branch camera control (UCPE and Plücker mixing), a two-stage generation…

Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, et al.
Published
May 2026
Citations
13
Code
Not linked
27

arXiv.org

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation

Causal Forcing++ is a scalable pipeline for real-time interactive video generation that distills bidirectional diffusion models into few-step autoregressive (AR) students. It targets frame-wise autoregression with 1–2 sampling steps, a regime where existing initialization strategies fail: ODE distillation with a bidirectional teacher is architecturally…

Min Zhao, Hongzhou Zhu, Kaiwen Zheng, Zihan Zhou, et al.
Published
May 2026
Citations
11
Code
905 stars
28

NVIDIA

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

AnyFlow is a video diffusion distillation framework that enables any-step generation by learning flow-map transitions between arbitrary time pairs, unlike consistency models that degrade with more sampling steps. It uses a two-stage pipeline: forward flow map training (with interpolated timestep conditioning, guidance-fused training, and adaptive loss…

Yuchao Gu, Guian Fang, Yuxin Jiang, Weijia Mao, et al.
Published
May 2026
Citations
6
Code
406 stars
29

arXiv.org

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

CollabVR is a closed-loop framework for video reasoning that couples a Vision-Language Model (VLM) with a Video Generation Model (VGM) at step-level granularity. It addresses two failure modes of VGMs: long-horizon drift and mid-clip simulation errors. The VLM plans the immediate next action, verifies the generated clip, and folds the diagnosis into the…

Joowon Kim, Seungho Shin, Joonhyung Park, Eunho Yang
Published
May 2026
Citations
1
Code
9 stars
30

arXiv.org

Stream-T1: Test-Time Scaling for Streaming Video Generation

Stream-T1 is a Test-Time Scaling (TTS) framework for streaming video generation, addressing the high costs and lack of temporal guidance in existing diffusion-based TTS methods. It leverages chunk-level synthesis and few denoising steps to reduce computational overhead. The framework comprises three components: Stream-Scaled Noise Propagation, which…

Yijing Tu, Shaojin Wu, Mengqi Huang, Wenchuan Wang, et al.
Published
May 2026
Citations
2
Code
37 stars
31

arXiv.org

Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation

Stream-R1 is a framework for distilling autoregressive streaming video diffusion models, addressing limitations in existing distribution matching distillation (DMD) methods that treat all rollouts, frames, and pixels equally. It introduces two concepts: Inter-Reliability (varying reliability of supervision across rollouts) and Intra-Perplexity (varying…

Bin Wu, Mengqi Huang, Shaojin Wu, Weinan Jia, et al.
Published
May 2026
Citations
2
Code
54 stars
32

ACM Transactions on Graphics

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

UniVidX is a unified multimodal framework for versatile video generation that repurposes video diffusion model (VDM) priors to handle diverse tasks within a single model. It addresses limitations of existing approaches that train separate models for fixed input-output mappings, ignoring cross-modal correlations. UniVidX introduces three key designs:…

Houyuan Chen, Hong Li, Xianghao Kong, Tianrui Zhu, et al.
Published
May 2026
Citations
0
Code
249 stars
33

arXiv.org

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

HERMES++ is a unified driving world model that integrates 3D scene understanding and future geometry prediction within a single framework. It uses a Bird's-Eye View (BEV) representation to consolidate multi-view spatial information, making it compatible with Large Language Models (LLMs). The model introduces LLM-enhanced world queries for knowledge…

Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, et al.
Published
Apr 2026
Citations
3
Code
69 stars
34

arXiv.org

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

This roadmap paper argues that visual generation must evolve from appearance synthesis to intelligent visual generation, grounded in structure, dynamics, and causal relations. It proposes a five-level taxonomy—Atomic, Conditional, In-Context, Agentic, and World-Modeling Generation—to organize progress from passive rendering to interactive, world-aware…

Keming Wu, Zuhao Yang, Kaichen Zhang, Shizun Wang, et al.
Published
Apr 2026
Citations
4
Code
128 stars
35

arXiv.org

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

CoInteract is an end-to-end framework for speech-driven human-object interaction (HOI) video synthesis, conditioned on a person reference image, a product reference image, text prompts, and speech audio. It addresses two common failures in diffusion models: structural instability in hands and faces, and physically implausible contact (e.g., hand-object…

Xiangyang Luo, Xiaozhe Xin, Tao Feng, Xu Guo, et al.
Published
Apr 2026
Citations
0
Code
165 stars
36

arXiv.org

Seedance 2.0: Advancing Video Generation for World Complexity

Seedance 2.0, released by ByteDance in early February 2026, is a native multimodal audio-video generation model that supports text, image, audio, and video inputs. It generates 4-15 second clips at 480p/720p, with a Fast version for low latency. The model excels in real-world complexity, multimodal reference and editing, high-fidelity binaural audio, and…

Team Seedance, De Chen, Liyang Chen, Xin Chen, et al.
Published
Apr 2026
Citations
78
Code
Not linked
37

arXiv.org

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

OmniShow is an end-to-end framework for Human-Object Interaction Video Generation (HOIVG), synthesizing videos conditioned on text, reference images, audio, and pose. It introduces Unified Channel-wise Conditioning to inject reference images and pose via channel concatenation with pseudo-frame tokens, and Gated Local-Context Attention for precise…

Donghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li, et al.
Published
Apr 2026
Citations
3
Code
464 stars
38

arXiv.org

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

The paper introduces NUMINA, a training-free framework for improving numerical alignment in text-to-video diffusion models, which often fail to generate the correct number of objects specified in prompts. NUMINA uses an identify-then-guide paradigm: first, it identifies prompt-layout inconsistencies by selecting discriminative self- and cross-attention…

Zhengyang Sun, Yu Chen, Xin Zhou, Xiaofan Li, et al.
Published
Apr 2026
Citations
0
Code
68 stars
39

arXiv.org

LPM 1.0: Video-based Character Performance Model

LPM 1.0 is a video-based character performance model that generates identity-consistent conversational videos in real time. It addresses the 'performance trilemma'—the challenge of achieving high expressiveness, real-time inference, and long-horizon identity stability simultaneously. The system comprises a 17B-parameter Diffusion Transformer (Base LPM)…

Ailing Zeng, Casper Yang, Chauncey Ge, Eddie Zhang, et al.
Published
Apr 2026
Citations
5
Code
361 stars
40

arXiv.org

Generative World Renderer

The paper introduces a large-scale dataset for generative world rendering, curated from two AAA games (Cyberpunk 2077 and Black Myth: Wukong) to address the domain gap in inverse and forward rendering. The dataset includes over 4 million frames at 720p/30 FPS, with synchronized RGB and five G-buffer channels (depth, normals, albedo, metallic, roughness),…

Zheng-Hui Huang, Zhixiang Wang, Jiaming Tan, Ruihan Yu, et al.
Published
Apr 2026
Citations
1
Code
709 stars
41

arXiv.org

VOID: Video Object and Interaction Deletion

VOID is a video object removal framework that generates physically plausible counterfactual videos when an object is removed, addressing limitations of existing methods that only handle photometric effects like shadows. It uses a two-pass approach: first, a video diffusion model (CogVideoX) synthesizes a counterfactual trajectory guided by a quadmask,…

Saman Motamed, William Harvey, Benjamin Klein, Luc Van Gool, et al.
Published
Apr 2026
Citations
4
Code
2K stars
42

Google DeepMind

VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward

VGGRPO is a framework for geometry-aware post-training of video diffusion models, addressing geometric drift and unstable camera motion. It introduces a Latent Geometry Model (LGM) that stitches video diffusion latents to a geometry foundation model (e.g., Any4D) via a lightweight connector, enabling direct prediction of 4D scene geometry (camera poses,…

Zhaochong An, Orest Kupyn, Théo Uscidda, Andrea Colaco, et al.
Published
Mar 2026
Citations
14
Code
Not linked
43

arXiv.org

ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling

ShotStream is a novel causal multi-shot video generation architecture that enables interactive storytelling and real-time synthesis at 16 FPS on a single GPU. It reformulates multi-shot generation as a next-shot prediction task, allowing users to guide narratives via streaming prompts. The method first fine-tunes a text-to-video model into a bidirectional…

Yawen Luo, Xiaoyu Shi, Junhao Zhuang, Yutian Chen, et al.
Published
Mar 2026
Citations
14
Code
177 stars
44

arXiv.org

Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

The paper introduces Hybrid Memory, a paradigm for video world models that requires maintaining static background consistency while tracking dynamic subjects during out-of-view intervals. The authors construct HM-World, a large-scale dataset of 59K high-fidelity clips with 17 scenes, 49 subjects, and designed exit-entry events, and propose HyDRA, a memory…

Kaijin Chen, Dingkang Liang, Xin Zhou, Yikang Ding, et al.
Published
Mar 2026
Citations
11
Code
269 stars
45

arXiv.org

Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells

Lingshu-Cell is a masked discrete diffusion model (MDDM) for generative modeling of single-cell transcriptomics, introduced by Alibaba DAMO Academy. It models transcriptomic state distributions across ~18,000 genes without prior gene selection, operating directly in a discrete token space compatible with sparse, non-sequential scRNA-seq data. The model…

Han Zhang, Guo-Hua Yuan, Chaohao Yuan, Tingyang Xu, et al.
Published
Mar 2026
Citations
3
Code
Not linked
46

arXiv.org

WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG

WildWorld is a large-scale, action-conditioned world modeling dataset automatically collected from the AAA game Monster Hunter: Wilds. It contains over 108 million frames with more than 450 actions (movement, attacks, skill casting) and per-frame annotations including character skeletons, world states, camera poses, and depth maps. The dataset addresses…

Zhen Li, Zian Meng, Shuwei Shi, Wenshuo Peng, et al.
Published
Mar 2026
Citations
6
Code
420 stars
47

arXiv.org

Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

Omni-WorldBench is a new benchmark for evaluating the interactive response capabilities of video-based world models, addressing the gap left by existing benchmarks that focus on visual fidelity or static 3D reconstruction. It comprises Omni-WorldSuite, a set of 1,068 prompts with initial frames and optional camera trajectories, organized into three…

Meiqi Wu, Zhixin Cai, Fufangchen Zhao, Xiaokun Feng, et al.
Published
Mar 2026
Citations
5
Code
107 stars
48

arXiv.org

Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model

daVinci-MagiHuman is an open-source audio-video generative foundation model for human-centric generation, jointly producing synchronized video and audio via a single-stream Transformer that processes text, video, and audio in a unified token sequence using self-attention only. This design avoids multi-stream complexity and supports multilingual generation…

SII-GAIR, Sand. ai, :, Ethan Chern, et al.
Published
Mar 2026
Citations
12
Code
2.1K stars
49

arXiv.org

SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing

SAMA is a framework for instruction-guided video editing that factorizes the task into semantic anchoring and motion modeling. It uses Semantic Anchoring to predict semantic tokens from sparse anchor frames, enabling instruction-aware structural planning, and Motion Alignment, which pre-trains the backbone on motion-centric pretext tasks (cube inpainting,…

Xinyao Zhang, Wenkai Dong, Yuxin Song, Bo Fang, et al.
Published
Mar 2026
Citations
3
Code
207 stars
50

arXiv.org

3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model

The paper introduces 3DreamBooth, a framework for 3D-aware video customization that generates view-consistent videos of a subject from a few multi-view reference images. It addresses the limitation of existing subject-driven video generation methods that treat subjects as 2D entities, lacking 3D geometry priors. The framework comprises two components:…

Hyun-kyu Ko, Jihyeon Park, Younghyun Kim, Dongheok Park, et al.
Published
Mar 2026
Citations
0
Code
60 stars