The year/Topics/Video and spatial AI

Topic area

Video and spatial AI

Every collection across video and spatial ai.

Papers
164
Research labs
4
Official code
135

101150 of 164 papers in this topic area

101

International Conference on 3D Vision

CaricatureGS: Exaggerating 3D Gaussian Splatting Faces With Gaussian Curvature

CaricatureGS introduces a method for creating photorealistic, controllable 3D caricature avatars by combining curvature-based geometric deformation with 3D Gaussian Splatting (3DGS). The pipeline starts with a multiview video, extracts a FLAME mesh, and solves a curvature-weighted Poisson equation to produce an exaggerated mesh. To train the 3DGS,…

Eldad Matmon, Amit Bracha, Noam Rotstein, Ron Kimmel
Published
Jan 2026
Citations
0
Code
Not linked
102

arXiv.org

DreamID-V:Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer

The paper introduces DreamID-V, a Diffusion Transformer (DiT)-based framework for high-fidelity video face swapping (VFS). It addresses the gap between image face swapping (IFS) and VFS by proposing a data pipeline, SyncID-Pipe, which pre-trains an Identity-Anchored Video Synthesizer (IVS) to generate synthetic videos, combined with IFS models to create…

Xu Guo, Fulong Ye, Xinghui Li, Pengqi Tu, et al.
Published
Jan 2026
Citations
7
Code
668 stars
103

arXiv.org

Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation

Avatar Forcing is a framework for real-time interactive head avatar generation that models user-avatar interactions using diffusion forcing. It addresses two key challenges: real-time motion generation under causal constraints and learning expressive reactions without labeled data. The framework processes multimodal user inputs (audio and motion) with low…

Taekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon, et al.
Published
Jan 2026
Citations
13
Code
342 stars
104

arXiv.org

NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos

NeoVerse is a 4D world model that reconstructs dynamic 4D Gaussian Splatting (4DGS) from monocular videos in a feed-forward, pose-free manner, enabling novel-trajectory video generation and downstream applications. It addresses scalability limitations of prior methods by avoiding expensive multi-view data and offline preprocessing. Key innovations include…

Yuxue Yang, Lue Fan, Ziqi Shi, Junran Peng, et al.
Published
Jan 2026
Citations
28
Code
647 stars
105

arXiv.org

Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation

The paper addresses visual ungrounded hallucinations in Multimodal Large Language Models (MLLMs), which over-rely on language priors when processing counterfactual videos that defy common sense. To mitigate this, the authors introduce DualityForge, a framework using diffusion-based controllable video editing to transform real-world videos into…

Zhe Huang, Hao Wen, Aiming Hao, Bingze Song, et al.
Published
Dec 2025
Citations
6
Code
55 stars
106

arXiv.org

LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation

The paper introduces LiveTalk, a real-time multimodal interactive video diffusion system. It addresses the high inference cost of diffusion models by distilling a bidirectional, many-step model into a causal, 4-step autoregressive one. The authors identify that the leading on-policy distillation method, Self Forcing, suffers from training instability and…

Ethan Chern, Zhulin Hu, Bohao Tang, Jiadi Su, et al.
Published
Dec 2025
Citations
7
Code
331 stars
107

arXiv.org

Yume-1.5: A Text-Controlled Interactive World Generation Model

Yume1.5 is a framework for generating interactive, continuous virtual worlds from a single image or text prompt, with keyboard-based control for person and camera movement. It addresses limitations in existing video diffusion models, such as limited generalizability, high latency, and insufficient text control. The framework introduces three core…

Xiaofeng Mao, Zhen Li, Chuanhao Li, Xiaojie Xu, et al.
Published
Dec 2025
Citations
52
Code
681 stars
108

arXiv.org

SemanticGen: Video Generation in Semantic Space

SemanticGen is a novel video generation framework that operates in a compact semantic space rather than directly in the VAE latent space. It uses a two-stage process: first, a diffusion model generates compressed semantic video features (using Qwen-2.5-VL as the semantic encoder) that define the global layout; second, another diffusion model generates VAE…

Jianhong Bai, Xiaoshi Wu, Xintao Wang, Xiao Fu, et al.
Published
Dec 2025
Citations
7
Code
Not linked
109

Volume 1

LongVideoAgent: Multi-Agent Reasoning with Long Videos

LongVideoAgent is a multi-agent framework for long-video question answering. A master LLM coordinates a grounding agent to localize question-relevant segments and a vision agent to extract targeted visual observations. The master agent plans with a step limit and is trained with reinforcement learning (GRPO) to encourage concise, correct, and efficient…

Runtao Liu, Ziyi Liu, Jiaqi Tang, Yue Ma, et al.
Published
Dec 2025
Citations
21
Code
126 stars
110

Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers

MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation

MACE-Dance is a music-driven dance video generation framework using a cascaded Mixture-of-Experts (MoE) design, decoupling the task into a Motion Expert and an Appearance Expert. The Motion Expert generates 3D SMPL motion from music using a Diffusion Model with a BiMamba-Transformer hybrid architecture and Guidance-Free Training (GFT), achieving…

Kaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng, et al.
Published
Dec 2025
Citations
9
Code
108 stars
111

arXiv.org

Region-Constraint In-Context Generation for Instructional Video Editing

ReCo is a novel framework for instruction-based video editing that uses in-context generation with region constraints. It concatenates source and target videos for joint denoising and introduces two regularization terms: latent-space regularization increases latent discrepancy in editing regions while reducing it in non-editing areas, and attention-space…

Zhongwei Zhang, Fuchen Long, Wei Li, Zhaofan Qiu, et al.
Published
Dec 2025
Citations
14
Code
174 stars
112

Independent research

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

InsertAnywhere is a framework for video object insertion (VOI) that addresses limitations in 4D scene understanding and optical effects. It uses a two-stage pipeline: first, a 4D-aware mask generation module reconstructs the video into a 4D scene, allowing users to anchor an object's 3D pose in one frame, then propagates it via scene flow tracking to…

Hoiyeong Jin, Hyojin Jang, Junha Hyung, Jeongho Kim, et al.
Published
Dec 2025
Citations
0
Code
89 stars
113

arXiv.org

Kling-Omni Technical Report

Kling-Omni is a generalist generative framework from Kuaishou Technology that unifies video generation, editing, and reasoning into a single end-to-end system. It introduces Multi-modal Visual Language (MVL) as an interaction paradigm, combining text, images, and videos into a unified representation. The architecture includes a Prompt Enhancer (PE) based…

Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du, et al.
Published
Dec 2025
Citations
41
Code
Not linked
114

arXiv.org

TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times

TurboDiffusion is a video generation acceleration framework that achieves 100–200× end-to-end speedup while maintaining video quality. It combines four main techniques: low-bit SageAttention for attention acceleration, Sparse-Linear Attention (SLA) for sparse attention, rCM for step distillation, and W8A8 quantization for linear layers. Training involves…

Jintao Zhang, Kaiwen Zheng, Kai Jiang, Haoxu Wang, et al.
Published
Dec 2025
Citations
28
Code
3.6K stars
115

arXiv.org

MMGR: Multi-Modal Generative Reasoning

The paper introduces MMGR (Multi-Modal Generative Reasoning), a benchmark suite to evaluate the reasoning capabilities of video and image generation models across five core abilities: Physical, Logical, 3D Spatial, 2D Spatial, and Temporal reasoning. It comprises three domains: Abstract Reasoning (Maze, Sudoku, ARC-AGI, Math), Embodied Navigation (four…

Zefan Cai, Haoyi Qiu, Tianyi Ma, Haozhe Zhao, et al.
Published
Dec 2025
Citations
9
Code
Not linked
116

arXiv.org

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

WorldPlay is a real-time interactive world model that generates streaming 720p video at 24 FPS while maintaining long-term geometric consistency. It addresses the trade-off between speed and memory in existing methods. The model uses three key components: Dual Action Representation combining discrete keys and continuous camera poses for robust control;…

Wenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu, et al.
Published
Dec 2025
Citations
97
Code
1.6K stars
117

arXiv.org

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

LongVie 2 is an end-to-end autoregressive framework for controllable ultra-long video generation, extending pretrained diffusion backbones (Wan2.1-I2V-14B) into a video world model. It is trained in three progressive stages: (1) multi-modal guidance integrating dense (depth maps) and sparse (point maps) control signals via a ControlNet-style architecture…

Jianxiong Gao, Zhaoxi Chen, Xian Liu, Junhao Zhuang, et al.
Published
Dec 2025
Citations
6
Code
335 stars
118

Independent research

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?

VideoASMR-Bench is a new benchmark for evaluating AI-generated ASMR videos, focusing on fine-grained audio-visual perception and sensory immersion. It includes 1,500 real ASMR videos from social media and 2,235 synthetic videos from nine video generation models (VGMs) under four settings. The benchmark introduces an adversarial evaluation framework where…

Jiaqi Wang, Weijia Wu, Yi Zhan, Rui Zhao, et al.
Published
Dec 2025
Citations
2
Code
23 stars
119

arXiv.org

StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation

StereoWorld is an end-to-end diffusion-based framework that converts monocular videos into high-fidelity stereo videos by adapting a pretrained video generator. It conditions the model on the left-view video and uses a geometry-aware regularization combining disparity and depth supervision to ensure 3D structural fidelity. A spatio-temporal tiling scheme…

Ke Xing, Xiaojie Jin, Longfei Li, Yuyang Yin, et al.
Published
Dec 2025
Citations
2
Code
Not linked
120

arXiv.org

Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance

Wan-Move is a framework for motion-controllable video generation that enhances existing image-to-video (I2V) models without adding auxiliary modules. It represents object motion using dense point trajectories, which are transferred into latent space and used to replicate first-frame features along each trajectory, creating a motion-aware condition feature…

Ruihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang, et al.
Published
Dec 2025
Citations
37
Code
650 stars
121

arXiv.org

Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform

Visionary is an open, web-native platform for real-time rendering of 3D Gaussian Splatting (3DGS) and meshes, built on WebGPU and ONNX. It introduces a Gaussian Generator contract, a standardized ONNX I/O interface that allows plug-and-play integration of various 3DGS algorithms (e.g., MLP-based 3DGS, 4DGS, neural avatars) without modifying the renderer.…

Yuning Gong, Yifei Liu, Yifan Zhan, Muyao Niu, et al.
Published
Dec 2025
Citations
2
Code
517 stars
122

arXiv.org

EgoX: Egocentric Video Generation from a Single Exocentric Video

EgoX is a novel framework that generates egocentric (first-person) videos from a single exocentric (third-person) video input. It leverages a pretrained video diffusion model (Wan 2.1 14B) with lightweight LoRA adaptation, avoiding the need for additional inputs like multiple views or initial frames. The method lifts the exocentric video into a 3D point…

Taewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park, et al.
Published
Dec 2025
Citations
6
Code
741 stars
123

arXiv.org

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

Live Avatar is an algorithm-system co-designed framework enabling real-time, streaming, and infinite-length audio-driven avatar generation using a 14-billion-parameter diffusion model. It addresses two key challenges: long-horizon consistency and the real-time-fidelity trade-off. The algorithm side uses a two-stage pipeline (Diffusion Forcing pretraining…

Yubo Huang, Hailong Guo, Fangtai Wu, Weiqiang Wang, et al.
Published
Dec 2025
Citations
0
Code
2.3K stars
124

arXiv.org

MultiShotMaster: A Controllable Multi-Shot Video Generation Framework

MultiShotMaster is a framework for controllable multi-shot video generation, extending a pretrained single-shot text-to-video model. It introduces two RoPE variants: Multi-Shot Narrative RoPE, which applies phase shifts at shot boundaries for flexible shot arrangement while preserving narrative order, and Spatiotemporal Position-Aware RoPE, which enables…

Qinghe Wang, Xiaoyu Shi, Baolu Li, Weikang Bian, et al.
Published
Dec 2025
Citations
28
Code
174 stars
125

arXiv.org

What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards

The paper introduces NewtonRewards, a physics-grounded post-training framework for video generation that enforces Newton's laws of motion using verifiable rewards. It addresses the issue that video diffusion models often produce visually realistic but physically implausible motion. The method uses optical flow as a proxy for velocity and visual features…

Minh-Quan Le, Yuanzhi Zhu, Vicky Kalogeiton, Dimitris Samaras
Published
Nov 2025
Citations
18
Code
16 stars
126

arXiv.org

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

LongVT is an end-to-end agentic framework that enables large multimodal models (LMMs) to reason over long videos by interleaving multimodal Chain-of-Tool-Thought (iMCoTT) with native video cropping tool calls. It mimics human global-to-local viewing: the model first skims the video, then invokes a crop_video tool to inspect specific temporal windows, and…

Zuhao Yang, Sudong Wang, Kaichen Zhang, Keming Wu, et al.
Published
Nov 2025
Citations
51
Code
259 stars
127

arXiv.org

V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models

V-ReasonBench is a benchmark for evaluating reasoning in generative video models under the Chain-of-Frame paradigm, where the final frame represents the model's answer. It covers four reasoning dimensions: structured problem-solving (arithmetic, code execution, Sudoku, Tic-Tac-Toe), spatial cognition (shape fitting, visual symmetry, color connection),…

Yang Luo, Xuanlei Zhao, Baijiong Lin, Lingting Zhu, et al.
Published
Nov 2025
Citations
17
Code
36 stars
128

Meta AI

SAM 3D: 3Dfy Anything in Images

SAM 3D is a generative model for 3D object reconstruction from a single image, predicting geometry, texture, and layout. It excels in natural images with occlusion and clutter, using a human- and model-in-the-loop pipeline to create large-scale 3D annotation data. The model uses a multi-stage training framework: synthetic pretraining on 2.7M meshes…

SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, et al.
Published
Nov 2025
Citations
197
Code
7.2K stars
129

arXiv.org

First Frame Is the Place to Go for Video Content Customization

This paper introduces FFGo, a lightweight add-on for video generation models that enables multi-reference video content customization without architectural changes or large-scale fine-tuning. The authors discover that pre-trained video models treat the first frame as a conceptual memory buffer, storing visual entities for later reuse. FFGo leverages this…

Jingxi Chen, Zongxia Li, Zhichao Liu, Guangyao Shi, et al.
Published
Nov 2025
Citations
8
Code
194 stars
130

arXiv.org

Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks

This paper introduces VR-Bench, a benchmark for evaluating the reasoning abilities of video generation models through maze-solving tasks. It comprises 7,920 procedurally generated videos across five maze types (Regular, Irregular, 3D, Trapfield, Sokoban) with varying difficulty and textures. The authors propose a 'reasoning via video' paradigm, where…

Cheng Yang, Haiyuan Wan, Yiran Peng, Xin Cheng, et al.
Published
Nov 2025
Citations
15
Code
2 stars
131

arXiv.org

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

Part-X-MLLM is a native 3D multimodal large language model that unifies diverse 3D tasks by formulating them as programs in a structured, executable grammar. Given an RGB point cloud and a natural language prompt, the model autoregressively generates a token sequence encoding part-level bounding boxes, semantic descriptions, and edit commands. This…

Chunshi Wang, Junliang Ye, Yunhan Yang, Yang Li, et al.
Published
Nov 2025
Citations
5
Code
119 stars
132

arXiv.org

Depth Anything 3: Recovering the Visual Space from Any Views

Depth Anything 3 (DA3) is a model that predicts spatially consistent geometry from any number of images, with or without known camera poses. It uses a single plain transformer (e.g., vanilla DINOv2) as backbone, with an input-adaptive cross-view self-attention mechanism and a dual-DPT head that jointly outputs depth and ray maps. A depth-ray representation…

Haotong Lin, Sili Chen, Junhao Liew, Donny Y. Chen, et al.
Published
Nov 2025
Citations
499
Code
6.1K stars
133

arXiv.org

PAN: A World Model for General, Interactable, and Long-Horizon World Simulation

PAN is a general, interactable, and long-horizon world model that predicts future world states via video simulation conditioned on history and natural language actions. It uses the Generative Latent Prediction (GLP) architecture, combining an autoregressive LLM-based latent dynamics backbone (Qwen2.5-VL-7B) with a video diffusion decoder (Wan2.1-T2V-14B)…

PAN Team, Jiannan Xiang, Yi Gu, Zihan Liu, et al.
Published
Nov 2025
Citations
33
Code
Not linked
134

arXiv.org

Time-to-Move: Training-Free Motion Controlled Video Generation via Dual-Clock Denoising

Time-to-Move (TTM) is a training-free, plug-and-play framework for motion- and appearance-controlled video generation using image-to-video (I2V) diffusion models. It uses crude reference animations (e.g., cut-and-drag or depth-based reprojection) as motion cues, adapting SDEdit's noise injection to video. To preserve appearance, it anchors generation to…

Assaf Singer, Noam Rotstein, Amir Mann, Ron Kimmel, et al.
Published
Nov 2025
Citations
11
Code
370 stars
135

arXiv.org

Visual Spatial Tuning

The paper introduces Visual Spatial Tuning (VST), a framework to enhance the spatial perception and reasoning abilities of Vision-Language Models (VLMs) without adding specialized 3D encoders. VST comprises two datasets: VST-P, with 4.1 million samples across 19 tasks covering single-image, multi-image, and video scenarios, and VST-R, with 135K samples for…

Rui Yang, Ziyu Zhu, Yanwei Li, Jingjia Huang, et al.
Published
Nov 2025
Citations
64
Code
201 stars
136

arXiv.org

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

The paper proposes 'Thinking with Video', a new paradigm using video generation models like Sora-2 for multimodal reasoning, addressing limitations of text- and image-based paradigms. The authors introduce VideoThinkBench, a benchmark with vision-centric tasks (eyeballing puzzles, visual puzzles, ARC-AGI-2, mazes) and text-centric tasks (GSM8K, MATH, MMLU,…

Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, et al.
Published
Nov 2025
Citations
27
Code
316 stars
137

arXiv.org

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

UniAVGen is a unified framework for human-centric joint audio and video generation, addressing limitations in existing methods like poor lip synchronization and semantic inconsistency. It uses a dual-branch architecture with two parallel Diffusion Transformers (DiTs) for video and audio, enabling a cohesive cross-modal latent space. The core innovation is…

Guozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng, et al.
Published
Nov 2025
Citations
28
Code
57 stars
138

arXiv.org

Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

Open-o3-Video is a non-agent framework that integrates explicit spatio-temporal evidence into video reasoning by highlighting key timestamps, objects, and bounding boxes. The authors construct two datasets, STGR-CoT-30k and STGR-RL-36k, combining existing temporal and spatial grounding resources with 5.9k newly annotated spatio-temporal samples. They adopt…

Jiahao Meng, Xiangtai Li, Haochen Wang, Yue Tan, et al.
Published
Oct 2025
Citations
43
Code
159 stars
139

arXiv.org

World-in-World: World Models in a Closed-Loop World

The paper introduces World-in-World, the first open benchmark for evaluating generative world models (WMs) in a closed-loop embodied setting, moving beyond visual quality to task success. It provides a unified online planning strategy and a standardized action API to integrate diverse WMs into four embodied tasks: Active Recognition, Image-Goal Navigation,…

Jiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu, et al.
Published
Oct 2025
Citations
35
Code
181 stars
140

arXiv.org

NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks

Nano3D is a training-free framework for 3D object editing that supports removal, addition, and replacement tasks without requiring masks. It integrates FlowEdit into the TRELLIS pipeline to perform localized edits guided by front-view renderings, and introduces region-aware merging strategies (Voxel/Slat-Merge) to preserve structural fidelity in unedited…

Junliang Ye, Shenghao Xie, Ruowen Zhao, Zhengyi Wang, et al.
Published
Oct 2025
Citations
27
Code
179 stars
141

AAAI Conference on Artificial Intelligence

ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints

The paper introduces ImagerySearch, a test-time search strategy for text-to-video generation that adapts to prompts with long-distance semantic relationships, which are rare in training data and cause performance degradation in imaginative scenarios. ImagerySearch dynamically adjusts the inference search space (SaDSS) and reward function (AIR) based on the…

Meiqi Wu, Jiashu Zhu, Xiaokun Feng, Chubin Chen, et al.
Published
Oct 2025
Citations
12
Code
56 stars
142

arXiv.org

FlashWorld: High-quality 3D Scene Generation within Seconds

FlashWorld is a generative model that creates 3D scenes from a single image or text prompt in seconds, being 10-100x faster than previous methods while achieving superior rendering quality. It shifts from the conventional multi-view-oriented (MV-oriented) paradigm to a 3D-oriented approach that directly produces 3D Gaussian representations during…

Xinyang Li, Tengfei Wang, Zixiao Gu, Shengchuan Zhang, et al.
Published
Oct 2025
Citations
26
Code
834 stars
143

arXiv.org

StreamingVLM: Real-Time Understanding for Infinite Video Streams

StreamingVLM is a framework for real-time understanding of infinite video streams, addressing the limitations of full attention (quadratic cost, poor long-video performance) and sliding window methods (coherence loss or high latency). It maintains a compact KV cache with attention sinks, a short vision window, and a long text window, using contiguous RoPE…

Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He, et al.
Published
Oct 2025
Citations
73
Code
1.1K stars
144

arXiv.org

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

VideoCanvas is a unified framework for arbitrary spatio-temporal video completion, where users specify patches at any spatial location and timestamp, and the model generates a coherent video. The paper identifies that causal video VAEs compress multiple frames into a single latent slot, creating temporal ambiguity, and that zero-padding in video mode…

Minghong Cai, Qiulin Wang, Zongli Ye, Wenze Liu, et al.
Published
Oct 2025
Citations
4
Code
68 stars
145

arXiv.org

Paper2Video: Automatic Video Generation from Scientific Papers

The paper introduces Paper2Video, the first benchmark of 101 research papers paired with author-created presentation videos, slides, and speaker metadata, along with four evaluation metrics: Meta Similarity, PresentArena, PresentQuiz, and IP Memory. It also proposes PaperTalker, a multi-agent framework that generates presentation videos from papers,…

Zeyu Zhu, Kevin Qinghong Lin, Mike Zheng Shou
Published
Oct 2025
Citations
24
Code
2.3K stars
146

arXiv.org

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation

Self-Forcing++ is a method for long-horizon video generation that extends autoregressive diffusion models beyond the training horizon of their teacher models. It addresses quality degradation from error accumulation by generating long self-rollouts (up to 100 seconds), re-injecting noise into these degraded sequences (backward noise initialization), and…

Justin Cui, Jie Wu, Ming Li, Tao Yang, et al.
Published
Oct 2025
Citations
156
Code
268 stars
147

NVIDIA

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

SANA-Video is a small diffusion model for efficient, high-resolution (up to 720×1280) and minute-long video generation, deployable on RTX 5090 GPUs. It uses a Linear DiT with linear attention (O(N) complexity) and a constant-memory KV cache for block linear attention, enabling long videos with fixed memory. Training cost is 12 days on 64 H100 GPUs (1% of…

Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu, et al.
Published
Sep 2025
Citations
82
Code
8.7K stars
148

NVIDIA

LongLive: Real-time Interactive Long Video Generation

LONGLIVE is a frame-level autoregressive (AR) framework for real-time, interactive long video generation, addressing efficiency and quality challenges in diffusion and AR models. It introduces KV-recache to refresh cached states with new prompts for smooth, adherent prompt switches; streaming long tuning to enable train-long-test-long alignment; and…

Shuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao, et al.
Published
Sep 2025
Citations
189
Code
Not linked
149

Google DeepMind

Video models are zero-shot learners and reasoners

This paper investigates whether generative video models, like Veo 3, can act as zero-shot learners and reasoners for general-purpose vision tasks, similar to how LLMs transformed NLP. The authors analyzed 18,384 generated videos across 62 qualitative and 7 quantitative tasks, finding that Veo 3 can solve tasks it wasn't explicitly trained for, including…

Thaddäus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, et al.
Published
Sep 2025
Citations
194
Code
Not linked
150

arXiv.org

OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models

OmniInsert is a unified framework for mask-free video insertion, allowing users to insert single or multiple reference subjects into a source video based on a text prompt. It addresses three key challenges: data scarcity, subject-scene equilibrium, and insertion harmonization. To tackle data scarcity, the authors propose InsertPipe, a data pipeline with…

Jinshu Chen, Xinghui Li, Xu Bai, Tianxiang Ma, et al.
Published
Sep 2025
Citations
9
Code
162 stars