The year/Topics/Video and spatial AI

Topic area

Video and spatial AI

Every collection across video and spatial ai.

Papers
164
Research labs
4
Official code
135

151164 of 164 papers in this topic area

151

arXiv.org

OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling

OmniWorld is a large-scale, multi-domain, multi-modal dataset for 4D world modeling, introduced by Shanghai AI Lab and ZJU. It comprises a self-collected OmniWorld-Game synthetic dataset (96K clips, 18.5M frames, 214+ hours) and curated public datasets from robot, human, and internet domains, totaling over 600K sequences and 300M frames. OmniWorld provides…

Yang Zhou, Yifan Wang, Jianjun Zhou, Wenzheng Chang, et al.
Published
Sep 2025
Citations
43
Code
489 stars
152

arXiv.org

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

HuMo is a unified framework for Human-Centric Video Generation (HCVG) that enables collaborative control from text, reference images, and audio. It addresses two challenges: data scarcity and difficulty in coordinating sub-tasks of subject preservation and audio-visual sync. To overcome data scarcity, HuMo constructs a high-quality dataset with paired…

Liyang Chen, Tianxiang Ma, Jiawei Liu, Bingchuan Li, et al.
Published
Sep 2025
Citations
45
Code
1.3K stars
153

arXiv.org

3D and 4D World Modeling: A Survey

This survey provides the first comprehensive review of 3D and 4D world modeling, addressing the lack of standardized definitions and the fragmented literature that often focuses on 2D generative methods. It establishes precise definitions and a hierarchical taxonomy categorizing methods into video-based (VideoGen), occupancy-based (OccGen), and LiDAR-based…

Lingdong Kong, Yu Yang, Jianbiao Mei, Youquan Liu, et al.
Published
Sep 2025
Citations
65
Code
960 stars
154

arXiv.org

From Editor to Dense Geometry Estimator

FE2E is a framework that adapts a pre-trained image editing model, Step1X-Edit, for monocular dense geometry prediction (depth and normal estimation). The authors argue that editing models, unlike text-to-image generators, possess inherent structural priors that make them more suitable for image-to-image tasks. They introduce three key adaptations: a…

JiYuan Wang, Chunyu Lin, Lei Sun, Rongying Liu, et al.
Published
Sep 2025
Citations
25
Code
243 stars
155

arXiv.org

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding

The paper introduces ELV-Halluc, the first benchmark for evaluating Semantic Aggregation Hallucination (SAH) in long videos. SAH occurs when models correctly perceive frame-level semantics but misattribute them across events, a problem that intensifies with semantic complexity. The benchmark uses event-by-event videos (average 672.4 seconds) and…

Hao Lu, Jiahao Wang, Yaolun Zhang, Ruohui Wang, et al.
Published
Aug 2025
Citations
10
Code
11 stars
156

arXiv.org

Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation

Droplet3D addresses 3D data scarcity by leveraging commonsense priors from videos for 3D generation. The authors introduce Droplet3D-4M, a large-scale dataset of 4 million 3D models, each with an 85-frame 360-degree orbital rendering video and dense multi-view-level text captions averaging 260 words. They also present Droplet3D, a generative model…

Xiaochuan Li, Guoguang Du, Runze Zhang, Liang Jin, et al.
Published
Aug 2025
Citations
2
Code
43 stars
157

arXiv.org

MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds

MeshCoder is a framework that reconstructs 3D objects from point clouds into editable Blender Python scripts. It introduces a set of expressive Blender Python APIs capable of modeling complex geometries beyond simple primitives, including primitives, translation, bridge loops, boolean operations, and arrays. A large-scale paired object-code dataset was…

Bingquan Dai, Li Ray Luo, Qihong Tang, Jie Wang, et al.
Published
Aug 2025
Citations
12
Code
501 stars
158

IEEE International Conference on Computer Vision

LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

LongSplat is a framework for novel view synthesis from casually captured long videos without known camera poses. It jointly optimizes camera poses and 3D Gaussian Splatting (3DGS) to address pose drift, inaccurate geometry initialization, and memory limitations. Key components include incremental joint optimization, a pose estimation module using learned…

Chin-Yang Lin, Cheng Sun, Fu-En Yang, Min-Hung Chen, et al.
Published
Aug 2025
Citations
28
Code
799 stars
159

arXiv.org

4DNeX: Feed-Forward 4D Generative Modeling Made Easy

4DNeX is the first feed-forward framework for generating 4D (dynamic 3D) scene representations from a single image. It fine-tunes a pretrained video diffusion model (Wan2.1) to generate a unified 6D video representation (RGB and XYZ sequences) jointly, enabling efficient end-to-end image-to-4D generation. To address data scarcity, the authors constructed…

Zhaoxi Chen, Tianqi Liu, Long Zhuo, Jiawei Ren, et al.
Published
Aug 2025
Citations
31
Code
840 stars
160

arXiv.org

ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing

ToonComposer is a generative model that unifies the traditionally separate inbetweening and colorization stages of cartoon production into a single post-keyframing stage. Built on the DiT-based video foundation model Wan 2.1, it uses a sparse sketch injection mechanism for precise control from keyframe sketches and a spatial low-rank adapter (SLRA) to…

Lingen Li, Guangzhi Wang, Zhaoyang Zhang, Yaowei Li, et al.
Published
Aug 2025
Citations
8
Code
584 stars
161

arXiv.org

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

The paper introduces M3-Agent, a multimodal agent framework with long-term memory that processes real-time video and audio to build episodic and semantic memories, organized in an entity-centric multimodal graph. It uses reinforcement learning for multi-turn reasoning and iterative memory retrieval. The authors also present M3-Bench, a long-video question…

Lin Long, Yichen He, Wentao Ye, Yiyuan Pan, et al.
Published
Aug 2025
Citations
64
Code
1.4K stars
162

arXiv.org

Matrix-3D: Omnidirectional Explorable 3D World Generation

Matrix-3D is a framework for generating omnidirectional, explorable 3D worlds from a single image or text prompt. It uses panoramic representations to overcome the limited field of view of perspective-based methods. The pipeline first generates a panorama image, then a trajectory-guided panoramic video using a diffusion model conditioned on scene mesh…

Zhongqi Yang, Wenhang Ge, Yuqi Li, Jiaqi Chen, et al.
Published
Aug 2025
Citations
28
Code
777 stars
163

AAAI Conference on Artificial Intelligence

Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation

Omni-Effects is a unified framework for generating spatially controllable visual effects (VFX) in videos, addressing limitations of existing per-effect LoRA training. It introduces two key innovations: LoRA-based Mixture of Experts (LoRA-MoE) to integrate diverse effects in a single model while mitigating cross-task interference, and Spatial-Aware Prompt…

Fangyuan Mao, Aiming Hao, Jintao Chen, Dongxia Liu, et al.
Published
Aug 2025
Citations
28
Code
175 stars
164

arXiv.org

LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation

LongVie is a framework for controllable ultra-long video generation, addressing temporal inconsistency and visual degradation in autoregressive generation. It identifies three key issues: separate noise initialization, independent control signal normalization, and single-modality guidance limitations. LongVie introduces unified noise initialization and…

Jianxiong Gao, Zhaoxi Chen, Xian Liu, Jianfeng Feng, et al.
Published
Aug 2025
Citations
19
Code
Not linked