The year/Topics/Video and spatial AI

Topic area

Video and spatial AI

Every collection across video and spatial ai.

Papers
164
Research labs
4
Official code
135

150 of 164 papers in this topic area

01

Independent research

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

VideoCoCo is an agentic dual-engine framework for physically consistent text-to-video generation. It uses executable Blender code as a process-level chain of thought. A coding agent synthesizes a Blender program from a text prompt, which is executed in a sandbox to produce a deterministic, low-fidelity spatiotemporal draft. A generative video engine then…

Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, et al.
Published
Jul 2026
Citations
0
Code
Not linked
02

Independent research

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

The paper DistillAlign revisits autoregressive video distillation from a distributional perspective. It argues that existing multi-stage pipelines, which separate initialization (e.g., ODE or consistency distillation) from DMD refinement, often have misaligned target distributions. Since DMD is mode-seeking, a good initialization must match the mode…

Jiaxing Li, Kai Zou, Cindy Zhou, Kaichen Huang, et al.
Published
Jul 2026
Citations
0
Code
97 stars
03

Independent research

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

ABot-World-0 is an action-conditioned video world model for real-time, long-horizon closed-loop interaction, deployable on a single NVIDIA RTX 5090 GPU. It uses raw keyboard inputs as a unified control interface for both scene roaming and third-person character control, with reference-character memory for identity consistency. The model is trained on…

Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, et al.
Published
Jul 2026
Citations
0
Code
1.8K stars
04

Independent research

Generative World Renderer at the Speed of Play

AlayaRenderer-Flash is a real-time generative world renderer that accelerates the offline AlayaRenderer from 0.56 FPS to 31.54 FPS, enabling interactive, prompt-controllable gameplay. It reformulates the original renderer into a few-step autoregressive streaming model with three key improvements: autoregressive generation over unbounded G-buffer streams,…

Guixu Lin, Zheng-Hui Huang, Siqi Yang, Ming-Hsuan Yang, et al.
Published
Jul 2026
Citations
0
Code
53 stars
05

Independent research

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

HOMIE is a framework for human-object centric video personalization (HOCVP) that unifies inter-subject (distinct subjects) and intra-subject (multiple references of the same subject) personalization. It addresses limitations of existing methods by integrating Multimodal Large Language Models (MLLMs) while preserving the text encoder, avoiding costly…

Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, et al.
Published
Jul 2026
Citations
0
Code
166 stars
06

Independent research

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

AlayaWorld is an interactive long-horizon video world model that generates 24-fps video at 540p and 720p, built on a 15B video diffusion transformer. It generates short latent chunks autoregressively under camera trajectories and switchable text prompts, using a bounded visual context that combines a persistent sink frame, compressed temporal history,…

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, et al.
Published
Jul 2026
Citations
0
Code
834 stars
07

Independent research

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

TimeLens2 introduces a generalist video temporal grounding model that predicts variable-cardinality sets of evidence intervals across diverse video lengths, domains, query forms, and viewpoints. It addresses two structural mismatches: unreliable long-video supervision and optimization that lacks interval-level geometry. The TimeLens2-93K dataset pipeline…

Yuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, et al.
Published
Jul 2026
Citations
0
Code
117 stars
08

Independent research

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

VideoChat3 is a fully open, efficient, and generalist video multimodal large language model (MLLM) with 4B parameters, designed to address limitations in existing open-source models: limited cross-domain generalization, high computational demands, and incomplete openness. It introduces two key architectural innovations: the Inflated 3D Vision Transformer…

Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, et al.
Published
Jul 2026
Citations
0
Code
377 stars
09

Independent research

Video Generation Models are General-Purpose Vision Learners

The paper introduces GenCeption, a general-purpose vision model that uses large-scale text-to-video generation as a pre-training paradigm. By repurposing a pre-trained video diffusion backbone (WAN 2.1) into a feed-forward model, GenCeption performs multiple vision tasks—depth, surface normal, camera pose, segmentation, and 3D keypoint estimation—steered…

Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, et al.
Published
Jul 2026
Citations
1
Code
Not linked
10

Independent research

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

LingBot-Video is a DiT-based video pretraining paradigm for embodied intelligence, introduced as the first large-scale open-source Mixture-of-Experts (MoE) video foundation model. It addresses domain mismatch in video generation by using a sparse MoE framework for better capacity-efficiency trade-off, a data profiling engine that augments internet videos…

Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, et al.
Published
Jul 2026
Citations
3
Code
908 stars
11

Independent research

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld is a full-stack, open-source framework for building interactive generative worlds, fine-tuned from LTX-2.3. It addresses four key challenges: control, consistency, stability, and runtime. For control, it combines a 3D cache rendered along the camera trajectory with AdaLN-style camera modulation, and supports prompt-driven actions via a…

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, et al.
Published
Jul 2026
Citations
1
Code
Not linked
12

Independent research

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

PixWorld is a unified pixel-space diffusion framework for 3D scene generation and reconstruction. It partitions multi-view inputs into clean and noisy subsets, processes them with a two-stream diffusion transformer, and decodes features into a pixel-aligned 3D Gaussian representation. The diffusion objective is applied directly on rendered images,…

Sensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, et al.
Published
Jul 2026
Citations
1
Code
239 stars
13

Independent research

Vidu S1: A Real-Time Interactive Video Generation Model

Vidu S1 is a real-time interactive video generation model that enables users to control digital characters via voice instructions during generation, supporting infinite-length video without blurring or drift. Built with TurboDiffusion and TurboServe, it outputs 540p video at up to 42 FPS on consumer GPUs. The model uses a three-stage training pipeline:…

Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, et al.
Published
Jul 2026
Citations
2
Code
240 stars
14

Independent research

Orca: The World is in Your Mind

Orca, developed by the Beijing Academy of Artificial Intelligence, is a general world foundation model that learns a unified world latent space from multimodal signals (vision and language) using Next-State-Prediction modeling. It employs two complementary learning paradigms: unconscious learning captures dense natural state transitions from continuous…

Yihao Wang, Yuheng Ji, Mingyu Cao, Yanqing Shen, et al.
Published
Jun 2026
Citations
0
Code
906 stars
15

Independent research

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

LiveEdit is a novel streaming video editing framework that performs causal, chunk-by-chunk editing with high fidelity and ultra-low latency. It addresses two core issues: attention distribution shift and spatial-temporal token redundancy. The method uses a three-stage distillation pipeline: Stage 1 tunes a bidirectional DiT for editing, Stage 2 transitions…

Xinyu Wang, Chongbo Zhao, Fangneng Zhan, Yue Ma
Published
Jun 2026
Citations
2
Code
147 stars
16

Independent research

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

DomainShuttle is a novel framework for open-domain subject-driven text-to-video (S2V) generation, addressing both in-domain (high subject fidelity) and cross-domain (flexible adaptation of subject-irrelevant features) scenarios. It introduces three key components: Domain-MoT, which decouples video and reference features and uses domain-aware AdaLN for…

Nan Chen, Yiyang Cai, Rongchang Xie, Junwen Pan, et al.
Published
Jun 2026
Citations
1
Code
165 stars
17

Independent research

Looped World Models

The paper introduces Looped World Models (LoopWM), the first looped transformer architecture for world modeling, addressing the tension between deep computation for faithful long-horizon simulation and the high cost and error accumulation of deep models. LoopWM iteratively refines latent environment states through a parameter-shared transformer block with…

Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang, Jinrui Zeng, et al.
Published
Jun 2026
Citations
1
Code
Not linked
18

Independent research

DreamX-World 1.0: A General-Purpose Interactive World Model

DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation, supporting camera navigation, revisits, and promptable events across photorealistic, game-style, and stylized domains. It uses a data engine combining Unreal Engine rendering, gameplay recordings, and real-world videos. The model…

DreamX Team, Yancheng Bai, Rui Chen, Xiangxiang Chu, et al.
Published
Jun 2026
Citations
5
Code
748 stars
19

NVIDIA

SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

SpatialClaw is a training-free framework that improves spatial reasoning in vision-language models (VLMs) by using code as the action interface. It maintains a persistent Python kernel pre-loaded with input frames and perception tools, allowing a VLM-backed agent to write and execute one code cell per step, inspect intermediate results (e.g., masks, depth…

Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su, et al.
Published
Jun 2026
Citations
3
Code
358 stars
20

Independent research

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

OmniDirector is a framework for cloning camera motion from reference videos to animate source images, supporting multi-shot sequences without requiring cross-paired training data. It introduces a 'camera grid' representation, which renders camera parameters as a grid motion video within an empty 3D scene, decoupling camera motion from content and enabling…

Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, et al.
Published
Jun 2026
Citations
5
Code
79 stars
21

Z.ai / GLM

SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

SCAIL-2 is an end-to-end framework for controlled character animation that bypasses intermediate representations like pose skeletons or masked backgrounds, which cause information loss. It directly concatenates driving videos to the sequence, allowing the model to capture all visual information. To address the lack of end-to-end data, the authors unify…

Wenhao Yan, Fengjia Guo, Zhuoyi Yang, Jie Tang
Published
Jun 2026
Citations
1
Code
1.1K stars
22

Independent research

Latent Spatial Memory for Video World Models

This paper introduces latent spatial memory, a persistent 3D cache for video world models that stores scene information directly in the diffusion latent space, avoiding the pixel-space round trip of RGB point-cloud memories. The authors propose Mirage, a framework that constructs the memory by lifting latent tokens into 3D via depth-guided back-projection…

Weijie Wang, Haoyu Zhao, Yifan Yang, Feng Chen, et al.
Published
Jun 2026
Citations
2
Code
297 stars
23

Independent research

ABot-Earth 0.5: Generative 3D Earth Model

ABot-Earth 0.5, developed by AMAP CV Lab (Alibaba Group), is a generative 3D framework that synthesizes vast, seamless 3D environments from geospatially referenced satellite imagery using a native 3D Gaussian Splatting (3DGS) representation. Trained on real-world urban reconstructions, it generates realistic geometry and textures at under 10 minutes per…

Ming Qian, Tianjian Ouyang, Mingchao Sun, Zijian Wang, et al.
Published
Jun 2026
Citations
1
Code
192 stars
24

NVIDIA

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA introduces Cosmos 3, a family of omnimodal world models that jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. It subsumes vision-language models, video generators, world simulators, and world-action models into a single framework, supporting flexible input-output…

NVIDIA, :, Aditi, Niket Agarwal, et al.
Published
Jun 2026
Citations
36
Code
Not linked
25

arXiv.org

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

YoCausal is a two-level benchmark for evaluating causal cognition in video diffusion models (VDMs), inspired by the Violation of Expectation paradigm. It uses temporally reversed real-world videos as counterfactual samples, avoiding synthetic data and the sim-to-real gap. Level 1 introduces the Reverse Surprise Index (RSI), measuring arrow-of-time…

You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, et al.
Published
May 2026
Citations
1
Code
36 stars
26

arXiv.org

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

minWM is a full-stack open-source framework for converting bidirectional text-to-video (T2V) or text-and-image-to-video (TI2V) diffusion foundation models into camera-controllable, few-step autoregressive (AR) world models for real-time interaction. The pipeline has two phases: first, fine-tuning the bidirectional model with camera control via PRoPE…

Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou, et al.
Published
May 2026
Citations
6
Code
766 stars
27

NVIDIA

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

The paper investigates whether vision-language models (VLMs) achieve spatial reasoning through structured 3D understanding or by exploiting statistical shortcuts in natural images. The authors introduce a representation-level analysis framework using contrastive pairs to measure how spatial axes (horizontal, vertical, depth) are organized in VLM…

Cheolhong Min, Jaeyun Jung, Daeun Lee, Hyeonseong Jeon, et al.
Published
May 2026
Citations
1
Code
16 stars
28

NVIDIA

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

The paper introduces Gamma-World, a generative multi-agent world model for interactive simulation that scales beyond two players. It addresses limitations of prior work like Solaris, which uses dense attention and learned per-slot identities, by proposing two key innovations. First, Simplex Rotary Agent Encoding extends 3D RoPE with an agent axis,…

Fangfu Liu, Kai He, Tianchang Shen, Tianshi Cao, et al.
Published
May 2026
Citations
2
Code
Not linked
29

arXiv.org

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

SpatialBench is a new benchmark for evaluating spatial foundation models across diverse domains, input densities, and model paradigms. It includes 19 datasets, 546 scenes, 41 model variants, and 6 paradigms, using a deterministic multi-density sampling protocol (single, sparse, medium, dense). Key findings reveal that full-context attention models achieve…

Haosong Peng, Hao Li, Jiaqi Chen, Yuhao Pan, et al.
Published
May 2026
Citations
0
Code
119 stars
30

arXiv.org

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation

WBENCH is a comprehensive multi-turn benchmark for evaluating interactive video world models across five dimensions: video quality, setting adherence, interaction adherence, consistency, and physics compliance. It contains 289 test cases and 1,058 interaction turns, covering diverse scenes, styles, subjects, and both first- and third-person perspectives.…

Kaining Ying, Hengrui Hu, Siyu Ren, Jiamu Li, et al.
Published
May 2026
Citations
7
Code
175 stars
31

arXiv.org

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

EvalVerse is a comprehensive evaluation framework for professional cinematic video generation that addresses the gap between basic prompt-following and true cinematic quality. It introduces a pipeline-aware taxonomy mirroring the filmmaking workflow (pre-production, production, post-production) with 3 stages, 7 aspects, 18 dimensions, 45 sub-dimensions,…

Songlin Yang, Haobin Zhong, Ruilin Zhang, Xiaotong Zhao, et al.
Published
May 2026
Citations
2
Code
Not linked
32

arXiv.org

TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation

TransitLM is a large-scale dataset of over 13 million transit route planning records from four Chinese cities (Beijing, Shanghai, Shenzhen, Chengdu), covering 120,845 stations and 13,666 lines. It is released as a continual pre-training corpus and benchmark data for three tasks: optimal route generation, preference-aware planning, and multi-route…

Hanyu Guo, Jiedong Yang, Chao Chen, Longfei Xu, et al.
Published
May 2026
Citations
0
Code
125 stars
33

arXiv.org

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

This paper introduces Grounded Personality Reasoning (GPR) and the MM-OCEAN benchmark to evaluate whether Multimodal Large Language Models (MLLMs) perceive personality through behavioral evidence or merely prejudge via superficial patterns. The benchmark includes 1,104 videos and 5,320 cue-grounding MCQs, built via a multi-agent human-collaborative…

Caixin Kang, Tianyu Yan, Sitong Gong, Mingfang Zhang, et al.
Published
May 2026
Citations
0
Code
11 stars
34

arXiv.org

PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects

PhysX-Omni is a unified framework for generating simulation-ready physical 3D assets covering rigid, deformable, and articulated objects. It introduces a novel template-based run-length encoding (RLE) geometry representation for Vision-Language Models, which directly encodes high-resolution 3D structures without special tokens or segmentation modules,…

Ziang Cao, Yinghao Liu, Haitian Li, Runmao Yao, et al.
Published
May 2026
Citations
1
Code
307 stars
35

NVIDIA

LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

LongLive-2.0 is an NVFP4-based parallel infrastructure for long video generation, co-designing training and inference. For training, it introduces Balanced SP, a sequence-parallel autoregressive (AR) training method that pairs clean-history and noisy-target temporal chunks on each GPU, enabling a natural teacher-forcing mask and SP-aware chunked VAE…

Yukang Chen, Luozhou Wang, Wei Huang, Shuai Yang, et al.
Published
May 2026
Citations
11
Code
2.5K stars
36

arXiv.org

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

The paper introduces MIGA, a training-free method for infinite-frame long video generation that builds on frame-level autoregressive frameworks like FIFO-Diffusion. MIGA addresses two key limitations: the training-inference gap and long-term consistency. It proposes a Two-Stage Training-Inference Alignment (TTA) mechanism that reduces the noise span of…

X. Feng, J. Zhu, M. Wu, C. Chen, et al.
Published
May 2026
Citations
1
Code
Not linked
37

arXiv.org

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

FashionChameleon is a real-time and interactive framework for human-garment video customization, enabling users to switch garments during generation while preserving motion coherence. It uses three key techniques: a Teacher Model with In-Context Learning trained on single-garment data to implicitly handle garment switching; Streaming Distillation with…

Quanjian Song, Yefeng Shen, Mengting Chen, Hao Sun, et al.
Published
May 2026
Citations
3
Code
256 stars
38

NVIDIA

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

SANA-WM is a 2.6B-parameter open-source world model for generating one-minute, 720p videos with precise 6-DoF camera control. It uses a hybrid linear diffusion transformer combining frame-wise Gated DeltaNet (GDN) and softmax attention for efficient long-context modeling, a dual-branch camera control (UCPE and Plücker mixing), a two-stage generation…

Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, et al.
Published
May 2026
Citations
13
Code
Not linked
39

arXiv.org

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation

Causal Forcing++ is a scalable pipeline for real-time interactive video generation that distills bidirectional diffusion models into few-step autoregressive (AR) students. It targets frame-wise autoregression with 1–2 sampling steps, a regime where existing initialization strategies fail: ODE distillation with a bidirectional teacher is architecturally…

Min Zhao, Hongzhou Zhu, Kaiwen Zheng, Zihan Zhou, et al.
Published
May 2026
Citations
11
Code
905 stars
40

NVIDIA

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

AnyFlow is a video diffusion distillation framework that enables any-step generation by learning flow-map transitions between arbitrary time pairs, unlike consistency models that degrade with more sampling steps. It uses a two-stage pipeline: forward flow map training (with interpolated timestep conditioning, guidance-fused training, and adaptive loss…

Yuchao Gu, Guian Fang, Yuxin Jiang, Weijia Mao, et al.
Published
May 2026
Citations
6
Code
406 stars
41

arXiv.org

When Vision Speaks for Sound

The paper identifies a 'Clever Hans effect' in video-capable multimodal LLMs, where models appear to understand audio but actually rely on visual cues to hallucinate or infer sounds without verifying the audio stream. This is demonstrated across open-source and closed-source models. To systematically study this, the authors introduce THUD, a diagnostic…

Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu, Rui Cai, et al.
Published
May 2026
Citations
2
Code
98 stars
42

arXiv.org

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

CollabVR is a closed-loop framework for video reasoning that couples a Vision-Language Model (VLM) with a Video Generation Model (VGM) at step-level granularity. It addresses two failure modes of VGMs: long-horizon drift and mid-clip simulation errors. The VLM plans the immediate next action, verifies the generated clip, and folds the diagnosis into the…

Joowon Kim, Seungho Shin, Joonhyung Park, Eunho Yang
Published
May 2026
Citations
1
Code
9 stars
43

arXiv.org

HumanNet: Scaling Human-centric Video Learning to One Million Hours

HumanNet is a one-million-hour human-centric video corpus designed to scale embodied learning by capturing how humans interact with the physical world. It includes both first-person and third-person perspectives, covering fine-grained activities, human-object interactions, tool use, and long-horizon behaviors across diverse environments. The dataset…

Yufan Deng, Daquan Zhou
Published
May 2026
Citations
6
Code
281 stars
44

arXiv.org

Stream-T1: Test-Time Scaling for Streaming Video Generation

Stream-T1 is a Test-Time Scaling (TTS) framework for streaming video generation, addressing the high costs and lack of temporal guidance in existing diffusion-based TTS methods. It leverages chunk-level synthesis and few denoising steps to reduce computational overhead. The framework comprises three components: Stream-Scaled Noise Propagation, which…

Yijing Tu, Shaojin Wu, Mengqi Huang, Wenchuan Wang, et al.
Published
May 2026
Citations
2
Code
37 stars
45

arXiv.org

Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation

Stream-R1 is a framework for distilling autoregressive streaming video diffusion models, addressing limitations in existing distribution matching distillation (DMD) methods that treat all rollouts, frames, and pixels equally. It introduces two concepts: Inter-Reliability (varying reliability of supervision across rollouts) and Intra-Perplexity (varying…

Bin Wu, Mengqi Huang, Shaojin Wu, Weinan Jia, et al.
Published
May 2026
Citations
2
Code
54 stars
46

ACM Transactions on Graphics

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

UniVidX is a unified multimodal framework for versatile video generation that repurposes video diffusion model (VDM) priors to handle diverse tasks within a single model. It addresses limitations of existing approaches that train separate models for fixed input-output mappings, ignoring cross-modal correlations. UniVidX introduces three key designs:…

Houyuan Chen, Hong Li, Xianghao Kong, Tianrui Zhu, et al.
Published
May 2026
Citations
0
Code
249 stars
47

arXiv.org

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

HERMES++ is a unified driving world model that integrates 3D scene understanding and future geometry prediction within a single framework. It uses a Bird's-Eye View (BEV) representation to consolidate multi-view spatial information, making it compatible with Large Language Models (LLMs). The model introduces LLM-enhanced world queries for knowledge…

Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan, et al.
Published
Apr 2026
Citations
3
Code
69 stars
48

arXiv.org

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

This roadmap paper argues that visual generation must evolve from appearance synthesis to intelligent visual generation, grounded in structure, dynamics, and causal relations. It proposes a five-level taxonomy—Atomic, Conditional, In-Context, Agentic, and World-Modeling Generation—to organize progress from passive rendering to interactive, world-aware…

Keming Wu, Zuhao Yang, Kaichen Zhang, Shizun Wang, et al.
Published
Apr 2026
Citations
4
Code
128 stars
49

arXiv.org

RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

RADIO-ViPE is an online, calibration-free semantic SLAM system that processes raw monocular RGB video to produce geometry-aware, open-vocabulary 3D grounding. It tightly couples multi-modal embeddings from agglomerative foundation models (RADIO/RADSeg) with geometric information within a dense bundle adjustment framework, using a factor graph with…

Zaid Nasser, Mikhail Iumanov, Tianhao Li, Maxim Popov, et al.
Published
Apr 2026
Citations
0
Code
138 stars
50

arXiv.org

ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

The paper introduces ReVSI, a benchmark for evaluating vision-language models' (VLMs) 3D spatial reasoning, addressing validity issues in the existing VSI-Bench. Two key pitfalls are identified: annotation-to-video ground-truth drift (errors from point-cloud-based annotations) and scene-observability mismatch (questions unanswerable under sparse frame…

Yiming Zhang, Jiacheng Chen, Jiaqi Tan, Yongsen Mao, et al.
Published
Apr 2026
Citations
7
Code
83 stars