Paper 2602.23152
The Trinity of Consistency as a Defining Principle for General World Models
- Published
- Feb 2026
- Research lab
- Independent
- Citations
- 4
- GitHub
- Not linked
01 In brief
Summary
This paper proposes that a General World Model must be grounded in the Trinity of Consistency: Modal Consistency (semantic interface), Spatial Consistency (geometric basis), and Temporal Consistency (causal engine).
The authors systematically review the evolution of multimodal learning from specialized modules to unified architectures, arguing that dissolving barriers between these dimensions is necessary for world simulation.
They introduce CoW-Bench, a benchmark with 1,485 samples across 18 sub-tasks, evaluating video generation models and Unified Multimodal Models (UMMs) under a unified protocol.
Main results show closed-source image models (e.g., GPT-image-1.5) outperform open-source video generators, with temporal control and cross-consistency tasks (e.g., Maze-2D) being major bottlenecks.
The paper concludes that consistency is not optional but the criterion of existence for a world model, distinguishing texture synthesizers from true simulators, and proposes a future 'Prompt-as-Action' paradigm for interactive world models.
- The Trinity of Consistency comprises Modal, Spatial, and Temporal consistency as orthogonal yet synergistic constraints.
- CoW-Bench evaluates 18 sub-tasks across single and cross-consistency dimensions, using atomic checks and a 0-2 scoring scale.
- GPT-image-1.5 achieves the best overall performance (85.62), while open-source video models lag significantly.
- Key failures include constraint backoff, identity drift, and poor performance on navigation-style tasks like Maze-2D.
- The paper advocates for a paradigm shift from 'Vector-as-Action' and 'Key-as-Action' to 'Prompt-as-Action' for future world models.
02 From the paper
Abstract
The construction of World Models capable of learning, simulating, and reasoning about objective physical laws constitutes a foundational challenge in the pursuit of Artificial General Intelligence. Recent advancements represented by video generation models like Sora have demonstrated the potential of data-driven scaling laws to approximate physical dynamics, while the emerging Unified Multimodal Model (UMM) offers a promising architectural paradigm for integrating perception, language, and reasoning. Despite these advances, the field still lacks a principled theoretical framework that defines the essential properties requisite for a General World Model. In this paper, we propose that a World Model must be grounded in the Trinity of Consistency: Modal Consistency as the semantic interface, Spatial Consistency as the geometric basis, and Temporal Consistency as the causal engine. Through this tripartite lens, we systematically review the evolution of multimodal learning, revealing a trajectory from loosely coupled specialized modules toward unified architectures that enable the synergistic emergence of internal world simulators. To complement this conceptual framework, we introduce CoW-Bench, a benchmark centered on multi-frame reasoning and generation scenarios. CoW-Bench evaluates both video generation models and UMMs under a unified evaluation protocol. Our work establishes a principled pathway toward general world models, clarifying both the limitations of current systems and the architectural requirements for future progress.