Paper 2605.28816
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
- Published
- May 2026
- Research lab
- NVIDIA
- Citations
- 2
- GitHub
- Not linked
01 In brief
Summary
The paper introduces Gamma-World, a generative multi-agent world model for interactive simulation that scales beyond two players.
It addresses limitations of prior work like Solaris, which uses dense attention and learned per-slot identities, by proposing two key innovations.
First, Simplex Rotary Agent Encoding extends 3D RoPE with an agent axis, representing agents as vertices of a regular simplex in rotary angle space.
This provides distinct, permutation-symmetric identities without learned embeddings, enabling zero-shot scaling to more agents.
Second, Sparse Hub Attention uses learnable hub tokens to mediate cross-agent communication, reducing cross-agent attention cost from quadratic to linear in the number of agents.
The model is trained in three stages: a bidirectional teacher, a causal student with Diffusion Forcing, and a distilled few-step generator for real-time inference.
Experiments in multiplayer Minecraft environments show Gamma-World improves video fidelity, action controllability, and inter-agent consistency over baselines, and generalizes from two to four players without retraining.
It also demonstrates applicability to real-world robotic coordination tasks.
The model achieves 24 FPS streaming rollouts with KV caching, and ablations validate the design choices, including the number of hub tokens and training stage comparisons.
02 From the paper
Abstract
World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many generated environments require multi-agent interaction: multiple players, robots, or embodied agents act simultaneously within a shared space. Scaling world models to such settings requires a principled multi-agent design: agents should remain independently controllable, permutation-symmetric, and support efficient inference while maintaining consistency across time and perspectives. In this paper, we present our generative multi-agent world model for interactive simulation. It introduces Simplex Rotary Agent Encoding, a parameter-free extension of 3D RoPE that represents agents as vertices of a regular simplex in rotary angle space. This gives each agent a distinct phase while making all agents permutation-equivalent, enabling scalable agent identity without learned per-slot identities or a fixed agent ordering. To avoid dense all-to-all attention across agents, we further propose Sparse Hub Attention, where learnable hub tokens mediate token interaction across agents, reducing cross-agent attention cost from quadratic to linear in the number of agents. For real-time rollout, we distill a full-context diffusion teacher into a causal student that generates temporal blocks sequentially with KV caching, enabling action-responsive generation at 24 FPS. Experiments in multiplayer virtual environments show that our model improves video fidelity, action controllability, and inter-agent consistency over slot-based and dense-attention baselines, while generalizing from two to four players without additional training.