The year/Independent research

Paper 2605.12178

Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics

Published
May 2026
Research lab
Independent
Citations
0
GitHub
Not linked

01 In brief

Summary

The paper investigates whether enterprise systems need learned world models, arguing that runtime discovery of configurable transition dynamics is more robust than offline training.

It introduces CascadeBench, a benchmark for enterprise cascade prediction, and enterprise discovery agents that retrieve business rules at inference time.

Experiments show offline-trained world models perform well in-distribution but degrade under deployment shift, while discovery agents remain robust by grounding predictions in the active instance.

The study uses Enterprise Gym, a live platform environment, to collect 27,243 transition samples across 64 worlds.

Results indicate that prompting without rules fails, fine-tuning helps only in-distribution, and runtime retrieval recovers cross-instance accuracy, though it does not fully match oracle performance, especially for Tier 3 execution-inferred dynamics.

02 From the paper

Abstract

World models enable agents to anticipate the effects of their actions by internalizing environment dynamics. In enterprise systems, however, these dynamics are often defined by tenant-specific business logic that varies across deployments and evolves over time, making models trained on historical transitions brittle under deployment shift. We ask a question the world-models literature has not addressed: when the rules can be read at inference time, does an agent still need to learn them? We argue, and demonstrate empirically, that in settings where transition dynamics are configurable and readable, runtime discovery complements offline training by grounding predictions in the active system instance. We propose enterprise discovery agents, which recover relevant transition dynamics at runtime by reading the system's configuration rather than relying solely on internalized representations. We introduce CascadeBench, a reasoning-focused benchmark for enterprise cascade prediction that adopts the evaluation methodology of World of Workflows on diverse synthetic environments, and use it together with deployment-shift evaluation to show that offline-trained world models can perform well in-distribution but degrade as dynamics change, whereas discovery-based agents are more robust under shift by grounding their predictions in the current instance. Our findings suggest that, in configurable enterprise environments, agents should not rely solely on fixed internalized dynamics, but should incorporate mechanisms for discovering relevant transition logic at runtime.