The year/Independent research

Paper 2607.02501

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

Published
Jul 2026
Research lab
Independent
Citations
1
GitHub
121 stars

01 In brief

Summary

Embodied.cpp is a portable C++ inference runtime for embodied AI models, addressing the fragmented deployment of vision-language-action (VLA) models and world-action models (WAMs) on heterogeneous edge devices.

It identifies three key runtime requirements: multi-rate execution, latency-first batch-1 inference, and extensible embodied interfaces.

The runtime is organized into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters.

Evaluations show HY-VLA achieves 100.0% task success on RoboTwin, and pi0.5 achieves 91.0% success.

A preliminary WAM benchmark on a LingBot-VA Transformer block reduces memory from 312.2 MiB to 88.1 MiB with Q4_K quantization, keeping MAE below 3.3e-2 and cosine similarity above 0.9997.

The runtime supports heterogeneous hardware, robots, and simulators through one backend abstraction.

02 From the paper

Abstract

Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are designed mainly for request-response serving and therefore do not satisfy the runtime contract of embodied deployment: multi-rate execution inside closed-loop control, latency-first batch-1 inference on heterogeneous hardware, and extensible embodied interfaces beyond fixed token I/O. We present Embodied$.$cpp, a portable C++ inference runtime for embodied models. Based on an architectural analysis of representative VLA models and WAMs, Embodied$.$cpp captures a shared execution path and organizes it into five layers: input adapters, sequence builders, backbone execution, head plugins, and deployment adapters. The runtime provides modular multi-rate execution, latency-first fused inference, and extensible operator and I/O support, enabling deployment across heterogeneous devices, robots, and simulators through one backend abstraction. We evaluate Embodied$.$cpp on two VLA models, HY-VLA and pi0.5, and on a preliminary WAM benchmark using a LingBot-VA Transformer block. The VLA deployments achieve successful closed-loop execution with 100.0% and 91.0% task success rates, respectively. The WAM benchmark reduces block memory from 312.2 MiB to 88.1 MiB. These results show that Embodied$.$cpp improves deployment efficiency while preserving high accuracy across diverse embodied model architectures.