RLMesh puts a network boundary between models and environments, so each side keeps its own language and dependencies. Managed runs optimize inference across rollouts, so the GPU stays saturated.
Built for evaluation, and for any model ↔ environment inference workload, including data collection across simulators and the real world.
Fully managed, end to end. RLMesh handles scheduling, scaling, and recovery, so your evaluations just run.
Open-source framework
Any language on either side of the wire
A model and an environment don't have to share a language, or a dependency stack, to run together. RLMesh keeps each in its own process, talking over one wire, so neither is locked to the other's interpreter or versions. Python and Rust work today; C and C++ are planned on the same wire.
env processowns the simulator
1fromrlmeshimportclass EnvServer(env, address: str) Serves a Gymnasium-style environment over the RLMesh wire. Clients in any language and any value backend can connect.
2importgymnasiumasgym
3
4env = gym.make("CartPole-v1")
5class EnvServer(env, address: str) Serves a Gymnasium-style environment over the RLMesh wire. Clients in any language and any value backend can connect.(env, "127.0.0.1:5555").def serve(self) -> None Binds the address and serves reset/step until shutdown.()
~
~
~
~
valuesbackend-agnostic; clients choose how to decode
obs →← action
obs ↓action ↑
client processowns the model
1fromrlmesh.numpyimportclass RemoteEnv(address: str) Gymnasium-compatible client for a served environment. Observations decode as NumPy arrays.
2
3env = class RemoteEnv(address: str) Gymnasium-compatible client for a served environment. Observations decode as NumPy arrays.("127.0.0.1:5555")
4obs: np.ndarray Decoded client-side as NumPy arrays. On the wire it is a space-typed value, so the env process never needs this stack., info = env.def reset(self, *, seed: int | None = None) -> tuple[np.ndarray, dict] Starts an episode on the served environment and returns the first observation.()
5
6action: int Sampled from the mirrored Discrete space. It crosses the wire as a space-typed value with the same type in every backend. = env.action_space: Space Mirrors the served environment's action space. Spaces convert across the wire..sample()
7obs: np.ndarray Decoded client-side as NumPy arrays. On the wire it is a space-typed value, so the env process never needs this stack., reward, terminated, truncated, info = env.def step(self, action) -> tuple[np.ndarray, float, bool, bool, dict] One environment step across the process boundary. The observation arrives as NumPy arrays.(action: int Sampled from the mirrored Discrete space. It crosses the wire as a space-typed value with the same type in every backend.)
~
~
values
Serve from Python and connect from Rust, or reverse it. RLMesh keeps endpoints, spaces, and values on one wire contract.
Every environment reports observations in its own shape, and a generalist policy can't reshape its inputs for each one. The old way, you hand-write an adapter to reconcile every model with every environment: 4 envs means 4 adapters, 10 means 10. RLMesh makes the environment declare its observation and action spaces once, so any model connects through that contract and runs against all of them.
+1 model → +N adapters vs +1 spec
Old way
+1 model = +N adapters
RLMesh
+1 model = +1 spec
Managed platform
Keep GPUs busy through the eval loop
GPU infrastructure today is built for workloads that stay fed: LLM serving runs off a constant queue of requests, and training holds its whole batch on the device. Robotics policy inference is fundamentally different. The model and the environment take turns, so every action waits on the next observation and the GPU sits idle between steps. That gap is what RLMesh closes: managed runs batch inference across concurrent rollouts, so the GPU stays busy on one batch while other environments step.
GPU timeline · who stays busyillustrative
Physical AI Inferenceidle between steps
env
model
Physical AI Inference on RLMeshbatched across rollouts
env
model
Same suite, same silicon, fewer idle gaps.
Open Source
The framework
Use RLMesh when simulator and model stacks need separate processes. Gymnasium-style environments get a network boundary, and value decoding stays explicit.
EnvServer and RemoteEnv across process boundaries
Sandboxed environments, each with its own dependency stack
Gymnasium-style environment contracts and space conversion
Run models × tasks on a connected Kubernetes cluster. RLMesh batches inference across rollouts, streams status through the control plane, and reconciles final results from artifacts.
Models × tasks expanded into workloads
Live episode results and logs through the control plane
Final results reconciled from object-storage artifacts