Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
-
Updated
Jul 24, 2026 - Go
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
A lightweight LLM inference runtime, evolving toward heterogeneous, stage-disaggregated serving.
Add a description, image, and links to the pd-disaggregation topic page so that developers can more easily learn about it.
To associate your repository with the pd-disaggregation topic, visit your repo's landing page and select "manage topics."