NVIDIA
ModelExpress Server
Container
NVIDIA
ModelExpress Server

Model Express is a Rust-based model cache management service designed to be deployed as a sidecar alongside existing inference solutions such as NVIDIA Dynamo.

  • ModelExpress-server is a container providing a model cache management service designed to be deployed as a sidecar alongside existing inference solutions such as NVIDIA Dynamo.

    Publisher
    NVIDIA
    Latest Tag0.5.1
    UpdatedAugust 20, 2026 UTC
    Compressed Size71.21 MB
    Multinode SupportNo
    Multi-Arch SupportYes