The deploy command deploys a compiled model to your GPU fleets, whether on-prem or from third-party GPU platforms.
Compiled models can yield up to 45x reductions in cold-start times, along with reduced latency and higher utilization.
See why.
Each deployment runs an OpenAI-compatible web server, which then
forwards requests to the compiled model.
Usage
Options
Deploying to Modal
Deploy a compiled model to Modal:
This command will create a lightweight app on Modal that runs the web server.
This command requires the modal package to be installed. Run pip install modal.
Deploying to Baseten
Deploy a compiled model to Baseten:
This command will create and deploy a lightweight service on Baseten that runs the web server.
This command requires the truss package to be installed and logged in. Run pip install truss, then truss login.
Deploying to Spheron
Deploy a compiled model to Spheron:
This command requires a Spheron API key in the SPHERON_API_KEY environment variable, and a --gpu.
Deploying to Your Own Servers
Deploy a compiled model to any server you can SSH into:
The web server listens on port 8000. Use --endpoint-url to tell Muna the public URL that routes to it.