> ## Documentation Index
> Fetch the complete documentation index at: https://docs.muna.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# muna deploy

> Deploy a compiled model onto a compute cloud.

The `deploy` command deploys a compiled model to your GPU fleets, whether on-prem or from third-party GPU platforms.
Compiled models can yield up to 45x reductions in cold-start times, along with reduced latency and higher utilization.
[See why](/compile/create#why-compile).

Each deployment runs an [OpenAI-compatible web server](https://github.com/muna-ai/muna-server), which then
forwards requests to the compiled model.

## Usage

```bash icon="terminal" theme={null}
$ muna deploy [OPTIONS] TAG --provider PROVIDER
```

| Argument | Description |
| :- | :- |
| `TAG` | Compiled model tag (e.g. `@google/gemma-4-26b-a4b-it`) |

## Options

| Option | Description |
| :- | :- |
| `--provider` | Cloud to deploy to: `modal`, `baseten`, `spheron`, or `baremetal`. Required. |
| `--gpu` | GPU: `a100`, `h100`, `h200`, or `b200`. |
| `--gpu-count` | Number of GPUs to request. |
| `--cpu` | Number of vCPUs to request. |
| `--memory` | Memory to request, in MB. |
| `--name` | Deployment name. |
| `--concurrency` | Maximum concurrent requests sent to an instance. |
| `--min-replicas` | Minimum replicas for autoscaling. |
| `--max-replicas` | Maximum replicas for autoscaling. |
| `--scaledown-window` | Autoscaling scale down window, in seconds. |
| `--ssh-host` | SSH target for `baremetal` deployments, e.g. `'root@1.2.3.4 -p 22 -i ~/.ssh/key'`. |
| `--endpoint-url` | Public base URL where the deployed server is reachable. Required for `baremetal`. |
| `--wait` | Wait until the deployment is complete. |

## Deploying to Modal

Deploy a compiled model to [Modal](https://modal.com):

```bash icon="terminal" theme={null}
# 🚀 Deploy a compiled LLM to Modal
$ muna deploy @google/gemma-4-26b-a4b-it --provider modal --gpu b200
```

This command will create a lightweight app on Modal that runs the web server.

<Tip>
  This command requires the `modal` package to be installed. Run `pip install modal`.
</Tip>

## Deploying to Baseten

Deploy a compiled model to [Baseten](https://baseten.co):

```bash icon="terminal" theme={null}
# 🚀 Deploy a compiled LLM to Baseten
$ muna deploy @google/gemma-4-26b-a4b-it --provider baseten --gpu b200
```

This command will create and deploy a lightweight service on Baseten that runs the web server.

<Tip>
  This command requires the `truss` package to be installed and logged in. Run `pip install truss`, then `truss login`.
</Tip>

## Deploying to Spheron

Deploy a compiled model to [Spheron](https://spheron.network):

```bash icon="terminal" theme={null}
# 🚀 Deploy a compiled LLM to Spheron
$ muna deploy @google/gemma-4-26b-a4b-it --provider spheron --gpu b200
```

<Tip>
  This command requires a Spheron API key in the `SPHERON_API_KEY` environment variable, and a `--gpu`.
</Tip>

## Deploying to Your Own Servers

Deploy a compiled model to any server you can SSH into:

```bash icon="terminal" theme={null}
# 🚀 Deploy a compiled LLM to your own GPU server
$ muna deploy @google/gemma-4-26b-a4b-it \
    --provider baremetal \
    --ssh-host "root@1.2.3.4 -p 22 -i ~/.ssh/key" \
    --endpoint-url https://gemma.example.com
```

The web server listens on port `8000`. Use `--endpoint-url` to tell Muna the public URL that routes to it.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.