Skip to main content
Our hosted endpoint is local inference under the hood: each GPU server runs the Muna SDK behind an OpenAI-compatible web server. Using the SDK directly peels back that layer, so you can run the same compiled models on the CPU, GPU, or neural processor of your choosing. Learn more.

Installing Muna

We provide SDKs for common development frameworks:
Muna uses access keys to authenticate all requests. Head to your developer settings and generate an access key.
Most of our client SDKs are open-source. Star them on GitHub!
Keep your access key secret! Do not embed it in client-side code that you ship to users. See our security model for more information.

Choosing a Client

The SDK exposes three ways to run a model:

Using the Muna Client

Run any compiled model with the raw muna.predictions API.

Using the OpenAI Client

Chat, embeddings, speech, and transcriptions, with the OpenAI interface.

Using the Anthropic Client

The Messages API, with the Anthropic interface.