This page reads the live catalog, so it always matches what the endpoint serves and what you are billed.
Prices are in USD per million tokens.
Chat Models
Embedding Models
Use the tag in the Model column as the model in your request. Don’t see the model you need?
Ask in our community Slack.
Prompt Caching
Coding agents resend the whole conversation on every turn, so most of the input in an agent session
is a prefix the model has already seen. Muna caches these prefixes automatically and bills them at the
cached input rate. There is nothing to configure: no cache keys and no cache_control markers.
In practice, about 80% of the input tokens in agent sessions on Muna are served from cache. That is
why the cached input rate matters more to your bill than the input rate.
Caching works best when the start of the prompt stays the same between turns. Agents already do this.
In your own apps, put stable content like the system prompt and tool definitions first.
Cold Starts
Models load on demand. If a model has not been used recently, the first request waits a few seconds
while it loads. Requests after that are served at full speed.