> ## Documentation Index
> Fetch the complete documentation index at: https://docs.muna.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Models

> Every model on the inference endpoint, with live prices.

export const LiveModels = ({kind}) => {
  const GATEWAY_URL = "https://inference.muna.ai";
  const API_URL = "https://api.muna.ai/v1";
  const [rows, setRows] = useState(null);
  const [failed, setFailed] = useState(false);
  const usd = value => {
    if (value === undefined || value === null) return "—";
    return new Intl.NumberFormat("en-US", {
      style: "currency",
      currency: "USD",
      minimumFractionDigits: 2,
      maximumFractionDigits: 4
    }).format(value);
  };
  const tokens = value => {
    if (!value) return "—";
    return value >= 1024 ? `${Math.round(value / 1024)}K` : `${value}`;
  };
  const kindOf = (model, pricing) => {
    if (pricing?.kind === "tokens" && pricing.outputPerMillion === undefined) return "embedding";
    if (!model.max_tokens) return "embedding";
    return "chat";
  };
  useEffect(() => {
    const load = async () => {
      try {
        const [models, endpoints] = await Promise.all([fetch(`${GATEWAY_URL}/v1/models`).then(r => r.json()), fetch(`${API_URL}/endpoints?limit=100`).then(r => r.json()).catch(() => ({
          data: []
        }))]);
        const pricing = new Map((endpoints.data ?? []).map(e => [e.tag, e.pricing]));
        const list = (models.data ?? []).map(model => ({
          tag: model.id,
          kind: kindOf(model, pricing.get(model.id)),
          context: model.max_input_tokens,
          pricing: pricing.get(model.id)
        }));
        list.sort((a, b) => a.kind.localeCompare(b.kind) || a.tag.localeCompare(b.tag));
        setRows(list);
      } catch {
        setFailed(true);
      }
    };
    load();
  }, []);
  if (failed) return <p>
        The live catalog is unavailable right now. <a href={`${GATEWAY_URL}/v1/models`}>/v1/models</a> lists
        every model on the endpoint.
      </p>;
  if (!rows) return <p>Loading the live model catalog…</p>;
  const shown = kind ? rows.filter(row => row.kind === kind) : rows;
  const chat = kind !== "embedding";
  return <table>
      <thead>
        <tr>
          <th>Model</th>
          {chat && <th>Context</th>}
          <th>Input / 1M</th>
          {chat && <th>Cached input / 1M</th>}
          {chat && <th>Output / 1M</th>}
        </tr>
      </thead>
      <tbody>
        {shown.map(row => <tr key={row.tag}>
            <td><code>{row.tag}</code></td>
            {chat && <td>{tokens(row.context)}</td>}
            <td>{usd(row.pricing?.inputPerMillion)}</td>
            {chat && <td>{usd(row.pricing?.cachedInputPerMillion)}</td>}
            {chat && <td>{usd(row.pricing?.outputPerMillion)}</td>}
          </tr>)}
      </tbody>
    </table>;
};

This page reads the live catalog, so it always matches what the endpoint serves and what you are billed.
Prices are in USD per million tokens.

## Chat Models

<LiveModels kind="chat" />

## Embedding Models

<LiveModels kind="embedding" />

<Tip>
  Use the tag in the **Model** column as the `model` in your request. Don't see the model you need?
  Ask in [our community Slack](https://muna.ai/slack).
</Tip>

## Prompt Caching

Coding agents resend the whole conversation on every turn, so most of the input in an agent session
is a prefix the model has already seen. Muna caches these prefixes automatically and bills them at the
**cached input** rate. There is nothing to configure: no cache keys and no `cache_control` markers.

In practice, about 80% of the input tokens in agent sessions on Muna are served from cache. That is
why the cached input rate matters more to your bill than the input rate.

<Tip>
  Caching works best when the start of the prompt stays the same between turns. Agents already do this.
  In your own apps, put stable content like the system prompt and tool definitions first.
</Tip>

## Cold Starts

Models load on demand. If a model has not been used recently, the first request waits a few seconds
while it loads. Requests after that are served at full speed.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.