> ## Documentation Index
> Fetch the complete documentation index at: https://docs.muna.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Inference Backends

> Configuring how the compiler lowers your model for inference.

Muna's compiler supports specifying metadata, allowing you to configure the compiler or provide additional information.

<AccordionGroup>
  <Accordion title="TensorRT Inference Metadata">
    Use the `TensorRTInferenceMetadata` metadata type to compile a PyTorch [`nn.Module`](https://docs.pytorch.org/docs/stable/generated/torch.nn.Module.html) to [TensorRT](https://developer.nvidia.com/tensorrt):

    ```py ai.py icon="python" focus={1,5-8,12-20} theme={null}
    from muna.beta import TensorRTInferenceMetadata
    from torch import randn, Tensor
    from torch.nn import Module

    # Given a PyTorch model...
    model: Module = ...
    # With some example arguments...
    example_args: list[Tensor] = [randn(1, 3, 224, 224)]

    @compile(
        ...,
        metadata=[
            # Use TensorRT for model inference
            TensorRTInferenceMetadata(
                model=model,
                model_args=example_args,
                cuda_arch="sm_100",
                precision="int4"
            )
        ]
    )
    def predict() -> None:
        pass
    ```

    <Note>
      The TensorRT inference backend is only available on Linux and Windows devices with compatible Nvidia GPUs.
    </Note>

    <Tip>
      We are working on adding support for consumer RTX GPUs with [TensorRT for RTX](https://developer.nvidia.com/blog/nvidia-tensorrt-for-rtx-introduces-an-optimized-inference-ai-library-on-windows/).
    </Tip>

    ### Target CUDA Architectures

    TensorRT engines must be compiled for specific target CUDA architectures. Below are CUDA architectures that our compiler supports:

    | CUDA Architecture | GPU Family |
    | :- | :- |
    | `sm_80` | Ampere (e.g. A100) |
    | `sm_86` | Ampere |
    | `sm_87` | Ampere |
    | `sm_89` | Ada Lovelace (e.g. L40S) |
    | `sm_90` | Hopper (e.g. H100) |
    | `sm_100` | Blackwell (e.g. B200) |

    ### TensorRT Inference Precision

    TensorRT allows for specifying the inference engine's precision. Below are supported precision modes:

    | Precision | Notes |
    | :- | :- |
    | `fp32` | 32-bit single precision inference. |
    | `fp16` | 16-bit half precision inference. |
    | `int8` | 8-bit quantized integer inference. |
  </Accordion>

  <Accordion title="OnnxRuntime Inference Metadata">
    Use the `OnnxRuntimeInferenceMetadata` metadata type to compile a PyTorch [`nn.Module`](https://docs.pytorch.org/docs/stable/generated/torch.nn.Module.html) for inference with [ONNXRuntime](https://onnxruntime.ai/):

    ```py ai.py icon="python" focus={1,5-8,12-18} theme={null}
    from muna.beta import OnnxRuntimeInferenceMetadata
    from torch import randn, Tensor
    from torch.nn import Module

    # Given a PyTorch model...
    model: Module = ...
    # With some example arguments...
    example_args: list[Tensor] = [randn(1, 3, 224, 224)]

    @compile(
        ...,
        metadata=[
            # Use ONNXRuntime for model inference
            OnnxRuntimeInferenceMetadata(
                model=model,
                model_args=example_args
            )
        ]
    )
    def predict() -> None:
        pass
    ```
  </Accordion>

  <Accordion title="OnnxRuntime Inference Session Metadata">
    Use the `OnnxRuntimeInferenceSessionMetadata` metadata type to compile an OnnxRuntime [`InferenceSession`](https://onnxruntime.ai/docs/api/python/api_summary.html#inferencesession):

    ```py ai.py icon="python" focus={1,4-6,10-16} theme={null}
    from muna.beta import OnnxRuntimeInferenceSessionMetadata
    from onnxruntime import InferenceSession

    # Given an ONNXRuntime inference session...
    model_path = "/path/to/model.onnx"
    session = InferenceSession(model_path)

    @compile(
        ...,
        metadata=[
            # Use ONNXRuntime for model inference
            OnnxRuntimeInferenceSessionMetadata(
                session=session,
                model_path=model_path
            )
        ]
    )
    def predict(...) -> None:
        pass
    ```

    <Warning>
      The ONNX model file must exist at the provided `model_path` **within the compiler sandbox**.
    </Warning>
  </Accordion>

  <Accordion title="CoreML Inference Metadata">
    Use the `CoreMLInferenceMetadata` metadata type to compile a PyTorch
    [`nn.Module`](https://docs.pytorch.org/docs/stable/generated/torch.nn.Module.html) to
    [CoreML](https://developer.apple.com/documentation/coreml):

    ```py ai.py icon="python" focus={1,5-8,12-18} theme={null}
    from muna.beta import CoreMLInferenceMetadata
    from torch import randn, Tensor
    from torch.nn import Module

    # Given a PyTorch model...
    model: Module = ...
    # With some example arguments...
    example_args: list[Tensor] = [randn(1, 3, 224, 224)]

    @compile(
        ...,
        metadata=[
            # Use CoreML for model inference
            CoreMLInferenceMetadata(
                model=model,
                model_args=example_args
            )
        ]
    )
    def predict() -> None:
        pass
    ```

    <Note>
      The CoreML inference backend is only available on iOS, macOS, and visionOS devices.
    </Note>
  </Accordion>

  <Accordion title="Llama.cpp Inference Metadata">
    Use the `LlamaCppInferenceMetadata` metadata type to compile a [`Llama`](https://github.com/abetlen/llama-cpp-python)
    instance:

    ```py llm.py icon="python" focus={9-15} theme={null}
    from muna.beta import LlamaCppInferenceMetadata
    from llama_cpp import Llama

    # Given an LLM
    llm = Llama(...)

    @compile(
        ...,
        metadata=[
            # Specify Llama.cpp inference metadata
            LlamaCppInferenceMetadata(
                model=llm,
                backends=["cuda"]
            )
        ]
    )
    def predict() -> None:
        pass
    ```

    ### Llama.cpp Hardware Backends

    Llama.cpp supports several hardware backends to accelerate model inference.
    Below are targets that are currently supported by Muna:

    | Backend | Notes |
    | :- | :- |
    | `cuda` | [Nvidia CUDA backend](https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md#cuda). Linux only. |
  </Accordion>

  <Accordion title="ExecuTorch Inference Metadata">
    Use the `ExecuTorchInferenceMetadata` metadata type to compile a PyTorch [`nn.Module`](https://docs.pytorch.org/docs/stable/generated/torch.nn.Module.html) for inference with [ExecuTorch](https://docs.pytorch.org/executorch/stable/index.html):

    ```py ai.py icon="python" focus={1,5-8,12-19} theme={null}
    from muna.beta import ExecuTorchInferenceMetadata
    from torch import randn, Tensor
    from torch.nn import Module

    # Given a PyTorch model...
    model: Module = ...
    # With some example arguments...
    example_args: list[Tensor] = [randn(1, 3, 224, 224)]

    @compile(
        ...,
        metadata=[
            # Use ExecuTorch for model inference
            ExecuTorchInferenceMetadata(
                model=model,
                model_args=example_args,
                backend="xnnpack"
            )
        ]
    )
    def predict() -> None:
        pass
    ```

    <Note>
      The ExecuTorch inference backend is only available on Android.
    </Note>

    ### ExecuTorch Hardware Backends

    ExecuTorch supports several [hardware backends](https://docs.pytorch.org/executorch/stable/backends-overview.html) to
    accelerate model inference. Below are targets that are currently supported by Muna:

    | Backend | Notes |
    | :- | :- |
    | `xnnpack` | [XNNPACK CPU backend](https://docs.pytorch.org/executorch/stable/backends-xnnpack.html). Always enabled. |
    | `vulkan` | [Vulkan GPU backend](https://docs.pytorch.org/executorch/stable/backends-vulkan.html). Only supported on Android. |
  </Accordion>

  <Accordion title="LiteRT Inference Metadata">
    Use the `LiteRTInferenceMetadata` metadata type to compile a PyTorch [`nn.Module`](https://docs.pytorch.org/docs/stable/generated/torch.nn.Module.html) for inference with [LiteRT](https://ai.google.dev/edge/litert):

    ```py ai.py icon="python" focus={1,5-8,12-18} theme={null}
    from muna.beta import LiteRTInferenceMetadata
    from torch import randn, Tensor
    from torch.nn import Module

    # Given a PyTorch model...
    model: Module = ...
    # With some example arguments...
    example_args: list[Tensor] = [randn(1, 3, 224, 224)]

    @compile(
        ...,
        metadata=[
            # Use LiteRT for model inference
            LiteRTInferenceMetadata(
                model=model,
                model_args=example_args
            )
        ]
    )
    def predict() -> None:
        pass
    ```
  </Accordion>

  <Accordion title="TensorFlow Lite Interpreter Metadata">
    Use the `TFLiteInterpreterMetadata` metadata type to compile a TensorFlow Lite
    [`Interpreter`](https://ai.google.dev/edge/api/tflite/python/tf/lite/Interpreter):

    ```py ai.py icon="python" focus={1,4-6,10-16} theme={null}
    from muna.beta import TFLiteInterpreterMetadata
    from tensorflow import lite

    # Given a TFLite interpreter...
    model_path = "/path/to/model.tflite"
    interpreter = lite.Interpreter(model_path)

    @compile(
        ...,
        metadata=[
            # Use TensorFlow Lite for model inference
            TFLiteInterpreterMetadata(
                interpreter=interpreter,
                model_path=model_path
            )
        ]
    )
    def predict(...) -> None:
        pass
    ```

    <Warning>
      The TensorFlow Lite model file must exist at the provided `model_path` **within the compiler sandbox**.
    </Warning>
  </Accordion>

  <Accordion title="QNN Inference Metadata">
    Use the `QnnInferenceMetadata` metadata type to compile a PyTorch [`nn.Module`](https://docs.pytorch.org/docs/stable/generated/torch.nn.Module.html) to a [Qualcomm QNN](https://docs.qualcomm.com/bundle/publicresource/topics/80-63442-50/introduction.html?product=1601111740009302) context binary:

    ```py ai.py icon="python" focus={1,5-8,12-20} theme={null}
    from muna.beta import QnnInferenceMetadata
    from torch import randn, Tensor
    from torch.nn import Module

    # Given a PyTorch model...
    model: Module = ...
    # With some example arguments...
    example_args: list[Tensor] = [randn(1, 3, 224, 224)]

    @compile(
        ...,
        metadata=[
            # Use QNN for model inference
            QnnInferenceMetadata(
                model=model,
                model_args=example_args,
                backend="gpu",
                quantization=None
            )
        ]
    )
    def predict() -> None:
        pass
    ```

    <Note>
      The QNN inference backend is only available on Android and Windows devices with Qualcomm processors.
    </Note>

    ### QNN Hardware Backends

    QNN requires that a hardware device `backend` is specified ahead of time. Below are supported backends:

    | Backend | Notes |
    | :- | :- |
    | `cpu` | Reference `aarch64` CPU backend. |
    | `gpu` | Adreno GPU backend, accelerated by OpenCL. |
    | `htp` | Hexagon NPU backend. |

    <Info>
      Learn more about [QNN hardware backends](https://docs.qualcomm.com/bundle/publicresource/topics/80-63442-50/backend.html?product=1601111740009302).
    </Info>

    ### QNN Model Quantization

    When using the `htp` backend, you **must** specify a model `quantization` mode as the Hexagon NPU only supports
    running integer-quantized models. Below are supported quantization modes:

    | Quantization | Notes |
    | :- | :- |
    | `w8a8` | Weights and activations are quantized to `uint8`. |
    | `w8a16` | Weights are quantized to `uint8` while activations are quantized to `uint16`. |
    | `w4a8` | Weights are quantized to `uint4` while activations are quantized to `uint8`. |
    | `w4a16` | Weights are quantized to `uint4` while activations are quantized to `uint16`. |
  </Accordion>

  <Accordion title="OpenVINO Inference Metadata">
    Use the `OpenVINOInferenceMetadata` metadata type to compile a PyTorch [`nn.Module`](https://docs.pytorch.org/docs/stable/generated/torch.nn.Module.html) to [OpenVINO](https://docs.openvino.ai/2025/index.html) IR:

    ```py ai.py icon="python" focus={1,5-8,12-18} theme={null}
    from muna.beta import OpenVINOInferenceMetadata
    from torch import randn, Tensor
    from torch.nn import Module

    # Given a PyTorch model...
    model: Module = ...
    # With some example arguments...
    example_args: list[Tensor] = [randn(1, 3, 224, 224)]

    @compile(
        ...,
        metadata=[
            # Use OpenVINO for model inference
            OpenVINOInferenceMetadata(
                model=model,
                model_args=example_args
            )
        ]
    )
    def predict() -> None:
        pass
    ```

    At runtime, the OpenVINO IR will be used for inference with the [OpenVINO toolkit](https://github.com/openvinotoolkit/openvino).

    <Note>
      The OpenVINO inference backend is only available on Linux and Windows `x86_64` devices with Intel processors.
    </Note>
  </Accordion>

  <Accordion title="IREE Inference Metadata">
    Use the `muna.beta.IREEInferenceMetadata` metadata type to compile a PyTorch [`nn.Module`](https://docs.pytorch.org/docs/stable/generated/torch.nn.Module.html) for inference with [IREE](https://iree.dev/):

    ```py ai.py icon="python" focus={1,5-8,12-19} theme={null}
    from muna.beta import IREEInferenceMetadata
    from torch import randn, Tensor
    from torch.nn import Module

    # Given a PyTorch model...
    model: Module = ...
    # With some example arguments...
    example_args: list[Tensor] = [randn(1, 3, 224, 224)]

    @compile(
        ...,
        metadata=[
            # Use IREE for model inference
            IREEInferenceMetadata(
                model=model,
                model_args=example_args,
                backend="vulkan"
            )
        ]
    )
    def predict() -> None:
        pass
    ```

    <Note>
      The IREE inference backend is only available on Android devices.
    </Note>

    ### IREE HAL Target Backends

    IREE supports several HAL target backends that
    the `model` can be compiled against. Below are targets that are currently supported by Muna:

    | Target | Notes |
    | :- | :- |
    | `vulkan` | [Vulkan GPU backend](https://iree.dev/guides/deployment-configurations/gpu-vulkan/). Only supported on Android. |
  </Accordion>

  <Accordion title="MIGraphX Inference Metadata">
    *Coming soon* 🤫.
  </Accordion>
</AccordionGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.