Skip to main content
Muna’s compiler supports specifying metadata, allowing you to configure the compiler or provide additional information.
Use the TensorRTInferenceMetadata metadata type to compile a PyTorch nn.Module to TensorRT:
ai.py
The TensorRT inference backend is only available on Linux and Windows devices with compatible Nvidia GPUs.
We are working on adding support for consumer RTX GPUs with TensorRT for RTX.

Target CUDA Architectures

TensorRT engines must be compiled for specific target CUDA architectures. Below are CUDA architectures that our compiler supports:

TensorRT Inference Precision

TensorRT allows for specifying the inference engine’s precision. Below are supported precision modes:
Use the OnnxRuntimeInferenceMetadata metadata type to compile a PyTorch nn.Module for inference with ONNXRuntime:
ai.py
Use the OnnxRuntimeInferenceSessionMetadata metadata type to compile an OnnxRuntime InferenceSession:
ai.py
The ONNX model file must exist at the provided model_path within the compiler sandbox.
Use the CoreMLInferenceMetadata metadata type to compile a PyTorch nn.Module to CoreML:
ai.py
The CoreML inference backend is only available on iOS, macOS, and visionOS devices.
Use the LlamaCppInferenceMetadata metadata type to compile a Llama instance:
llm.py

Llama.cpp Hardware Backends

Llama.cpp supports several hardware backends to accelerate model inference. Below are targets that are currently supported by Muna:
Use the ExecuTorchInferenceMetadata metadata type to compile a PyTorch nn.Module for inference with ExecuTorch:
ai.py
The ExecuTorch inference backend is only available on Android.

ExecuTorch Hardware Backends

ExecuTorch supports several hardware backends to accelerate model inference. Below are targets that are currently supported by Muna:
Use the LiteRTInferenceMetadata metadata type to compile a PyTorch nn.Module for inference with LiteRT:
ai.py
Use the TFLiteInterpreterMetadata metadata type to compile a TensorFlow Lite Interpreter:
ai.py
The TensorFlow Lite model file must exist at the provided model_path within the compiler sandbox.
Use the QnnInferenceMetadata metadata type to compile a PyTorch nn.Module to a Qualcomm QNN context binary:
ai.py
The QNN inference backend is only available on Android and Windows devices with Qualcomm processors.

QNN Hardware Backends

QNN requires that a hardware device backend is specified ahead of time. Below are supported backends:
Learn more about QNN hardware backends.

QNN Model Quantization

When using the htp backend, you must specify a model quantization mode as the Hexagon NPU only supports running integer-quantized models. Below are supported quantization modes:
Use the OpenVINOInferenceMetadata metadata type to compile a PyTorch nn.Module to OpenVINO IR:
ai.py
At runtime, the OpenVINO IR will be used for inference with the OpenVINO toolkit.
The OpenVINO inference backend is only available on Linux and Windows x86_64 devices with Intel processors.
Use the muna.beta.IREEInferenceMetadata metadata type to compile a PyTorch nn.Module for inference with IREE:
ai.py
The IREE inference backend is only available on Android devices.

IREE HAL Target Backends

IREE supports several HAL target backends that the model can be compiled against. Below are targets that are currently supported by Muna:
Coming soon 🤫.