TensorRT Inference Metadata
TensorRT Inference Metadata
Use the
TensorRTInferenceMetadata metadata type to compile a PyTorch nn.Module to TensorRT:ai.py
The TensorRT inference backend is only available on Linux and Windows devices with compatible Nvidia GPUs.
Target CUDA Architectures
TensorRT engines must be compiled for specific target CUDA architectures. Below are CUDA architectures that our compiler supports:TensorRT Inference Precision
TensorRT allows for specifying the inference engine’s precision. Below are supported precision modes:OnnxRuntime Inference Metadata
OnnxRuntime Inference Metadata
Use the
OnnxRuntimeInferenceMetadata metadata type to compile a PyTorch nn.Module for inference with ONNXRuntime:ai.py
OnnxRuntime Inference Session Metadata
OnnxRuntime Inference Session Metadata
Use the
OnnxRuntimeInferenceSessionMetadata metadata type to compile an OnnxRuntime InferenceSession:ai.py
CoreML Inference Metadata
CoreML Inference Metadata
Llama.cpp Inference Metadata
Llama.cpp Inference Metadata
ExecuTorch Inference Metadata
ExecuTorch Inference Metadata
Use the
ExecuTorchInferenceMetadata metadata type to compile a PyTorch nn.Module for inference with ExecuTorch:ai.py
The ExecuTorch inference backend is only available on Android.
ExecuTorch Hardware Backends
ExecuTorch supports several hardware backends to accelerate model inference. Below are targets that are currently supported by Muna:LiteRT Inference Metadata
LiteRT Inference Metadata
TensorFlow Lite Interpreter Metadata
TensorFlow Lite Interpreter Metadata
Use the
TFLiteInterpreterMetadata metadata type to compile a TensorFlow Lite
Interpreter:ai.py
QNN Inference Metadata
QNN Inference Metadata
Use the
QnnInferenceMetadata metadata type to compile a PyTorch nn.Module to a Qualcomm QNN context binary:ai.py
The QNN inference backend is only available on Android and Windows devices with Qualcomm processors.
QNN Hardware Backends
QNN requires that a hardware devicebackend is specified ahead of time. Below are supported backends:Learn more about QNN hardware backends.
QNN Model Quantization
When using thehtp backend, you must specify a model quantization mode as the Hexagon NPU only supports
running integer-quantized models. Below are supported quantization modes:OpenVINO Inference Metadata
OpenVINO Inference Metadata
Use the At runtime, the OpenVINO IR will be used for inference with the OpenVINO toolkit.
OpenVINOInferenceMetadata metadata type to compile a PyTorch nn.Module to OpenVINO IR:ai.py
The OpenVINO inference backend is only available on Linux and Windows
x86_64 devices with Intel processors.IREE Inference Metadata
IREE Inference Metadata
Use the
muna.beta.IREEInferenceMetadata metadata type to compile a PyTorch nn.Module for inference with IREE:ai.py
The IREE inference backend is only available on Android devices.
IREE HAL Target Backends
IREE supports several HAL target backends that themodel can be compiled against. Below are targets that are currently supported by Muna:MIGraphX Inference Metadata
MIGraphX Inference Metadata
Coming soon 🤫.