Skip to main content
Muna’s native client, muna.predictions, runs any compiled model locally.

Creating Predictions

Streaming Predictions

Muna supports consuming the partial results of an inference request as they are made available by the compiled model:

Consuming Inference Streams

Streaming in Muna is designed to fully separate how a compiled function is implemented from how the function might be consumed. Consider these two functions:
Here are the results of creating vs. streaming each function at runtime:
In this case, the single result is returned:
In this case, the Muna client will consume all partial results yielded by the function then return the very last one:
In this case, the Muna client will return an inference stream with the single result returned by the function:
In this case, the Muna client will provide an inference stream containing all partial results yielded by the function:
You can choose how to consume a compiled function depending on what works best for your user experience. You don’t have to care about the underlying function!

Using Inference Values

Muna supports a fixed set of value types for inference input and output values:
Muna supports the following floating-point numbers:
In languages that don’t support fixed-size floating point scalars, the data type for floating point values defaults to float32. Use a tensor constructor to explicitly specify the data type.
Support for half-precision floating point scalars float16 is planned for the future depending on language support.
Muna supports floating point vectors (i.e. one-dimensional floating point tensors):
Although Muna supports input vectors, compiled models will always output either scalars or Tensor instances—never plain vectors.
Muna supports floating point tensors:
Muna supports several signed and unsigned integer scalars:
When integer scalars are passed to compiled models, the data type defaults to int32. Use a tensor constructor to explicitly specify the data type.
Muna supports integer vectors (i.e. one-dimensional integer tensors) of the aforementioned integer types:
Although Muna supports input vectors, compiled models will always output either scalars or Tensor instances—never plain vectors.
Muna supports integer tensors:
Unsigned integer tensors are not supported in our Android client because of missing language support in Java.
Muna supports boolean scalars:
Muna supports boolean vectors (i.e. one-dimensional boolean tensors):
Although Muna supports input vectors, compiled models will always output either scalars or Tensor instances—never plain vectors.
Muna supports boolean tensors:
Muna assumes that boolean values are 1 byte.
Muna supports string values:
Muna supports lists of values, each with potentially different types:
Input list values must be JSON-serializable.
Muna supports dictionary values:
Input dictionary values must be JSON-serializable.
Muna supports images, represented as raw pixel buffers with 8 bytes per pixel and interleaved by channel. Muna supports three pixel buffer formats:Some client SDKs provide Image utility types for working with images:
Muna supports binary blobs:
Because Muna’s security model prohibits file system access, binary input values are always fully read into memory before being passed to the compiled model.To run inference on large files, consider mapping the file into memory using mmap or your environment’s equivalent.

Specifying Inference Acceleration

Muna lets you choose which processor runs inference on the local device, per-request. Specify an acceleration when creating or streaming a prediction:

Specifying the Acceleration

Below are the currently supported acceleration specifiers:

Specifying the Device

Some Muna clients allow you to specify the acceleration device used to run inference. Our clients expose this field as an untyped integer or pointer. The underlying type depends on the current operating system:
Currently unsupported.
The device is reinterpreted as a Metal device with type id<MTLDevice>.
The device is reinterpeted as a pointer to a CUDA device ID or HIP device ID with type int*.
The device is reinterpreted as a Metal device with type id<MTLDevice>.
The device is reinterpreted as a WebGPU device with type GPUDevice.
The device is reinterpreted as a DirectX 12 device with type ID3D12Device*.
You should never specify the inference device unless you know what the hell you’re doing.
The inference device is merely a hint. Setting a device does not guarantee that all or any operation in the compiled function will actually use that acceleration device.