Detailed description of the requested feature
Alpamayo Example mentions that "only FP16 is supported for Alpamayo export in this release (v0.9.0)".
We would like to know if NVIDIA has plans to support quantized Alpamayo model export, such as FP8 or NVFP4, in the future.
We also noticed alpamayo-recipes repository provides a quantization script. We are wondering if we will be able to convert these quantized models using TensorRT Edge LLM for inference on edge devices.
Timeline
If there are plans to release this, any roadmap information would be very helpful.
Target hardware/use case
Hardware: DGX Spark (Also, is it possible to run TensorRT Edge LLM on DRIVE AGX Orin?)
CUDA: 13.0
Use case: Running quantized Alpamayo on edge devices.
Detailed description of the requested feature
Alpamayo Example mentions that "only FP16 is supported for Alpamayo export in this release (v0.9.0)".
We would like to know if NVIDIA has plans to support quantized Alpamayo model export, such as FP8 or NVFP4, in the future.
We also noticed alpamayo-recipes repository provides a quantization script. We are wondering if we will be able to convert these quantized models using TensorRT Edge LLM for inference on edge devices.
Timeline
If there are plans to release this, any roadmap information would be very helpful.
Target hardware/use case
Hardware: DGX Spark (Also, is it possible to run TensorRT Edge LLM on DRIVE AGX Orin?)
CUDA: 13.0
Use case: Running quantized Alpamayo on edge devices.