Support Matrix — TensorRT-LLM

nvidia.github.io

Tin mới

NVIDIA Ada Lovelace Architecture

TensorRT-LLM optimizes the performance of a range of well-known models on NVIDIA GPUs. The following sections provide a list of supported GPU architectures as well as important features implemented in TensorRT-LLM.

Baichuan/Baichuan2

LLaMA/LLaMA 2/LLaMA 3/LLaMA 3.1

Phi-1.5/Phi-2/Phi-3

Qwen/Qwen1.5/Qwen2/Qwen3

NVIDIA GB200 NVL72

NVIDIA Blackwell Architecture

NVIDIA Grace Hopper Superchip

NVIDIA Hopper Architecture

NVIDIA Ada Lovelace Architecture

NVIDIA Ampere Architecture