Quick Start Guide — TensorRT LLM

nvidia.github.io

Tin mới

TensorRT LLM - Home

TensorRT LLM - Home

Installation Guide

Supported Hardware

Generate text asynchronously

Generate text in streaming

Distributed LLM Generation

Generate text with guided decoding

Control generated text using logits processor

Generate text with multiple LoRA adapters

Speculative Decoding

KV Cache Connector

KV Cache Offloading

Runtime Configuration Examples

Sampling Techniques Showcase

LMCache KV Cache Connector

Run LLM-API with pytorch backend on Slurm

Run trtllm-bench with pytorch backend on Slurm

Run trtllm-serve with pytorch backend on Slurm

VisualGen Examples

Online Serving Examples

Aiperf Client For Multimodal

Curl Chat Client For Multimodal

Curl Completion Client

Curl Responses Client

Deepseek R1 Reasoning Parser

OpenAI Chat Client

OpenAI Chat Client for Multimodal

OpenAI Completion Client

Openai Completion Client For Lora

OpenAI Completion Client with JSON Schema

OpenAI Responses Client

Prometheus Metrics

Dynamo K8s Example

Deployment Guide for Nemotron v3 (Ultra & Super) on TensorRT LLM - Blackwell & Hopper Hardware

Deployment Guide for DeepSeek R1 on TensorRT LLM - Blackwell & Hopper Hardware

Deployment Guide for Llama3.3 70B on TensorRT LLM - Blackwell & Hopper Hardware

Deployment Guide for Llama4 Scout 17B on TensorRT LLM - Blackwell & Hopper Hardware

Deployment Guide for GPT-OSS on TensorRT-LLM - Blackwell Hardware

Deployment Guide for Qwen3 on TensorRT LLM - Blackwell & Hopper Hardware

Deployment Guide for Qwen3.8 MoE and Qwen3.5 MoE on TensorRT LLM - Blackwell Hardware

Deployment Guide for Kimi K2 Thinking on TensorRT LLM - Blackwell

Deployment Guide for Kimi K3 on TensorRT LLM - Blackwell

Deployment Guide for GLM-5 on TensorRT LLM - Blackwell Hardware

Deployment Guide for MiniMax-M3 on TensorRT LLM

CPU Affinity configuration in TensorRT LLM

Visual Generation (Beta)

Adding a New Model

Run benchmarking with trtllm-serve

LLM API Introduction

GuidedDecodingParams