Skip to main content

Prerequisites

Before starting, ensure the following:
  • NVIDIA Jetson AGX Orin Devkit is set up with JetPack 6.1 or later.
  • CUDA Toolkit and cuDNN are installed.
  • Verify that the Jetson AGX Orin is in high-performance mode:

Installing and running SGLang with Jetson Containers

Clone the jetson-containers github repository:
Run the installation script:
Build the container image:
Run the container:
Or you can also manually run a container with this command:

Running Inference

Launch the server:
The quantization and limited context length (--dtype half --context-length 8192) are due to the limited computational resources in Nvidia jetson kit. A detailed explanation can be found in Server Arguments. After launching the engine, refer to Chat completions to test the usability.

Structured output with XGrammar

Please refer to SGLang doc structured output.
Thanks to the support from Nurgaliyev Shakhizat, Dustin Franklin and Johnny Núñez Cano.

References