Skip to main content

Runtime

Environment VariableDefaultDescription
SGLANG_EXTERNAL_MODEL_PACKAGEnot setInstalled package that registers out-of-tree diffusion pipelines and component models. The package is imported once in every process.
SGLANG_DIFFUSION_PLATFORM_OVERRIDEnot setSelect cpu, cuda, rocm, mps, npu, musa, or an installed sglang.multimodal_gen.platforms entry-point name. XPU remains automatic-only. When unset, SGLang Diffusion detects installed platform plugins before trying built-in platforms.
SGLANG_PLUGINSnot setComma-separated allowlist for installed sglang.multimodal_gen.plugins hooks. When unset, all discovered diffusion hooks load.
SGLANG_DIFFUSION_TARGET_DEVICEcudaTarget device for inference (cuda, rocm, xpu, npu, musa, mps, cpu)
SGLANG_DIFFUSION_ATTENTION_BACKENDnot setOverride attention backend via env var (e.g. fa, torch_sdpa, sage_attn)
SGLANG_DIFFUSION_ATTENTION_CONFIGnot setPath to attention backend configuration file (JSON/YAML)
SGLANG_DIFFUSION_STAGE_LOGGINGfalseEnable per-stage timing logs
SGLANG_DIFFUSION_SERVER_DEV_MODEfalseEnable dev-only HTTP endpoints for debugging
SGLANG_DIFFUSION_TORCH_PROFILER_DIRnot setDirectory for torch profiler traces (absolute path). Enables profiling when set
SGLANG_DIFFUSION_CACHE_ROOT~/.cache/sgl_diffusionRoot directory for cache files
SGLANG_DIFFUSION_CONFIG_ROOT~/.config/sgl_diffusionRoot directory for configuration files
SGLANG_DIFFUSION_LOGGING_LEVELINFODefault logging level
SGLANG_DIFFUSION_IPC_A2AtrueEnable CUDA-IPC all-to-all for eligible same-host, peer-accessible TP1 + two-rank Ulysses groups. Unsupported topology falls back to NCCL; set 0 to force NCCL.
SGLANG_DIFFUSION_IPC_A2A_TIMEOUT_MS10000Peer-wait timeout in milliseconds. This is a hang backstop, not a per-step budget; raise it only for known long stalls such as layerwise offload.
SGLANG_DIFFUSION_IPC_A2A_MAX_BUFFERS16Maximum cached IPC staging-buffer shape pairs. Cap this for multi-resolution serving to bound permanent staging memory.
SGLANG_USE_RUNAI_MODEL_STREAMERtrueUse Run:AI model streamer for model loading
SGLANG_KITCHEN_INT8_MAX_ROWS8192Max activation rows per kitchen_int8 fused GEMM. Set 0 to disable row splitting.
SGLANG_KITCHEN_INT8_MIN_SPLIT_N8192Minimum output features before kitchen_int8 splits large-M GEMMs. Narrow outputs stay as a single call.

Platform-Specific

Apple MPS

Environment VariableDefaultDescription
SGLANG_USE_MLXnot setSRT only: enables the MLX serving backend. It has no effect on SGLang Diffusion, which uses PyTorch MPS.

ROCm (AMD GPUs)

Environment VariableDefaultDescription
SGLANG_USE_ROCM_VAEfalseUse AITer GroupNorm in VAE for improved performance on ROCm
SGLANG_USE_ROCM_CUDNN_BENCHMARKfalseEnable MIOpen auto-tuning for VAE conv layers on ROCm

Quantization

Environment VariableDefaultDescription
SGLANG_DIFFUSION_FLASHINFER_FP4_GEMM_BACKENDnot setOptional FlashInfer FP4 GEMM backend override for diffusion NVFP4. When unset, SGLang defaults to flashinfer_trtllm.
SGLANG_DIFFUSION_ENABLE_W8A8_FP8_GEMMfalseExperimental opt-in for fused W8A8 FP8 GEMM in diffusion weight-only FP8 linears. When disabled, FP8 weights are dequantized to the compute dtype before matmul. Enabling this dynamically quantizes activations to FP8 and may change output quality.
SGLANG_DIFFUSION_ENABLE_MXFP8_ATTENTIONfalseEnable Ascend MXFP8 FA for supported non-causal self-attention layers. Applies to online MXFP8Config and offline ModelSlim W8A8_MXFP8 checkpoints. Unsupported calls continue to use the regular attention path.
SGLANG_DIFFUSION_MXFP8_FA_HEAD_CHUNK_SIZE4Maximum number of attention heads processed by each Ascend MXFP8 FA call. Smaller chunks may improve performance for large workloads but add kernel launches; the optimal value depends on the model and input shape. Set to 0 to disable head splitting.

Caching Acceleration

These variables configure caching acceleration for Diffusion Transformer (DiT) models. SGLang supports multiple caching strategies - see caching documentation for an overview.

Cache-DiT Configuration

See cache-dit documentation for detailed configuration.
Environment VariableDefaultDescription
SGLANG_CACHE_DIT_ENABLEDfalseEnable Cache-DiT acceleration
SGLANG_CACHE_DIT_FN1First N blocks to always compute
SGLANG_CACHE_DIT_BN0Last N blocks to always compute
SGLANG_CACHE_DIT_WARMUP4Warmup steps before caching
SGLANG_CACHE_DIT_RDT0.24Residual difference threshold
SGLANG_CACHE_DIT_MC3Max continuous cached steps
SGLANG_CACHE_DIT_TAYLORSEERfalseEnable TaylorSeer calibrator
SGLANG_CACHE_DIT_TS_ORDER1TaylorSeer order (1 or 2)
SGLANG_CACHE_DIT_DMDfalseEnable the DMD (Dynamic Mode Decomposition) calibrator (mutually exclusive with TaylorSeer)
SGLANG_CACHE_DIT_DMD_HISTORY6DMD snapshot window length (5-6 typical; needs >= 4 uniformly spaced snapshots, otherwise DMD falls back to TaylorSeer)
SGLANG_CACHE_DIT_DMD_RANK0DMD SVD truncation rank (0 = automatic)
SGLANG_CACHE_DIT_DMD_RIDGE1e-8DMD Tikhonov regularization added to the inverted singular values
SGLANG_CACHE_DIT_DMD_SVD_PRECISIONmediumDMD SVD precision (low/medium/high)
SGLANG_CACHE_DIT_SCM_PRESETnoneSCM preset (none/slow/medium/fast/ultra)
SGLANG_CACHE_DIT_SCM_POLICYdynamicSCM caching policy
SGLANG_CACHE_DIT_SCM_COMPUTE_BINSnot setCustom SCM compute bins
SGLANG_CACHE_DIT_SCM_CACHE_BINSnot setCustom SCM cache bins

Cache-DiT Secondary Transformer

For dual-transformer models (e.g., Wan2.2 with high/low-noise experts), these variables configure caching for the secondary transformer. Each falls back to its primary counterpart if not set.
Environment VariableDefaultDescription
SGLANG_CACHE_DIT_SECONDARY_FN(from primary)First N blocks to always compute
SGLANG_CACHE_DIT_SECONDARY_BN(from primary)Last N blocks to always compute
SGLANG_CACHE_DIT_SECONDARY_WARMUP(from primary)Warmup steps before caching
SGLANG_CACHE_DIT_SECONDARY_RDT(from primary)Residual difference threshold
SGLANG_CACHE_DIT_SECONDARY_MC(from primary)Max continuous cached steps
SGLANG_CACHE_DIT_SECONDARY_TAYLORSEER(from primary)Enable TaylorSeer calibrator
SGLANG_CACHE_DIT_SECONDARY_TS_ORDER(from primary)TaylorSeer order (1 or 2)
SGLANG_CACHE_DIT_SECONDARY_DMD(from primary)Enable the DMD calibrator (mutually exclusive with TaylorSeer)
SGLANG_CACHE_DIT_SECONDARY_DMD_HISTORY(from primary)DMD snapshot window length
SGLANG_CACHE_DIT_SECONDARY_DMD_RANK(from primary)DMD SVD truncation rank (0 = automatic)
SGLANG_CACHE_DIT_SECONDARY_DMD_RIDGE(from primary)DMD Tikhonov regularization term
SGLANG_CACHE_DIT_SECONDARY_DMD_SVD_PRECISION(from primary)DMD SVD precision (low/medium/high)

Cloud Storage

These variables configure S3-compatible cloud storage for automatically uploading generated images and videos.
Environment VariableDefaultDescription
SGLANG_CLOUD_STORAGE_TYPEnot setSet to s3 to enable cloud storage
SGLANG_S3_BUCKET_NAMEnot setThe name of the S3 bucket
SGLANG_S3_ENDPOINT_URLnot setCustom endpoint URL (for MinIO, OSS, etc.)
SGLANG_S3_REGION_NAMEus-east-1AWS region name
SGLANG_S3_ACCESS_KEY_IDnot setAWS Access Key ID
SGLANG_S3_SECRET_ACCESS_KEYnot setAWS Secret Access Key

CUDA Crash Debugging

These variables enable kernel API logging and optional input/output dumps around diffusion CUDA kernel call boundaries. They are useful when tracking down CUDA crashes such as illegal memory access, device-side assert, or shape mismatches in custom kernels.
Environment VariableDefaultDescription
SGLANG_KERNEL_API_LOGLEVEL0Controls crash-debug kernel API logging. 1 logs API names, 3 logs tensor metadata, 5 adds tensor statistics, and 10 also writes dump snapshots.
SGLANG_KERNEL_API_LOGDESTstdoutDestination for crash-debug kernel API logs. Use stdout, stderr, or a file path. %i is replaced with the process PID.
SGLANG_KERNEL_API_DUMP_DIRsglang_kernel_api_dumpsOutput directory for level-10 kernel API dumps. %i is replaced with the process PID.
SGLANG_KERNEL_API_DUMP_INCLUDEnot setComma-separated wildcard patterns for kernel API names to include in level-10 dumps.
SGLANG_KERNEL_API_DUMP_EXCLUDEnot setComma-separated wildcard patterns for kernel API names to exclude from level-10 dumps.