Skip to main content

Overview

Allows hardware vendors and developers to extend SGLang without modifying the main repository code. The framework provides two plugin types, both discovered via Python’s standard setuptools entry_points:
Plugin TypeEntry Point GroupsPurpose
Hardware Platform Pluginsglang.srt.platforms
sglang.multimodal_gen.platforms
Register a custom hardware platform (device operations, KV cache pools, attention backends, graph capture, compilation backends, etc.)
General Pluginsglang.srt.plugins
sglang.multimodal_gen.plugins
Inject hooks (before/after/around/replace) into any function/method, or replace entire classes

Principles

  • Non-intrusive: Built-in platforms remain the fallback when no OOT platform activates.
  • Install-time discovery: Plugins are discovered from Python entry points after installation.
  • Environment variable control: SGLANG_PLATFORM selects an SRT platform, SGLANG_DIFFUSION_PLATFORM_OVERRIDE selects a diffusion platform, and SGLANG_PLUGINS filters general plugins in either hook group.

Current scope

The platform plugin system targets out-of-tree (OOT) hardware platforms. Diffusion support is experimental and covers the seams documented below, not every device-specific branch.

Architecture

Runtime-specific platform interfaces

SRT and diffusion have separate platform base classes and platform identity types in sglang.srt.platforms and sglang.multimodal_gen.runtime.platforms. A package supporting both runtimes should define separate platform classes and use each runtime’s own PlatformEnum.OOT value. Each platform entry point resolves to a zero-argument activation callback:
Return a fully qualified class name when the provider can run, or None otherwise. Keep activation import-safe: do not access current_platform or initialize runtime or device state. Put required backend setup in the platform’s init_backend() method and reserve general plugins for hooks that the platform interface cannot express. Explicit selection enumerates entry-point metadata and imports only the selected callback. Automatic selection invokes installed platform callbacks to determine which provider is active. Diffusion reserves cpu, cuda, rocm, xpu, mps, npu, and musa. Automatic discovery also rejects duplicate entry-point names; explicit selection validates only the selected name and does not import unrelated providers.

SRT selection

current_platform is a lazy singleton in sglang.srt.platforms. On first access it resolves the active platform through the following priority chain:

Diffusion selection

SGLang Diffusion resolves its platform in this order:
  1. When SGLANG_DIFFUSION_PLATFORM_OVERRIDE names cpu, cuda, rocm, mps, npu, or musa, select that built-in platform without hardware probing. XPU remains automatic-only, preserving the existing selector behavior.
  2. When it contains another name, load only the matching sglang.multimodal_gen.platforms entry point. An unknown name or a callback that returns None is an error.
  3. When it is unset, activate installed diffusion platform plugins. No active plugin continues to built-in detection, one selects that plugin, and multiple active plugins raise an error that asks you to set the selector. An activation callback that raises aborts startup rather than falling back, so a broken vendor runtime cannot silently run the job on a built-in platform.
  4. Try built-in platforms in order: MPS, XPU, ROCm, CUDA, NPU, MUSA, then CPU.
The existing override variable is therefore the single explicit selector for supported built-in aliases and OOT entry-point names. Selection resolves lazily, the first time anything in a process touches current_platform, so it needs no call site and happens in every process automatically. SGLang Diffusion also records which distribution supplied the selected platform, and skips the hooks of every other installed platform package.

Required platform initialization vs. optional hooks

The two mechanisms have different failure semantics, and it matters which one you use: Anything your hardware needs in order to be correct belongs on the Platform subclass, so a broken platform cannot silently serve. Reach for a general plugin only when the Platform interface has no seam for what you need — and please report that gap. A plugin shipped in the selected platform’s own distribution is treated as part of that platform’s contract: a failure to load it, to run its callback, or to apply any hook it registered aborts startup rather than leaving the platform half-initialized. Plugins from any other installed package stay best-effort, so a broken third party cannot take the server down. An explicit SGLANG_PLUGINS allowlist can disable any general plugin, including one from the selected platform package; required hardware setup therefore belongs in init_backend().

Plugin Loading Flow

Each runtime has a process-local hook registry. SRT retains its single load_plugins() activation step. Diffusion separates registration (load_plugins()) from target resolution (apply_plugin_hooks()), because resolving a dotted hook target can import that target’s entire module graph. Both runtimes honor SGLANG_PLUGINS. The loader is called at these SRT entry points:
Call SiteProcessTiming
cli/serve.py serve()MainBefore prepare_server_args()
launch_server.py mainMainBefore prepare_server_args()
engine.py _launch_subprocesses()MainBefore server_args.check_server_args()
scheduler.py run_scheduler_process()SubprocessBefore Scheduler() construction
Note: Diffusion plugin registration and hook application each run once per process. Spawned subprocesses start from a blank interpreter and establish their own hooks — nothing the parent patched survives the spawn boundary. SRT’s load_plugins() performs both phases in one call.
Diffusion launchers call apply_plugin_hooks(), which first performs registration if necessary and then resolves and patches targets. Scheduler children use a stricter lifecycle in runtime/managers/worker_bootstrap.py:
HTTP children likewise finish plugin registration and hook application before importing runtime.launch_server or materializing ServerArgs. This matters for class replacements and other hooks whose effects cannot be retroactively applied to classes or registrations created while importing the server module graph. The bootstrap module is the mp.Process target and imports no diffusion runtime module at module scope. ServerArgs is serialized inside ServerArgsPayload, so the multiprocessing unpickler cannot import pipeline configuration modules before the target starts. Workers always use a local spawn context, independent of an embedding application’s global multiprocessing setting. This ordering makes backend initialization precede plugin callback imports, hook-target resolution, ServerArgs materialization, and worker imports. A platform activation module necessarily loads before its own init_backend(); activation modules must therefore remain import-safe. Prefer platform methods and registries for required behavior. Each activation phase runs once per process behind a lock: a re-entrant call from a plugin callback returns and lets the outer call finish, another thread waits for it, and a failure is terminal. A callback must not hand activation to a second thread and join it — that deadlocks.

Offline scripts and the spawn boundary

spawn re-executes the launching script’s module scope in every child before it unpickles the target’s arguments, so an offline script’s own imports run ahead of that child’s platform initialization and ServerArgsPayload cannot help. The supported script layout is:
DiffGenerator on the sglang.multimodal_gen facade is a lazy proxy, so binding it at module scope costs no diffusion import and the child still reaches initialize_current_platform() with a clean module table. Module scope also stays open to envs and runtime.platforms, including runtime.platforms.plugins. These modules remain import-safe so a plugin can subclass Platform and register hooks before backend initialization. Everything else, including SamplingParams and PipelineConfig from the same facade, belongs inside the if __name__ == "__main__": guard or inside the function that uses it. A child that finds runtime modules already imported names them in a warning: hook application can still patch those modules, but classes and registrations created while importing them are already past reach.

Plugin Type 1: Hardware Platform Plugin

Description

A hardware platform plugin registers an SRT SRTPlatform subclass, a diffusion Platform subclass, or both. The selected class tells that runtime how to interact with a specific hardware backend.

SRT quick start

1. Create a minimal package:
2. pyproject.toml:
3. __init__.py — activation function:
4. device.py — device mixin:
5. platform.py — SRT platform:
6. Install and verify:

SRT platform interface reference

Identity Queries (from DeviceMixin)

MethodDefaultDescription
is_cuda()Based on _enumWhether this is an NVIDIA CUDA platform
is_rocm()Based on _enumWhether this is an AMD ROCm platform
is_npu()Based on _enumWhether this is a Huawei NPU platform
is_cpu()Based on _enumWhether this is a CPU-only platform
is_xpu()Based on _enumWhether this is an Intel XPU platform
is_musa()Based on _enumWhether this is a Moore Threads MUSA platform
is_cuda_alike()CUDA+ROCM+MUSATrue if the hardware supports CUDA-like APIs
is_out_of_tree()True for OOTAutomatically detected based on _enum = PlatformEnum.OOT

Device Operations (from DeviceMixin)

Methods annotated [Active] are called by SGLang core through current_platform — OOT implementations take effect immediately. Methods annotated [Planned] are reserved interfaces — SGLang core still uses hardcoded calls (e.g. torch.cuda.empty_cache()). OOT implementations will NOT take effect until the core is migrated in a future PR.
MethodDefaultStatusDescription
get_device(local_rank)raise NotImplementedErrorPlannedReturn torch.device for a given local rank
set_device(device)raise NotImplementedErrorPlannedSet the current device
get_device_name(device_id)raise NotImplementedErrorPlannedGet human-readable device name
get_device_uuid(device_id)raise NotImplementedErrorPlannedGet unique device identifier
get_device_capability(device_id)raise NotImplementedErrorPlannedGet DeviceCapability(major, minor). None if N/A
empty_cache()passPlannedRelease cached device memory
synchronize()passPlannedSynchronize device operations
get_device_total_memory(device_id)raise NotImplementedErrorActiveGet total device memory in bytes
get_available_memory(device_id)raise NotImplementedErrorPlannedReturn (free_bytes, total_bytes)
get_current_memory_usage(device)raise NotImplementedErrorActiveGet current peak memory usage in bytes
is_pin_memory_available(device=None)FalseActiveWhether pinned host memory is available for a target device
get_torch_distributed_backend_str()raise NotImplementedErrorPlannedDistributed backend string (e.g. “nccl”, “hccl”)
get_communicator_class()NonePlannedPlatform-specific communicator class
inference_mode()torch.inference_mode(True)PlannedReturn inference mode context manager
seed_everything(seed)Set random/np/torch seedsPlannedSet random seeds for reproducibility
verify_quantization(quant)passPlannedValidate quantization method support
get_cpu_architecture()Auto-detect x86/armPlannedDetect CPU architecture (CpuArchEnum)

Types (from DeviceMixin)

TypeDescription
PlatformEnumEnumeration of platform types: CUDA, ROCM, CPU, XPU, MUSA, NPU, TPU, MPS, OOT, UNSPECIFIED
CpuArchEnumCPU architecture: X86, ARM, UNSPECIFIED
DeviceCapabilityNamedTuple(major, minor) with comparison support. Methods: as_version_str(), to_int()

Capability Flags (from SRTPlatform)

MethodDefaultDescription
support_cuda_graph()FalseWhether device graph capture is supported (plain CUDA graph)
support_piecewise_cuda_graph()FalseWhether piecewise CUDA graph (torch.compile backend) is supported
supports_fp8()FalseWhether FP8 quantization is supported

Subsystem Factory Methods (from SRTPlatform)

MethodDefaultDescription
get_default_attention_backend()raise NotImplementedErrorDefault attention backend name
get_graph_runner_cls()raise NotImplementedErrorGraph Runner class
get_mha_kv_pool_cls()raise NotImplementedErrorMHA KV cache pool class
get_mla_kv_pool_cls()raise NotImplementedErrorMLA KV cache pool class
get_dsa_kv_pool_cls()raise NotImplementedErrorDSA KV cache pool class (DeepSeek V3.2)
get_paged_allocator_cls()raise NotImplementedErrorPaged allocator class
get_quantization_config(quantization)raise NotImplementedErrorReturn hardware-specific quantization config for the specific quantization scheme, raise an error if not supported or return None to use the default config.
get_piecewise_backend_cls()raise NotImplementedErrorPiecewise compilation backend class
get_compile_backend(mode)“inductor”Compilation backend string
get_dispatch_key_name()“native”BaseFusedOp (fused-op) dispatch key name

Lifecycle Hooks (from SRTPlatform)

MethodInvocation TimingPurpose
apply_server_args_defaults(server_args)After ServerArgs parsing, in post_initSet platform-specific defaults
init_backend()In each worker, before model constructionOne-time backend initialization

Platform and plugin environment variables

VariableDescription
SGLANG_PLATFORMSelect the platform plugin by entry_point name (e.g. kunlun, demo_cuda). When set, only the named plugin’s activate() is called (front-loading filter) — other plugins are not touched. Additionally, general plugins (sglang.srt.plugins) from unselected platform packages are automatically skipped to avoid importing their dependencies. Required when multiple plugins would activate. Errors if the name is not found or if the plugin’s hardware is unavailable.
SGLANG_PLUGINSComma-separated whitelist of general plugin names to load from either hook group. It also filters automatic SRT platform discovery when SGLANG_PLATFORM is unset; explicit platform selection ignores it.

Add diffusion support

A package may support SRT, SGLang Diffusion, or both. Diffusion uses a zero-argument Platform subclass and separate platform and hook entry points:
Keep activation import-safe and return the fully qualified class name only when the backend is available:
The referenced class must be zero-argument constructible:
Set dispatch_key to the PyTorch dispatcher key used by direct torch.library registrations. Keep it separate from get_dispatch_key_name(), which selects CustomOp implementations. Bring up the remaining contract in dependency order:
  1. Implement get_device_name(), get_device_total_memory(), and get_available_gpu_memory() before constructing ServerArgs.
  2. Implement get_device() and get_local_torch_device() before worker binding.
  3. Configure distributed initialization, attention selection, and custom-op implementations for supported workloads.
  4. Run an end-to-end workload and audit model-specific kernels, compilation, and remaining device-family branches.
Register custom-op implementations from init_backend(), which runs once in each worker before any pipeline module is constructed:
init_backend() runs at most once per worker process, before the worker implementation is imported. Raising aborts startup and the failed initialization is not retried in that process, because registrations and other backend side effects may be only partially reversible. Custom-op dispatch is resolved when each operation is constructed, after backend initialization has completed. The selected callable therefore remains stable when the operation is compiled instead of changing on the first compiled call. The function receives the operation instance before its normal arguments. Registration matches the exact operation class and the value returned by get_dispatch_key_name(). Without a registration, SGLang looks for a matching forward_<key>() implementation on the operation, then falls back to forward_oot(), whose base implementation calls forward_native(). Returning "cuda", for example, lets a CUDA-compatible OOT platform reuse operation-specific forward_cuda() implementations without reporting CUDA platform identity. The sglang.multimodal_gen.plugins entry point remains for seams the Platform interface does not cover. Its callbacks run in launchers and workers before scheduler or model construction, so use a BEFORE or AROUND hook on GPUWorker.init_device_and_model() for rank-aware initialization. Hooks from platform distributions other than the selected one are skipped. Because these ship in your platform’s distribution, a failure at any stage — load, callback, or hook application — aborts startup, the same as init_backend().

Diffusion platform contract

Active interfaces are called through current_platform. Compatibility interfaces are retained adapters; unlisted Platform methods are not a stable OOT contract. Install the package, restart Python to refresh entry-point metadata, and verify explicit selection:

Diffusion platform limitations

  • SRT and diffusion require separate platform classes.
  • Only one external platform can be active per process; set the runtime selector if multiple callbacks activate.
  • get_all_to_all_communicator_cls() controls only all_to_all_4D(), not every collective, graph-capture, or synchronization path; get_device_communicator_cls() remains for compatibility.
  • Compile settings affect only build_torch_compile_kwargs() callers; audit static @torch.compile decorators and other direct compile paths.
  • Explicit attention selectors accept only AttentionBackendEnum names. Returning a custom backend works when none is selected; selecting it by name requires a hook or downstream patch.
  • Device-family branches remain outside the interface. Each supported workload needs a native fallback or an actionable unsupported-feature error.
  • The diffusion PlatformEnum.OOT identifies an external provider. Override built-in identity predicates only after auditing every enabled branch.

Plugin Type 2: General Plugin

Description

General function plugins inject behavior into SRT or diffusion without requiring a custom platform. The two runtimes have separate entry-point groups and hook registries:
Register an SRT plugin under sglang.srt.plugins and a diffusion plugin under sglang.multimodal_gen.plugins, and import the hook API from the matching module above. Each runtime applies only its own registry, so a hook registered through the other runtime’s API never runs. The examples below use SRT. The diffusion hook API lives alongside platform discovery in sglang.multimodal_gen.runtime.platforms.plugins; sglang.multimodal_gen.plugins is its entry-point group, not a Python module path. Use cases include:
  • Observability: Add logging, metrics, and tracing to any function
  • Behavior modification: Modify function arguments or return values
  • Performance profiling: Add timing to critical functions
  • A/B testing: Replace implementations at runtime

Quick Start

1. Create a minimal package:
2. pyproject.toml:
3. __init__.py — register hooks:
4. Install and run:

Hook Types

HookRegistry supports four hook types:
Hook TypeSignatureDescription
BEFOREfn(*args, **kwargs) -> (args, kwargs) | NoneRuns before the original. Return None to keep args unchanged, or (args, kwargs) to modify.
AFTERfn(result, *args, **kwargs) -> new_result | NoneRuns after the original. Return None to keep result, or a new value to replace.
AROUNDfn(original_fn, *args, **kwargs) -> resultWraps the original. You must call original_fn yourself. Full control over execution.
REPLACEfn(*args, **kwargs) -> result or classReplace the original function or class entirely. For class targets, pass a replacement class directly — it is substituted via setattr preserving isinstance()/issubclass() semantics.
Note: Only REPLACE accepts a class as the hook. Passing a class to BEFORE/AFTER/AROUND raises TypeError at registration time.

Registration API

Hooks can be registered using the imperative API or the decorator API:

Hook Target Resolution

Target paths use fully-qualified dotted notation. Both formats are supported:
  • Dotted: sglang.srt.managers.scheduler.Scheduler.__init__
  • Entry-points style: sglang.srt.managers.scheduler:Scheduler.__init__ (colon treated as dot)

Common SRT Hook Targets

TargetDescription
sglang.srt.server_args.ServerArgs.add_cli_argsAdd custom CLI arguments
sglang.srt.server_args.ServerArgs.post_initModify ServerArgs after parsing
sglang.srt.server_args.ServerArgs.check_server_argsAdd/relax validation
sglang.srt.managers.scheduler.Scheduler.initCustom scheduler state
sglang.srt.managers.scheduler.Scheduler.get_next_batch_to_runCustom scheduling policy
sglang.srt.managers.scheduler.Scheduler.run_batchProfiling / inspection
sglang.srt.managers.scheduler.Scheduler.process_batch_resultCustom metrics
sglang.srt.managers.tp_worker.TpModelWorker.initCustom worker state
sglang.srt.managers.tp_worker.TpModelWorker.forward_batch_generationForward pass wrapping

File Reference

FileDescription
sglang/srt/platforms/device_mixin.pyDeviceMixin base class and SRT platform identity types
sglang/srt/platforms/interface.pySRTPlatform base class (extends DeviceMixin)
sglang/srt/platforms/init.pycurrent_platform lazy singleton + discovery logic
sglang/multimodal_gen/runtime/platforms/interface.pyDiffusion Platform base class
sglang/multimodal_gen/runtime/platforms/init.pyDiffusion current_platform lazy singleton and built-in fallback order
sglang/multimodal_gen/runtime/platforms/plugins.pyDiffusion plugin registration, explicit hook-application phase, and hook registry
sglang/multimodal_gen/runtime/managers/worker_bootstrap.pyImport-neutral process specifications and spawn targets that initialize the backend before resolving runtime hooks
sglang/srt/plugins/init.pyload_plugins() + load_plugins_by_group()
sglang/srt/plugins/hook_registry.pyHookRegistry, HookType, plugin_hook decorator