Overview
Allows hardware vendors and developers to extend SGLang without modifying the main repository code. The framework provides two plugin types, both discovered via Python’s standardsetuptools entry_points:
| Plugin Type | Entry Point Groups | Purpose |
|---|---|---|
| Hardware Platform Plugin | sglang.srt.platformssglang.multimodal_gen.platforms | Register a custom hardware platform (device operations, KV cache pools, attention backends, graph capture, compilation backends, etc.) |
| General Plugin | sglang.srt.pluginssglang.multimodal_gen.plugins | Inject hooks (before/after/around/replace) into any function/method, or replace entire classes |
Principles
- Non-intrusive: Built-in platforms remain the fallback when no OOT platform activates.
- Install-time discovery: Plugins are discovered from Python entry points after installation.
- Environment variable control:
SGLANG_PLATFORMselects an SRT platform,SGLANG_DIFFUSION_PLATFORM_OVERRIDEselects a diffusion platform, andSGLANG_PLUGINSfilters general plugins in either hook group.
Current scope
The platform plugin system targets out-of-tree (OOT) hardware platforms. Diffusion support is experimental and covers the seams documented below, not every device-specific branch.Architecture
Runtime-specific platform interfaces
SRT and diffusion have separate platform base classes and platform identity types insglang.srt.platforms and sglang.multimodal_gen.runtime.platforms. A package supporting both runtimes should define separate platform classes and use each runtime’s own PlatformEnum.OOT value.
Each platform entry point resolves to a zero-argument activation callback:
None otherwise. Keep activation import-safe: do not access current_platform or initialize runtime or device state. Put required backend setup in the platform’s init_backend() method and reserve general plugins for hooks that the platform interface cannot express.
Explicit selection enumerates entry-point metadata and imports only the selected callback. Automatic selection invokes installed platform callbacks to determine which provider is active.
Diffusion reserves cpu, cuda, rocm, xpu, mps, npu, and musa. Automatic discovery also rejects duplicate entry-point names; explicit selection validates only the selected name and does not import unrelated providers.
SRT selection
current_platform is a lazy singleton in sglang.srt.platforms. On first access it resolves the active platform through the following priority chain:
Diffusion selection
SGLang Diffusion resolves its platform in this order:- When
SGLANG_DIFFUSION_PLATFORM_OVERRIDEnamescpu,cuda,rocm,mps,npu, ormusa, select that built-in platform without hardware probing. XPU remains automatic-only, preserving the existing selector behavior. - When it contains another name, load only the matching
sglang.multimodal_gen.platformsentry point. An unknown name or a callback that returnsNoneis an error. - When it is unset, activate installed diffusion platform plugins. No active plugin continues to built-in detection, one selects that plugin, and multiple active plugins raise an error that asks you to set the selector. An activation callback that raises aborts startup rather than falling back, so a broken vendor runtime cannot silently run the job on a built-in platform.
- Try built-in platforms in order: MPS, XPU, ROCm, CUDA, NPU, MUSA, then CPU.
current_platform, so it needs no call site and happens in every process automatically. SGLang Diffusion also records which distribution supplied the selected platform, and skips the hooks of every other installed platform package.
Required platform initialization vs. optional hooks
The two mechanisms have different failure semantics, and it matters which one you use:
Anything your hardware needs in order to be correct belongs on the
Platform subclass, so a broken platform cannot silently serve. Reach for a general plugin only when the Platform interface has no seam for what you need — and please report that gap.
A plugin shipped in the selected platform’s own distribution is treated as part of that platform’s contract: a failure to load it, to run its callback, or to apply any hook it registered aborts startup rather than leaving the platform half-initialized. Plugins from any other installed package stay best-effort, so a broken third party cannot take the server down. An explicit SGLANG_PLUGINS allowlist can disable any general plugin, including one from the selected platform package; required hardware setup therefore belongs in init_backend().
Plugin Loading Flow
Each runtime has a process-local hook registry. SRT retains its singleload_plugins() activation step. Diffusion separates registration
(load_plugins()) from target resolution (apply_plugin_hooks()), because
resolving a dotted hook target can import that target’s entire module graph.
Both runtimes honor SGLANG_PLUGINS.
The loader is called at these SRT entry points:
| Call Site | Process | Timing |
|---|---|---|
cli/serve.py serve() | Main | Before prepare_server_args() |
launch_server.py main | Main | Before prepare_server_args() |
engine.py _launch_subprocesses() | Main | Before server_args.check_server_args() |
scheduler.py run_scheduler_process() | Subprocess | Before Scheduler() construction |
Note: Diffusion plugin registration and hook application each run once per process. Spawned subprocesses start from a blank interpreter and establish their own hooks — nothing the parent patched survives the spawn boundary. SRT’s load_plugins() performs both phases in one call.
apply_plugin_hooks(), which first performs registration
if necessary and then resolves and patches targets. Scheduler children use a
stricter lifecycle in runtime/managers/worker_bootstrap.py:
runtime.launch_server or materializing ServerArgs. This matters for
class replacements and other hooks whose effects cannot be retroactively applied
to classes or registrations created while importing the server module graph.
The bootstrap module is the mp.Process target and imports no diffusion runtime
module at module scope. ServerArgs is serialized inside ServerArgsPayload, so
the multiprocessing unpickler cannot import pipeline configuration modules
before the target starts. Workers always use a local spawn context, independent
of an embedding application’s global multiprocessing setting.
This ordering makes backend initialization precede plugin callback imports,
hook-target resolution, ServerArgs materialization, and worker imports. A
platform activation module necessarily loads before its own init_backend();
activation modules must therefore remain import-safe. Prefer platform methods
and registries for required behavior.
Each activation phase runs once per process behind a lock: a re-entrant call
from a plugin callback returns and lets the outer call finish, another thread
waits for it, and a failure is terminal. A callback must not hand activation to
a second thread and join it — that deadlocks.
Offline scripts and the spawn boundary
spawn re-executes the launching script’s module scope in every child before
it unpickles the target’s arguments, so an offline script’s own imports run
ahead of that child’s platform initialization and ServerArgsPayload cannot
help. The supported script layout is:
DiffGenerator on the sglang.multimodal_gen facade is a lazy proxy, so
binding it at module scope costs no diffusion import and the child still reaches
initialize_current_platform() with a clean module table. Module scope also
stays open to envs and runtime.platforms, including runtime.platforms.plugins.
These modules remain import-safe so a plugin can subclass Platform and register
hooks before backend initialization. Everything else, including SamplingParams and PipelineConfig
from the same facade, belongs inside the if __name__ == "__main__": guard or
inside the function that uses it. A child
that finds runtime modules already imported names them in a warning: hook
application can still patch those modules, but classes and registrations
created while importing them are already past reach.
Plugin Type 1: Hardware Platform Plugin
Description
A hardware platform plugin registers an SRTSRTPlatform subclass, a diffusion Platform subclass, or both. The selected class tells that runtime how to interact with a specific hardware backend.
SRT quick start
1. Create a minimal package:pyproject.toml:
__init__.py — activation function:
device.py — device mixin:
platform.py — SRT platform:
SRT platform interface reference
Identity Queries (from DeviceMixin)
| Method | Default | Description |
|---|---|---|
is_cuda() | Based on _enum | Whether this is an NVIDIA CUDA platform |
is_rocm() | Based on _enum | Whether this is an AMD ROCm platform |
is_npu() | Based on _enum | Whether this is a Huawei NPU platform |
is_cpu() | Based on _enum | Whether this is a CPU-only platform |
is_xpu() | Based on _enum | Whether this is an Intel XPU platform |
is_musa() | Based on _enum | Whether this is a Moore Threads MUSA platform |
is_cuda_alike() | CUDA+ROCM+MUSA | True if the hardware supports CUDA-like APIs |
is_out_of_tree() | True for OOT | Automatically detected based on _enum = PlatformEnum.OOT |
Device Operations (from DeviceMixin)
Methods annotated [Active] are called by SGLang core throughcurrent_platform— OOT implementations take effect immediately. Methods annotated [Planned] are reserved interfaces — SGLang core still uses hardcoded calls (e.g.torch.cuda.empty_cache()). OOT implementations will NOT take effect until the core is migrated in a future PR.
| Method | Default | Status | Description |
|---|---|---|---|
get_device(local_rank) | raise NotImplementedError | Planned | Return torch.device for a given local rank |
set_device(device) | raise NotImplementedError | Planned | Set the current device |
get_device_name(device_id) | raise NotImplementedError | Planned | Get human-readable device name |
get_device_uuid(device_id) | raise NotImplementedError | Planned | Get unique device identifier |
get_device_capability(device_id) | raise NotImplementedError | Planned | Get DeviceCapability(major, minor). None if N/A |
empty_cache() | pass | Planned | Release cached device memory |
synchronize() | pass | Planned | Synchronize device operations |
get_device_total_memory(device_id) | raise NotImplementedError | Active | Get total device memory in bytes |
get_available_memory(device_id) | raise NotImplementedError | Planned | Return (free_bytes, total_bytes) |
get_current_memory_usage(device) | raise NotImplementedError | Active | Get current peak memory usage in bytes |
is_pin_memory_available(device=None) | False | Active | Whether pinned host memory is available for a target device |
get_torch_distributed_backend_str() | raise NotImplementedError | Planned | Distributed backend string (e.g. “nccl”, “hccl”) |
get_communicator_class() | None | Planned | Platform-specific communicator class |
inference_mode() | torch.inference_mode(True) | Planned | Return inference mode context manager |
seed_everything(seed) | Set random/np/torch seeds | Planned | Set random seeds for reproducibility |
verify_quantization(quant) | pass | Planned | Validate quantization method support |
get_cpu_architecture() | Auto-detect x86/arm | Planned | Detect CPU architecture (CpuArchEnum) |
Types (from DeviceMixin)
| Type | Description |
|---|---|
PlatformEnum | Enumeration of platform types: CUDA, ROCM, CPU, XPU, MUSA, NPU, TPU, MPS, OOT, UNSPECIFIED |
CpuArchEnum | CPU architecture: X86, ARM, UNSPECIFIED |
DeviceCapability | NamedTuple(major, minor) with comparison support. Methods: as_version_str(), to_int() |
Capability Flags (from SRTPlatform)
| Method | Default | Description |
|---|---|---|
support_cuda_graph() | False | Whether device graph capture is supported (plain CUDA graph) |
support_piecewise_cuda_graph() | False | Whether piecewise CUDA graph (torch.compile backend) is supported |
supports_fp8() | False | Whether FP8 quantization is supported |
Subsystem Factory Methods (from SRTPlatform)
| Method | Default | Description |
|---|---|---|
get_default_attention_backend() | raise NotImplementedError | Default attention backend name |
get_graph_runner_cls() | raise NotImplementedError | Graph Runner class |
get_mha_kv_pool_cls() | raise NotImplementedError | MHA KV cache pool class |
get_mla_kv_pool_cls() | raise NotImplementedError | MLA KV cache pool class |
get_dsa_kv_pool_cls() | raise NotImplementedError | DSA KV cache pool class (DeepSeek V3.2) |
get_paged_allocator_cls() | raise NotImplementedError | Paged allocator class |
get_quantization_config(quantization) | raise NotImplementedError | Return hardware-specific quantization config for the specific quantization scheme, raise an error if not supported or return None to use the default config. |
get_piecewise_backend_cls() | raise NotImplementedError | Piecewise compilation backend class |
get_compile_backend(mode) | “inductor” | Compilation backend string |
get_dispatch_key_name() | “native” | BaseFusedOp (fused-op) dispatch key name |
Lifecycle Hooks (from SRTPlatform)
| Method | Invocation Timing | Purpose |
|---|---|---|
apply_server_args_defaults(server_args) | After ServerArgs parsing, in post_init | Set platform-specific defaults |
init_backend() | In each worker, before model construction | One-time backend initialization |
Platform and plugin environment variables
| Variable | Description |
|---|---|
SGLANG_PLATFORM | Select the platform plugin by entry_point name (e.g. kunlun, demo_cuda). When set, only the named plugin’s activate() is called (front-loading filter) — other plugins are not touched. Additionally, general plugins (sglang.srt.plugins) from unselected platform packages are automatically skipped to avoid importing their dependencies. Required when multiple plugins would activate. Errors if the name is not found or if the plugin’s hardware is unavailable. |
SGLANG_PLUGINS | Comma-separated whitelist of general plugin names to load from either hook group. It also filters automatic SRT platform discovery when SGLANG_PLATFORM is unset; explicit platform selection ignores it. |
Add diffusion support
A package may support SRT, SGLang Diffusion, or both. Diffusion uses a zero-argumentPlatform subclass and separate platform and hook entry points:
dispatch_key to the PyTorch dispatcher key used by direct torch.library registrations. Keep it separate from get_dispatch_key_name(), which selects CustomOp implementations. Bring up the remaining contract in dependency order:
- Implement
get_device_name(),get_device_total_memory(), andget_available_gpu_memory()before constructingServerArgs. - Implement
get_device()andget_local_torch_device()before worker binding. - Configure distributed initialization, attention selection, and custom-op implementations for supported workloads.
- Run an end-to-end workload and audit model-specific kernels, compilation, and remaining device-family branches.
init_backend(), which runs once in each worker before any pipeline module is constructed:
init_backend() runs at most once per worker process, before the worker implementation is imported. Raising aborts startup and the failed initialization is not retried in that process, because registrations and other backend side effects may be only partially reversible. Custom-op dispatch is resolved when each operation is constructed, after backend initialization has completed. The selected callable therefore remains stable when the operation is compiled instead of changing on the first compiled call.
The function receives the operation instance before its normal arguments. Registration matches the exact operation class and the value returned by get_dispatch_key_name(). Without a registration, SGLang looks for a matching forward_<key>() implementation on the operation, then falls back to forward_oot(), whose base implementation calls forward_native(). Returning "cuda", for example, lets a CUDA-compatible OOT platform reuse operation-specific forward_cuda() implementations without reporting CUDA platform identity.
The sglang.multimodal_gen.plugins entry point remains for seams the Platform interface does not cover. Its callbacks run in launchers and workers before scheduler or model construction, so use a BEFORE or AROUND hook on GPUWorker.init_device_and_model() for rank-aware initialization. Hooks from platform distributions other than the selected one are skipped. Because these ship in your platform’s distribution, a failure at any stage — load, callback, or hook application — aborts startup, the same as init_backend().
Diffusion platform contract
Active interfaces are called throughcurrent_platform. Compatibility interfaces are retained adapters; unlisted Platform methods are not a stable OOT contract.
Install the package, restart Python to refresh entry-point metadata, and verify explicit selection:
Diffusion platform limitations
- SRT and diffusion require separate platform classes.
- Only one external platform can be active per process; set the runtime selector if multiple callbacks activate.
get_all_to_all_communicator_cls()controls onlyall_to_all_4D(), not every collective, graph-capture, or synchronization path;get_device_communicator_cls()remains for compatibility.- Compile settings affect only
build_torch_compile_kwargs()callers; audit static@torch.compiledecorators and other direct compile paths. - Explicit attention selectors accept only
AttentionBackendEnumnames. Returning a custom backend works when none is selected; selecting it by name requires a hook or downstream patch. - Device-family branches remain outside the interface. Each supported workload needs a native fallback or an actionable unsupported-feature error.
- The diffusion
PlatformEnum.OOTidentifies an external provider. Override built-in identity predicates only after auditing every enabled branch.
Plugin Type 2: General Plugin
Description
General function plugins inject behavior into SRT or diffusion without requiring a custom platform. The two runtimes have separate entry-point groups and hook registries:sglang.srt.plugins and a diffusion plugin under
sglang.multimodal_gen.plugins, and import the hook API from the matching
module above. Each runtime applies only its own registry, so a hook registered
through the other runtime’s API never runs. The examples below use SRT.
The diffusion hook API lives alongside platform discovery in
sglang.multimodal_gen.runtime.platforms.plugins;
sglang.multimodal_gen.plugins is its entry-point group, not a Python module path.
Use cases include:
- Observability: Add logging, metrics, and tracing to any function
- Behavior modification: Modify function arguments or return values
- Performance profiling: Add timing to critical functions
- A/B testing: Replace implementations at runtime
Quick Start
1. Create a minimal package:pyproject.toml:
__init__.py — register hooks:
Hook Types
HookRegistry supports four hook types:
| Hook Type | Signature | Description |
|---|---|---|
| BEFORE | fn(*args, **kwargs) -> (args, kwargs) | None | Runs before the original. Return None to keep args unchanged, or (args, kwargs) to modify. |
| AFTER | fn(result, *args, **kwargs) -> new_result | None | Runs after the original. Return None to keep result, or a new value to replace. |
| AROUND | fn(original_fn, *args, **kwargs) -> result | Wraps the original. You must call original_fn yourself. Full control over execution. |
| REPLACE | fn(*args, **kwargs) -> result or class | Replace the original function or class entirely. For class targets, pass a replacement class directly — it is substituted via setattr preserving isinstance()/issubclass() semantics. |
Note: OnlyREPLACEaccepts a class as the hook. Passing a class toBEFORE/AFTER/AROUNDraisesTypeErrorat registration time.
Registration API
Hooks can be registered using the imperative API or the decorator API:Hook Target Resolution
Target paths use fully-qualified dotted notation. Both formats are supported:- Dotted:
sglang.srt.managers.scheduler.Scheduler.__init__ - Entry-points style:
sglang.srt.managers.scheduler:Scheduler.__init__(colon treated as dot)
Common SRT Hook Targets
| Target | Description |
|---|---|
sglang.srt.server_args.ServerArgs.add_cli_args | Add custom CLI arguments |
sglang.srt.server_args.ServerArgs.post_init | Modify ServerArgs after parsing |
sglang.srt.server_args.ServerArgs.check_server_args | Add/relax validation |
sglang.srt.managers.scheduler.Scheduler.init | Custom scheduler state |
sglang.srt.managers.scheduler.Scheduler.get_next_batch_to_run | Custom scheduling policy |
sglang.srt.managers.scheduler.Scheduler.run_batch | Profiling / inspection |
sglang.srt.managers.scheduler.Scheduler.process_batch_result | Custom metrics |
sglang.srt.managers.tp_worker.TpModelWorker.init | Custom worker state |
sglang.srt.managers.tp_worker.TpModelWorker.forward_batch_generation | Forward pass wrapping |
File Reference
| File | Description |
|---|---|
sglang/srt/platforms/device_mixin.py | DeviceMixin base class and SRT platform identity types |
sglang/srt/platforms/interface.py | SRTPlatform base class (extends DeviceMixin) |
sglang/srt/platforms/init.py | current_platform lazy singleton + discovery logic |
sglang/multimodal_gen/runtime/platforms/interface.py | Diffusion Platform base class |
sglang/multimodal_gen/runtime/platforms/init.py | Diffusion current_platform lazy singleton and built-in fallback order |
sglang/multimodal_gen/runtime/platforms/plugins.py | Diffusion plugin registration, explicit hook-application phase, and hook registry |
sglang/multimodal_gen/runtime/managers/worker_bootstrap.py | Import-neutral process specifications and spawn targets that initialize the backend before resolving runtime hooks |
sglang/srt/plugins/init.py | load_plugins() + load_plugins_by_group() |
sglang/srt/plugins/hook_registry.py | HookRegistry, HookType, plugin_hook decorator |
