Skip to main content
Run an SRT server for prompt expansion and a diffusion server for generation. The diffusion HTTP frontend calls SRT’s /v1/chat/completions once before constructing the generation request. The enhancer does not replace the diffusion model’s text encoder or run inside its GPU workers.
Prompt enhancement is opt-in per request. It changes conditioning and can change intent or quality; it is not a lossless acceleration feature. Evaluate the chosen enhancer and template on your workload before enabling it.

Deploy two services

This example uses Linux with NVIDIA CUDA and two GPUs, one for each server. Install SGLang for SRT and SGLang Diffusion for diffusion. Separate environments and hosts are supported; the diffusion host needs HTTP access to SRT. Start a text-only enhancer:
Create enhancer.json on the diffusion host:
Start diffusion in another terminal:
Keep the SRT port separate from diffusion’s HTTP and internal ports; diffusion reserves --port + 1 for its broker. For separate hosts, bind SRT to its service interface with --host and update base_url. Restrict network access and configure SRT authentication. Do not expose an unauthenticated enhancer publicly. A shared enhancer can serve multiple diffusion servers; size its concurrency and memory independently. Send a generation request:
Omit enhance_prompt or set it to false to use the original prompt without calling SRT. With the OpenAI Python SDK, pass extra_body={"enhance_prompt": True} to images.generate(). The same switch works with /v1/images/edits and /v1/videos, including multipart requests (-F 'enhance_prompt=true'). Image responses expose the rewritten text as data[].revised_prompt; video responses use revised_prompt. Video creation waits for enhancement before returning the queued job. All n outputs share one rewritten prompt. Include enhancement in client-observed end-to-end latency. The response’s inference_time_s measures diffusion execution, not the preceding SRT call.

Configure the enhancer

The configuration file is read at server startup. Its supported fields are: Default generation options are max_tokens=512 and temperature=0. generation_kwargs cannot replace the model, messages, single-output or non-streaming settings, or enable tools. Choose a non-thinking model or configure its supported chat template to suppress reasoning; only message content is used as the diffusion prompt. The enhancer receives a JSON text message containing task (image, image_edit, or video) and prompt. It returns the replacement prompt as text. Seed, negative prompt, references and diffusion sampling options are not rewritten. The diffusion frontend reuses HTTP connections and makes no enhancer calls during synthetic warmup. Timeouts return HTTP 504. Upstream failures, empty responses and truncated completions return HTTP 502 before diffusion is queued. There is no automatic retry or fallback to the original prompt. An unconfigured server rejects enhance_prompt=true with HTTP 400.

Choose a model combination

There is no enhancer or diffusion model allowlist in this integration. Protocol compatibility does not establish quality: use a template appropriate for the target checkpoint and verify prompt adherence. For a VLM, deploy an image-capable SRT model and set include_images=true in its configuration. Uploaded images are sent as data URLs, so the SRT server does not need access to the diffusion host’s filesystem. HTTP image URLs must be reachable from SRT. Text-only models should leave this option disabled. Audio and video references are not forwarded to the enhancer; it cannot inspect their content. Image edits forward the supplied images; video creation forwards the primary image reference, not model-specific reference bundles. For Ideogram 4, instruct the enhancer to produce the checkpoint’s JSON caption schema, including high_level_description and compositional_deconstruction. Set the matching JSON schema through generation_kwargs.response_format. JSON text is passed through unchanged, not converted to prose. Valid JSON alone does not guarantee the correct caption schema or equivalent output to a hosted service’s expander. This integration applies to the three HTTP image/video endpoints above, including their normal parallelism, caching and disaggregated diffusion paths. It does not add enhancement to offline sglang generate, realtime streaming, mesh or action endpoints. For those workflows, call SRT in your application first and pass its returned text as the diffusion prompt. For reproducibility, store the rewritten prompt and replay it with enhance_prompt=false; a fixed diffusion seed does not fix the enhancer’s output.

ERNIE-Image’s checkpoint-provided enhancer

ERNIE-Image also has a separate, existing PE integration. By default, diffusion loads its native Ministral3 enhancer in-process:
Use use_pe=false to skip it for a request. For a memory-constrained deployment, --layerwise-offload-components pe streams the native PE decoder from CPU. To serve the checkpoint-provided PE through SRT, use a local checkpoint’s PE directory with its tokenizer and retain the ERNIE-specific protocol:
--pe-server-url uses the checkpoint’s PE tokenizer/template and SRT’s /generate; it is not the generic chat-completion integration. Keep using it when you need ERNIE’s trained PE behavior. If you instead use --prompt-enhancer-config, send both enhance_prompt=true and use_pe=false. For LongCat, use enable_prompt_rewrite=false with the generic enhancer. These request flags skip native rewriting, not loading its model components. For Ascend setup, see models with AR stages.