/v1/chat/completions once before
constructing the generation request. The enhancer does not replace the
diffusion model’s text encoder or run inside its GPU workers.
Deploy two services
This example uses Linux with NVIDIA CUDA and two GPUs, one for each server. Install SGLang for SRT and SGLang Diffusion for diffusion. Separate environments and hosts are supported; the diffusion host needs HTTP access to SRT. Start a text-only enhancer:enhancer.json on the diffusion host:
--port + 1 for its broker. For separate hosts, bind SRT to its service
interface with --host and update
base_url. Restrict network access and configure SRT authentication. Do not
expose an unauthenticated enhancer publicly. A shared enhancer can serve
multiple diffusion servers; size its concurrency and memory independently.
Send a generation request:
enhance_prompt or set it to false to use the original prompt without
calling SRT. With the OpenAI Python SDK, pass extra_body={"enhance_prompt": True}
to images.generate().
The same switch works with /v1/images/edits and /v1/videos, including
multipart requests (-F 'enhance_prompt=true'). Image responses expose the
rewritten text as data[].revised_prompt; video responses use revised_prompt.
Video creation waits for enhancement before returning the queued job. All n
outputs share one rewritten prompt.
Include enhancement in client-observed end-to-end latency. The response’s
inference_time_s measures diffusion execution, not the preceding SRT call.
Configure the enhancer
The configuration file is read at server startup. Its supported fields are:
Default generation options are
max_tokens=512 and temperature=0.
generation_kwargs cannot replace the model, messages, single-output or
non-streaming settings, or enable tools. Choose a non-thinking model or configure
its supported chat template to suppress reasoning; only message content is
used as the diffusion prompt.
The enhancer receives a JSON text message containing task (image,
image_edit, or video) and prompt. It returns the replacement prompt as
text. Seed, negative prompt, references and diffusion sampling options are not
rewritten. The diffusion frontend reuses HTTP connections and makes no enhancer
calls during synthetic warmup.
Timeouts return HTTP 504. Upstream failures, empty responses and truncated
completions return HTTP 502 before diffusion is queued. There is no automatic
retry or fallback to the original prompt. An unconfigured server rejects
enhance_prompt=true with HTTP 400.
Choose a model combination
There is no enhancer or diffusion model allowlist in this integration. Protocol compatibility does not establish quality: use a template appropriate for the target checkpoint and verify prompt adherence.
For a VLM, deploy an image-capable SRT model and set
include_images=true in
its configuration. Uploaded images are sent as data URLs, so the SRT server
does not need access to the diffusion host’s filesystem. HTTP image URLs must
be reachable from SRT. Text-only models should leave this option disabled.
Audio and video references are not forwarded to the enhancer; it cannot inspect
their content. Image edits forward the supplied images; video creation forwards
the primary image reference, not model-specific reference bundles.
For Ideogram 4, instruct the enhancer
to produce the checkpoint’s JSON caption schema, including
high_level_description and compositional_deconstruction. Set the matching
JSON schema through generation_kwargs.response_format. JSON text is passed
through unchanged, not converted to prose. Valid JSON alone does not guarantee
the correct caption schema or equivalent output to a hosted service’s expander.
This integration applies to the three HTTP image/video endpoints above,
including their normal parallelism, caching and disaggregated diffusion paths.
It does not add enhancement to offline sglang generate, realtime streaming,
mesh or action endpoints. For those workflows, call SRT in your application
first and pass its returned text as the diffusion prompt. For reproducibility,
store the rewritten prompt and replay it with enhance_prompt=false; a fixed
diffusion seed does not fix the enhancer’s output.
ERNIE-Image’s checkpoint-provided enhancer
ERNIE-Image also has a separate, existing PE integration. By default, diffusion loads its native Ministral3 enhancer in-process:use_pe=false to skip it for a request. For a memory-constrained deployment,
--layerwise-offload-components pe streams the native PE decoder from CPU.
To serve the checkpoint-provided PE through SRT, use a local checkpoint’s PE
directory with its tokenizer and retain the ERNIE-specific protocol:
--pe-server-url uses the checkpoint’s PE tokenizer/template and SRT’s
/generate; it is not the generic chat-completion integration. Keep using it
when you need ERNIE’s trained PE behavior. If you instead use
--prompt-enhancer-config, send both enhance_prompt=true and use_pe=false.
For LongCat, use enable_prompt_rewrite=false with the generic enhancer.
These request flags skip native rewriting, not loading its model components.
For Ascend setup, see models with AR stages.