Skip to main content

Ming-Image integration status

Native support covers inclusionAI/Ming-Image-0.1-Design (generation and single-image editing) and inclusionAI/Ming-Image-0.1-Design-Layer (ordered RGBA layers). The checkpoints are not interchangeable. No remote modeling code is used; checkpoint directories do not need conversion to a Diffusers layout. H200 full-checkpoint checks include 2048-square generation, 1024-square four-layer decomposition, and repeated 512-square HTTP generation and editing. The HTTP checks cover cold starts, request warmup, layerwise offload, opt-in VAE tiling, and one or two sequential generations per request, requiring identical pixels across repeated requests with fixed inputs. Two-H200 functional checks cover DiT TP2, Ulysses2, Ring2 with FlashAttention, CFG parallelism for Design-Layer, encoder folding, and spatial VAE decoding. Single-GPU checks cover DiT and encoder layerwise offload, Cache-DiT, SageAttention, actual breakable CUDA graph replay, and dynamic LoRA load/remove with a synthetic adapter. These checks do not establish quality equivalence for lossy optimizations or bit-identical output across parallel topologies. Quantized checkpoints and other GPU families remain unverified. Run the opt-in HTTP regression tests on a CUDA host with sufficient memory:
See the Ming-Image cookbook for launch commands, request parameters, and model-specific limits.

Qwen-Image 2.1 integration status

Qwen-Image 2.1 has a separate native pipeline for text-to-image and optional reference-image conditioning. Full-checkpoint validation on H200 includes 1024×1024, 40-step generation and editing. Reference comparisons use Transformers 4.57.3, BF16 and matching Torch SDPA settings with cuDNN SDPA disabled. The native encoder explicitly preserves that version’s pre-final-norm hidden state and vision position interpolation semantics. Full-checkpoint functional checks on up to two H200 GPUs cover DiT TP, Ulysses, Ring + FA, CFG parallel, DiT layerwise offload, FA/Sage, parallel tiled VAE decode, Cache-DiT, breakable CUDA graphs, and multiple reference images and outputs. One H200 with resident weights also passed HTTP generation and editing, including multiple outputs and reference images. These are functional checks, not a broad quality evaluation of lossy optimizations. Additional full-checkpoint checks on B200 cover native encoder TP, combined TP2 × Ulysses2 and TP2 × Ring2 + FA on four GPUs, and --layerwise-offload-components all. Two-GPU VAE spatial_shard decode passed with TP2, CFG parallelism, and all-component offload. These use 512px/4-step editing, plus a 1024px/4-step comparison of spatial and full VAE decode. Single-GPU 1024px/40-step generation, editing, and transparent PNG output also passed at source revision d629d234f1. Two-B200 component regressions assert actual encoder weight sharding and compare text and image conditioning with a single-rank reference. VAE regressions cover both decoder structures and even/odd latent heights in FP32 and FP64. FP64 spatial decode matches the full decoder; BF16 full-checkpoint output is not bit-exact. In the tested image, spatial versus full decode gave RGBA PSNR 60.69 dB; this is one sample, not a general quality guarantee. Request-local prefix KV reuse and new-request graph replay have CUDA regression tests, including batched targets and independent variable-length prefixes. Compatible text-to-image requests support opt-in dynamic batching; image-edit requests remain separate, while multiple outputs within one request are supported. See the cookbook’s batching guidance for deployment settings and floating-point reproducibility limits. The deployment picker marks only its exact tested HTTP combinations as verified, including H200, B200, RTX PRO 6000 96GB, RTX 5090, and RTX 4090. CLI-only combinations remain Unverified in the picker. On 1eab5de5990, one RTX PRO 6000 Blackwell Server Edition passed repeated 1024px/40-step HTTP generation, editing, transparent generation, and transparent input editing with resident weights or DiT layerwise offload and Torch SDPA. Repeated requests and corresponding outputs across placements matched RGBA pixels; transparent outputs retained alpha from 0 to 255. Its SM120 RoPE, normalization, fusion-gate, and model regressions also passed. This check does not establish multi-GPU or quantized-checkpoint coverage on that platform. Disaggregated encoder, denoiser, and decoder roles passed full-checkpoint HTTP generation, variable-length prompts, image editing, and CFG on three B200 GPUs, using Mooncake’s same-host TCP transport. At 512px/4 steps, their RGBA pixels matched the monolithic server exactly. Multi-host RDMA and multi-rank roles are not covered by that check. Online FP8 for the DiT, encoder, and both together passed 1024px/40-step HTTP generation and editing on one B200; both together also passed transparent PNG generation. Serialized E4M3FN FP8 component folders with scalar weight scales and dynamic activations passed generation, editing, and transparent output for each component and both together. The combined export also passed separate checks with TP2 plus encoder folding and with single-GPU all-component layerwise offload. The 476 loaded quantized matrices and scales matched their checkpoint values. Quantization and TP change numerical results. Revision f1f3366c7c fixes vision rotary initialization rounding during CPU/GPU transfer. In 1024px/40-step B200 checks, all-component layerwise offload matched resident generation, editing, and transparent RGBA pixels exactly for both native precision and the combined serialized export. Revision 81c8c550fa additionally preserves FP8 weights and FP32 buffers during whole-encoder CPU offload; separate checks of that mode also matched resident pixels. Resident output was unchanged by these fixes. See the cookbook for sample quality metrics. Native-name Q4_0 GGUF files passed 1024px/40-step B200 HTTP generation, editing, and transparent output for DiT-only, encoder-only, and both together. The 476 packed matrices matched their exported bytes. The combined export also passed TP2 with encoder folding; separate all-component layerwise and whole-encoder CPU offload checks matched resident RGBA pixels exactly. Quantization and TP change output pixels, including alpha. These checks cover the documented private exports, not arbitrary community checkpoints. Other GGUF types and hardware remain unverified for this model. Calibrated ModelOpt-format NVFP4 component directories also passed the same six-configuration B200 HTTP matrix: each component, both together, two offload modes, and TP2 with encoder folding. All 476 packed matrices and their scales matched the exports after runtime layout transforms. Both offload modes matched resident RGBA pixels exactly. NVFP4 and TP2 change image and alpha values; this private export uses max calibration on only six requests, without SVDQuant or AWQ. See the cookbook for measured errors. Native FP4 requires Blackwell, and this model’s NVFP4 exports remain unverified on RTX PRO 6000 and RTX 5090. Synthetic Diffusers-format LoRA passed dynamic loading, merging, and removal on one B200 and TP2 with encoder folding; both removal paths restored the base image exactly. Trained-adapter quality remains unverified. Platform latency measurements and their exact workloads are in the cookbook. Breakable graphs replay only matching warmup signatures, including condition-prefix length; other shapes run eagerly. VAE tiling is opt-in and can change numerical results. Do not infer compatibility from the older Qwen-Image row. Use the public Qwen/Qwen-Image-2.1 checkpoint; no Hugging Face token is required. See the Qwen-Image 2.1 cookbook for checkpoint layout and usage. The standard two-GPU E2E suite includes qwen_image21_t2i_tp2 with TP2, 1024px/40-step generation, two requests, and image consistency checks. The additional opt-in HTTP cases check repeated generation, editing, and real RGBA alpha output from a local checkpoint:
The module skips without the model variable; only editing skips when the image variable is absent. Credentials stay in the runner environment.

Supported model inventory

Use a listed checkpoint as --model-path with sglang generate or sglang serve. This registry-backed list contains known public entry points; family detection may also support compatible local directories. Open the linked Cookbook recipe for launch commands, optimizations, adapters, and model-specific notes.