Qwen-Image 2512 (20B MMDiT, bf16) text-to-image, strong at typography and lettering, SFW only. Renders on the RunPod VIDEO endpoint and bills the video ledger while returning a still (heavy-image, A.02.05): the bf16 transformer alone is 40.86 GB and its Qwen2.5-VL 7B encoder 16.58 GB, which no 24-32 GB image card holds. Translated from Comfy-Org/workflow_templates@aaac56dd templates/image_qwen_Image_2512.json with its Enable 4 Steps LoRA switch off (the template's default): UNETLoader, CLIPLoader type qwen_image, VAELoader, ModelSamplingAuraFlow shift 3.1, KSampler euler/simple at 50 steps and CFG 4, the template's Chinese negative prompt, EmptySD3LatentImage at 1328x1328; the aspect options are the seven sizes the template's Aspect Ratios note lists. The template loads the fp8_e4m3fn transformer and fp8_scaled encoder; this catalog loads the bf16 files on the volume (D-V5). The Lightning 4-step LoRA the switch would add has no manifest row and is not offered.
Tags: heavy-imagerunpod-serverlessvolume-backedtext-to-imagetypographyqwen-image
Inputs (7)#
The typed parameter surface callers bind when they request this workflow. Enum options and numeric bounds are the values the workflow document declares.
promptstringrequireddefault Urban alleyway at dusk. Tall, statuesque high-fashion model striding elegantly, mid distant full body shot from an angular perspective, cinematic/editorial with bold contrasts and tactile materials. They wear a rose-gold metallic trench coat with deconstructed elements over a black long-sleeved turtleneck with subtle texture; paired with forest-green pleated pants with raw hems and a soft texture. Long braided dark hair, medium complexion. They carry a vibrant yellow designer handbag with geometric details and a structured silhouette. White architectural sneakers with bold geometric cutouts. Bold, high-contrast, tactile, urban-grit meets high-fashion impact, extreme clarity, extreme layering, post-processing with transparent light-transmitting ultra-smooth high-definition film effect, removing all noise and grain, removing all blur, removing all vintage feel, removing all roughness, drawn with 32K pixel precision, unparalleled fine line drawing of every single detail, the entire image like a brand new photograph, photorealisticnegative_promptstringdefault 低分辨率,低画质,肢体畸形,手指畸形,画面过饱和,蜡像感,人脸无细节,过度光滑,画面具有AI感。构图混乱。文字模糊,扭曲stepsintegerdefault 50min 1max 60cfgfloatdefault 4.0min 1.0max 10.0seedintegerdefault -1aspectenumdefault square_1328x1328square_1328x1328landscape_16x9_1664x928portrait_9x16_928x1664landscape_4x3_1472x1104portrait_3x4_1104x1472landscape_3x2_1584x1056portrait_2x3_1056x1584batch_sizeintegerdefault 1min 1max 1ComfyUI node graph (10)#
The executable ComfyUI prompt graph: 10 nodes across 9 distinct node classes, wired by 10 data dependencies. Nodes tinted green come from a custom node pack this workflow declares; the rest are ComfyUI core / baked-community classes.
Nodes (10)#
1UNETLoadercoreunet_name = qwen_image_2512_bf16.safetensorsweight_dtype = defaultMODEL2CLIPLoadercoreclip_name = qwen_2.5_vl_7b.safetensorstype = qwen_imagedevice = defaultCLIP3VAELoadercorevae_name = qwen_image_vae.safetensorsVAE4ModelSamplingAuraFlowcoremodel = ◂ node 1 · out[0]shift = 3.1MODEL5CLIPTextEncodecoretext = {{constructed_prompt}} tmplclip = ◂ node 2 · out[0]CONDITIONING6CLIPTextEncodecoretext = {{negative_prompt}} tmplclip = ◂ node 2 · out[0]CONDITIONING7EmptySD3LatentImagecorewidth = {{aspect_map[aspect].width}} tmplheight = {{aspect_map[aspect].height}} tmplbatch_size = {{batch_size}} tmplLATENT8KSamplercoremodel = ◂ node 4 · out[0]seed = {{seed}} tmplsteps = {{steps}} tmplcfg = {{cfg}} tmplsampler_name = eulerscheduler = simplepositive = ◂ node 5 · out[0]negative = ◂ node 6 · out[0]latent_image = ◂ node 7 · out[0]denoise = 1.0LATENT9VAEDecodecoresamples = ◂ node 8 · out[0]vae = ◂ node 3 · out[0]IMAGE10SaveImagecoreimages = ◂ node 9 · out[0]filename_prefix = isis/qwen-txt2imgPrompt construction#
template{base_prompt}Variables (1)#
base_prompt{{prompt}} tmplParameter banks (2)#
The prompt / configuration lookup tables this workflow keys into from its inputs — the vocabulary that turns a style / palette / preset selection into graph parameters.
aspect_map (7)#
square_1328x1328{"width": 1328, "height": 1328}landscape_16x9_1664x928{"width": 1664, "height": 928}portrait_9x16_928x1664{"width": 928, "height": 1664}landscape_4x3_1472x1104{"width": 1472, "height": 1104}portrait_3x4_1104x1472{"width": 1104, "height": 1472}landscape_3x2_1584x1056{"width": 1584, "height": 1056}portrait_2x3_1056x1584{"width": 1056, "height": 1584}requires_families (1)#
qwen-imageModels & dependencies#
Models required (3)#
qwen_image_2512_bf16.safetensorsqwen_2.5_vl_7b.safetensorsqwen_image_vae.safetensorsOutput contract#
What a successful run of this workflow returns.
typeimageformatpngTaxonomy & routing#
How the control plane classifies this workflow — from the committed workflow-taxonomy-registry.json. It drives the consistency / control surface the agentic director can exercise over the workflow.
assetFamilyportrait-hero-image-bundleoutputPackageProfileimage-single-profilecontrolModalitiesmodel-locksampler-scheduler-lockseed-lockprompt-template-lockconsistencyDimensionsidentitylightinglensnotesQwen-Image 2512 text-to-image with typography, a still rendered on the RunPod video endpoint (heavy-image), SFW only.