Tongyi-MAI · Apache 2.0
Small model, long batches

Every prompt you have, through
Z-Image

Z-Image is the text-to-image model Alibaba's Tongyi-MAI team open-sourced in November 2025: six billion parameters, Apache 2.0 weights, and a distilled Turbo build that reaches a finished image in eight sampling steps. It occupies the bottom rung of the price list on BulkImagen, which is what makes a genuinely long bulk z-image run worth setting up.

One per line — each line is a separate prompt.

0 prompts

Z-Image generates from text prompts only, so there's no reference-image upload.

QwenZ-Image
Will generate 0 images · 0 credits

Why a six-billion-parameter model is on this list

Z-Image is not the largest model BulkImagen carries and is not trying to be. Its entire argument is ratio — how much a small network gives back for the compute it asks for — which is a different argument from most of the models here.

Measured against much larger weights

Artificial Analysis put Z-Image Turbo at the top of its Image Arena among open-weight text-to-image models when it appeared, ahead of FLUX.2 [dev] and Qwen-Image. The Z-Image technical report makes the same case from the other direction: an 87.4% G+S Rate on roughly a fifth of FLUX.2 dev's 32B parameter count.

Bilingual lettering, with numbers behind it

Text is where the Z-Image paper puts its sharpest results — 0.8671 word accuracy on CVTG-2K, and 0.935 English against 0.936 Chinese on LongText-Bench, reported ahead of GPT-Image-1 and Qwen-Image. If the batch carries captions in either language, that benchmark is the reason to look at Z-Image at all.

Photoreal, in eight steps

The Z-Image model card and the paper both lead with photorealism, and the distilled Turbo weights get there in eight sampling steps rather than the usual dozens. The open release fits on a 16GB card, but nothing on this page asks you to own one — the batch runs on our side.

What Z-Image is set up to do here

Aspect ratios, the resolution tier, prompt length and the credit rate for Z-Image, pulled out of the running configuration rather than written into a paragraph where they would slowly go out of date.

Qwen

Z-Image

Made by
Alibaba
Credits / image
1
Resolutions
1K
Aspect ratios
5
Reference images
Text only
Prompt limit
1,000 chars

Z-Image reads as having no reference-image support on purpose — the FAQ below explains what that means before you plan a batch around it.

Three steps, and no reference uploads

Z-Image reads words and nothing else, so this workflow is shorter than the one on the reference-driven model pages. Describe, frame, collect.

  1. 01

    Put everything into words

    There is no image input, so any detail you want has to live in the prompt itself. The Z-Image paper is also candid that six billion parameters buy limited world knowledge, so describe a niche subject rather than assuming the model recognises the name.

  2. 02

    Choose the framing and the volume

    Aspect ratio and how many images each prompt returns — that is the whole form, because this page is pinned to Z-Image and it exposes one resolution tier. The credit total appears on the confirmation screen before the first call goes out.

  3. 03

    Skim, shortlist, export

    Results arrive as they finish. The Z-Image Turbo model card rates its output diversity as low, so if rows come back near-identical the prompts were probably too close together — rewrite the wording rather than re-running the same line and hoping.

Batches that suit this particular model

A small, fast, text-strong model has a shape to it, and Z-Image is no exception. These are the jobs that fit that shape instead of fighting it.

Posters and UI mockups with real copy

Layouts carrying English or Chinese lettering are the case the Z-Image benchmarks were built around. Produce a whole set of screens or poster variants in one pass, then read the type on each before committing to any of them.

Photoreal people and places at volume

Photorealism is what Tongyi-MAI foregrounds about Z-Image, and quantity is what this page adds. Lifestyle frames, stock-style scenes and background plates come back as a set you can sift, not as one considered render at a time.

Teams already running the open weights

If you have Z-Image checked out locally, the same family is available here without a GPU queue in front of it. Draft and sort a prompt set against the hosted Turbo build, then take the shortlist back to your own environment.

Z-Image questions

Anything else? Email support@bulkimagen.com

  • Tongyi-MAI, a team inside Alibaba that sits apart from both the Qwen group and the Wan group. Z-Image Turbo was released as open source on 26 November 2025, and the technical report went up on arXiv the following month.
  • Yes. The Z-Image weights are published under Apache 2.0 on HuggingFace, ModelScope and GitHub, which is a permissive licence by the standards of models at this quality level. Running it through this page simply saves you hosting it.
  • The distilled Turbo build. Distillation is what brings Z-Image down to eight sampling steps; the trade-offs, both stated on the model card, are that Turbo takes no guidance setting and is not intended for fine-tuning.
  • No, and the attempt is refused before you are charged rather than after. The published Z-Image pipeline accepts a prompt, a size, a step count and a seed — there is no image input in it at all. Z-Image-Edit, the separately trained model that adds one, was still marked as coming soon in August 2026.
  • On the benchmarks in its paper, well for its size: 0.8671 word accuracy on CVTG-2K, plus 0.935 and 0.936 on the English and Chinese halves of LongText-Bench. Those are paper-reported figures, so a small batch of your own copy is still the test that counts.
  • Z-Image was trained bilingually, and its Chinese LongText-Bench score matches its English one, so Chinese prompts are a first-class input rather than a translation exercise bolted onto an English model.
  • Its own paper names the limit plainly: a 6B model buys efficiency at the cost of world knowledge, so obscure people, places and brand specifics are shaky. The Turbo card separately rates output diversity as low, which shows up as sameness across closely worded prompts.
  • Closed flagships like Nano Banana Pro and Flux 2 Pro go further on resolution ceilings, fine detail and reference-driven consistency. Z-Image trades those away for cost per render, which is why it is the one to reach for when the list is long.
  • Because a small distilled model costs the least to serve, and that lands Z-Image at the bottom of the credit table above. A run that would need a budget conversation on a flagship model is, here, just a run.
  • A CSV or an Excel sheet both work; the prompt column becomes the rows of the Z-Image batch. Useful when the list already exists somewhere as a spreadsheet and retyping it would be the slowest part of the job.
  • Mark the renders worth keeping in the Z-Image batch view and pull that selection down as one ZIP. Whatever you skip stays put, and any prompt that missed can be sent again on its own.

Point a long list at Z-Image

The cheapest slot on the platform, a Z-Image build with published bilingual text scores, and no reference uploads to prepare first. Paste the prompts, run them as one batch, and take the ZIP away.