AI image generators work by turning your words into numbers and then building a picture that matches them, using patterns learned from billions of captioned images. Stable Diffusion, whose makers published its full design, starts every picture as random static and cleans it up over dozens of steps. Below are those steps in plain words, the two main methods (diffusion and autoregressive), where the training pictures come from, and why the wording of your prompt changes the result so much.
How AI image generators work, step by step
For beginners asking how AI image generators work, Stable Diffusion is the easiest example, because its makers published the design and a public model card. Here is how AI image generation works in it, in five stages:
- Training on captioned images. Before anyone types a prompt, the model studies a huge set of pictures paired with text. Stable Diffusion v1-4 learned from subsets of LAION-2B (en), the English part of LAION-5B, a public collection of 5.85 billion image and text pairs.
- Turning your prompt into numbers. A text encoder converts your words into a list of numbers that capture their meaning. Stable Diffusion v1-4 uses OpenAI's CLIP ViT-L/14 encoder. CLIP was trained on 400 million image and caption pairs, learning to predict which caption belongs to which picture.
- Starting from noise. The picture begins as pure random static, like a TV tuned between channels.
- Removing noise, step by step. At each step the model predicts which part of the static to take away so the image looks a little more like your prompt. The prompt's numbers feed into every step. Hugging Face's Stable Diffusion pipeline runs 50 of these steps by default.
- Decoding to pixels. Stable Diffusion does steps 3 and 4 on a small, compressed version of the image, then a decoder expands the finished result into a full-size picture.
How does an AI art generator work?
An AI art generator works the same way whether the output is a photo, a watercolor or a cartoon. The style is just another part of the prompt the model learned to match: asking for "watercolor" or "35mm film photo" steers the same noise-removal process toward the look those words were paired with in training. If you can't name a style in words, a reference image is the other way to show it.
How diffusion turns noise into a picture
A diffusion model is a network trained to undo noise. In training, it sees real images with more and more static mixed in and learns to predict the noise that was added, so it can later run the process in reverse. Jonathan Ho, Ajay Jain and Pieter Abbeel showed in 2020 that this approach, which they called denoising diffusion probabilistic models, could produce high-quality images. The same family of models sits behind Stable Diffusion, and OpenAI used a diffusion model to draw the final image in DALL-E 2.
How does AI art work in simple terms?
Picture a sculptor who has studied millions of statues, handed a block of stone and a note that says "a cat on a windowsill". Each pass of the chisel removes a little of whatever doesn't look like a cat on a windowsill. In AI art, the block is random noise, the note is your prompt, and each pass is one denoising step. The model doesn't look up a stored cat picture; it carves a new one out of the noise.
Why it runs fast enough to use
Working on full-size pixels takes an enormous amount of computing. The 2021 latent diffusion paper by Robin Rombach and colleagues, the design behind Stable Diffusion, says training the strongest pixel-based diffusion models can take hundreds of GPU days. Their fix was to run diffusion on a compressed "latent" version of the image instead. Stable Diffusion's autoencoder shrinks each side of the image by a factor of 8 and keeps 4 values per spot instead of 3 colors. For a 512 x 512 px picture:
| Version of the image | Size | Numbers stored |
|---|---|---|
| Full picture | 512 x 512 px, 3 color channels | 786,432 (512 x 512 x 3) |
| Latent version | 64 x 64, 4 channels | 16,384 (64 x 64 x 4) |
That is 48 times fewer numbers to process at every step.
Guidance and seeds
Stable Diffusion tools expose a guidance scale, which sets how strictly the image follows your prompt. The technique behind it, classifier-free guidance, was published by Jonathan Ho and Tim Salimans in 2022: the model makes one prediction with your prompt and one without, and the difference between them is pushed further in your prompt's direction. Hugging Face's pipeline defaults to 7.5, and its docs note that a higher value follows the prompt more closely at some cost to image quality. The seed is the number that picks the starting noise. Keep it fixed and you can change one word of the prompt and compare the two results fairly. A negative prompt plugs into this same guidance step: its words take the place of the empty prompt in the second prediction.
How autoregressive image models work
An autoregressive model builds its output in sequence, working out each new piece from the ones it has already made, the same way a chatbot writes a reply word by word. OpenAI's GPT-4o image generation, which OpenAI introduced in ChatGPT on March 25, 2025, works this way. OpenAI's system card says that 4o image generation, unlike its earlier diffusion-based DALL-E models, works autoregressively and sits natively inside ChatGPT.
Because one model handles both the words and the picture, it can use what it knows from language. OpenAI says it trained the model on images and text together, and that it can keep track of 10 to 20 different objects in a prompt where other systems struggled with about 5 to 8. It can also take your own pictures as input and return a new one based on them, which the system card lists as a new capability.
Diffusion vs autoregressive image generators
| Feature | Diffusion | Autoregressive |
|---|---|---|
| How the picture is built | A whole noisy image is refined over many steps | The image is predicted piece by piece, in order |
| Well-known examples | Stable Diffusion, SDXL, DALL-E 2 | GPT-4o image generation |
| How the prompt steers it | A separate text encoder feeds every step | The same model reads the prompt and makes the image |
| Common controls | Steps, guidance scale, seed, negative prompt | Follow-up requests in plain language |
| Weak spots named by the makers | Hands, legible text, placing objects correctly | Cropping, dense small text, editing precision |
You can use either kind without knowing which one is inside. In both, the things you control are the prompt, the size, and any reference images.
What came before diffusion
Before diffusion, many image models were generative adversarial networks, or GANs, an idea introduced by Ian Goodfellow and colleagues in 2014. A GAN trains two networks against each other: a generator makes images, and a discriminator tries to tell them apart from real training images. Each gets better by beating the other. The text-to-image tools named in this guide are diffusion or autoregressive models, not GANs.
Where the training images come from
Most training images come from the public web. LAION-5B, the collection Stable Diffusion's training set came from, was built from Common Crawl, a public archive of web pages, by pulling out images and the alt-text written for them, then filtering the pairs with CLIP. Its creators say they released it to open up research on large image and text models.
What goes into training shapes what comes out. The Stable Diffusion v1-4 model card says its captions are mostly English, so the model tends to treat white and Western cultures as the default and does noticeably worse with prompts in other languages. A 2023 study led by Nicholas Carlini found that diffusion models sometimes memorize individual training images and can reproduce them, including photos of real people and trademarked logos.
Whether training on copyrighted images without permission is legal is still being decided. In a May 2025 pre-publication report, the US Copyright Office noted that dozens of lawsuits over it are pending in US courts, and said the answer under fair use depends on which works were used, where they came from, and what the model is used for. For what this means when you use AI images yourself, read whether AI images can be used commercially. This is general information, not legal advice.
Why the prompt matters so much
The prompt is the main thing steering the picture, so anything you leave out becomes the model's choice. "A dog in a park" leaves the breed, the season, the light, the angle and the style open, and the model fills each gap with whatever its training pictures made most likely.
A useful prompt covers what the picture shows, where it is, how it's lit, the camera angle or art style, and any words to print in double quotes. These two are written for Image Generator, and you can paste either one into another text-to-image tool to start from:
Photo of a golden retriever pup sitting in fallen maple leaves in a city park, late afternoon sun behind it, shot at eye level with a 50mm lens, soft background blur, warm autumn colors.Flat illustration of a lemonade stand on a sunny sidewalk with a hand-painted sign that reads "Fresh Lemonade" in bold letters, two paper cups on the counter, bright yellow and teal colors, simple shapes, no other text.For a full method, read the guide to prompting an AI image generator, or copy a starting point from the AI image prompt library.
Why AI images still get things wrong
The model makers' own documents list the same weak spots, and they follow from how the models learn.
- Hands. Stability AI's SDXL paper lists human hands as its first limitation, because hands vary so much from photo to photo. The guide to why AI is bad at hands covers the causes and fixes.
- Text. Stable Diffusion v1-4's model card admits the model "cannot render legible text", and even GPT-4o image generation lists dense small text as a limit. See why AI messes up text in images.
- Placing and counting objects. The same model card uses "A red cube on top of a blue sphere" as an example of a prompt it struggles with.
- Faces and people. The v1-4 card warns that faces, and people as a whole, may not come out right.
Editing images, not just making them
Text-to-image is only one way to use these models. Image-to-image starts from a picture you supply instead of pure noise, which is how a photo gets restyled as a painting; see what image-to-image AI is. Inpainting regenerates only a selected area and leaves the rest alone, which is how an eraser removes a stranger from a photo; see what inpainting is.
Make an image in Image Generator
Toybox AI's Image Generator turns a description into one image, with optional photos to guide it. It works in your browser on a phone or computer.
- Open Image Generator and sign in with your free Toybox account.
- In "Describe your image", write the subject, setting, light, angle or style, and any words to print in double quotes.
- Optional: add up to 3 "Reference photos" to guide the style, subject or layout.
- Pick a Size: Square 1:1, Portrait 3:4, Landscape 4:3, Story 9:16 or Wide 16:9.
- For a design with words, open More options and switch on "Enhanced quality" (currently 45 extra credits).
- Tap Create image. A result usually takes about 30 seconds.
- Zoom in and check hands, faces and any text before you use it.
Each Image Generator result currently costs 50 credits. New accounts get 20 free credits, which covers text tools but not an image, since AI image tools start at 35 credits. Reference photos guide the result but don't guarantee an exact likeness, so check faces against your originals. See current prices for plans and top-ups.