Image-to-image AI starts from a picture you already have, not an empty canvas. You give it a photo, a sketch or a screenshot, usually with a short text prompt, and it returns a new picture that keeps part of the original and changes the rest. It's how a phone photo becomes an oil painting, or a doodle becomes a storybook scene. Below is what the term means, how the process works, the setting that controls how much changes, real uses with prompts, and where it falls short.
What is image-to-image AI?
An image-to-image AI model takes an image as input and produces a new image related to it. The term covers three things that work differently, and knowing which one a tool uses explains its results.
Image-to-image generation
In diffusion models such as Stable Diffusion, image-to-image generation means starting the usual noise-removal process from your picture instead of from pure static. Hugging Face's Diffusers documentation describes it step by step: your image is encoded into a compressed version, noise is added to it, the model removes that noise while following your prompt, and a decoder turns the result back into a picture. This is the classic form of image-to-image AI generation, often shortened to "img2img".
Reference-guided generation
Chat-style image models take your photo as a reference and follow a written instruction about it. OpenAI's system card for GPT-4o image generation, published March 25, 2025, lists image-to-image transformation as a new capability: the model can accept one or several pictures and return a new picture based on them or altered from them. Here there's no strength slider. You control how much changes with words, such as "keep the layout and colors, change only the season".
What is image-to-image translation?
Image-to-image translation is the research name for teaching a model to convert one kind of picture into another. A 2016 paper by Phillip Isola, Jun-Yan Zhu, Tinghui Zhou and Alexei Efros, released with software called pix2pix, showed one general approach doing several of these conversions: label maps into photos, edge drawings into objects, and black and white images into color. In that approach each conversion is trained on its own examples, while prompt-based tools handle many jobs in one model.
How image-to-image generation works
The key setting in diffusion image-to-image is strength: how much noise gets added to your picture before the model redraws it. In the Diffusers Stable Diffusion image-to-image pipeline, strength runs from 0 to 1 and defaults to 0.8. At 1, the noise is at its maximum and your image is essentially ignored.
| Strength | What you get | Good for |
|---|---|---|
| Low | Close to your original, with light changes in texture and color | Touch-ups, mild style shifts |
| Middle | The same layout with new details and style | Photo to painting, sketch to illustration |
| High | Loosely based on your original's composition | Using a photo only as a rough starting point |
| 1.0 | Your image is essentially ignored | Starting over; text to image does this directly |
The trade-off behind the slider was laid out in the 2021 SDEdit paper by Chenlin Meng and colleagues: the method adds noise to your input, then removes it, and the amount of noise balances realism against faithfulness to what you gave it. Their inputs included hand-drawn colored strokes, which the method turned into realistic images. Strength also changes how many denoising steps run, so a low-strength edit finishes faster.
Diffusion tools usually pair strength with guidance, a second slider that controls how literally the result follows the prompt. Hugging Face suggests high strength with high guidance for the most creative freedom, and low values of both for a result that stays close to your picture and loose on the prompt.
How image to image differs from text to image and editing
| Question | Text to image | Image to image | Photo editing with AI |
|---|---|---|---|
| Starts from | Random noise and a prompt | Your picture and a prompt | Your picture and an instruction |
| What carries over | Nothing | Layout, colors, subject or style | Everything except the part you change |
| Typical request | "A cabin in snowy woods at dusk" | "Turn this photo into a watercolor" | "Remove the power lines above the roof" |
| When to use it | You have nothing to start from | You like the picture's bones but want a new look | The picture is right except for one thing |
The last column overlaps with inpainting, where only a selected area is redrawn; the guide to what inpainting is covers it. For how the noise-removal process itself works, see how AI image generators work.
What image-to-image AI is used for
Here are four everyday jobs, each with a prompt written for Toybox AI's Image Generator. Add your photo under "Reference photos" first, then paste the prompt.
Restyle a photo
Turn a photo into a painting, a drawing or another art style while keeping its composition. For painting styles, print sizes and gift ideas, see the guide to making a painting from your photo.
Turn my reference photo of our front porch into a loose watercolor painting, with soft washes of color, visible paper texture and relaxed brushwork on the plants. Keep the door, the rocking chair and the pumpkins in the same places, and keep the warm fall colors.Put a subject in a new scene
Keep the product, pet or object from your photo and change everything around it. This is the reference-guided kind of image-to-image, and the guide to the AI image generator with a reference image goes deeper on keeping a character or product consistent.
Place the speckled blue ceramic mug from my reference photo on a rustic wooden cafe table by a rainy window, with steam rising from it. The mug should look just as it does in the photo, with the same speckles, handle and glaze. Soft gray daylight, shallow depth of field, eye-level angle.Turn a sketch into a finished picture
A rough drawing fixes the layout, and the model fills in detail. Research tools such as ControlNet, published in 2023 by Lvmin Zhang, Anyi Rao and Maneesh Agrawala, go further by letting diffusion models follow edges, depth maps, body poses or segmentation maps taken from an input image. For drawings, Toybox also has Sketch to Life, which renders a drawing in a look you pick, such as Photo, Watercolor or Oil painting.
Turn my pencil sketch of a treehouse into a detailed storybook illustration. Keep the ladder, the round window and the shape of the tree where they are in the sketch. Add green leaves, a rope swing and warm evening light, in a soft colored-pencil style.Colorize or restore
Adding color to a black and white photo is one of the original image-to-image translation tasks from the pix2pix paper. A dedicated colorizer is a better fit than a general generator here, because it is built to keep the photo's content as it is. Toybox's Image Colorizer, for example, adds color only and aims to keep the composition and the people as they were. It doesn't repair scratches or tears.
Tips for better image-to-image results
- Say what to keep before what to change. "Keep the layout, the people and the colors, change only the season to winter" gives two clear instructions.
- Change one thing per run. A new style, a new background and a new outfit at once leaves little of your original. Make one change, then run again on the result.
- Start from a clear image. A sharp, well-lit photo or a sketch with clean lines carries over better than a blurry screenshot.
- Match the output shape to your picture. A tall photo forced into a wide frame has to be cropped or filled in. Pick the size closest to your original.
- Expect text on the original to garble. Words on signs, shirts and labels are redrawn along with everything else, and the guide on why AI messes up text explains why.
Limits and rights
Image-to-image AI redraws your picture rather than copying it, so details drift. Faces can come back looking like a sibling, small logos change shape, and the number of windows on a house can change. Check the result against your original before you use it anywhere it has to be accurate.
Rights matter as well. Use your own photos, or ones you have permission for, and don't put real people into images without their consent, which Toybox's Terms of Service prohibit. On copyright, a January 2025 report from the US Copyright Office found that human authors keep copyright in their own work that is perceptible in an AI output, such as your drawing when it shows through in the result, while prompts alone don't give enough control to count as authorship. The guide to whether AI images can be used commercially covers the rest. This is general information, not legal advice.
Guide a new image with your own photos
Image Generator in Toybox AI takes a text description plus up to 3 reference photos, and uses them to guide the style, subject or layout of one new image.
- Open Image Generator. You'll need a free account to create anything.
- Under "Reference photos", add up to 3 images. You can mix photos of different people or things, such as a pet and a place.
- In "Describe your image", say what to keep from the photos and what to change.
- Pick a Size close to your photo's shape: Square 1:1, Portrait 3:4, Landscape 4:3, Story 9:16 or Wide 16:9.
- Tap Create image, then give it about 30 seconds.
- Check faces, logos and small details against your photos, and run it again with a clearer instruction if anything drifted.
Image Generator currently costs 50 credits a picture, or 95 with "Enhanced quality" switched on. Toybox doesn't promise an exact likeness from reference photos, so check faces closely. To change one thing in a photo while keeping the rest, Smart Editor (Pro plan or higher) edits a single image from a written instruction. See current prices.