An AI image generator with a reference image makes a new picture from two inputs: a photo you upload and a sentence describing what you want. The photo supplies the subject, the style or the layout, and your words supply everything else. In Toybox AI, Image Generator takes up to 3 reference photos per image, and Image Generator Lite takes up to 3 photos for cheaper 1K drafts. Below are eight jobs a reference photo handles, five prompts that say what to keep, where the photo stops helping, and how ChatGPT, Gemini and Midjourney compare.
What you can make with a reference image
A reference photo works best when one thing in it has to carry over: a face, a pet, a toy, a product or a look. These are the jobs it handles:
- Restyle a photo. Turn a vacation photo into a watercolor or a family snapshot into a picture book scene. This is what "image to image" means: an AI image generator using an existing image keeps its subject and layout and takes the new style from your words.
- Put yourself somewhere new. Upload a clear photo and describe the new setting, such as a mountain trail or a 1970s living room. The guide to making AI images of yourself covers faces in more depth.
- Combine two photos. Add a photo of each person, or a person and a place, and describe how they meet. The combine two photos page walks through that job.
- Paint a pet portrait. A photo gives the AI your dog's or cat's real markings and ear shape to work from, details that are hard to put into words. See the pet portrait page for styles.
- Reuse an original character. Add the same photo of a mascot or plush toy in every run to place it in new scenes for a story or a shop, and compare each result with the photo.
- Show a design on a product. A photo of your artwork plus "this design printed on a heather gray t-shirt" gives a quick preview, though colors and fine lines in the print can drift from your file.
- Borrow a layout or a color scheme. Toybox's upload box says the photos guide "the style, subject or layout", so a room photo you like can set the palette for a new one.
- Then and now pictures. A photo of you today and one from childhood can go in the same run, for example sitting side by side on a porch.
How reference photos guide the result
Every AI image generator that uses reference images reads the photo and the words together, but tools differ in how closely they follow the photo. Midjourney's docs say its image prompts are inspiration for a new image rather than a copy, and OpenAI's image generation docs say its model can struggle to keep recurring characters and brand elements consistent. Google's Nano Banana guide gives the prompt shape that makes the borrowing clear: the photos, then a relationship instruction (what to take from each), then the new scenario.
The words matter as much as the photo. Midjourney's guide says the text should spell out the whole finished picture rather than tell the model how to change the reference. So write "my dog, painted as a Renaissance oil portrait in a velvet jacket", not "make this photo fancier".
Prompts that say what to take from each photo
Each of these runs in Image Generator. Add the photos first with "Add reference photos", in the order the prompt names them.
A pet portrait from one photo:
A soft watercolor of the cat from my photo asleep on a sunny windowsill, curled around a potted fern, with soft washes of green and gold and white paper showing at the edges. Keep her orange tabby stripes, white paws and the notch in her left ear from the photo.Two people photographed separately, in one picture:
The woman from my first photo and the man from my second photo sitting together on a wooden porch swing at sunset, laughing, with a glass of lemonade each. Warm golden light, candid photo style, a slightly soft background of a green lawn. Keep both faces, hairstyles and glasses true to their own photos.An original toy character in a new scene:
The knitted gray elephant toy from my photo as the hero of a picture book page: it stands at the edge of a sunny meadow holding a tiny red kite, with wildflowers around its feet. Soft gouache illustration style. Keep its button eyes, short trunk and striped scarf from the photo.A style borrowed from one photo, applied to a new subject:
Use the colors and light of my photo of the beach at dusk as the style for a new image: a small blue fishing boat pulled up on the sand, with the same pink and orange sky and soft haze. Wide shot, calm and quiet.A design shown on a product:
The illustration from my photo as a screen print on a heather gray crewneck t-shirt, spread flat on a pale wooden floor with soft daylight from a window. Center the print on the chest and keep its colors as close to my file as possible.Tips for better results from a reference photo
- Use a sharp photo with one clear subject. A well-lit, front-facing photo of one person or pet gives the model the most to work from. A group shot leaves it guessing which face you mean.
- Crop to what matters. Midjourney's docs recommend cropping a reference to match the shape of the final image. Crop out strangers, clutter and a busy background before you upload.
- Number the photos in the prompt. "The dog from my first photo" and "the sofa in my second photo" tell the model which job each one has, the relationship instruction Google's guide describes.
- Name the details you check first. Google's Nano Banana guide tells you to state plainly which details must not change. Write "white chest patch" or "round red glasses" even though they are visible in the photo.
- Describe the whole new scene. Light, setting and style come from your words, so a prompt that only says "make it cooler" leaves most of the picture to chance.
- Pick the size to match the job. Image Generator offers Square 1:1, Portrait 3:4, Landscape 4:3, Story 9:16 and Wide 16:9. Lite adds shapes such as Social 4:5 and Cinema 21:9.
Limits to know
- Likeness isn't guaranteed. Toybox doesn't promise an exact likeness: the AI uses your photos as a guide, so faces can come out different. Try a sharper, front-facing photo, name the features that matter ("dimple on the left cheek"), or run it again. For a headshot from a selfie, Portrait Pro is built for that job and aims to keep your face recognizable.
- Three photos per image. Image Generator and Image Generator Lite each take up to 3. To add a fourth element, make the picture with three, then use the result as a reference in the next run.
- It makes a new picture; it doesn't edit yours. Even with a reference, the whole image is generated fresh, so small details shift. To change one thing in an existing photo, such as a shirt color, use Smart Editor on a Pro plan, which edits one photo from a plain-language instruction.
- Products and printed text can drift. Logos, labels and small lettering may change between the photo and the result. For a shop listing where the item must stay the same, Product Mockup places your exact product into the scene you pick without redesigning it, and every label still needs a proofread.
Other AI image generators that take reference images
The big general image tools also take a reference photo. As of September 2026:
| Tool | How you add a photo | How many | Free option |
|---|---|---|---|
| ChatGPT Images | Tap + and choose "Add photos & files", then describe the change | No fixed number; OpenAI says it depends on image size and text | All plans; Free image generation is limited and slower |
| Gemini app | Upload one or more photos in the chat and describe the new image | Multiple images in one request | Nano Banana 2 without a paid plan; uploaded photos need age 18 or older; 1K downloads without a plan |
| Midjourney | Image Prompts, Style References, or the Edit model on V8.1 and V8.2 | Up to 4 references in the Edit model | No free trial on the website; Basic is $10 a month |
| Toybox Image Generator | "Add reference photos" | Up to 3 | New accounts get 20 free credits; an image currently costs 50 |
| Toybox Image Generator Lite | "Add photos to edit" | Up to 3 | An image currently costs 35 credits |
ChatGPT accepts PNG, JPEG and non-animated GIF uploads of up to 20 MB each. Toybox takes JPG or PNG photos up to 10 MB each. For the idea behind all of these, read what image to image AI is.
How to use a reference image in Image Generator
- Go to Image Generator and sign in, or create a free account and verify your email.
- Tap Add reference photos and upload up to 3 JPG or PNG photos, up to 10 MB each.
- In Describe your image, name what to take from each photo ("the dog from my first photo"), then describe the new scene, the light and the style.
- Under Size, pick Square, Portrait, Landscape, Story or Wide. For a design with words, open More options and switch on Enhanced quality (currently 45 extra credits).
- Tap Create image. Image tools usually finish in about 30 seconds.
- Compare the face, markings and any text with your photo. On a Pro plan, type a fix into "Want to change something?" under the result; otherwise adjust the prompt and create it again.
Results you don't save stay available for 24 hours, so save a picture to your gallery before you close the tab. For plans and credit top-ups, see Toybox pricing. For help with the words half of the prompt, the guide to writing image prompts step by step walks through one idea from vague to usable.