An AI photo editor can combine two photos into one by reading both and drawing a single new picture with what you asked for from each: two people on one bench, or your dog on the sofa from another shot. In Toybox AI, Image Generator does this from a written description plus up to 3 reference photos. Room Guest starts from your scene photo and adds one person, pet or thing to it. Smart Editor edits one image at a time, so it can't merge two.
What an AI photo editor can make when you combine two photos
Before you upload anything, write one sentence that says what the finished picture shows. These are the combine jobs each Toybox tool handles:
- Two people who have never been in one photo. Your photo and a friend's photo from overseas become one picture of you both on a park bench, made in Image Generator.
- You now and you as a kid. A current photo plus a childhood photo, described as one instant-camera shot of the two of you side by side. The AI Polaroid prompt page has wording for this trend.
- Three siblings from three separate photos. Image Generator's 3 reference slots cover it, so the three of you can stand together on a front porch you describe.
- A missing person added to a real group photo. Room Guest starts from the actual group photo and places one person in it. The guide on adding someone to a photo covers matching size, light and shadows.
- A pet in the family picture. Room Guest takes a photo of the pet, or a description such as "our gray tabby curled up on the armchair".
- Your product in a setting from another photo. A photo of your mug plus a photo of your kitchen counter, described as one shot of the mug on that counter. For preset scenes such as Studio (a clean white background) or Flat lay, Product Mockup is built for it.
- A person in front of a place from another photo. Background Swapper's "Use my picture" takes a background photo you upload and builds the new scene from it, starting from your photo of the person. The scene is redrawn by AI, so check landmarks and signs against your original.
Tips for better combined photos
The AI only knows what your photos show and what your description says. These habits give it less to guess:
- Name each subject by what it looks like. Write "the woman with short gray hair and round glasses" and "the man in the blue plaid shirt" rather than relying on the order you uploaded the photos.
- Describe one scene, not two photos. Say where they are, the time of day, the light, how far the camera is, and what each person is doing.
- Use sharp, front-facing photos. A face cropped from a distant group shot gives the AI very little to copy. A clear, well-lit photo of each person works better.
- Pick the size first. Landscape (4:3) or Wide (16:9) gives two people room side by side, Portrait (3:4) suits a standing pair, and Square (1:1) fits a profile picture.
- Say what must stay the same. Write it out, such as "keep the sofa's color and the striped pillows as they are", and check that part first when the result comes back.
- Keep the contact simple. Hugs, held hands and arms around shoulders mean inventing hands and arms. Standing or sitting side by side comes out more believable.
These descriptions go in Image Generator's "Describe your image" box, with the photos added through "Add reference photos":
Combine the two people from my reference photos into one photo: the woman with short gray hair and round glasses and the man in the blue plaid shirt, sitting side by side on a wooden porch swing at golden hour, both smiling at the camera. Natural light, a relaxed eye-level photo, a soft green yard behind them.My golden retriever lying on the blue sofa shown in the other reference photo, curled up on the cushions with its head on the armrest. Keep the sofa's color and the striped pillows as they are. Soft afternoon light from a window on the left, a cozy living room photo.Put the speckled white ceramic mug from my reference photos on the oak kitchen counter shown in the other photo, next to a small potted herb, in soft morning light. Keep the mug's shape, color and glaze the same. A clean, realistic product photo from a slightly raised angle.Limits to know
Combining photos with AI is a redraw, not a cut and paste. Four limits matter, each with a way around it:
- Faces and products can come out different. Image Generator uses your photos as a guide rather than copying them, so a face may not look exactly like the person, and a logo, label or glaze may drift from your product photo. Use the sharpest photos you have. When every face must stay exactly as photographed, make a collage instead.
- Mismatched light doesn't always blend. A flash-lit indoor selfie and a sunny beach photo carry different light. Say which light you want in the description, or choose photos taken in similar light. In Room Guest, the "Enhanced blending" switch (Pro plan, currently 20 extra credits) is labeled for cleaner edges and better lighting and size.
- Each run takes a limited number of inputs. Image Generator stops at 3 reference photos per image, and Room Guest places a single subject each time. For a second person in Room Guest, upload your first result as the scene and add them next.
- Smart Editor can't merge photos. It edits one image per run, either an upload or one of your own Toybox results. Use it after combining to fix a detail, like "put a shadow under the dog on the cushion" (Pro plan or higher).
Collage or AI merge: which to use
A collage puts photos next to each other and changes nobody. An AI merge makes one scene and redraws the people in it. Pick by what matters more, exact faces or a single scene:
| Situation | Better choice | Why |
|---|---|---|
| Every face must stay exactly as photographed | Collage | Nothing gets redrawn |
| Two people should look like they were together | Image Generator | It draws one scene around both |
| One person is missing from a real group photo | Room Guest | It starts from the real photo and adds one person |
| A person should stand in front of a real place | Background Swapper with "Use my picture" | The subject stays, and the new scene is built from your background photo |
| A before-and-after or a side-by-side comparison | Collage | Viewers need to see both originals |
To make a collage on Android, pick up to 6 photos in the Google Photos app, then tap Create, then Collage. Tap Templates to choose a layout, and Borders to set the width, corner roundness and color. Google Photos needs at least 3 GB of RAM and Android 8.0 or later for this, and its collages don't let you layer one photo over another. On an iPhone, touch and hold a person in a photo and tap Copy Subject to paste a cutout into any app with layers, if you'd rather combine by hand.
How to combine two photos in Image Generator
Image Generator turns your description and up to three reference photos into one new image. It runs in the browser, so a phone works as well as a laptop.
- Sign in to Image Generator (a free account is fine).
- Tap "Add reference photos" and choose both photos, JPG or PNG, up to 10 MB each. A third photo fits too, if the scene needs one.
- In "Describe your image", name each person or thing by what they look like and describe the one scene you want.
- Under "Size", choose Landscape or Wide for two people side by side, or Portrait for a standing pair.
- If the picture includes words or a logo, switch on "Enhanced quality" under "More options" (currently 45 extra credits).
- Tap "Create image" and allow about 30 seconds.
- Compare every face with your originals. Pro plans can send a follow-up through "Want to change something?", such as "make his beard shorter, like in his photo". Free plan users rewrite the description with the fix and run it once more.
Image Generator is open to Free plan accounts, at a current 50 credits per image, or 95 with "Enhanced quality". The 20 free credits you get at sign-up fall short of one image (current prices). Only upload photos of people who have agreed to it, because Toybox's Terms of Service require the explicit consent of any real person you generate an image of. For more on how reference photos guide a result, read about the AI image generator with a reference image.