Getting Started with Nano Banana Photo Editing — How to Write Prompts for Background Swaps, Removing Bystanders, and Colorizing
Pour-over coffee has a basic logic to it: to get a good cup out of a good batch of beans, you don't just dump all the water in at once — you pour in stages, the early pour pulls out the aroma, the middle pour pulls out the sweetness, the late pour controls the bitterness, each stage handled on its own, and the layers stack up at the end. Editing photos with Nano Banana works the same way — split the effect you want into layers and write them out separately, and the result comes out far steadier than trying to cram every requirement into one sentence.
Background swaps: describe the subject and the background separately
The easiest way for a background swap to go wrong isn't the background itself — it's the lighting on the subject not matching the new background, so it looks like a cutout pasted on top. It's worth spelling out two things in the prompt: first, explicitly say the original subject should stay unchanged (person, pose, outfit details all stay put); second, the description of the new background needs to include the light direction and color temperature — something like "keep the subject's original pose and outfit details, replace with a warm indoor setting, light source coming in from the left." Actively write the lighting information into the prompt instead of a vague "change to an indoor background" — the more specific the detail, the lower the odds the front notes and back notes end up clashing.
This matters even more when handling background swaps for e-commerce product shots — the product's own reflections and shadows need to match the logic of the new background's light source, or the product ends up looking like it's floating above the background — instantly fake.
Removing bystanders: the key is in the "fill," not the "delete"
For removing bystanders, a lot of people assume writing "remove the pedestrian in the background" is enough — and yes, that phrase does trigger a removal — but what really determines the quality of the result is what fills in the empty space afterward. It's best to add a line emphasizing background continuity in the prompt, something like "remove the pedestrian from the background, keep the ground texture and the building lines in the distance continuous" — spell out the "fill logic" too, so the model knows what visual information should continue into the gap. Otherwise you often get a patch of color that doesn't match the surrounding style — like an over-extracted pour that comes out bitter: you can taste something's off, but can't say exactly where.
For scenes with multiple bystanders, it's better to process them in batches — handle the one or two most noticeable ones first, check how stable the result is, then keep going. Asking it to remove an entire crowd in one go makes it much harder to pin down which step went wrong if something does.
Colorizing old photos: don't rush to make it "vivid"
The most common way old-photo colorization goes wrong is stuffing the prompt with too many words like "vivid, saturated, colorful," which ends up giving the old photo a heavily digital-looking coat of color, completely at odds with the period texture of the photo itself — like dumping syrup into what should be a mellow pour-over: sure, it's sweeter now, but all the character of the bean is gone.
A steadier approach is to emphasize preserving the material and period feel — something like "colorize this black-and-white photo naturally, preserve the original grain texture and period atmosphere, moderate color saturation, avoid over-saturation." Actively write the requirement for restraint into the prompt instead of letting the model freewheel on saturation. For details like skin tone and clothing material, you can add a specific reference in the prompt, like "skin tone naturally warm, avoid looking overly pale" — the more specific you get, the lower the chance of drifting off course.
A suggested step-by-step order
- First confirm whether this edit is a single request or a compound one (doing a background swap and removing bystanders at once makes it much harder to tell which step went wrong if something does)
- For compound requests, split them into two separate passes — confirm the first step is clean before moving to the second, same as staged pouring: one stage at a time controls the result better than one long boil
- For every step, actively spell out the lighting, material, and what needs to be preserved — don't count on the model to guess what effect you're picturing
- When you're not happy with the result, first figure out whether the description wasn't specific enough, or whether the request itself needed to be split more finely — the two problems have different fixes
More fine-grained local processing, like facial retouching, follows a similar logic — the more clearly the detail is spelled out, the steadier the result. See How to Fix Facial Distortion for the region-by-region approach it covers. Overall, Nano Banana isn't hard to pick up for these kinds of needs — what's hard is breaking down the effect you have in your head into a concrete text description. That step is like picking beans and dialing in water temperature — it takes practice to build a feel for it; no single tutorial gets you there in one read. For more everyday photo-editing use cases, check the photo editing category.
