AI Image Remix Workflow: Subject + Background + Style
There are two powerful ways to create this type of AI image:
Create Stunning AI Images Without Writing Complex Prompts
There are two powerful ways to create this type of AI image:
Method 1 — Whisk:
Combine three reference images:
Person + Background + Style → New Image
Method 2 — Google Flow + Gemini Omni:
Give the AI your reference images and explain what you want using a natural-language prompt.
The goal is the same: use visual references to control the final image instead of trying to describe everything from scratch.
PART 1 — Create the Image Using Google Whisk
What is Whisk?
Whisk is Google's image-remixing workflow designed around image inputs rather than long text prompts.
The original Whisk workflow allows you to provide:
- Subject — the person/object you want
- Scene — the environment/background
- Style — the visual aesthetic you want
Google describes Whisk as a creative exploration tool rather than a traditional pixel-perfect image editor. It uses Gemini to understand/caption the reference images and then uses Google's image-generation model to create the result. (blog.google)
Think of Whisk as:
Person + Background + Style = New AI Image
Step 1 — Choose Your Subject
The Subject is the person, product, object or character that you want to appear in the final image.
For example:
Your portrait → Subject
Use a relatively clear photograph where the person is visible.
Good Subject Images
- Clear face
- Good lighting
- Person clearly visible
- Minimal obstruction
- Useful body/pose visibility
- High-resolution image
The subject image doesn't necessarily need to have the same background as your desired final image.
Step 2 — Choose Your Background / Scene
The second image defines where you want the subject to appear.
For example:
Santorini street → Scene
You can use:
- Travel photographs
- Architecture
- Beaches
- Cafés
- Hotels
- Streets
- Offices
- Landscapes
- Studios
- Any environment you want
The key idea is that the Scene reference gives the AI visual information about the environment.
Step 3 — Choose Your Style
This is the part most beginners misunderstand.
Style does NOT mean the subject.
It means:
How should the final image look?
For example, you could use a Pinterest image showing:
- Cinematic photography
- 35mm film photography
- Editorial photography
- High-fashion photography
- Candid photography
- Golden-hour photography
- Luxury photography
- Vintage photography
- Moody photography
- Futuristic photography
For example:
Cinematic portrait → Style
You are essentially telling the AI:
"Take the visual language of this image and apply it to my creation."
Step 4 — Upload Your Three Images
Your workflow should now look like:
IMAGE 1
Subject
Your person/object
↓
IMAGE 2
Scene
Your desired environment
↓
IMAGE 3
Style
Your desired aesthetic
Google's original Whisk interface was specifically designed around these three image inputs. (blog.google)
Step 5 — Remix the Images
Once you've uploaded the three references, you can ask Whisk to remix them.
A simple instruction can be:
"Remix these three images into one cohesive image."
Or:
"Create a new image using the subject from the first image, the scene from the second image, and the visual style from the third image."
The second version is better when you want more control.
Step 6 — Generate and Iterate
Don't expect the first generation to always be perfect.
Generate several variations and look for:
- Facial consistency
- Body proportions
- Clothing
- Background integration
- Lighting
- Composition
- Style consistency
If something is wrong, change the relevant input or instruction and generate again.
Important limitation
Whisk is designed for creative remixing, not pixel-perfect compositing.
Google specifically notes that Whisk captures the essence of the reference rather than making an exact replica. That means characteristics such as hairstyle, proportions or other details can sometimes change. (blog.google)
So don't think:
"Copy this exact person and paste her into this exact background."
Think:
"Use this person, this environment and this visual aesthetic as references to create a new image."
PART 2 — Doing the Same Thing in Google Flow + Gemini Omni
Google has been moving Whisk's image-generation capabilities into Flow, making Flow a broader creative workspace for generating, editing and animating images and videos. Google announced that Whisk and ImageFX capabilities were being brought directly into Flow. (blog.google)
Today, Flow includes models and tools such as Nano Banana for image generation/editing and Gemini Omni for multimodal video creation and editing. (Google Labs)
The Big Difference
With Whisk, the workflow is primarily:
Show the AI what you want.
With Flow + Gemini Omni, the workflow becomes:
Show the AI what you want + tell it what to do.
So instead of relying only on:
Subject + Scene + Style
you can provide:
References + Prompt + Instructions
Example: Recreating Our Workflow in Flow
Let's say we have:
Reference 1
Your photograph
Purpose: Person / Subject
Reference 2
A beautiful Mediterranean street
Purpose: Background / Scene
Reference 3
A cinematic Pinterest photograph
Purpose: Style
Now upload these references into the Flow workflow.
Then give Gemini Omni a natural-language instruction.
Example Prompt
Use the person from the first reference image as the main subject. Place her naturally into the environment shown in the second reference image. Use the third reference image as the visual style reference, including its cinematic lighting, color grading, photographic mood and overall aesthetic. Create a photorealistic image that combines these three references into one cohesive scene. Preserve the subject's recognizable facial features and overall appearance while naturally adapting the lighting and perspective to the new environment.
This is essentially the same creative idea as Whisk, but you're explicitly explaining the relationship between the references.
Why Gemini Omni Makes This More Powerful
Gemini Omni in Flow is designed to work with different kinds of inputs and lets you iterate conversationally. Google describes Omni as combining Gemini's intelligence with Google's generative-media models, including improved character consistency and the ability to blend real-world inspiration with generated content. (blog.google)
This means you can start with your references and then continue refining the result.
For example:
First instruction
Put the subject in the new environment using the cinematic style reference.
Then:
Make the lighting warmer.
Then:
Move the subject slightly to the right.
Then:
Make the camera angle slightly lower.
Then:
Make the image look more like a 35mm photograph.
Instead of starting over every time, you're iteratively directing the creation.
PART 3 — Whisk vs Flow + Gemini Omni
| Whisk-style workflow | Flow + Gemini Omni | |
|---|---|---|
| Main approach | Image references | Images + natural language |
| Subject reference | ✅ | ✅ |
| Background reference | ✅ | ✅ |
| Style reference | ✅ | ✅ |
| Long prompt required | ❌ Usually not | ✅ Useful |
| Conversational refinement | Limited | ✅ |
| Image generation | ✅ | ✅ |
| Video generation | Not its primary purpose | ✅ |
| Video editing | Not its primary purpose | ✅ |
| Best for | Fast image remixing | Image + video workflow |
| Creative control | Simple | More detailed |
Flow's current image tools can also blend elements from multiple reference images, while Gemini Omni is designed for more multimodal, conversational creation and editing. (blog.google)
PART 4 — The Simple Mental Model
Whenever you want to create an AI image using references, think about these three questions:
1. WHO?
Subject
Who or what should appear?
↓
2. WHERE?
Scene
Where should they appear?
↓
3. HOW?
Style
How should the final image look?
↓
RESULT
New AI-generated image
Example Workflow
Your photograph
PERSON
↓
Santorini photograph
BACKGROUND
↓
Cinematic film photograph
STYLE
↓
AI
REMIX / GENERATE
↓
Final image
You + Santorini + cinematic photography
PART 5 — How to Find Good Style References
Pinterest is particularly useful for finding style references.
Instead of searching:
"Beautiful AI images"
search for the photographic aesthetic you want.
Cinematic
cinematic portrait photography
Film
35mm film photography portrait
Editorial
fashion editorial photography
Luxury
luxury fashion editorial photography
Candid
candid lifestyle photography
Vintage
vintage film photography aesthetic
Golden hour
golden hour portrait photography
Flash
direct flash fashion photography
Moody
moody cinematic portrait photography
Futuristic
futuristic fashion editorial photography
The important thing is that you're looking for an image that communicates lighting, composition, color, camera treatment and mood.
The Complete Formula
For Whisk
SUBJECT IMAGE
SCENE IMAGE
STYLE IMAGE
↓
REMIX
↓
NEW AI IMAGE
For Flow + Gemini Omni
SUBJECT IMAGE
SCENE IMAGE
STYLE IMAGE
NATURAL-LANGUAGE INSTRUCTION
↓
GENERATE / EDIT
↓
NEW IMAGE OR VIDEO
The Most Important Takeaway
You don't always need to start with a blank text box and write a huge AI prompt.
You can show the AI what you mean.
Instead of describing:
"A woman standing on a Mediterranean cliffside with pink flowers, cinematic lighting, warm color grading, 35mm film look..."
you can provide:
Person 📸 + Background 📸 + Style 📸
and let the AI interpret those visual references.
That's the core idea behind this workflow.
Three visual references can become the starting point for an entirely new AI-generated image.
Current Google terminology note
For your content, I'd call the tool “Whisk” when you're specifically demonstrating the Subject/Scene/Style image-remixing workflow. But for a current tutorial, also mention that Google has brought Whisk's capabilities into Flow, so users may encounter this workflow inside Flow rather than as a completely separate product. (blog.google)
And if you're demonstrating Gemini Omni, describe it as the more advanced reference + instruction workflow rather than saying that Omni is simply “Whisk with prompts.” Google positions Omni as a broader multimodal creation/editing system inside Flow. (blog.google)
Different style references :
| Style you want | Pinterest search |
|---|---|
| 🖤 Luxury / premium | luxury editorial fashion photography |
| 🎬 Cinematic | cinematic portrait photography still |
| 💼 Corporate premium | luxury corporate portrait photography |
| 📰 Magazine | Vogue editorial photography aesthetic |
| 📸 Candid | candid street style photography editorial |
| 🎞️ Film look | 35mm film editorial photography |
| 🌃 Night / neon | cinematic neon portrait photography |
| ☀️ Warm lifestyle | warm golden hour lifestyle photography |
| 🤍 Minimal | minimalist fashion photography studio |
| 🏙️ Urban | urban fashion editorial photography |
| 🇮🇳 Indian luxury | Indian fashion editorial photography |
| 💎 High fashion | high fashion editorial photography |
| 🎨 Artistic | fine art portrait photography editorial |
| 📷 Flash photography | direct flash fashion editorial photography |
| 🕯️ Moody | moody dark editorial portrait photography |
| 🌸 Soft feminine | soft feminine editorial photography |
| 🧊 Clean futuristic | futuristic fashion editorial photography |
| 🏎️ Luxury lifestyle | luxury lifestyle editorial photography |
| 📱 Social-media aesthetic | Instagram fashion editorial photography |
| 🎥 Movie still | cinematic movie still portrait photography |