Blog/AI Image Remix Workflow: Subject + Background + Style
@onlygrowthtalks·
24 June 2026

AI Image Remix Workflow: Subject + Background + Style

There are two powerful ways to create this type of AI image:

Create Stunning AI Images Without Writing Complex Prompts

There are two powerful ways to create this type of AI image:

Method 1 — Whisk:

Combine three reference images:

Person + Background + Style → New Image

Method 2 — Google Flow + Gemini Omni:

Give the AI your reference images and explain what you want using a natural-language prompt.

The goal is the same: use visual references to control the final image instead of trying to describe everything from scratch.

PART 1 — Create the Image Using Google Whisk

What is Whisk?

Whisk is Google's image-remixing workflow designed around image inputs rather than long text prompts.

The original Whisk workflow allows you to provide:

  • Subject — the person/object you want
  • Scene — the environment/background
  • Style — the visual aesthetic you want

Google describes Whisk as a creative exploration tool rather than a traditional pixel-perfect image editor. It uses Gemini to understand/caption the reference images and then uses Google's image-generation model to create the result. (blog.google)

Think of Whisk as:

Person + Background + Style = New AI Image

Step 1 — Choose Your Subject

The Subject is the person, product, object or character that you want to appear in the final image.

For example:

Your portrait → Subject

Use a relatively clear photograph where the person is visible.

Good Subject Images

  • Clear face
  • Good lighting
  • Person clearly visible
  • Minimal obstruction
  • Useful body/pose visibility
  • High-resolution image

The subject image doesn't necessarily need to have the same background as your desired final image.

Step 2 — Choose Your Background / Scene

The second image defines where you want the subject to appear.

For example:

Santorini street → Scene

You can use:

  • Travel photographs
  • Architecture
  • Beaches
  • Cafés
  • Hotels
  • Streets
  • Offices
  • Landscapes
  • Studios
  • Any environment you want

The key idea is that the Scene reference gives the AI visual information about the environment.

Step 3 — Choose Your Style

This is the part most beginners misunderstand.

Style does NOT mean the subject.

It means:

How should the final image look?

For example, you could use a Pinterest image showing:

  • Cinematic photography
  • 35mm film photography
  • Editorial photography
  • High-fashion photography
  • Candid photography
  • Golden-hour photography
  • Luxury photography
  • Vintage photography
  • Moody photography
  • Futuristic photography

For example:

Cinematic portrait → Style

You are essentially telling the AI:

"Take the visual language of this image and apply it to my creation."

Step 4 — Upload Your Three Images

Your workflow should now look like:

IMAGE 1

Subject

Your person/object

IMAGE 2

Scene

Your desired environment

IMAGE 3

Style

Your desired aesthetic

Google's original Whisk interface was specifically designed around these three image inputs. (blog.google)

Step 5 — Remix the Images

Once you've uploaded the three references, you can ask Whisk to remix them.

A simple instruction can be:

"Remix these three images into one cohesive image."

Or:

"Create a new image using the subject from the first image, the scene from the second image, and the visual style from the third image."

The second version is better when you want more control.

Step 6 — Generate and Iterate

Don't expect the first generation to always be perfect.

Generate several variations and look for:

  • Facial consistency
  • Body proportions
  • Clothing
  • Background integration
  • Lighting
  • Composition
  • Style consistency

If something is wrong, change the relevant input or instruction and generate again.

Important limitation

Whisk is designed for creative remixing, not pixel-perfect compositing.

Google specifically notes that Whisk captures the essence of the reference rather than making an exact replica. That means characteristics such as hairstyle, proportions or other details can sometimes change. (blog.google)

So don't think:

"Copy this exact person and paste her into this exact background."

Think:

"Use this person, this environment and this visual aesthetic as references to create a new image."

PART 2 — Doing the Same Thing in Google Flow + Gemini Omni

Google has been moving Whisk's image-generation capabilities into Flow, making Flow a broader creative workspace for generating, editing and animating images and videos. Google announced that Whisk and ImageFX capabilities were being brought directly into Flow. (blog.google)

Open Google Flow

Today, Flow includes models and tools such as Nano Banana for image generation/editing and Gemini Omni for multimodal video creation and editing. (Google Labs)

The Big Difference

With Whisk, the workflow is primarily:

Show the AI what you want.

With Flow + Gemini Omni, the workflow becomes:

Show the AI what you want + tell it what to do.

So instead of relying only on:

Subject + Scene + Style

you can provide:

References + Prompt + Instructions

Example: Recreating Our Workflow in Flow

Let's say we have:

Reference 1

Your photograph

Purpose: Person / Subject

Reference 2

A beautiful Mediterranean street

Purpose: Background / Scene

Reference 3

A cinematic Pinterest photograph

Purpose: Style

Now upload these references into the Flow workflow.

Then give Gemini Omni a natural-language instruction.

Example Prompt

Use the person from the first reference image as the main subject. Place her naturally into the environment shown in the second reference image. Use the third reference image as the visual style reference, including its cinematic lighting, color grading, photographic mood and overall aesthetic. Create a photorealistic image that combines these three references into one cohesive scene. Preserve the subject's recognizable facial features and overall appearance while naturally adapting the lighting and perspective to the new environment.

This is essentially the same creative idea as Whisk, but you're explicitly explaining the relationship between the references.

Why Gemini Omni Makes This More Powerful

Gemini Omni in Flow is designed to work with different kinds of inputs and lets you iterate conversationally. Google describes Omni as combining Gemini's intelligence with Google's generative-media models, including improved character consistency and the ability to blend real-world inspiration with generated content. (blog.google)

This means you can start with your references and then continue refining the result.

For example:

First instruction

Put the subject in the new environment using the cinematic style reference.

Then:

Make the lighting warmer.

Then:

Move the subject slightly to the right.

Then:

Make the camera angle slightly lower.

Then:

Make the image look more like a 35mm photograph.

Instead of starting over every time, you're iteratively directing the creation.

PART 3 — Whisk vs Flow + Gemini Omni

Whisk-style workflowFlow + Gemini Omni
Main approachImage referencesImages + natural language
Subject reference
Background reference
Style reference
Long prompt required❌ Usually not✅ Useful
Conversational refinementLimited
Image generation
Video generationNot its primary purpose
Video editingNot its primary purpose
Best forFast image remixingImage + video workflow
Creative controlSimpleMore detailed

Flow's current image tools can also blend elements from multiple reference images, while Gemini Omni is designed for more multimodal, conversational creation and editing. (blog.google)

PART 4 — The Simple Mental Model

Whenever you want to create an AI image using references, think about these three questions:

1. WHO?

Subject

Who or what should appear?

2. WHERE?

Scene

Where should they appear?

3. HOW?

Style

How should the final image look?

RESULT

New AI-generated image

Example Workflow

Your photograph

PERSON

Santorini photograph

BACKGROUND

Cinematic film photograph

STYLE

AI

REMIX / GENERATE

Final image

You + Santorini + cinematic photography

PART 5 — How to Find Good Style References

Pinterest is particularly useful for finding style references.

Instead of searching:

"Beautiful AI images"

search for the photographic aesthetic you want.

Cinematic

cinematic portrait photography

Film

35mm film photography portrait

Editorial

fashion editorial photography

Luxury

luxury fashion editorial photography

Candid

candid lifestyle photography

Vintage

vintage film photography aesthetic

Golden hour

golden hour portrait photography

Flash

direct flash fashion photography

Moody

moody cinematic portrait photography

Futuristic

futuristic fashion editorial photography

The important thing is that you're looking for an image that communicates lighting, composition, color, camera treatment and mood.

The Complete Formula

For Whisk

SUBJECT IMAGE

SCENE IMAGE

STYLE IMAGE

REMIX

NEW AI IMAGE

For Flow + Gemini Omni

SUBJECT IMAGE

SCENE IMAGE

STYLE IMAGE

NATURAL-LANGUAGE INSTRUCTION

GENERATE / EDIT

NEW IMAGE OR VIDEO

The Most Important Takeaway

You don't always need to start with a blank text box and write a huge AI prompt.

You can show the AI what you mean.

Instead of describing:

"A woman standing on a Mediterranean cliffside with pink flowers, cinematic lighting, warm color grading, 35mm film look..."

you can provide:

Person 📸 + Background 📸 + Style 📸

and let the AI interpret those visual references.

That's the core idea behind this workflow.

Three visual references can become the starting point for an entirely new AI-generated image.

Current Google terminology note

For your content, I'd call the tool “Whisk” when you're specifically demonstrating the Subject/Scene/Style image-remixing workflow. But for a current tutorial, also mention that Google has brought Whisk's capabilities into Flow, so users may encounter this workflow inside Flow rather than as a completely separate product. (blog.google)

And if you're demonstrating Gemini Omni, describe it as the more advanced reference + instruction workflow rather than saying that Omni is simply “Whisk with prompts.” Google positions Omni as a broader multimodal creation/editing system inside Flow. (blog.google)

Different style references :

Style you wantPinterest search
🖤 Luxury / premiumluxury editorial fashion photography
🎬 Cinematiccinematic portrait photography still
💼 Corporate premiumluxury corporate portrait photography
📰 MagazineVogue editorial photography aesthetic
📸 Candidcandid street style photography editorial
🎞️ Film look35mm film editorial photography
🌃 Night / neoncinematic neon portrait photography
☀️ Warm lifestylewarm golden hour lifestyle photography
🤍 Minimalminimalist fashion photography studio
🏙️ Urbanurban fashion editorial photography
🇮🇳 Indian luxuryIndian fashion editorial photography
💎 High fashionhigh fashion editorial photography
🎨 Artisticfine art portrait photography editorial
📷 Flash photographydirect flash fashion editorial photography
🕯️ Moodymoody dark editorial portrait photography
🌸 Soft femininesoft feminine editorial photography
🧊 Clean futuristicfuturistic fashion editorial photography
🏎️ Luxury lifestyleluxury lifestyle editorial photography
📱 Social-media aestheticInstagram fashion editorial photography
🎥 Movie stillcinematic movie still portrait photography