What you'll find in this article
- Why I started testing visual AI on advertising creatives — the practical context
- The 4 main tools compared: Midjourney v7, Flux 1.1 Pro, GPT Image 1.5 and Imagen 4
- Display test: how AI images behave against Google RDA requirements
- Shopping test: precise white background, product materials and the AI metadata problem
- Performance Max test: how to generate variety of consistent assets with less effort
- The multi-tool workflow that cut creative production times from days to hours
Why I started testing visual AI on creatives
The starting point was practical, not theoretical. I was managing Google Ads campaigns for a client with a limited creative budget — no photographer, no set, a product catalog with low-quality in-house photography. The classic problem: sufficient media budget, zero creative budget.
The concrete question I asked myself was this: do AI tools for Google Ads — starting with Midjourney — produce images that pass Google Ads review, reach the correct minimum dimensions for the placements I care about, and aren't penalized by the algorithm for some reason I don't yet know?
The short answer is yes — but with precise constraints that depend on the ad format. Display, Shopping and Performance Max have different logics that require different approaches for ad creatives, and no single AI tool excels across all three.
I started with Midjourney because it was the one I knew best. Over time I added Flux for cases where compositional precision mattered more than aesthetics, GPT Image for quick variants, and Google's Imagen 4 when working on products with complex materials. The formats tested cover the entire ecosystem: Google Display ads, Shopping and Performance Max. A year of experimentation gave me enough observational data to write this article with something concrete to say.
The 4 tools tested — characteristics and limitations for Google Ads
🎨 Midjourney v7
Strength: unrivaled artistic quality and lifestyle imagery. The images have an aesthetic that stands out — dramatic lighting, refined compositions, high visual coherence. For Display campaigns where you want to stop the scroll, it is still the best.
Critical limitation: it interprets prompts aesthetically rather than following them literally. If you specify "product centered, white background 80%", Midjourney adds unrequested atmospheric details. It has no public API for automation.
✓ Excellent for Display⚙️ Flux 1.1 Pro
Strength: prompt control. When you write "product bottom left, diagonal light from top right, dark gradient background" — it executes it. It is the tool for precise technical briefs. Available via API with batch generation for automated pipelines.
Critical limitation: slightly clinical and "over-sharpened" aesthetic — a trained eye recognizes it as generative AI. For Shopping where precision matters it is perfect; for Display where emotional immersion counts it is less convincing.
✓ Excellent for Shopping💬 GPT Image 1.5
Strength: understanding of text in images and conversational prompt following. It is the successor to DALL-E 3 — faster, text in images finally reliable. The conversational flow in ChatGPT allows refining through dialogue instead of rewriting the prompt from scratch.
Critical limitation: overall quality of complex scenes still below Midjourney and Flux. Excellent for quick variants of an existing asset, less so for creating high-quality master assets from scratch.
✓ Excellent for PMax variants🔬 Imagen 4
Strength: material rendering. Glass, metal, glossy packaging, liquids — Imagen 4 is an AI-powered engine that reproduces them with accuracy surpassing all competitors on benchmark. For isolated product photography it is the absolute best in 2026.
Critical limitation: available only via Vertex AI (Google Cloud) — no consumer interface. Requires technical setup. Scenes with people regress in quality compared to product-only photography.
✓ Excellent for product shotsTest 1 — Google Display: lifestyle vs product, and the 20% rule
The first scenario I tested systematically concerns Responsive Display Ads. The question: do images generated by Midjourney pass Google Ads review and how do they behave against technical and editorial requirements?
Review result: no automatic disapproval for "AI image". Google Ads has no detection system for the production method — it evaluates content, not origin. Midjourney images that comply with requirements are approved exactly like photographs.
The problem I encountered most often is different: Midjourney tends to generate images with a very rich atmosphere — textures, gradients, background details — that visually seem excellent but, when Google overlays headline and ad copy in the assembled ad, produce a cluttered result. The 20% text rule applies to the text I put inside the image, not to the text Google adds on top — but total visual clutter is a real problem.
The prompt that gave me the most reliable results for Display with Midjourney:
The --style raw flag reduces Midjourney's excessive artistic interpretations and brings the output closer to editorial photography. no text, no overlay in the prompt is not an absolute guarantee — Midjourney sometimes
ignores negatives — but significantly lowers the frequency of inadvertently generated text.
Test 2 — Google Shopping: white background, materials and the metadata problem
Shopping is the territory where differences between tools become most pronounced — and where disapproval risks are highest.
The fundamental requirement for Shopping is white or neutral background, product clearly visible, no overlaid text. Flux 1.1 Pro is the most suitable tool for this use case — its ability to follow precise compositions allows specifying exactly where the product should be and what percentage of the frame the background should occupy.
The prompt I use for Shopping with Flux:
With Flux the result is predictable and repeatable — a fundamental characteristic when you need to produce images for hundreds of SKUs with stylistic consistency. Midjourney with the same prompt produces aesthetically more beautiful but less predictable images — background that turns pearl grey, shadows that appear, compositions that shift. For Shopping this is unacceptable.
The Imagen 4 case for materials: when working on products with complex materials — a transparent glass bottle, a watch with a steel case, packaging with a glossy finish — Imagen 4 produces markedly superior results. The rendering of reflective and transparent surfaces by Imagen 4 is documented as the best among available models in 2026. The problem is access: Vertex AI requires a Google Cloud account and a non-trivial technical configuration.
How to remove metadata before uploading: in Photoshop — File → File Info → delete all. Alternatively, export with "Save for Web (legacy)" which removes IPTC metadata by default. Command-line tools such as exiftool -all= image.jpg are the most scalable solution for file batches.
Test 3 — Performance Max: asset variety with less effort
Performance Max is the format where visual AI delivers the most concrete advantage — not on the quality of the individual asset, but on the speed of producing the necessary variety. PMax requires up to 20 images per asset group in three different formats (landscape, square, portrait). Producing 20 quality images with traditional photography is expensive and slow. With visual AI it is a matter of hours.
The workflow I have developed for PMax:
Step 1 — Master asset with Midjourney. I generate 3-5 high-quality images representing the main creative themes of the campaign. I invest time in the prompt and variants until I have assets I am truly satisfied with. These are the images that will receive the most traffic in the first weeks before the system accumulates performance data.
Step 2 — Variants with GPT Image 1.5. I load the best master asset and request variations — different angle, different light intensity, subject in a different position, slightly varied background. GPT Image 1.5 in ChatGPT's conversational flow allows doing this quickly without having to rewrite the prompt from scratch each time. In 30 minutes I can have 10-15 variants of the master.
Step 3 — Format adaptation with Canva or Photoshop. AI images are generated primarily in one format. Adapting to landscape 1200×628, square 1200×1200 and portrait 960×1200 requires an editing step. For portrait formats in particular — which many skip and which block access to Discover and Shorts — it is essential to ensure the subject is centered before cropping.
One thing I have observed that is worth stating explicitly: the PMax asset rating system (Low/Good/Best) measures relative statistical Google Ads performance — it does not discriminate between AI images and photographic images. I have had Midjourney images rated Best and photographic images rated Low in the same asset group. The measure is CTR and conversion rate, not the production method.
What visual AI still can't do well
It would be dishonest to present only the advantages. There are cases where visual AI is not yet the right answer, and being clear about them avoids costly mistakes.
Consistency of the specific product. If you have a physical product with a precise design — exact brand colors, specific shape, details that identify the product — generative AI tools struggle to reproduce it faithfully without visual references. Midjourney's Character Reference function and image prompting allow getting close, but the result is rarely identical to the real product. For Shopping campaigns where Google compares the image with the product declared in the feed, this can be a concrete problem.
Hands and complex text. In 2026 models have improved significantly, but human hands and complex text in images remain weak points. GPT Image 1.5 has improved text in images compared to DALL-E 3, but for critical headlines in the image (not in the ad layout) it is still more reliable to add text in post with dedicated tools.
People with specific characteristics. Generating a person of a specific age, ethnicity or physical characteristic with consistency across different variants is still an iterative and imprecise process. For campaigns requiring representation of specific people or testimonials, photography remains superior.
Brand consistency across long campaigns. Maintaining the same consistent visual style across dozens of assets produced at different times requires a disciplined prompt management system — a library of prompts with approved templates. Without this, after 6-8 weeks of creative refresh the catalog becomes visually incoherent. This is one aspect of Google Ads management that AI does not yet automate reliably.
The multi-tool workflow that works
After a year of experimentation, the workflow I use in practice is not "one tool for everything" — it is a specialized chain where each tool does what it does best.
1. Brief and prompt templates. Before generating any image, I write a document with the visual specifications of the campaign — color palette, photographic style, main subject, composition, what to avoid. From this I create the prompt templates I use across all tools. The written brief forces clarity and reduces subsequent iteration.
2. Midjourney for master assets. 3-5 main images per campaign, produced with care. I invest time here because these assets set the visual tone of the entire campaign and will receive the greatest distribution in the first weeks.
3. GPT Image 1.5 for variants. From the best master, I generate 8-12 variants through conversational editing. "Change the light from side to top", "shift the subject slightly to the left", "make the background darker". Each variant is a PMax asset or a candidate for A/B testing in Display.
4. Flux for Shopping assets. Separately, I generate the product images for Shopping with Flux — same product, white background, precise composition. I never mix lifestyle styles with Shopping images in the same prompt or the same tool.
5. Cleanup before uploading. Removal of IPTC metadata from all files. Dimension check (minimum 500×500 for Shopping, 1200×628 for Display landscape). Visual check of the subject in the central 80% for Smart Cropping. No overlaid text.
6. Upload and monitoring. Upload to Merchant Center and your Google Ads account. After 14 days, first check of asset rating in PMax and disapprovals in Merchant Center. After 4-6 weeks, refresh the 2-3 lowest-performing assets with new variants.