OpenAI retired DALL-E 3 in May 2026. Google is sunsetting Imagen 4 in August. Midjourney is on V8.1 while most users are still defaulted to V7. If you picked an AI image generator eighteen months ago and haven't looked since, the tool you settled on has probably shipped two major model versions and rewritten its pricing page at least once.
This guide is for developers and builders who need to pick a tool (or an API) for a real project, not scroll through a hype thread. We looked at eight generators - four consumer apps, four API-first platforms - across the things that actually matter day to day: how well they follow a prompt, whether text renders correctly inside the image, what a thousand images actually costs, and whether you can hit an API endpoint instead of clicking through a web UI.
What Changed in 2026
Three shifts explain why the leaderboard looks different than it did in 2024.
Text rendering got solved, then commoditized. Two years ago, every model mangled text inside images - a sign that read "OPEN 9AM-5PM" would come out as garbled symbols more often than not. Ideogram broke that ceiling first, and now GPT Image 2, Google's Nano Banana 2, and Recraft all handle short text passably. If your use case is logos, posters, or UI mockups, this is the single biggest capability gap between tools that closed this year.
Pricing moved from flat subscriptions to metered tokens. OpenAI, Google, and Black Forest Labs all price image generation the way they price LLM tokens - by resolution and quality tier, not a flat "$X per image." A 1024x1024 GPT Image 2 output can cost anywhere from $0.006 to $0.21 depending on the quality setting you pass. That's a 35x spread on the same endpoint, which means the pricing page alone doesn't tell you what you'll actually pay.
Commercial liability became a selling point. Adobe Firefly is still the only major generator offering full IP indemnification, because Firefly's training data is licensed stock and public domain content rather than scraped web images. For agencies and brands shipping client work, that single fact outweighs every benchmark score.
Quick Comparison
| Tool | Best For | Free Tier | Starting Price | Text-in-Image |
|---|---|---|---|---|
| Midjourney | Artistic, stylized imagery | No | $10/mo | Weak (~30-40%) |
| GPT Image 2 (OpenAI) | Prompt accuracy, editing, general use | Yes (rate-limited) | $20/mo or pay-per-token API | Good |
| Nano Banana 2 (Google) | Photorealism, world knowledge, high-volume API | Yes (AI Studio) | ~$0.02-0.15/image API | Good |
| Adobe Firefly | Commercially indemnified assets | Yes (25 credits/mo) | $9.99/mo | Moderate |
| Ideogram 4 | Logos, posters, readable text, open-weight | Yes (10/day) | $7/mo or $0.03/MP API | Best (0.97 OCR score) |
| Flux.2 (Black Forest Labs) | Product shots, open-weight self-hosting | Via third-party hosts | $0.014-0.07/MP API (5 variants) | Moderate |
| Recraft V3 | Vector art, icons, brand kits | Yes (50/day) | $10/mo | Good (sized text) |
| Leonardo AI | Game/concept art, generous credits | Yes (150 tokens/day) | $12/mo | Moderate |
Midjourney

Midjourney is still the reference point for "looks like art, not like AI." It has no API in the traditional REST sense - you generate through the web app or a Discord bot - which makes it the odd one out on this list for developers, but nobody else touches its output on texture, lighting, and painterly composition.
The model line moved to V8.1 in April 2026 (faster generation, 2K support, better small-detail retention, an opt-in Raw mode), though V7 remains the default most users land on in the web app. Prompt adherence closed a lot of ground on the competition this generation, but Midjourney is still noticeably behind Ideogram and GPT Image 2 on rendering legible text - budget on the 30-40% range if a sign or label needs to read correctly.
Pricing: Basic $10/mo (3.3 fast GPU hours) · Standard $30/mo (15 hours, unlimited Relax mode) · Pro $60/mo (30 hours) · Mega $120/mo (60 hours). Annual billing cuts each tier by 20%. There's no free trial - Midjourney pulled it years ago after abuse - so budget the $10 minimum just to test it.
Verdict: The pick when the output needs to look hand-crafted rather than merely correct - concept art, hero images, mood boards. Skip it if you need an API or reliable text rendering.
GPT Image 2 (OpenAI)

OpenAI shipped GPT Image 2 on April 21, 2026, and retired DALL-E 3 the same cycle - DALL-E 2 and 3 were pulled from the API entirely on May 12. If you have code still pointing at dall-e-3, it's already broken; this is the only image model OpenAI ships now.
The model is genuinely good at following multi-part instructions ("a diagram with three labeled arrows pointing left") and at iterative editing - describe a change to an existing image and it applies just that change rather than regenerating the whole scene. 4K output is in beta. It's also directly reachable from ChatGPT, so non-developers on your team can use the exact same model without touching the API.
Pricing is token-based, not per-image, which catches people off guard:
Pricing: $8/million image input tokens · $2/million cached image input tokens · $30/million image output tokens · $5/million text input tokens. In practice a 1024x1024 image runs about $0.006 at low quality, $0.053 at medium, and $0.211 at high. ChatGPT Plus ($20/mo) or Pro ($200/mo) get you the model through the chat UI with generous but rate-limited use; the free ChatGPT tier includes it with tighter limits.
Bashcurl https://api.openai.com/v1/images/generations \ -H "Authorization: Bearer $OPENAI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-image-2", "prompt": "a minimalist flat-vector logo for a coffee roastery, warm brown palette", "size": "1024x1024", "quality": "medium" }'
Verdict: The safest default for general-purpose generation and editing. The token pricing means you should always pass an explicit quality value in production - leaving it unset at "auto" can silently push a batch job into the $0.21/image tier.
Nano Banana 2 (Google)

Google's standalone Imagen line is being wound down - Imagen 4 shuts off on August 17, 2026 - in favor of image generation built directly into the Gemini model family, branded Nano Banana. Google DeepMind shipped Nano Banana 2 (model ID gemini-3.1-flash-image) on February 26, 2026, rolling out across the Gemini app, Search, AI Studio, the Gemini API, Vertex AI, and Google Ads. Google's own framing is that it brings "the advanced world knowledge, quality and reasoning you love in Nano Banana Pro, at lightning-fast speed" - real-time web search grounding, precise text rendering and translation inside images, and subject consistency across up to five characters and 14 objects in one generation.
That makes for three tiers in the current lineup, and it's worth knowing which one an API call is actually hitting:
- Nano Banana 2 Lite (
gemini-3.1-flash-lite-image) - shipped June 30, 2026 as the fastest, cheapest tier (about four seconds per image), and now replaces the original Nano Banana, which Google has moved to legacy status. - Nano Banana 2 (
gemini-3.1-flash-image) - the general-purpose default: Pro-level quality at Flash speed, 512px to 4K output. - Nano Banana Pro (
gemini-3-pro-image) - the ceiling tier for tasks that need maximum factual accuracy, brand consistency, or complex localization, at a higher per-image cost.
Bashcurl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-image:generateContent?key=$GEMINI_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "contents": [{"parts": [{"text": "product photo of a stainless steel water bottle on a marble countertop, studio lighting"}]}] }'
Pricing: Nano Banana 2 Lite $0.034/image (1K) · Nano Banana 2 $0.045 (512px) / $0.067 (1K) / $0.101 (2K) / $0.151 (4K) · Nano Banana Pro $0.134 (1K-2K) / $0.24 (4K). Batch jobs get a flat 50% discount across every tier if you can tolerate a 24-hour processing window. Free tier available through Google AI Studio for testing.
Verdict: Nano Banana 2 Lite is the cheapest credible option for high-volume, API-only workloads; Nano Banana 2 is the right default for everyday generation and editing; reach for Nano Banana Pro only when a generation genuinely needs the extra factual accuracy. Migrate off imagen-4 model IDs before August 17 or your integration silently breaks.
Adobe Firefly

Adobe Firefly isn't chasing the top of any quality benchmark - it's chasing the one thing legal and compliance teams actually ask about. Firefly is trained exclusively on Adobe Stock, licensed content, and public domain material, which is why it's the only major generator offering full commercial copyright indemnification. If a client's brand assets need to survive a legal review, this is the deciding factor before image quality even enters the conversation.
Where Firefly really pays off is inside the Creative Cloud workflow: Generative Fill in Photoshop, text effects in Illustrator, and the standalone web app all point at the same model, so designers already in Adobe tools don't have to context-switch to a separate generator and re-import results.
Pricing: Free tier with 25 generative credits/month · Standard $9.99/mo (2,000 premium credits) · Pro $19.99/mo (4,000 credits) · Premium $199.99/mo (50,000 credits). Standard generation features (text-to-image, Generative Fill) are unlimited on paid plans - credits are only consumed by premium features like video and lip sync. Enterprise API access runs $0.02-0.10/image with roughly a $1,000/month minimum.
Verdict: Choose Firefly when indemnification and Creative Cloud integration matter more than chasing the sharpest model. Everyone else on this list will out-render it on raw quality.
Ideogram 4

Ideogram 4 shipped in June 2026 and changed two things at once: it's meaningfully better at text, and it's now open-weight - Ideogram's own framing is "the first frontier-grade open-weight text-to-image foundation model," with weights downloadable to fine-tune and self-host rather than locked behind an API. That puts it in direct competition with Flux.2 [dev] for anyone who wants a top-tier model on their own hardware.
On the capability side, Ideogram 4 scores 0.97 on the X-Omni OCR benchmark for in-image text accuracy (competing models cluster around 0.85), renders natively at 2K with background transparency, and accepts a structured JSON layout schema with bounding-box coordinates - you can specify exactly where a headline or logo mark should sit rather than hoping the model interprets your prompt correctly. It currently ranks #1 among open-weight models on the DesignArena leaderboard.
Bashcurl -X POST https://api.ideogram.ai/v1/ideogram-4.0/generate \ -H "Api-Key: $IDEOGRAM_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "retro diner sign that reads OPEN 24 HOURS, neon style", "rendering_speed": "DEFAULT" }'
Pricing: Free tier with 10 prompts/day (~40 images) · Basic $7/mo annual (400 priority generations) · Plus $15-20/mo (1,000 priority) · Pro $42-60/mo (3,000+ priority). API is pay-per-megapixel with no subscription required: Turbo $0.03/MP, Balanced $0.06/MP, Quality $0.10/MP, plus a flat $0.03 if you use prompt expansion. Commercial self-hosting requires a license from Ideogram scaled to usage.
Verdict: If the brief includes any readable text - a sign, a label, a headline - start here before trying anything else on this list. Now also worth a look if you specifically wanted an open-weight model, since it's the newest entrant in that category alongside Flux.2.
Flux.2 (Black Forest Labs)

Flux, from Black Forest Labs (a team of ex-Stability AI researchers), is the model family cloud generators get benchmarked against, and it's the only one on this list that spans fully open-weight to proprietary-flagship in a single lineup. FLUX.2 replaced FLUX.1 as the current generation, and it now ships as five distinct variants rather than one model, each with a different license and price point:
- FLUX.2 [max] - the top-tier flagship (released November 2025), with grounded generation that pulls real-time web context, up to 10 reference images for character/product consistency, and a #2 ranking on the Artificial Analysis text-to-image and editing leaderboards. API-only, priced from $0.07/megapixel.
- FLUX.2 [pro] - the everyday flagship: photorealistic output up to 4MP, strong prompt adherence, faster and cheaper than [max]. API-only, from $0.03/megapixel.
- FLUX.2 [flex] - exposes step count and guidance scale directly, built for text-heavy work like infographics and UI mockups where you need to trade speed for precision. API-only, $0.06/megapixel.
- FLUX.2 [dev] - the open-weight flagship, source-available under a non-commercial license (a commercial license is purchasable from BFL directly). This is the one to self-host.
- FLUX.2 [klein] - the fastest tier, generating and editing images in under a second, released as fully open-source (Apache 2.0) in 4B and 9B parameter sizes built for consumer hardware and real-time use. Priced from $0.014 for the first megapixel plus $0.001 per additional megapixel via the API.
All five are reachable through BFL's own API as well as third-party hosts like fal.ai, Replicate, and DeepInfra, which often shave a bit off BFL's list price.
Bashcurl -X POST https://api.bfl.ai/v1/flux-2-pro \ -H "x-key: $BFL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "isometric illustration of a cloud server rack, flat colors", "width": 1024, "height": 1024 }'
Pricing: [max] $0.07/MP · [flex] $0.06/MP · [pro] $0.03/MP · [klein] $0.014 (first MP) + $0.001/MP after. [dev] and [klein] are the two you can self-host - [dev] needs a commercial license for paid use, [klein] is Apache 2.0. Self-hosting FLUX.2 [dev] needs as little as 8GB VRAM on a quantized GGUF build - see our guide to self-hosting open-source image generators for the full setup.
Verdict: Start with [pro] for general use, move to [max] only when a generation needs grounded web context or more than 4-5 reference images, and reach for [flex] specifically for typography-heavy layouts. [dev] and [klein] are the pair to know about if you want the option to self-host later at zero marginal cost.
Recraft V3

Recraft V3 does something none of the models above can: clean, scalable vector output - actual SVG paths, not a raster image that looks vector-ish. For logos, icon sets, and brand assets that need to scale to a billboard or shrink to a favicon without artifacts, this is the only generator on the list built for the job. It also handles text at any size within the composition, from a small caption to a large headline, without the legibility collapsing the way it does in most diffusion models.
Pricing: Free plan with 50 daily credits (public images, personal use only) · $10/mo for 1,000 credits with commercial rights and private generation · Pro Premium tier at 8,400 credits/month (roughly 210 raster or 105 vector images). API pricing is $0.04 per raster image and $0.08 per vector image.
Verdict: Skip this one entirely unless you specifically need vector output - for raster images its quality is good but not class-leading. For logo and icon work, nothing else here competes.
Leonardo AI

Leonardo AI leans into the game development and concept art crowd, with fine-tuned models for character sheets, environment art, and texture generation that general-purpose tools don't specialize in. It's also, credit-for-credit, the most generous of the subscription tools here.
Pricing: Free tier gives 150 tokens/day (roughly 10-15 images on the first-party Lucid Origin model) · Essential $12/mo (8,500 credits, ~1,400-2,100 images) · Premium $30/mo (25,000 credits, ~4,100-6,250 images) · Ultimate $60/mo (60,000 credits). Paid subscribers own full commercial IP rights on generations. Annual billing saves roughly 20% across every paid tier.
Verdict: The best free-to-cheap option if you're generating a high volume of game assets, concept art, or iterative variations and don't need best-in-class photorealism or text rendering.
How to Choose
Need an API, not a UI? GPT Image 2, Google's Nano Banana 2, Flux.2, and Ideogram all have straightforward REST APIs. Midjourney does not - rule it out immediately if the workflow needs to be programmatic.
Text has to render correctly? Start with Ideogram 4. If you're already committed to another model's ecosystem, GPT Image 2 and Nano Banana 2 are the next-best options for short text.
Client work needs legal cover? Adobe Firefly is the only model with full commercial indemnification. Nothing else on this list makes that guarantee, regardless of what the terms of service imply.
Generating at real volume? Nano Banana 2 Lite's $0.034/image floor (and roughly half that on batch) and Flux.2 [klein]'s $0.014-per-megapixel API pricing are built for scale in a way flat subscriptions aren't. Do the math on your monthly volume before defaulting to a subscription tier - past a few thousand images a month, per-image API pricing usually wins.
Need vectors, not rasters? Recraft V3 is the only real option. Everything else outputs raster images regardless of what the prompt asks for.
Want to self-host and pay nothing per image? Flux.2 [dev] and [klein] are the strongest open-weight options - BFL licenses self-hosting in tiers (Builder for [klein] only, up to Enterprise for the full model family with permissive commercial use). Ideogram 4 is a newer open-weight alternative if text rendering matters more than raw photorealism. See the self-hosting guide for VRAM requirements and setup.
Working the Output Into a Real Pipeline
Generating the image is the easy part - what happens after is where most of these workflows actually break. If you're pulling generations into a web app, run them through the image compressor before shipping - a raw 4K output from GPT Image 2 or Nano Banana 2 is often 5-10x larger than it needs to be for a browser. For anywhere you need multiple sizes, the image resizer handles the crop-and-scale step without round-tripping through a design tool.
If you went the Recraft route for vector logos, the SVG to PNG converter covers you for any surface that still needs a raster fallback, and the favicon generator turns that same asset into a full favicon set in one pass. Converting generated assets to a modern format is worth doing by default - the JPG to WebP converter typically cuts file size 25-35% with no visible quality loss.
Two more that come up constantly in practice: if you're pulling a brand palette out of a generated hero image to drive the rest of a site's design tokens, the color palette generator extracts it directly from the image. And if a generated asset needs a visible watermark before it goes out for client review, the image watermark tool does it client-side without uploading the file anywhere.
Conclusion
GPT Image 2 for general-purpose default use, Ideogram 4 the moment text needs to render correctly, Nano Banana 2 for real volume via API, Adobe Firefly when a client needs commercial indemnification, and Flux.2 if you want the option to self-host later. Pick based on the constraint that actually bites - API access, text accuracy, legal cover, or cost at volume - rather than whichever tool trends that week. All eight are good enough now that the differentiator is almost never raw quality.
Related reading: Best Open-Source AI Image Generators to Self-Host in 2026 covers the five open-weight models worth running on your own hardware, including full setup instructions. Best AI Video Generation Models covers the video-generation tools that share a lot of the same underlying model families.
