Added May 10, 2026
GPT-Image-2 is OpenAI's image generation model, first announced on April 21, 2026, and available through the API in early May 2026
Why: GPT-Image-2 is OpenAI's most capable image model to date, with notably better text-in-image accuracy. It is a natural choice for OpenAI API users who want image generation alongside text and audio in a single platform.
Added May 19, 2026
ComfyUI is an open-source node-based interface for diffusion models
Why: If you want to know exactly what happened between the prompt and the image, this is the tool that shows you. Every step is a node you can open, change and re-run, and a workflow someone else built arrives as a file you can load rather than a screenshot you have to reverse engineer. That reproducibility is why it became the format the rest of the field builds around.
New this month
Added Aug 8, 2026
OpenArt is a hosted generative image platform with a node-based workflow builder alongside a conventional prompt interface
Why: It is the shortest path from wanting a node workflow to having one running. ComfyUI asks you to bring a GPU and set it up; OpenArt hosts the compute and ships a template library, so the graph is something you edit rather than something you first have to stand up.
New this month
Added Aug 8, 2026
Invoke is an open-source generative image platform combining a unified canvas with inpainting, outpainting and layer control, plus a node editor for building repeatable workflows
Why: Its centre of gravity is the canvas rather than the graph, which suits the way a lot of real work happens: generate something, then keep editing regions of it. The node editor is there when a job needs repeating, instead of being the only way in.
Added May 19, 2026
Gemini Omni is a single Google model announced at I/O 2026 that can generate and reason across text, images, video, and audio from unified prompts
Why: Gemini Omni represents Google's push toward a single model for all media types. For teams building multimodal products, it simplifies architecture by replacing multiple specialized endpoints with one interface.
Added Jan 31, 2026
Flora is a collaborative AI design canvas that moves beyond the prompt box and into node-based workflow orchestration
Why: Flora is built for more than one person working on the same graph at the same time, which most node canvases are not. If the bottleneck in your work is handing a workflow to a colleague rather than the workflow itself, that is what it solves. ComfyUI gives you more control and Invoke gives you a better editing canvas, so pick this one for the collaboration.
Added May 15, 2026
FLUX
Why: FLUX.2 [max] continues the FLUX lineage of excellent prompt adherence and typography. It is a top choice for designers, advertisers, and developers who need reliable, high-quality image generation.
Added Jan 1, 2026
Generates high-quality photorealistic images from text prompts using Tongyi-MAI's Z-Image model with Single-Stream Diffusion Transformer (S3-DiT) architecture
Why: Ultra-fast photorealistic generation with superior bilingual text rendering, making it ideal for designs requiring text-in-image accuracy.
Added Jan 1, 2026
Generates high-quality images from text prompts using Alibaba's Tongyi Qianwen 20-billion parameter MMDiT model
Why: Top-performing open-source model with exceptional text rendering and advanced image editing capabilities, optimized for efficient deployment.
Added May 20, 2026
Recraft V4 is a design-focused image generation model from Recraft, released in 2026
Why: Recraft V4 is built for designers rather than casual prompt users. Its emphasis on brand consistency, vector output, and editable design assets makes it unique among image generation tools.
Added Jan 1, 2026
FLUX
Why: FLUX.2 represents the shift toward 'High-End Open Source.' We picked it because it matches Midjourney's aesthetic quality while offering the transparency and customizability that only an open-weight model can provide.
Added May 22, 2026
Nano Banana 2 is Google's fast text-to-image model, available in part through Fal's model hosting platform
Why: Nano Banana 2 fills the need for a lightning-fast diffusion-style model from a major lab. Its availability on Fal makes it easy for developers to drop into existing inference pipelines without managing their own GPU infrastructure.
Added Feb 5, 2026
Generates high-fidelity images from text prompts using OpenAI's GPT-Image 1
Why: OpenAI's flagship image generation model with state-of-the-art prompt following and detail preservation, representing the cutting edge of text-to-image quality.
Added Feb 5, 2026
Black Forest Labs' FLUX
Why: FLUX.1 [pro] is the 'Master Artist' for AI images. Most AI tools are bad at writing words inside pictures, but Flux is perfect at it. It's the best tool for designers who need high-quality posters, logos, and photos that look 100% real.
Added Feb 5, 2026
Generates images with adjustable inference steps and guidance scale using Flux 2 Flex model, featuring enhanced typography and text rendering capabilities
Why: Best control over generation parameters + superior text rendering, making it ideal for projects requiring precise control and accurate text in images.
Added Jun 26, 2025
BRIA FIBO is an 8B-parameter DiT text-to-image model trained on long structured JSON captions
Why: FIBO stands out for native JSON structured prompting and fully licensed training data, making it the strongest open-source choice for enterprises that need predictable, legally safe image generation.
Added Nov 11, 2025
BRIA FIBO Lite is a lightweight variant of the FIBO image generation pipeline
Why: FIBO Lite gives teams a FIBO-family option optimized for speed and data sovereignty, with a fully local deployment path that the full FIBO pipeline does not emphasize.
Added Dec 13, 2023
Imagen 2 is a diffusion-based text-to-image model developed by Google DeepMind
Added May 14, 2024
Open-source text-to-image diffusion transformer with fine-grained Chinese and English understanding, multi-turn prompt refinement, ControlNet, LoRA, and IP-Adapter support
Why: Leading open-source bilingual text-to-image model with strong Chinese prompt understanding and a rich ecosystem.
Added Sep 28, 2025
Native multimodal MoE image generation model with 80 billion total parameters and 13 billion active parameters
Why: The largest open-source image generation model, combining high parameter counts with efficient MoE inference.
Added Aug 1, 2024
Generates photorealistic and stylized images from text prompts with a major leap in realism, prompt adherence, and text rendering over the first Ideogram model
Why: Ideogram 2.0 was the release that made Ideogram a serious alternative to Midjourney for realistic, text-heavy marketing imagery before V3 arrived.
Added Sep 1, 2024
A speed-optimized variant of Ideogram 2
Why: Ideogram 2a gives creators a faster, cheaper way to produce the same text-in-image style when iteration speed matters more than pixel-perfect quality.
Added Mar 1, 2025
Kling Image 2
Added Jun 15, 2026
Luma Uni-1
Why: Uni-1.1 ties a reasoning model directly to pixel generation, making it unusually good at following brand references and complex creative direction in images.
Added Sep 26, 2023
Microsoft Copilot is the free consumer AI assistant formerly known as Bing Chat
Why: The free, broadly available Microsoft AI assistant that brings search, chat, and image generation into one cross-platform experience.
Added Oct 1, 2022
Microsoft Designer is a browser-based and mobile design app that generates images from text prompts, creates social graphics, and combines AI-generated visuals with templates
Why: Microsoft's free, template-driven AI design tool that pairs DALL-E image generation with practical layout tools.
Added Mar 13, 2024
Recraft V2 is the second-generation image generation model released by Recraft in March 2024
Why: Recraft V2 was the first generational upgrade that explicitly positioned Recraft as a designer-first model with strong style and anatomy control.
Added May 14, 2026
Recraft V4
Why: Recraft V4.1 is the current flagship model, offering more natural photorealism, refined illustration quality, and dedicated Utility and Vector variants for production design workflows.
Added Oct 1, 2024
Runway Frames is a dedicated image generation model from Runway, designed to create stylized images with strong consistency and to serve as the starting frame for video generations
Added Oct 20, 2022
Stable Diffusion 1
Added Jul 26, 2023
Stable Diffusion XL (SDXL) is a 2023 open-source text-to-image model that generates higher-quality, higher-resolution images than SD 1
Added Nov 28, 2023
Stable Diffusion XL Turbo is a distilled, fast variant of SDXL that can generate images in a single step or a few steps
Added Jun 12, 2024
Stable Diffusion 3 is a 2024 text-to-image model from Stability AI based on a Multimodal Diffusion Transformer architecture
Added Jun 12, 2024
Stable Diffusion 3 Medium is a 2B-parameter version of SD3 designed to run well on consumer GPUs
Added Oct 22, 2024
Stable Diffusion 3
Added Feb 5, 2026
Generates and edits images with context awareness for better coherence using Flux Kontext model
Why: Context-aware generation for more coherent results, making it superior for image editing and variation tasks requiring consistency.
Added Feb 5, 2026
Generates images from text with open-source flexibility and community support using Stable Diffusion 3
Why: Open-source standard with extensive customization options, making it the foundation for many custom image generation workflows.
Added Jul 7, 2026
Microsoft announced a family of MAI-branded models at Build 2026, including MAI-Thinking-1 for reasoning, MAI-Image-2
Why: The MAI family gives Microsoft a cohesive, enterprise-ready AI stack. For organizations already using Microsoft services, these models reduce friction by running inside familiar tools rather than requiring separate platforms.
Added Feb 5, 2026
Generates high-quality images from text prompts using Black Forest Labs' Flux 1 schnell (fast) variant
Why: Fastest Flux variant maintaining top-tier quality, perfect for workflows requiring speed without compromising on image fidelity.
Added Feb 5, 2026
Generates realistic, high-quality images from text prompts using Google's Imagen 3 model
Why: Google's flagship image generation model with state-of-the-art quality and photorealism, representing one of the best text-to-image systems available.
Added Feb 5, 2026
Generates long texts, vector art, and images in brand style using Recraft V3
Why: SOTA model excelling at vector art and brand consistency, making it unique for design workflows requiring precise style control and typography.
Added Feb 5, 2026
Generates high-quality images, posters, and logos with exceptional typography handling and realistic outputs
Why: Best-in-class typography rendering makes it the top choice for designs requiring text integration, logos, and marketing materials with readable text.
Added Feb 5, 2026
Generates high-quality images from text prompts using Black Forest Labs' Flux 1 development version
Why: Development version offering advanced control and experimental features, ideal for developers and power users requiring maximum customization.
Added Feb 5, 2026
Generates images optimized for quick, high-quality text rendering, making it suitable for creating marketing graphics with typography, UI mockups, and social media posts with captions
Why: Specialized for marketing graphics and text-heavy designs, making it the ideal choice for social media and UI mockup generation requiring readable text.
Added Feb 5, 2026
Generates images with multilingual text rendering and photorealism using a 6B parameter model optimized for deployment efficiency
Why: Unique multilingual text rendering capabilities make it essential for global marketing and content creation requiring text in multiple languages.
Added Feb 5, 2026
A 7B parameter multimodal model developed by ByteDance-Seed, capable of generating both text and images
Why: Unique multimodal capabilities combining text and image generation with editing, making it versatile for complex content creation workflows requiring multiple modalities.
Added Feb 5, 2026
Generates photorealistic images from text prompts using Black Forest Labs' Flux model enhanced with Realism LoRA (Low-Rank Adaptation)
Why: Unique photorealistic variant of Flux with LoRA fine-tuning, offering specialized realism capabilities that complement the base Flux models for professional photography-style generation.
Added Feb 5, 2026
Generates images from text prompts using Black Forest Labs' Flux model with LoRA (Low-Rank Adaptation) support for custom style fine-tuning
Why: LoRA-enabled Flux variant offering customizable style fine-tuning, making it ideal for specialized use cases requiring consistent character generation or specific artistic styles.
New this month
Added Aug 4, 2026
FLUX 3 is Black Forest Labs' multimodal foundation model, announced 23 July 2026
Why: The first credible attempt to collapse image, video and audio generation into a single model rather than a pipeline of separate ones, from the team behind the most widely self-hosted open image models. Access is the catch: video is gated early-access and the open-weight release has not shipped, so treat availability as limited until FLUX 3 Dev lands.