Added May 10, 2026
GPT-Image-2 is OpenAI's image generation model, first announced on April 21, 2026, and available through the API in early May 2026. It improves prompt adherence, text rendering, and photorealism compared to earlier DALL-E generations.
Why: GPT-Image-2 is OpenAI's most capable image model to date, with notably better text-in-image accuracy. It is a natural choice for OpenAI API users who want image generation alongside text and audio in a single platform.
Added May 19, 2026
ComfyUI is an open-source node-based interface for diffusion models. Instead of a prompt box, a generation is a graph: loaders, samplers, conditioning, upscalers and masks wired together, each step inspectable and re-runnable. Workflows save as JSON and can be shared, which is why most published Stable Diffusion and FLUX pipelines circulate as ComfyUI graphs.
Why: If you want to know exactly what happened between the prompt and the image, this is the tool that shows you. Every step is a node you can open, change and re-run, and a workflow someone else built arrives as a file you can load rather than a screenshot you have to reverse engineer. That reproducibility is why it became the format the rest of the field builds around.
Added Aug 8, 2026
OpenArt is a hosted generative image platform with a node-based workflow builder alongside a conventional prompt interface. Workflows can be built from templates or from scratch, run on hosted compute, and published for other people to reuse.
Why: It is the shortest path from wanting a node workflow to having one running. ComfyUI asks you to bring a GPU and set it up; OpenArt hosts the compute and ships a template library, so the graph is something you edit rather than something you first have to stand up.
Added Aug 8, 2026
Invoke is an open-source generative image platform combining a unified canvas with inpainting, outpainting and layer control, plus a node editor for building repeatable workflows. It can be self-hosted or used through a commercial hosted tier aimed at studios.
Why: Its centre of gravity is the canvas rather than the graph, which suits the way a lot of real work happens: generate something, then keep editing regions of it. The node editor is there when a job needs repeating, instead of being the only way in.
Added May 19, 2026
Gemini Omni is a single Google model announced at I/O 2026 that can generate and reason across text, images, video, and audio from unified prompts. It is designed to reduce the need for separate modality-specific models by handling generation and understanding in one architecture.
Why: Gemini Omni represents Google's push toward a single model for all media types. For teams building multimodal products, it simplifies architecture by replacing multiple specialized endpoints with one interface.
Added Jan 31, 2026
Flora is a collaborative AI design canvas that moves beyond the prompt box and into node-based workflow orchestration. It allows designers to build complex, repeatable AI creative engines by connecting different 'Nodes', such as Sketch-to-Image, ControlNet, and multi-model refinement layers. Unlike traditional AI tools, Flora is built for teams, offering real-time collaborative spaces where multiple creators can design and iterate on the same AI canvas simultaneously. It represents the shift from simple prompting to professional AI design systems.
Why: Flora is built for more than one person working on the same graph at the same time, which most node canvases are not. If the bottleneck in your work is handing a workflow to a colleague rather than the workflow itself, that is what it solves. ComfyUI gives you more control and Invoke gives you a better editing canvas, so pick this one for the collaboration.
Added May 15, 2026
FLUX.2 [max] is Black Forest Labs' flagship image generation model. It builds on the FLUX architecture with improved prompt adherence, anatomy, text rendering, and aesthetic quality for professional image creation.
Why: FLUX.2 [max] continues the FLUX lineage of excellent prompt adherence and typography. It is a top choice for designers, advertisers, and developers who need reliable, high-quality image generation.
Added Jan 1, 2026
Generates high-quality photorealistic images from text prompts using Tongyi-MAI's Z-Image model with Single-Stream Diffusion Transformer (S3-DiT) architecture. Produces images in seconds with exceptional detail, lighting, and texture control. Features three variants: Z-Image-Turbo for ultra-fast generation, Z-Image-Base for community fine-tuning, and Z-Image-Edit for precise image editing. Excels at bilingual text rendering, accurately generating both Chinese and English text within images with commercial-grade quality.
Why: Ultra-fast photorealistic generation with superior bilingual text rendering, making it ideal for designs requiring text-in-image accuracy.
Added Jan 1, 2026
Generates high-quality images from text prompts using Alibaba's Tongyi Qianwen 20-billion parameter MMDiT model. Excels at complex text rendering with commercial-grade quality, supporting multi-line layouts and paragraph-level text generation in both Chinese and English. Provides advanced image editing capabilities including style transfer, object insertion/removal, and detail enhancement. Ranks first in multiple public benchmark tests, surpassing similar open-source models with superior prompt understanding and visual quality.
Why: Top-performing open-source model with exceptional text rendering and advanced image editing capabilities, optimized for efficient deployment.
Added May 20, 2026
Recraft V4 is a design-focused image generation model from Recraft, released in 2026. It specializes in brand-consistent visuals, vector graphics, illustrations, and marketing assets with precise style control.
Why: Recraft V4 is built for designers rather than casual prompt users. Its emphasis on brand consistency, vector output, and editable design assets makes it unique among image generation tools.
Added Jan 1, 2026
FLUX.2 Pro is the definitive answer to closed-source image generators like Midjourney. Developed by Black Forest Labs (the original creators of Stable Diffusion), it represents the pinnacle of high-fidelity, open-weight image generation. It features a massive 12B parameter 'Flow' architecture that produces photorealistic textures, perfect human anatomy, and industry-leading text rendering. Unlike its competitors, FLUX is built for the open-source community, supporting LoRA training, ControlNet, and local deployment, allowing creators to maintain full control over their artistic style and data.
Why: FLUX.2 represents the shift toward 'High-End Open Source.' We picked it because it matches Midjourney's aesthetic quality while offering the transparency and customizability that only an open-weight model can provide.
Added May 22, 2026
Nano Banana 2 is Google's fast text-to-image model, available in part through Fal's model hosting platform. It is optimized for speed and low cost, making it suitable for real-time and high-volume image generation applications.
Why: Nano Banana 2 fills the need for a lightning-fast diffusion-style model from a major lab. Its availability on Fal makes it easy for developers to drop into existing inference pipelines without managing their own GPU infrastructure.
Added Feb 5, 2026
Generates high-fidelity images from text prompts using OpenAI's GPT-Image 1.5 model with exceptional prompt adherence and detail preservation. Maintains accurate composition, realistic lighting, and fine-grained details across diverse styles and subjects for production-ready image outputs. Represents OpenAI's latest advancement in image generation with superior prompt understanding, detail accuracy, and visual quality. Suitable for professional workflows requiring high-fidelity outputs with precise prompt control.
Why: OpenAI's flagship image generation model with state-of-the-art prompt following and detail preservation, representing the cutting edge of text-to-image quality.
Added Feb 5, 2026
Black Forest Labs' FLUX.1 [pro] is a state-of-the-art image generation model that outperforms almost everything in prompt adherence, human anatomy, and complex text rendering within images.
Why: FLUX.1 [pro] is the 'Master Artist' for AI images. Most AI tools are bad at writing words inside pictures, but Flux is perfect at it. It's the best tool for designers who need high-quality posters, logos, and photos that look 100% real.
Added Feb 5, 2026
Generates images with adjustable inference steps and guidance scale using Flux 2 Flex model, featuring enhanced typography and text rendering capabilities. Provides fine-tuned control over generation parameters for balancing quality, speed, and style. Allows users to adjust inference steps for speed/quality trade-offs and guidance scale for prompt adherence. Superior text rendering makes it ideal for designs requiring readable text, logos, and typography-heavy graphics.
Why: Best control over generation parameters + superior text rendering, making it ideal for projects requiring precise control and accurate text in images.
Added Oct 2, 2024
FLUX.1.1 [pro] is a refined version of the top-tier FLUX.1 [pro] model from Black Forest Labs, released in October 2024. It improves prompt adherence, realism, and generation speed while maintaining the same high-quality output and text rendering capabilities.
Added Oct 2, 2024
FLUX.1 Canny is a control-oriented Black Forest Labs model that uses Canny edge maps to guide image generation and structure-preserving edits. It is useful for maintaining pose, composition, and object outlines while changing styles or content.
Added Oct 2, 2024
FLUX.1 Depth is a control model from Black Forest Labs that uses depth maps to guide new image generation or editing. It preserves the spatial structure of a scene while allowing changes to objects, lighting, and style.
Added Jun 1, 2025
FLUX.2 [schnell] is the fastest open-weights FLUX.2 variant, designed for 4-8 step local inference on consumer hardware. It retains strong prompt adherence and text rendering while being freely available for local and commercial use.
Added Jun 1, 2025
FLUX.2 [dev] is an open-weight FLUX.2 model from Black Forest Labs, released for non-commercial and commercial research. It offers a strong balance of quality and efficiency, making it the base for many fine-tunes and community LoRAs.
Added Jun 26, 2025
BRIA FIBO is an 8B-parameter DiT text-to-image model trained on long structured JSON captions. It turns short prompts into detailed structured schemas and generates images with precise, reproducible control over composition, lighting, camera, and color. It is also available in an image-to-image 'Inspire' mode and is trained entirely on licensed data for commercial safety.
Why: FIBO stands out for native JSON structured prompting and fully licensed training data, making it the strongest open-source choice for enterprises that need predictable, legally safe image generation.
Added Nov 11, 2025
BRIA FIBO Lite is a lightweight variant of the FIBO image generation pipeline. It pairs an open-source FIBO-VLM bridge with a smaller FIBO Lite model to enable rapid inference and fully local, on-prem deployment for privacy-critical environments.
Why: FIBO Lite gives teams a FIBO-family option optimized for speed and data sovereignty, with a fully local deployment path that the full FIBO pipeline does not emphasize.
Added Dec 13, 2023
Imagen 2 is a diffusion-based text-to-image model developed by Google DeepMind. It emphasizes photorealism, accurate text rendering, and flexible aspect ratios, and is available through Vertex AI and the Google AI Studio image API.
Added May 14, 2024
Open-source text-to-image diffusion transformer with fine-grained Chinese and English understanding, multi-turn prompt refinement, ControlNet, LoRA, and IP-Adapter support. Available via Hugging Face, Diffusers, ComfyUI, and a web demo.
Why: Leading open-source bilingual text-to-image model with strong Chinese prompt understanding and a rich ecosystem.
Added Sep 28, 2025
Native multimodal MoE image generation model with 80 billion total parameters and 13 billion active parameters. It unifies multimodal understanding and generation within an autoregressive framework, supports image editing and multi-image fusion, and is released as the largest open-source image generation model.
Why: The largest open-source image generation model, combining high parameter counts with efficient MoE inference.
Added Aug 1, 2024
Generates photorealistic and stylized images from text prompts with a major leap in realism, prompt adherence, and text rendering over the first Ideogram model. It introduced four style presets—Design, Realistic, 3D, and Anime—along with custom aspect ratios and color palette controls.
Why: Ideogram 2.0 was the release that made Ideogram a serious alternative to Midjourney for realistic, text-heavy marketing imagery before V3 arrived.
Added Sep 1, 2024
A speed-optimized variant of Ideogram 2.0 that trades a small amount of fidelity for much faster generation and lower credit cost, making it ideal for quickly iterating on concepts and drafts.
Why: Ideogram 2a gives creators a faster, cheaper way to produce the same text-in-image style when iteration speed matters more than pixel-perfect quality.
Added Mar 1, 2025
Kling Image 2.0 is Kling AI's image generation model, optimized for high-quality text-to-image and image-to-image generation with strong style and composition control.
Added Jun 15, 2026
Luma Uni-1.1 is a multimodal reasoning model that understands intention, follows reference images, and generates or edits images with style and brand consistency. It supports text-to-image, image-to-image, and multi-reference generation, and ranks highly in human preference benchmarks for overall quality, style and editing, and reference-based generation.
Why: Uni-1.1 ties a reasoning model directly to pixel generation, making it unusually good at following brand references and complex creative direction in images.
Added Sep 26, 2023
Microsoft Copilot is the free consumer AI assistant formerly known as Bing Chat. It answers questions, summarizes web pages, drafts text, generates images, and supports voice conversations across Windows, the web, and mobile apps.
Why: The free, broadly available Microsoft AI assistant that brings search, chat, and image generation into one cross-platform experience.
Added Oct 1, 2022
Microsoft Designer is a browser-based and mobile design app that generates images from text prompts, creates social graphics, and combines AI-generated visuals with templates. It is also the home of Bing Image Creator and integrates with Word, PowerPoint, and Microsoft Photos.
Why: Microsoft's free, template-driven AI design tool that pairs DALL-E image generation with practical layout tools.
Added Mar 13, 2024
Recraft V2 is the second-generation image generation model released by Recraft in March 2024. It was built for professional designers, emphasizing style consistency, vector and raster output, anatomical accuracy, and brand-controlled visuals compared to the first generation.
Why: Recraft V2 was the first generational upgrade that explicitly positioned Recraft as a designer-first model with strong style and anatomy control.
Added May 14, 2026
Recraft V4.1 is Recraft's latest image generation model, released in May 2026. It improves photorealism, short-prompt understanding, and illustration quality, and ships with Standard, Pro, Vector, and Utility variants for different creative and production needs.
Why: Recraft V4.1 is the current flagship model, offering more natural photorealism, refined illustration quality, and dedicated Utility and Vector variants for production design workflows.
Added Oct 1, 2024
Runway Frames is a dedicated image generation model from Runway, designed to create stylized images with strong consistency and to serve as the starting frame for video generations.
Added Oct 20, 2022
Stable Diffusion 1.5 is the landmark open-source latent diffusion model released by Stability AI in 2022. It established the open image generation ecosystem and remains the base for countless fine-tunes, LoRAs, and ControlNet models.
Added Jul 26, 2023
Stable Diffusion XL (SDXL) is a 2023 open-source text-to-image model that generates higher-quality, higher-resolution images than SD 1.5. It uses a two-stage base-plus-refiner pipeline and is widely used for production image workflows.
Added Nov 28, 2023
Stable Diffusion XL Turbo is a distilled, fast variant of SDXL that can generate images in a single step or a few steps. It is optimized for low-latency applications and real-time interactive generation.
Added Jun 12, 2024
Stable Diffusion 3 is a 2024 text-to-image model from Stability AI based on a Multimodal Diffusion Transformer architecture. It improved text rendering, composition, and prompt adherence over SDXL.
Added Jun 12, 2024
Stable Diffusion 3 Medium is a 2B-parameter version of SD3 designed to run well on consumer GPUs. It offers a balance of SD3 quality and efficiency, making it accessible for local creators.
Added Oct 22, 2024
Stable Diffusion 3.5 Large is an 8B-parameter open-weights model in the SD 3.5 family, offering the highest quality and best prompt adherence of the 3.5 series for demanding image generation tasks.
New this month
Added Sep 10, 2026
ChatGPT Images 2.5 is OpenAI's latest image generation model announced September 8, 2026. Delivers sharper detail, more natural lighting and texture, better preservation of people and products in reference photos, and 50% faster generation than Images 2.0. Reliable editing across multi-turn conversations. Available in ChatGPT (all tiers), Codex, and API with two variants: Flare (high-quality, low-latency default) and Sunburst (premium control for production workflows).
Why: Represents significant iterative improvement in image quality, speed, and editing reliability. Two API variants show sophisticated tuning for different use cases. 50% speed improvement is substantial for production workflows.
Added Feb 5, 2026
Generates and edits images with context awareness for better coherence using Flux Kontext model. Understands image context and relationships to produce more coherent variations, edits, and style transfers with improved consistency. Advanced context understanding enables the model to maintain visual relationships, preserve important elements, and create coherent edits that respect the original image's context. Ideal for image editing, variations, and style transfer tasks requiring consistency.
Why: Context-aware generation for more coherent results, making it superior for image editing and variation tasks requiring consistency.
Added Feb 5, 2026
Generates images from text with open-source flexibility and community support using Stable Diffusion 3.5 model. Provides extensive customization options, community models, LoRA support, and self-hosting capabilities for complete workflow control. Latest version of the Stable Diffusion ecosystem with improved quality, better prompt understanding, and enhanced capabilities. Supports local deployment, API access, and extensive community ecosystem with thousands of custom models and tools.
Why: Open-source standard with extensive customization options, making it the foundation for many custom image generation workflows.
Added Jul 7, 2026
Microsoft announced a family of MAI-branded models at Build 2026, including MAI-Thinking-1 for reasoning, MAI-Image-2.5 for image generation and editing, MAI-Voice-2 for expressive text-to-speech, MAI-Transcribe-1.5 for speech-to-text, MAI-Code-1-Flash for coding in GitHub Copilot, and Scout as a workplace personal agent. They integrate tightly with Microsoft 365, Azure, and GitHub.
Why: The MAI family gives Microsoft a cohesive, enterprise-ready AI stack. For organizations already using Microsoft services, these models reduce friction by running inside familiar tools rather than requiring separate platforms.
New this month
Added Sep 6, 2026
ChatGPT Images 2.0 is OpenAI's updated image generation model integrated into ChatGPT. Improves upon prior versions with better detail rendering, faster generation speed (up to 4x faster), and notably improved text rendering in images. Can generate images from text prompts and edit existing photos with more precision. Ranked second globally in text-to-image generation and image editing (behind previous versions of its own gpt-image-2 model). Available to ChatGPT users with Plus/Pro subscriptions.
Why: Shows iterative improvement in text-to-image quality and speed. Better text rendering is significant for design and creative use. Demonstrates practical advancement in generative image quality.
New this month
Added Sep 8, 2026
Google Pics is an AI-powered design platform that simplifies graphic design by replacing traditional UI tools with natural language prompts, similar to generative image models. The tool automatically handles layout, typography, color coordination, and asset selection based on user descriptions. Makes professional-quality design accessible to non-designers within seconds, directly challenging Canva's dominance in accessible design and competing with Adobe's creative suite.
Why: Represents Google's entry into the accessible design space. Demonstrates shift from tool-based to prompt-based creative workflows. Early indication of how generative AI will reshape design software.
Added Feb 5, 2026
Generates high-quality images from text prompts using Black Forest Labs' Flux 1 schnell (fast) variant. Provides the same exceptional image quality as Flux 1 with significantly faster inference times, making it ideal for rapid iteration and high-volume image generation workflows. Optimized architecture enables fast generation while maintaining the superior quality and prompt adherence of the base Flux 1 model. Perfect balance of speed and quality for production workflows requiring rapid image generation.
Why: Fastest Flux variant maintaining top-tier quality, perfect for workflows requiring speed without compromising on image fidelity.
Added Feb 5, 2026
Generates realistic, high-quality images from text prompts using Google's Imagen 3 model. Produces photorealistic images with exceptional detail, proper composition, and accurate prompt understanding. Supports complex scene descriptions and maintains consistency across various artistic styles. Represents Google DeepMind's latest advancement in image generation with superior photorealism, detail accuracy, and scene understanding. Suitable for professional workflows requiring high-fidelity, photorealistic outputs.
Why: Google's flagship image generation model with state-of-the-art quality and photorealism, representing one of the best text-to-image systems available.
Added Feb 5, 2026
Generates long texts, vector art, and images in brand style using Recraft V3. Recognized as state-of-the-art in image generation with exceptional performance on Hugging Face's Text-to-Image Benchmark. Excels at anatomy depiction, prompt understanding, and aesthetic quality, surpassing competitors like Midjourney and OpenAI. Specialized capabilities in vector art generation, brand style consistency, and typography make it unique for design workflows requiring precise style control and readable text in images.
Why: SOTA model excelling at vector art and brand consistency, making it unique for design workflows requiring precise style control and typography.
Added Feb 5, 2026
Generates high-quality images, posters, and logos with exceptional typography handling and realistic outputs. Optimized for both commercial and creative use, with improved realism and understanding of complex text layouts. Capable of generating legible text within images, a feature that sets it apart from other text-to-image models. Latest version (V3) represents improvements in typography accuracy, text readability, and design quality, making it ideal for marketing materials, logos, and text-heavy designs.
Why: Best-in-class typography rendering makes it the top choice for designs requiring text integration, logos, and marketing materials with readable text.
Added Feb 5, 2026
Generates high-quality images from text prompts using Black Forest Labs' Flux 1 development version. Provides advanced control and customization options for developers and power users, with access to experimental features and fine-tuning capabilities for specialized use cases. Development version offers extended parameter control, experimental generation modes, and advanced customization options not available in standard versions. Ideal for developers building custom applications, researchers experimenting with generation parameters, and power users requiring maximum control.
Why: Development version offering advanced control and experimental features, ideal for developers and power users requiring maximum customization.
Added Feb 5, 2026
Generates images optimized for quick, high-quality text rendering, making it suitable for creating marketing graphics with typography, UI mockups, and social media posts with captions. A 7B parameter model designed for efficient deployment and fast iteration in design workflows. Specialized architecture optimized for text-heavy designs, enabling rapid generation of marketing materials, UI mockups, and social media content with readable text. Efficient model size allows for fast deployment and cost-effective generation.
Why: Specialized for marketing graphics and text-heavy designs, making it the ideal choice for social media and UI mockup generation requiring readable text.
Added Feb 5, 2026
Generates images with multilingual text rendering and photorealism using a 6B parameter model optimized for deployment efficiency. Excels at creating multilingual marketing assets and text-heavy social content with proper text rendering across multiple languages and scripts. Unique capability to render text accurately in multiple languages and writing systems, making it essential for global marketing campaigns and international content creation. Combines multilingual text rendering with photorealistic image generation for comprehensive global content workflows.
Why: Unique multilingual text rendering capabilities make it essential for global marketing and content creation requiring text in multiple languages.
Added Feb 5, 2026
A 7B parameter multimodal model developed by ByteDance-Seed, capable of generating both text and images. Supports text-to-image generation, image-to-image editing, and image understanding in a unified framework. Provides versatile capabilities for content creation and image manipulation workflows. Multimodal architecture enables seamless integration of text and image generation with editing capabilities, making it ideal for complex content creation workflows requiring multiple modalities in a single model.
Why: Unique multimodal capabilities combining text and image generation with editing, making it versatile for complex content creation workflows requiring multiple modalities.
Added Feb 5, 2026
Generates photorealistic images from text prompts using Black Forest Labs' Flux model enhanced with Realism LoRA (Low-Rank Adaptation). Combines the exceptional quality of Flux with specialized fine-tuning for realistic, lifelike image generation. Produces images with natural lighting, accurate textures, and authentic details suitable for professional photography-style outputs. LoRA fine-tuning enables specialized realism while maintaining Flux's superior base quality, making it ideal for projects requiring photorealistic outputs.
Why: Unique photorealistic variant of Flux with LoRA fine-tuning, offering specialized realism capabilities that complement the base Flux models for professional photography-style generation.
Added Feb 5, 2026
Generates images from text prompts using Black Forest Labs' Flux model with LoRA (Low-Rank Adaptation) support for custom style fine-tuning. Enables users to apply specialized LoRA models for specific artistic styles, character consistency, or domain-specific generation. Provides the flexibility to customize Flux's output while maintaining its high-quality base generation capabilities. LoRA support allows fine-tuning without retraining the entire model, enabling efficient customization for specialized use cases.
Why: LoRA-enabled Flux variant offering customizable style fine-tuning, making it ideal for specialized use cases requiring consistent character generation or specific artistic styles.
Added Aug 4, 2026
FLUX 3 is Black Forest Labs' multimodal foundation model, announced 23 July 2026. Unlike the FLUX.1 and FLUX.2 image models before it, FLUX 3 learns jointly across images, video and audio in a single unified architecture: it generates video with native synchronised audio, edits images, renders readable text, and, via a FLUX-mimic variant, predicts robot actions, all from the same weights. Video generation runs up to 20 seconds. At launch, video is available through a gated early-access programme, with image generation stated to follow and an open-weight FLUX 3 Dev backbone planned later.
Why: The first credible attempt to collapse image, video and audio generation into a single model rather than a pipeline of separate ones, from the team behind the most widely self-hosted open image models. Access is the catch: video is gated early-access and the open-weight release has not shipped, so treat availability as limited until FLUX 3 Dev lands.