Added May 10, 2026
GPT-Image-2 is OpenAI's image generation model, first announced on April 21, 2026, and available through the API in early May 2026. It improves prompt adherence, text rendering, and photorealism compared to earlier DALL-E generations.
Why: GPT-Image-2 is OpenAI's most capable image model to date, with notably better text-in-image accuracy. It is a natural choice for OpenAI API users who want image generation alongside text and audio in a single platform.
Added May 19, 2026
ComfyUI is an open-source node-based interface for diffusion models. Instead of a prompt box, a generation is a graph: loaders, samplers, conditioning, upscalers and masks wired together, each step inspectable and re-runnable. Workflows save as JSON and can be shared, which is why most published Stable Diffusion and FLUX pipelines circulate as ComfyUI graphs.
Why: If you want to know exactly what happened between the prompt and the image, this is the tool that shows you. Every step is a node you can open, change and re-run, and a workflow someone else built arrives as a file you can load rather than a screenshot you have to reverse engineer. That reproducibility is why it became the format the rest of the field builds around.
Added Aug 8, 2026
OpenArt is a hosted generative image platform with a node-based workflow builder alongside a conventional prompt interface. Workflows can be built from templates or from scratch, run on hosted compute, and published for other people to reuse.
Why: It is the shortest path from wanting a node workflow to having one running. ComfyUI asks you to bring a GPU and set it up; OpenArt hosts the compute and ships a template library, so the graph is something you edit rather than something you first have to stand up.
Added Aug 8, 2026
Invoke is an open-source generative image platform combining a unified canvas with inpainting, outpainting and layer control, plus a node editor for building repeatable workflows. It can be self-hosted or used through a commercial hosted tier aimed at studios.
Why: Its centre of gravity is the canvas rather than the graph, which suits the way a lot of real work happens: generate something, then keep editing regions of it. The node editor is there when a job needs repeating, instead of being the only way in.
Added May 19, 2026
Gemini Omni is a single Google model announced at I/O 2026 that can generate and reason across text, images, video, and audio from unified prompts. It is designed to reduce the need for separate modality-specific models by handling generation and understanding in one architecture.
Why: Gemini Omni represents Google's push toward a single model for all media types. For teams building multimodal products, it simplifies architecture by replacing multiple specialized endpoints with one interface.
Added Jan 31, 2026
Flora is a collaborative AI design canvas that moves beyond the prompt box and into node-based workflow orchestration. It allows designers to build complex, repeatable AI creative engines by connecting different 'Nodes', such as Sketch-to-Image, ControlNet, and multi-model refinement layers. Unlike traditional AI tools, Flora is built for teams, offering real-time collaborative spaces where multiple creators can design and iterate on the same AI canvas simultaneously. It represents the shift from simple prompting to professional AI design systems.
Why: Flora is built for more than one person working on the same graph at the same time, which most node canvases are not. If the bottleneck in your work is handing a workflow to a colleague rather than the workflow itself, that is what it solves. ComfyUI gives you more control and Invoke gives you a better editing canvas, so pick this one for the collaboration.
New
Added Oct 2, 2026
MAI-Image-2.6 is Microsoft AI's image model for text-to-image and image editing. The model card dates the 2.6 release to 14 August 2026 and the faster MAI-Image-2.6-Flash release to 4 September 2026. It is a diffusion model with a maximum of 2,359,296 pixels, equivalent to 1,536 by 1,536, and either side can run past 1,536 if the total stays inside that cap. The 4 September post adds multi-image reference editing, web grounding, and dynamic aspect ratios, and puts both models in public preview on Microsoft Foundry and in the MAI Playground. Weights are not released.
Why: On 2 October 2026 it is third on Artificial Analysis AA-Image-Editing v2.0, at 1,137 Elo, and fifth on AA-Image-T2I v2.0, at 1,149 Elo. Microsoft's 4 September post said it was first for editing and second for text-to-image as of that day. The board has moved. It is still the strongest non-OpenAI model on the editing board, at a representative price of $38.90 per 1,000 images.
Added May 15, 2026
FLUX.2 [max] is Black Forest Labs' flagship image generation model. It builds on the FLUX architecture with improved prompt adherence, anatomy, text rendering, and aesthetic quality for professional image creation.
Why: FLUX.2 [max] continues the FLUX lineage of excellent prompt adherence and typography. It is a top choice for designers, advertisers, and developers who need reliable, high-quality image generation.
New
Added Oct 2, 2026
Grok Imagine Image 2.0 is xAI's image model, generally available 7 August 2026 as Quality Mode on grok.com/imagine and in the Grok iOS and Android apps. The launch post describes instruction following for dense layouts and small text, a magic wand and segmentation for local edits, background removal, smart resize across aspect ratios, and multi-reference generation with up to five input images. The API id is grok-imagine-image-2.0, on POST /v1/images/generations and POST /v1/images/edits.
Why: It is the highest text-to-image model on Artificial Analysis outside OpenAI. On 2 October 2026, AA-Image-T2I v2.0 has it at 1,157 Elo, rank 4, inside a rank band of 4 to 5. Editing is tenth, at 1,107 Elo, inside a band of 7 to 13. The model page lists output at $0.04 per image. The leaderboard prices the setting it scored at $60 per 1,000 images.
New
Added Oct 2, 2026
Ideogram 4.5 is Ideogram's image model announced 30 September 2026, built for edits that do not drift. On the Precise Edit endpoint, pixels the instruction does not touch are copied from the source, and the result comes back at that image's own width and height. Ideogram's model page shows this on a 4,016 by 6,016 pixel source. The same model also generates from text. The API takes a prompt, an optional mask, and up to four reference images (three when a mask is set), with quality set to very_low, low, medium, or high. It is live on ideogram.ai and at POST /v2/image/generate/ideogram-4-5 and POST /v2/image/precise-edit/ideogram-4-5. Open weights were promised and have not shipped.
Why: This is the Ideogram release that is actually about editing. Precise Edit is a specific guarantee, copied pixels and original dimensions, not a slogan, and an independent board already has a score for the High tier. It does not lead that board. It is the current Ideogram to reach for when the job is a series of small, local changes.