ChatGPT Image Generation
Native image generation inside ChatGPT powered by ChatGPT Images 2.0 (gpt-image-2), with accurate text rendering, multi turn consistency, and in context learning from uploaded references.
What it is
Image generation built natively into ChatGPT. As of April 2026, ChatGPT Images 2.0 (model ID gpt-image-2) is the default image generation model, replacing GPT Image 1.5. Unlike the legacy DALL-E models (separate diffusion models), Images 2.0 generates images natively within the same model that handles text, enabling multi turn consistency and in context learning from uploads. It ships in two modes: Instant Mode (available on all tiers including Free) and Thinking Mode (Plus and above), which adds reasoning, web search, and character consistency across up to 8 images per prompt. Text rendering has improved dramatically, producing accurate and readable text in images (logos, signs, diagrams, infographics). Photorealistic quality has seen significant gains. Users can edit images conversationally by describing changes in chat, and the model applies targeted modifications while preserving the rest of the image. The model handles up to 10 to 20 objects per scene and supports transparent backgrounds. DALL-E 2 and DALL-E 3 were retired on May 12, 2026. The gpt-image-2 API is available to developers. All generated images include C2PA provenance metadata for content authenticity.
Who it is for
ChatGPT users who want image generation without switching to a separate tool
Marketers creating social media graphics with text overlays, logos, and branded visuals
Designers iterating on visual concepts through multi turn conversation with style consistency
Content creators needing quick illustrations with accurate, readable text in the image
Pros and cons
Pros
- +Integrated directly into ChatGPT, so there is no need to switch to a separate image generation tool
- +Dramatically improved text rendering produces accurate, readable text in images, a major leap over DALL-E 3 and most competing models
- +Conversational image editing lets you describe changes in chat and the model applies them without regenerating from scratch
- +Photorealistic quality has improved significantly with Images 2.0, and Thinking Mode adds reasoning and web search grounding for more complex compositions
- +Multi turn consistency keeps characters and styles visually coherent across multiple generations in the same conversation
- +In context learning from uploaded reference images lets you match brand styles, product photos, or illustration aesthetics
- +Free tier is available, so anyone can try image generation without paying
Cons
- −Quality can be inconsistent on complex scenes with many elements, especially beyond 10 to 20 objects
- −The gpt-image-2 API is available but billed per token, and Thinking Mode adds reasoning token overhead that can make complex prompts 3 to 5x more expensive than Instant Mode
- −Free tier is limited to Instant Mode only and US users see ads
- −Pro plans at $100/mo and $200/mo are expensive if image generation is your primary use case
- −DALL-E 2 and DALL-E 3 were retired on May 12, 2026, so any existing DALL-E integrations need migration to gpt-image-2
Pricing
- Images 2.0 Instant Mode
- GPT-5.3 Instant
- Text rendering supported
- Expanded daily image limits
- Images 2.0 Instant Mode
- Text rendering supported
- Images 2.0 Instant and Thinking Mode
- GPT-5.5 Thinking access
- Transparent backgrounds
- Multi turn consistency
- 5x Plus image generation limits
- GPT-5.5 Pro model access
- Images 2.0 Instant and Thinking Mode
- 20x Plus image generation limits
- Unlimited image generation
- Highest quality output
Prices change; check the official pricing →