10 September 2026

OpenAI GPT-Image 2.5: Measure approved-image cost, not generation speed

Flare promises up to 50% lower latency and Sunburst targets precise editing, but the production question is how many attempts it takes to reach an approved asset.

Virtual Arc · Editorial image

What launched

OpenAI released GPT-Image 2.5 on September 8 as two API models: Flare for faster, high-volume generation and Sunburst for precision-focused creation and editing. Both accept text and image inputs and are available through the Images API and the Responses API image tool.

OpenAI reports up to 50% lower latency for Flare versus GPT Image 2 and says multi-turn editing better preserves earlier changes. Published rates for both models are $5 per million text-input tokens, $8 per million image-input tokens and $30 per million image-output tokens.

Measure the right unit

The tempting migration case is speed, but production economics sit one layer higher. If the new model reaches approval in fewer edits, the team wins on compute, review time and queue latency; if it merely returns the first draft sooner, little changes.

Because OpenAI says the existing GPT Image 2 calculator does not estimate 2.5 token consumption, a spreadsheet based only on token rates can mislead. The unit we care about is cost per approved asset, segmented by generation, single edit and long edit chain.

Our rollout plan

We would replay a representative set of real jobs against the current model, Flare and Sunburst, with the same prompts and reference images. We would log token usage, wall-clock latency, failed edits, manual touch time and approval rate.

If Flare wins, it becomes the default route; Sunburst becomes a controlled fallback for brand-sensitive or composition-sensitive work. We would pin dated model snapshots and keep the existing path until the new route proves cheaper and at least as reliable.

Our take

Virtual Arc’s view is that GPT-Image 2.5 is worth testing now, but not because OpenAI says Flare cuts generation latency by up to 50%. The production win is the chance to reduce rejected iterations: if a model preserves the subject, composition and prior edits across several turns, teams spend less compute and less human review time reaching an approved asset. The rate card alone cannot prove that outcome. Flare and Sunburst share published token rates, and OpenAI warns that the GPT Image 2 calculator does not estimate 2.5 token consumption, so cost per image remains a workload question. We would run a shadow evaluation on real edit chains, record tokens, retries, elapsed time and human approval, pin the September 8 snapshots, then use Flare as the default and Sunburst only for precision-sensitive failures. We would not migrate the whole pipeline until cost per approved asset beats the current model.

Sources
  1. Introducing ChatGPT Images 2.5
  2. GPT-Image-2.5 Flare Model — OpenAI API
  3. GPT-Image-2.5 Sunburst Model — OpenAI API
  4. Exclusive: Hands on with ChatGPT's new image editor
  5. ChatGPT Images 2.5 is out — three new features tested

← All posts