Last verified: July 25, 2026
For the highest current broad image quality, start with GPT Image 2. For detailed typography, localization, brand-consistent edits, and Google-grounded visual work, test Nano Banana Pro. For photorealistic scenes at a much lower published per-image price, Seedream 5 Pro is the value challenger. There is no universal winner: a model that wins a blind text-to-image preference arena can still be the wrong choice for an e-commerce edit, a UI asset, or a high-volume creative pipeline.
This report combines current API documentation, the public Artificial Analysis Image Arena and one clearly labeled independent professional review. It does not claim that a public leaderboard reproduces every prompt, policy setting, API region, or gateway implementation. LumeAPI availability and price are checked separately from provider and leaderboard data.
Executive decision
- Best overall quality: GPT Image 2 high is first on the current Artificial Analysis text-to-image arena, at 1,337 Elo from 14,525 blind-vote samples.
- Best for high-value precise creative work: Nano Banana Pro is a premium model for world knowledge, localization, brand consistency and controlled editing; it is ninth on that arena at 1,217 Elo.
- Best photorealism/value challenger: Seedream 5 Pro is eighth at 1,229 Elo and costs $90 per 1,000 images in that arena's API-price normalization, versus $211 for GPT Image 2 high and $134 for Nano Banana Pro.
- Best low-cost current LumeAPI entry point: Qwen Image 2.0 is listed at $0.03 per image, and Seedream 5 at $0.035. Use a small acceptance-rate test before routing production traffic by price alone.
- Do not choose from one rank. Separate text rendering, photorealism, reference fidelity, edit control, safety policy, latency, resolution, and cost per accepted image.
What the public preference leaderboard says

Figure 1. Live public text-to-image leaderboard. The link is the source because its ranks and sample counts change continuously. It uses blind user preference votes; Elo is a relative preference rating, not a percentage of prompts completed.
| Model / setting | Arena Elo | 95% CI | Arena samples | Arena API-price reference | What that tells you |
|---|---|---|---|---|---|
| GPT Image 2 (high) | 1,337 | ±8 | 14,525 | $211 / 1k images | Current broad preference leader |
| Seedream 5.0 Pro | 1,229 | ±12 | 2,480 | $90 / 1k images | Strong frontier quality at lower normalized price |
| Nano Banana Pro | 1,217 | ±8 | 9,206 | $134 / 1k images | Premium Google image model; quality is task-dependent |
| FLUX.2 max | 1,193 | ±8 | 9,129 | $70 / 1k images | Competitive flexible alternative |
| Recraft V4.1 Utility Pro | 1,207 | ±8 | 7,729 | $210 / 1k images | Design-oriented specialist, not a low-cost default |
| Qwen Image 2.0 Pro | 1,170 | ±8 | 5,270 | $75 / 1k images | Relevant alternative for text and creative-generation testing |
The ranking is useful because it has real blind comparisons and confidence intervals. It is not a product requirements document. It does not tell you whether the output preserved a specific product SKU, met a regulated-content policy, kept a character consistent through ten revisions, or arrived quickly enough for an interactive UI. The listed price is also the leaderboard's standardized API reference, not a quote from LumeAPI.
Model capability and integration comparison
| Model | Best first use | Generation and editing | Resolution / control signal | Important limitation |
|---|---|---|---|---|
| GPT Image 2 | General-purpose premium generation, editing, and high-quality final assets | Text and image input; image output; OpenAI Images generation and edit endpoints | Flexible image sizes and high-fidelity image inputs | Highest benchmarked price in this table; test policy behavior and throughput |
| Nano Banana Pro (Gemini 3 Pro Image) | Brand assets, localized campaigns, knowledge-heavy compositions, precise revisions | Conversational multimodal image generation and editing | Google positions it for advanced localization, brand consistency and precision control; 1K/2K/4K pricing tiers | Higher cost than volume models; Google applies SynthID watermarking |
| Seedream 5 Pro | Photorealistic creative work and cost-conscious premium experiments | Text-to-image and provider-specific editing flows | Strong community/independent signals for photo realism | Less evidence that it is the best infographic or stylized-art choice |
| FLUX.2 max / pro | Alternative visual styles and teams wanting a BFL workflow | Provider-specific generation/edit controls | Max ranks 17th; Pro ranks 21st in the cited arena | Benchmark rank is below the three leaders; assess reference fidelity yourself |
| Recraft V4.1 Utility Pro | Vector-like design, utility artwork and graphic workflows | Design-oriented API workflow | Quality rank 11th in the arena | Its $210/1k reference price is near GPT Image 2 high |
| Qwen Image 2.0 | Low-cost challenger in a multi-model test set | Image generation through the selected provider route | LumeAPI lists 2K default output | Lower arena rank than the frontier leaders; test non-English text and edits for your exact briefs |
OpenAI documents GPT Image 2 as its state-of-the-art image model for generation and editing, supporting image input/output and the /v1/images/generations and /v1/images/edits endpoints. Google documents four Nano Banana API models and calls Pro the premium choice for complex visual tasks, with the highest world knowledge and advanced localization. Both statements describe product positioning; the comparative quality evidence comes from the independent arena, not from treating vendor claims as a head-to-head benchmark.
Sources: GPT Image 2 model documentation, Gemini image-generation documentation, and the Artificial Analysis leaderboard.
Price: normalize the decision around an accepted image
A $0.03 image is not cheaper if it needs five retries and still requires a designer rebuild. Conversely, the best-looking $0.21 image can be too expensive for pre-generating thousands of variants. Track the unit that matters: cost per accepted deliverable, including regeneration, edit passes, and human correction time.
The current public LumeAPI catalog, checked July 25, shows these callable image routes:
| LumeAPI model ID | Catalog price | Stated catalog capability | Practical use |
|---|---|---|---|
gpt-image-2-1k | $0.05/image | 1K, aspect-ratio size, image_urls for image-to-image | Default premium test lane |
gpt-image-2-2k | $0.08/image | 2K, aspect-ratio size, image_urls for image-to-image | Sharper final assets |
google/gemini-3-pro-image-preview | $0.134/image at 1K/2K; $0.24 at 4K | Nano Banana Pro, aspect-ratio size and image URLs | Complex high-value edits |
doubao-seedream-5.0 | $0.035/image | 2K default, aspect-ratio size | Photo-realistic/value test lane |
qwen-image-2.0 | $0.03/image | 2K default, aspect-ratio size | Low-cost challenger lane |
gemini-3.1-flash-image-preview | $0.05/image | 1K async image generation | Fast generalist draft lane |
The catalog's model IDs and behavior are the deployable facts for LumeAPI, while the vendors' native SDKs and feature sets can differ. In particular, do not assume that every native Google conversational feature, OpenAI streaming option, or provider editing parameter appears unchanged behind an OpenAI-compatible gateway. Check the relevant model documentation, make one authenticated staging request, and log the result before a rollout.
Official price vs LumeAPI price: the decision table
This is the missing comparison that should drive a real buying decision. The table below separates official native price from LumeAPI catalog price as of July 25, 2026. “Saving” is calculated as (native reference − LumeAPI price) / native reference. It is a per-generated-image comparison before prompt tokens, reference-image tokens, optional search grounding, retries, tax, credits, or human revision.
| Output route | Native/official price used | LumeAPI price | Saving | 1,000 images: native → LumeAPI | 10,000 images: native → LumeAPI | Price-source status |
|---|---|---|---|---|---|---|
| GPT Image 2 1K | $0.0584 | $0.050 | 14.4% | $58.40 → $50.00 | $584 → $500 | LumeAPI catalog official-reference field; OpenAI charges image output by tokens, so exact native cost also varies with request shape |
| GPT Image 2 2K | $0.170 | $0.080 | 52.9% | $170 → $80 | $1,700 → $800 | LumeAPI catalog official-reference field; verify with OpenAI's current calculator for your exact quality and dimensions |
| Gemini 3.1 Flash Image 1K | $0.067 | $0.050 | 25.4% | $67 → $50 | $670 → $500 | Google Gemini API standard image-output price |
| Gemini 3.1 Flash Image 2K | $0.101 | $0.080 | 20.8% | $101 → $80 | $1,010 → $800 | Google Gemini API standard image-output price |
| Nano Banana Pro 1K/2K | $0.134 | $0.134 | 0% | $134 → $134 | $1,340 → $1,340 | Google Gemini API standard image-output price |
| Seedream 5 | $0.150 | $0.035 | 76.7% | $150 → $35 | $1,500 → $350 | LumeAPI catalog official-reference field; native public price must be rechecked before a large commitment |
| Qwen Image 2.0 | $0.030 | $0.030 | 0% | $30 → $30 | $300 → $300 | LumeAPI catalog official-reference field |
What this table does — and does not — prove
The numbers do not prove that every LumeAPI route is cheaper. Nano Banana Pro and Qwen Image 2.0 currently match their stated reference rates. The clearest listed spread is Seedream 5 at 76.7%, followed by GPT Image 2 2K at 52.9%. GPT Image 2 1K is a modest 14.4% reduction, so choose it for output quality or edit success, not because the price gap alone is decisive.
For Google, the native comparison is directly cross-checkable. Google's current developer price page lists Gemini 3.1 Flash Image at $0.067 for 1K and $0.101 for 2K; it lists Nano Banana Pro at $0.134 for 1K/2K and $0.24 for 4K. Google also charges input tokens: Pro image input is approximately $0.0011 per image. Search grounding can add $14 per 1,000 queries after the shared free allowance. These items are outside a simple “per output image” card price.
For OpenAI, the correct native comparison is more conditional. GPT Image 2 bills text input, optional reference-image input, and generated image output tokens. The catalog's $0.0584 and $0.17 values are therefore useful stated reference points for its 1K and 2K routes, not a promise that every transparent-background or multi-reference edit has the same native bill. Use OpenAI's current image calculator before committing to a large native-provider budget.
For Seedream and Qwen, this report records the catalogue's current official-reference fields but does not claim an independently retrieved vendor invoice rate. Their direct providers, regions, credits, resolutions and commercial terms may differ. That distinction is deliberate: a gateway's price is actionable for a LumeAPI request; a separate vendor quote needs provider-side verification.
Sources: Google Gemini API pricing, Google Gemini 3 model guide, OpenAI GPT Image 2 documentation, and the current LumeAPI model catalog.
Cost per accepted image: the metric a production team should use
The card price is only the numerator. If a generated image is accepted without a material rework with probability a, then the first-pass cost per accepted image is:
cost per accepted image = price per generation / acceptance rateThis makes the routing threshold concrete. Assume GPT Image 2 2K costs $0.08 through LumeAPI and is accepted 70% of the time for your product-shot brief. Its first-pass cost per accepted image is $0.114. Seedream 5 at $0.035 only needs an acceptance rate above 30.6% to cost less per accepted image for the same task. It does not need to look equally good on every output; it needs to clear the acceptance threshold after your QA rule.
| Route and assumed acceptance rate | Cost / generated image | Cost / accepted image | When it wins |
|---|---|---|---|
| GPT Image 2 2K at 70% | $0.080 | $0.114 | A higher-quality benchmark for final assets |
| Seedream 5 at 50% | $0.035 | $0.070 | Wins on cost if half of photo-real outputs are usable |
| Gemini 3.1 Flash Image 1K at 65% | $0.050 | $0.077 | Strong draft/iteration lane if the output meets the brief |
| Nano Banana Pro at 90% | $0.134 | $0.149 | Rational only when control, localization or brand fidelity saves more human time |
| Qwen Image 2.0 at 45% | $0.030 | $0.067 | Cheap variant lane, provided QA can reject failures automatically |
These are scenario calculations, not benchmark outcomes. Add two real costs that teams routinely omit: a failed request that is still billed, and a human correction that turns a “pass” into a usable asset. For a four-variant ad experiment, record both generation attempts and the one image that actually reaches the campaign. For an image-editing workflow, record each reference upload and revision round as a separate cost center.
A deeper routing playbook by creative job
1. Product-commerce photos and lifestyle composites
Start with Seedream 5 and GPT Image 2 2K on the same 20 reference product briefs. Score: product shape preservation, logo/text integrity, contact shadows, material realism, and whether the result can enter a listing without manual retouch. Seedream's low listed $0.035 price makes it the economic challenger; GPT Image 2 is the quality-control lane. Route only after measuring a category-specific acceptance rate — cosmetics, furniture and transparent packaging behave differently.
2. Localized ads, readable text and structured infographics
Start with Nano Banana Pro and GPT Image 2. Google positions Nano Banana Pro for advanced localization, world knowledge, brand consistency and precision creative control. This does not mean every spelling is correct; evaluate native language, digit strings, legal disclaimer fidelity, logo distortion and layout overflow. A single incorrect price or translated claim can erase the apparent $0.05–$0.10 image saving.
3. High-volume creative exploration
Use Gemini 3.1 Flash Image 1K or Qwen Image 2.0 for first-pass variants. Keep the output brief deliberately narrow: one composition, one aspect ratio, a fixed product count, and one acceptance checklist. Promote the best candidates to GPT Image 2 2K or Nano Banana Pro only when a higher-resolution final is actually needed. This two-stage design prevents paying premium final-asset prices for 90% of drafts that will be discarded.
4. Image-to-image edits and brand consistency
Do not evaluate text-to-image and edit workflows with one score. Test 5, 10 and 20 reference-image edits separately. Measure retained object identity, unwanted global changes, text survival, background replacement quality and the number of repair turns. GPT Image 2 and Nano Banana Pro both expose image-input workflows, but provider-native capabilities and gateway parameter parity must be tested before you assume multi-turn behavior transfers.
5. Production launch controls
Keep an allowlist of model IDs, fixed dimension/aspect-ratio presets, an image-content policy, a per-job budget cap and an approval condition for public-facing assets. Log model, route, dimensions, prompt class, reference count, request ID, price, latency, retry count and final acceptance. The application should own the acceptance label; a gateway usage log cannot know whether a designer or customer accepted the picture.
Budget examples: use the right model at the right pipeline stage
| Monthly workload | Recommended routing starting point | LumeAPI generation budget before retries | Why |
|---|---|---|---|
| 10,000 concept drafts | Qwen Image 2.0 | ~$300 | Lowest listed card price among the reviewed live routes; quality gate required |
| 10,000 photo-real ecommerce attempts | Seedream 5 | ~$350 | The largest current catalog-reference reduction; compare against GPT Image 2 on acceptance rate |
| 10,000 premium 2K final attempts | GPT Image 2 2K | ~$800 | $900 below the stated $1,700 native-reference baseline |
| 1,000 brand-critical localized assets | Nano Banana Pro 1K/2K | ~$134 plus input/grounding where used | Price parity, so select for capability rather than a claimed discount |
| Two-stage: 10,000 drafts + 1,000 final upgrades | Qwen Image 2.0 then GPT Image 2 2K | ~$380 | Keeps premium output for the selected 10% rather than every trial |
A budget estimate should use attempts, not expected final deliverables. If your historical acceptance rate is 40% and you need 1,000 approved assets, plan approximately 2,500 generations plus a retry contingency. A lower card price can support more exploration; it does not make the production asset free.
What independent creative testing adds
See Contra Labs' four-model creative tournament charts and methodology
Figure 2. Independent four-model tournament source. Contra Labs used ten briefs, four models, blind round-robin rankings and ten creative professionals. This is a small study, so use it as directional evidence rather than a universal leaderboard.
A July 2026 Contra Labs study compared Seedream 5 Pro, ChatGPT Images 2.0, Nano Banana Pro and FLUX 2 across ten professional briefs. Across the tournament, ChatGPT Images led overall at a 35.9% first-place rate; Nano Banana Pro had 28.6% and Seedream 5 Pro had 28.0%. Seedream won photorealistic generation at 35.7% and was a close second in typography, but performed less well on stylized images and complex infographics.
This supports a useful routing hypothesis, not a blanket fact: route photo-real briefs to Seedream first, route infographics and exact layout demands to GPT Image 2 or Nano Banana Pro, then compare your own acceptance rate. The study did not test every model, used only ten briefs, and its "ChatGPT Images 2.0" name is not an API model-ID guarantee.
Which model should you start with?
| Your job | First model to test | Why | Required challenger |
|---|---|---|---|
| Marketing hero image or high-end final creative | GPT Image 2 2K | Strongest broad arena quality signal | Seedream 5 for cost; Nano Banana Pro for exact control |
| Product photo composite or natural lifestyle scene | Seedream 5 | Independent photorealism strength and low LumeAPI price | GPT Image 2 1K or 2K |
| Localized ad with readable language, brand rules, or world-aware elements | Nano Banana Pro | Google's explicit positioning for localization and consistency | GPT Image 2 |
| Thousands of rough concept variants | Qwen Image 2.0 or Gemini 3.1 Flash Image | Lower current catalog price | Seedream 5; measure acceptance rate |
| Diagram, polished graphic design, or utility art | Recraft V4.1 Utility Pro | Design-oriented specialist | GPT Image 2 and Nano Banana Pro |
| Existing image edit with a strict visual identity | GPT Image 2 and Nano Banana Pro A/B | Both expose image-input workflows and premium edit positioning | Seedream 5 if photo realism is the priority |
Avoid routing by prompt text alone. Add a request class such as photoreal_product, brand_localized, infographic, concept_draft, or image_edit, then save the selected model, requested dimensions, generation cost, latency, retry count and a human/automated acceptance label. This makes future routing evidence-based instead of a permanent guess.
A minimal multi-model evaluation plan
Create 20 to 50 briefs from real work, not generic "astronaut" prompts. Stratify them by the jobs above. For each model, use the same prompt, aspect ratio, reference images, seed/control values where available, and one pre-declared retry policy. Blind-rank outputs on: prompt satisfaction, text accuracy, reference fidelity, visual defects, brand compliance, and whether the asset can ship without material rework.
Calculate:
accepted-image rate = accepted outputs / all generated outputs
cost per accepted image = total generation cost / accepted outputs
revision burden = edit rounds + human correction minutesDo not turn a model on for user traffic only because it wins an internal visual vote. Test content moderation/refusals, image rights and watermark requirements, latency at your target concurrency, error retries, and your storage/deletion obligations. Synthetic imagery should never be presented as real documentary evidence.
Where LumeAPI fits
For this comparison, LumeAPI is useful when the product needs to evaluate models without maintaining separate wallets and request clients for every provider. The live catalog lists deployable routes for GPT Image 2, Nano Banana Pro, Seedream 5, Qwen Image 2.0 and Gemini 3.1 Flash Image. The same account, key and USD wallet can also cover supported text and video models.
The practical workflow is simple: keep a model allowlist, send the same approved briefs to two or three candidates through the image endpoint, and use usage logs plus your own acceptance field to compare cost and latency. Start from the multi-model API and the live model catalog. Never assume availability, price or a provider-native feature from this report; the catalog is the source of truth for the actual request you will send.
Final recommendation
Use GPT Image 2 as the quality benchmark and Nano Banana Pro as the control/brand-consistency benchmark. Add Seedream 5 as the photorealistic, low-cost challenger. For scale, let Qwen Image 2.0 or Gemini 3.1 Flash Image compete on accepted-image cost rather than on a global Elo alone. FLUX.2 and Recraft remain worthwhile specialist tests when style or design workflow matters more than a broad preference rank.
The best team setup is usually a small evaluated routing set, not a single permanent winner. One API entry point makes that experiment easier; the evidence that should decide it is your own acceptance rate, iteration burden and real delivered cost.
Sources and methodology
- Artificial Analysis Text-to-Image Leaderboard — live blind-preference Elo, confidence intervals, sample counts and normalized API-price references, checked July 25, 2026.
- OpenAI GPT Image 2 model documentation — model capabilities, image endpoints, modalities and rate-limit framework.
- Google Gemini image-generation documentation and Gemini API pricing — Nano Banana family, model IDs, SynthID disclosure and Pro price tiers.
- Contra Labs Seedream 5 Pro tournament — four-model, ten-brief, blind creative-professional comparison; included with its sample limitations.
- LumeAPI model catalog — current gateway model IDs, stated image request characteristics and catalog prices, checked July 25, 2026.
No billable provider request was made for this report. The original information gain is the cross-source decision matrix, normalized cost/acceptance methodology, current LumeAPI route audit, and explicit separation of independent, vendor and gateway evidence.