Every vendor page shows the best output. None publishes the miss rate, so buyers cannot plan capacity or budget. One published benchmark can: in July 2026 the strongest model tested passed a product-fidelity check on 29.0% of generations, meaning roughly two rejects for every keeper. Photoroom, a vendor, ran the benchmark on the hardest case, and it is still the most useful number available.
The relevant cost is per generation you can actually use.
This covers what the benchmark measured, what a reject rate does to your real cost, which failures recur, and how to build a process that expects them.
Key Takeaways
- The strongest model tested passed 29.0% of product-fidelity checks. The benchmark covered 850 products and 4,250 virtual-model generations across four models, published by Photoroom on 6 July 2026 (photoroom.com, accessed September 2026). It is vendor-run, and that matters.
- Pass rates clustered. 29.0%, 28.2% and 27.2% for the top three, with a fourth at 16.8%. That points to a limit of the whole category, and one weak model does not explain it.
- The vendor’s own fidelity layer raised the pass rate to 38.2%. Better, and still under two in five.
- Logo and text distortion was the single largest failure at 20.1% of generations. Then missing or changed elements at 12.5%, pattern changes at 11.4%, colour shifts at 8.1%.
- Reject rate multiplies your cost per usable asset. At a 29% pass rate, a $0.101 image is roughly $0.35 per keeper if rejects are charged.
- The benchmark tested the hardest case. Putting a real garment on a generated model at 2K is far harder than changing a background behind a product.
- Ask whether rejected generations are billed. That single term decides whether the reject rate costs you money or only time.
What did the benchmark actually measure?
A specific and demanding job: whether a product survived being placed on a virtual model.
The methodology, as published: 850 products spanning clothing, footwear, bags, jewellery and accessories, run through four image-editing models at 2K resolution, producing 4,250 virtual-model generations. Every generation was reviewed by at least three of ten trained annotators against a single question, whether it was still the same product. Published 6 July 2026 (photoroom.com, accessed September 2026).
Results by model: Nano Banana 2 at 29.0%, Nano Banana Pro at 28.2%, GPT Image 2 Medium at 27.2%, and FLUX.2 Klein 9B at 16.8%. Adding the publisher’s own fidelity layer to the top model raised it to 38.2%.
Photoroom, which ran this benchmark, sells a product-fidelity layer, and the result supports buying one. That alone is no reason to dismiss it. The methodology is disclosed, the sample is large, the annotation process is described, and no independent benchmark of comparable scale is public. It is the best number available and it should be read with its authorship in view.
The second caveat matters more for planning. This is the hardest case in the category. Placing a real garment with real logos and real patterns onto a generated human, at 2K, and requiring the garment to remain recognisably itself, is far harder than replacing a background behind a bottle. Do not read 29% as the pass rate for all AI imagery. Read it as the pass rate for the demanding end.
What does a reject rate do to your cost?
It divides your effective yield, and if rejects are billed it multiplies your price.
At a 29% pass rate you need roughly 3.4 generations per usable asset. The benchmark ran at 2K. Google lists Nano Banana 2 at $0.101 per 2K image and Nano Banana Pro at $0.134 (Gemini API pricing, accessed September 2026). At a 29% pass rate that is about $0.35 and $0.46 per keeper. Still cheap, and higher than the number on the pricing page.
For a catalogue the arithmetic compounds. A 200-image set at a 29% pass rate is about 680 generations. Each model costs a different amount, so the model you pick sets the ceiling. On DesignerBox the cost is shown before the run, so you can check a set against the plan allocation first: 500 credits a month on Basic, 1,000 on Pro, 2,500 on Premium. Pick the model before you plan the budget, and plan for the rejects too. The model list is on the models page.
Two questions follow, and they are the ones to put to any vendor:
- Are rejected generations charged? If yes, your cost per usable asset is the list price divided by your pass rate. If no, the reject rate costs time rather than money.
- What is the pass rate on my hardest product? Not on their demo. Reflective surfaces, fine patterns, small text and logos are where the failures concentrate.
The conversion arithmetic for any credit-based plan is in what an AI credit buys.
Which failures actually recur?
The published failure breakdown is the most useful part of the benchmark, because it tells you what to check rather than that something went wrong.
| Failure mode | Share of generations |
|---|---|
| Logo and text distortion | 20.1% |
| Missing or changed elements | 12.5% |
| Pattern and design changes | 11.4% |
| Colour shifts | 8.1% |
| Invalid virtual-model generations | 6.7% |
Logo and text distortion at 20.1% is the headline. Model vendors name text as a limit themselves: Google says Nano Banana Pro can still struggle with accurate spelling (deepmind.google, accessed September 2026), and OpenAI says GPT Image models can still struggle with precise text placement and clarity (OpenAI image generation guide, accessed September 2026). A garment with a printed logo carries text into every frame. If your products have printed branding, expect this to be your dominant failure and check it first.
Pattern changes at 11.4% matter for anyone selling prints, stripes or checks. A pattern that shifts scale or repeat is not the same garment, and it is the failure customers notice when the item arrives.
Colour shifts at 8.1% are the quietest and among the most damaging commercially, because a navy that renders black does not look wrong in isolation. It looks wrong next to the product.
These give you a review order rather than a general instruction to look carefully. Check text first, then missing elements, then pattern, then colour against the original.
How do you build a process that expects this?
Plan for the reject rate rather than hoping for a better one.
- Test your hardest product before you commit. Not the easiest. Run twenty generations on the item with the most printed branding and the finest pattern, and count what passes. That number is your planning rate, and it will differ from any published benchmark.
- Budget generations, not images. If you need 200 usable assets and your measured pass rate is 30%, plan for around 670 generations. Planning for 200 guarantees running out.
- Reject mechanically before judging aesthetically. Check the spec first: is the logo intact, is the pattern the same, is the colour right. Only then look at whether the image is good. Mixing the two makes the review slow and inconsistent.
- Fix the input before adding attempts. A better source photograph, a tighter reference and a clearer constraint move the pass rate more than regenerating does. Repeating a bad brief 20 times produces 20 rejects.
- Re-run the failed image on its own. When one result fails, regenerate that one. Re-running a whole set to fix one item is where the cost actually escalates.
- Use the easier job where you can. Changing a background behind a real product photograph is a materially easier task than generating a garment onto a model. Where the easier operation gets you the asset, use it.
Review and cost before the run
Using the easier job where you can is also a design decision in DesignerBox. A product photography workflow starts from the product photograph you add, so the product in the frame starts as your actual product. It does not remove the failures above. On-model work is the same hard case the benchmark measured. Virtual try-on and AI video start at the Premium plan, $75 a month billed monthly.
We keep our own account of what AI creative cannot do well for the same reason. Nobody in this category should claim a solved fidelity problem. The failure modes are measured, known and checkable, and a review process is required at volume. The mechanics of that review are in reviewing a bulk generation run, and the per-image checks are in AI product photo accuracy.
Measure the rate once, then reuse it. When you know the model and the review order that hold up for your hardest product, save them as a workflow. A saved workflow runs the same way on the next product. Three critic steps score the results of a run, and best-of-N keeps the best one. You still check the logo, the pattern and the colour yourself. The cost is shown before the run, so you know what the next product costs before you start. Batch is coming: it will run one workflow over a whole sheet of products. Start from a template and test it on your hardest product first.
FAQ
What is the pass rate for AI product image generation?
In a vendor-run benchmark published 6 July 2026, the strongest of four models tested passed a product-fidelity check on 29.0% of generations, with the others at 28.2%, 27.2% and 16.8%. It covered 850 products and 4,250 virtual-model generations at 2K. Photoroom ran it, and Photoroom sells a fidelity layer, which raised the top result to 38.2%.
How many AI generations do I need for one usable image?
At a 29% pass rate, roughly 3.4. That applies to the demanding case the vendor-run benchmark measured, which is placing a real garment on a generated model. Easier operations such as background replacement on a real product photograph pass more often. Measure your own rate on your hardest product rather than assuming a published figure.
Why do AI models distort logos and text on products?
Google and OpenAI both name text as a limit of their own image models, and it was the largest single failure in the vendor-run benchmark at 20.1% of generations. Any product with printed branding carries text into every frame, so brands with printed logos should expect this to be their dominant failure mode.
Does the reject rate affect what I pay?
Only if rejected generations are billed, which is the term to check. Where they are, your real cost is the list price divided by your pass rate: a $0.101 image at a 29% pass rate is about $0.35 per usable asset. Where rejects are not charged, the reject rate costs review time instead.
Can I trust a vendor-run benchmark?
Read it with its authorship in view rather than dismissing it. This one discloses its methodology, sample size and annotation process, and no independent benchmark of comparable scale is public. The result also supports the publisher selling a fidelity product, so treat the headline as directional and measure your own products.
How do I reduce the number of retries?
Improve the input before adding attempts. A better source photograph, a tighter reference image and explicit constraints move the pass rate more than regenerating does. Also choose the easier operation where it achieves the same asset, since editing a real product photograph is a materially easier task than generating the product from scratch.
Which products fail most often?
Those with printed logos or text, fine or repeating patterns, and colours that are hard to reproduce exactly. The published failure breakdown puts logo and text distortion at 20.1% of generations, missing or changed elements at 12.5%, pattern changes at 11.4% and colour shifts at 8.1%.
Sources
- Photoroom Product Fidelity Benchmark, published 6 July 2026, vendor-run (photoroom.com, accessed September 2026)
- Nano Banana 2 and Nano Banana Pro list prices per image: Google, Gemini API pricing, accessed September 2026
- Nano Banana Pro stated limits on spelling: Google DeepMind, Gemini image Pro, accessed September 2026
- GPT Image text placement limits: OpenAI, image generation guide, accessed September 2026
- DesignerBox plan allocations and the Premium gate for virtual try-on and AI video: DesignerBox pricing page (designerbox.ai/pricing), September 2026
Benchmark figures verified from the publisher’s own methodology page as of September 2026. The benchmark is vendor-run and measures virtual-model generation, the most demanding case in the category. Pass rates on easier operations differ. Individual results vary.