Skip to main content
Get started free

AI Models vs Real Models: How Fashion Brands Decide

Do AI models beat real models on conversion? What the most-cited study actually compared, and the shot-by-shot rule that decides it for your catalogue.

AI Models vs Real Models: How Fashion Brands Decide

No published test answers the AI-or-human question, and the numbers usually quoted for it measure something else. The most-cited study compared AI-generated garment designs against human-designed garments, not AI photography against studio photography. So decide on what you can actually observe: which gallery slot the shot fills, whether it matches the garment you ship, and who carries the disclosure risk.

That reframing matters because the question is usually asked at the wrong level. Brands ask “should we switch to AI models,” decide once, and apply the answer to a whole catalogue. The gallery does not work that way. A packshot, a fit shot, and a campaign hero each have to prove a different thing, and only one of them is really about a person. The order they move in is set by how much your existing images constrain each slot.

This covers what the research measured, why the published lifts cannot settle the argument, the slot-by-slot rule that replaces them, and what changed in EU law four days ago.

Key Takeaways

The 13% figure is not about photography. It comes from an Alibaba deployment that compared AI-generated garment designs against human-designed garments (arxiv.org/abs/2503.22182, KDD 2026 ADS track, accessed August 2026). The AI arm contained different products.

Triple-digit case-study lifts are a sample-size artefact. At a 2.5% baseline, detecting a 10% relative lift needs roughly 64,000 sessions per variant. Tests that report a 128% win are usually reading noise.

Decide per gallery slot, not per catalogue. The packshot, the scale shot, the fit shot, and the campaign hero each answer a different question. Only two of them depend on a person at all.

Platforms regulate accuracy, not provenance. Zalando permits AI-generated partner imagery and sets a quality floor with no disclosure rule (partner.zalando.com, updated May and June 2026). Walmart requires AI images to be “truthful, accurate and not misleading” (marketplacelearn.walmart.com, April 2026).

The disclosure risk is concentrated, not spread. EU AI Act Article 50 obligations applied from 2 August 2026. The marking duty falls on the provider of the generation system; the deepfake disclosure duty falls on you, and it turns on human likeness, not on product stills.

Fit is where AI on-model imagery costs you money. Size and fit drive the majority of apparel returns, and a flattering drape that the garment does not have converts once and comes back.

Produce both arms from the same source photo. If the control is last year’s studio shot and the variant is a fresh render, you compared two eras of your brand, not two treatments.

Do AI models convert better than real models?

There is no reliable general answer, and the effect on your catalogue is unmeasured until you measure it. The studies and case studies circulating in this category either compared a different variable or ran at sample sizes too small to detect the effect sizes that image treatments actually produce. Treat AI on-model imagery as a cost, coverage, and speed decision with an unknown conversion effect, rather than as a conversion tactic with a known return.

That is a less satisfying answer than the one most vendor pages give. It is also the only one the evidence supports.

What the most-cited study actually compared

One statistic anchors nearly every article on this topic: AI visuals delivering roughly 13% relative improvements to click-through and conversion, alongside a 7.9% drop in returns. The research is real and peer reviewed. It is routinely described as proof that AI fashion photography beats studio photography.

It is not that.

The paper is “Sell It Before You Make It: Revolutionizing E-Commerce with Personalized AI-Generated Items”, deployed at Alibaba and accepted to the KDD 2026 ADS track (arxiv.org/abs/2503.22182, accessed August 2026). Merchants described a garment in text. The system generated the design and a photorealistic image of it on a digital model, and the listing went live before the garment existed. Manufacturing started once orders arrived.

The comparison was AI-generated garment designs against human-designed garments. Different products, selected by a preference model trained on shopper behaviour. The photography was downstream of the variable that moved.

Quote that number to justify replacing your packshots and you have imported an effect size from an experiment that did not run your test. Your expected lift is unknown, and sizing a test around a borrowed 13% will size it wrong.

Why the published lifts cannot settle it

The other evidence on offer is vendor case studies, and they tend to report lifts of 128%, 150%, or 30%. Those numbers are not dishonest. They are unreadable, and the reason is arithmetic rather than motive.

Sample size comes from your baseline conversion rate and the smallest lift worth detecting. At a 2.5% baseline, detecting a 10% relative lift needs roughly 64,000 sessions per variant. A 30% lift needs around 7,800. Most single-brand case studies run on a fraction of that.

An underpowered test cannot return a small wrong answer. It returns a large one. A test with a few thousand sessions can only ever surface enormous swings, so the wins that get written up are precisely the ones least likely to survive a rerun. The full session counts each level of lift requires sit in the traffic math, and they are the first thing to check before you believe any published number in this category, including your own.

Image treatment effects on an established catalogue usually land in the single digits to low teens. That is the range hardest to measure and the range where the decision actually matters.

Once the conversion argument is set aside, the decision becomes tractable, because the gallery is not one thing. Each position answers a different shopper question, and the questions have different tolerance for a synthetic person.

Gallery slotWhat it must proveDepends on a person?Sensible default
Hero packshotThis is the item, accuratelyNoAI from your product photo
Scale referenceHow big it isWeaklyAI from your product photo
Detail cropFabric, stitching, hardwareNoPhotography of the real garment
On-model fitHow it hangs on a bodyYes, heavilyWhichever renders the garment honestly
Campaign heroWhat the brand meansYes, heavilyReal shoot
Variant coverageWhich colourway to pickNoAI from your product photo

Read the “depends on a person” column and the argument resolves itself. Four of the six slots barely involve a model. Those are the slots where AI generation is uncontroversial, cheap, and fast, and where most catalogues have gaps today. The two that do involve a person are where the real questions live, and they pull in opposite directions.

The practical consequence: most brands should not be choosing between AI and real models. They should be filling four slots with AI, shooting one with a camera, and thinking hard about the fifth.

Where AI on-model imagery is the stronger choice

Coverage you do not have at all. A catalogue missing the in-scale frame, the detail crop, and half its colourways loses more to those gaps than it will ever gain from optimising a treatment it already ships. Coverage beats optimisation when the gaps are this obvious.

Long-tail SKUs that will never justify a shoot. The economics of a studio day do not stretch to a 400-SKU tail. What a product photoshoot actually costs sets the baseline that decides where the line falls for your catalogue.

Colourway and variant permutations. The garment is identical and only the fabric colour changes. Reshooting that is waste.

Demand testing before manufacture. This is the one place the Alibaba result genuinely applies, because it is the thing that study measured.

Representation range. Showing a garment across more body types than a casting budget covers is a real gain, provided the drape stays honest for each one.

Where a real shoot still wins

Anything that carries brand meaning. A campaign hero is an argument about who the brand is. That is the slot with the least tolerance for approximation and the one where an audience is most primed to notice.

Fit fidelity on complex garments. Tailoring, drape, knitwear, and anything with structure are where generated on-model imagery is most likely to flatter. Fit is not a cosmetic problem. Size and fit drive the majority of apparel returns, so an image that improves the click and misrepresents the hang converts once and comes back. Run the accuracy checks before you ship on your hardest garment, not your easiest.

Texture that has to be legible. Where the purchase decision is the material itself, photograph the material.

Categories where the body is the product. Swimwear, lingerie, and activewear are judged on fit against a real body more than on styling.

The honest version of the tradeoff: AI on-model imagery is strongest where the garment is simple and the shot is functional, and weakest where the garment is structured and the shot is emotional. The consistency locks that hold a model steady from drop to drop matter more as you push further into the second category.

What platforms and regulators actually require

The rules in force do not ask whether pixels came from a camera. They ask whether the image matches what ships.

Zalando permits AI-generated partner imagery and constrains it on quality alone. Its image guidelines reject a shot that “is poor quality AI generation” and its video guidelines state that “low-quality AI video footage characterised by pixelation, motion blur, unnatural textures or anatomical inconsistencies is not permitted” (partner.zalando.com, updated May and June 2026). There is no disclosure requirement, and Zalando’s own Platform Rules v13 contain no AI clause at all. The company builds digital twins of real models itself (corporate.zalando.com, May 2025).

Walmart requires that “images generated by artificial intelligence must be truthful, accurate and not misleading” and does not permit stock photos in place of actual product images (marketplacelearn.walmart.com, April 2026). Etsy requires disclosure when the item itself is AI-created, and separately requires listing photos to be your own originals rather than renderings (etsy.com/legal/sellers, accessed August 2026).

The pattern is consistent. Fidelity to the product is the binding rule; provenance mostly is not. Full per-marketplace detail sits in what each marketplace requires for AI product images.

EU AI Act Article 50 transparency obligations applied from 2 August 2026, and the Commission’s final guidelines, adopted 20 July 2026, put photorealistic AI-generated people in scope even when no real individual is depicted (digital-strategy.ec.europa.eu, accessed August 2026). The machine-readable marking duty under Article 50(2) sits with the provider of the generation system. The deepfake disclosure duty sits with you as the deployer.

So the exposure is uneven, and it concentrates in one slot: the synthetic human. A packshot or a flat lay of your product triggers none of it. A photorealistic person wearing your garment needs a visible disclosure. Which assets in a drop trigger a label, and what a label costs, is covered in labeling AI-generated fashion images.

That is a real cost line on the on-model slot specifically, and it belongs in the comparison alongside the credit price.

Verify your own position against the regulation and your marketplace’s current policy before scaling a treatment across a catalogue. Do not settle a compliance question on a blog’s reading of it, this one included.

Produce both arms from the same product photo

Whatever you conclude, the test is only as clean as the assets. If your control is a two-year-old studio shot and your variant is a fresh render with a different crop and different lighting, you compared two eras of your brand.

Generate both arms from the same source photo of the real garment and change exactly one attribute. That is the practical argument for producing variants from your own product image rather than commissioning two separate shoots: the source is held constant by construction.

In DesignerBox that runs through Model Studio, where one product photo becomes the packshot, the styled scene, and the on-model shot from the same input. An image costs 5 credits, so a six-slot gallery is 30 credits per SKU. The free plan’s 112 credits cover 22 images, enough to build a first comparison set before committing. Virtual try-on and on-model generation sit on the Premium plan at $75 a month, which is the honest gate to know about before planning around them. Plan detail is on the pricing page.

Two further things worth knowing. A reusable fashion model persona is what keeps the same face and body across a drop instead of a new stranger per SKU, and where an AI fashion model actually comes from decides what you owe on disclosure. Model choice is not cosmetic either: the comparison of every image model on one product brief shows how differently they handle fabric on a body.

For the fit question specifically, what the research shows on virtual try-on accuracy covers what fit-aware generation changes and what it does not.

FAQ

Do AI fashion models convert better than real models?

No published evidence settles it. The most-cited study compared AI-generated garment designs against human-designed garments rather than AI photography against studio photography, and single-brand case studies typically run below the sample size needed to detect realistic image effects. Treat it as a cost and coverage decision with an unmeasured conversion effect.

Can I use AI models for my whole catalogue?

You can, but the gallery-slot table above suggests you should not want to. Packshots, scale shots, detail crops, and variant coverage are the slots where generation is strongest. Campaign heroes and fit-critical categories are where a real shoot still earns its cost.

Do I have to disclose AI-generated model images?

In the EU, a photorealistic AI-generated person wearing your garment does need disclosure. Article 50 obligations applied from 2 August 2026 and the Commission’s final guidelines put photorealistic synthetic people in scope even where no real individual is depicted. Packshots and flat lays are outside that trigger. Check the regulation and your marketplace’s policy directly for your own market.

Will marketplaces reject AI-generated product images?

Not for being AI-generated, in the cases verified here. Zalando permits AI imagery and rejects it only on quality. Walmart requires AI images to be truthful and accurate. Etsy requires listing photos to be your own originals rather than renderings, which is a fidelity rule rather than a provenance one.

Do AI model images increase returns?

They can, if the render flatters the garment’s drape or fit. Size and fit drive the majority of apparel returns, so an image that lifts click-through while misrepresenting the hang can raise returns enough to erase the gain. Track return rate on tested SKUs for a full return window before calling any winner. Generous returns policies also mask the effect, which is why generating a size range has to be read against returns rather than conversion alone.

How many sessions do I need to test AI against real model images?

At a 2.5% baseline conversion rate, detecting a 10% relative lift needs roughly 64,000 sessions per variant and a 30% lift needs about 7,800. Below those counts you cannot separate a result from noise. Run the arithmetic on your own baseline before you design the test.

What is the cheapest way to compare the two treatments?

Generate both arms from the same source photo so only one attribute differs, then screen them at the ad layer where impressions are purchasable rather than on a product page where traffic is not. Take the surviving two candidates to the PDP only if your traffic supports it.

Sources

All accessed August 2026.

  • What the 13% click-through and conversion figures and the 7.9% return reduction actually compared: “Sell It Before You Make It: Revolutionizing E-Commerce with Personalized AI-Generated Items”, KDD 2026 ADS track (arxiv.org/abs/2503.22182)
  • Zalando partner image and video guidelines, including the AI video footage policy and the AI generation quality rejection criteria (partner.zalando.com, pages updated May 2026 and June 2026)
  • Zalando digital twins of real models (corporate.zalando.com, May 2025)
  • Walmart Marketplace requirement that AI-generated images be truthful, accurate and not misleading (marketplacelearn.walmart.com, April 2026)
  • Etsy seller policy on AI disclosure for items and original listing photography (etsy.com/legal/sellers)
  • EU AI Act Article 50 transparency obligations, the 2 August 2026 application date, and the Commission’s final guidelines adopted 20 July 2026 putting photorealistic AI-generated people in scope regardless of whether a real individual is depicted (digital-strategy.ec.europa.eu)
  • DesignerBox credit costs, plan allocations and feature gating verified against live product configuration, August 2026

Research claims verified against arXiv:2503.22182, Zalando partner guidelines, Walmart Marketplace policy, Etsy seller policy and European Commission AI Act transparency guidance as of August 2026. Sample sizes computed with the standard two-proportion formula at 95% confidence and 80% power. Regulatory position varies by market and is not legal advice. Individual results vary.

Cristian

Head of Content at DesignerBox

Cristian covers AI product photography, video ad tools and model comparisons. He runs the same prompt and the same product across models, then publishes the output side by side, so you pick on evidence instead of marketing copy.

Follow along on Instagram at @designerboxai for campaign breakdowns.

Show it worn, without a casting call

Put a garment on a model from one flat photo. Keep the same face and body across a whole drop, and get a lookbook without booking a studio day.

Start free

Upload one product photo. Ship the whole campaign, without a photoshoot.