AI models for product photos are generated people who wear, hold or use your product, built from the product photo you already own. They exist to solve one problem: a shopper cannot judge size from a floating object on white. Baymard’s research names apparel, accessories such as bags, jewelry and watches, and cosmetics as the categories that need a human for scale.
Most guides on this topic are tool walkthroughs. They open with the assumption that you have already decided to put a person in the shot, then show you which button to press. That skips the only decision that matters, and it is the one that costs money when you get it wrong.
Two questions come before the tool. Which of your products actually need a person, and which part of a person do they need. Answer those and the brief writes itself. Skip them and you generate a full fashion model for a candle that needed a hand.
Key Takeaways
- The person is usually a hand, not a model. Full-body models are an apparel requirement. For skincare, drinkware, jewelry and small tech, a hand or a partial body answers the scale question at a fraction of the brief.
- Baymard names four categories that require a human. Apparel, accessories such as bags, jewelry and watches, and cosmetics. Everything else needs an in-scale reference, which a person is only one way to provide.
- 42% of shoppers try to judge size from images, and 28% of sites give them nothing to judge against (baymard.com, accessed August 2026).
- The platform decides whether the person is allowed in the first slot. Amazon requires a model on adult clothing main images and forbids one on shoes. The rule flips by category, not by taste.
- A hired hand model starts around $79 an hour at one production studio, before manicure and usage add-ons (soona.co/pro-services/hand-model, August 2026). That is the number an AI hand is competing with, not a studio day rate.
- A generated person in a catalog photo is not an endorsement. It becomes one when the person appears to give an opinion, which is where the FTC guides bite.
What are AI models for product photos?
An AI model for product photos is a generated person rendered wearing, holding or using your product, produced from your own product image rather than a text description. The output is a still that fills a gallery slot: a hand presenting a serum bottle, a wrist wearing a watch, a body carrying a bag. The product stays anchored to the real photograph. Only the person around it is generated.
That anchoring is the whole point. A model generated from a prompt alone produces a plausible bottle that is not your bottle, with a label that is not your label. A model generated from your photo produces your bottle, held. The difference decides whether the image can ship to a product page or only to a mood board.
Which products actually need a person in the frame
Baymard’s product page research is specific about this, and it is narrower than the category assumes. Products designed to be worn require the context of a human model to give the truest sense of the product: apparel, accessories such as bags, jewelry or watches, and cosmetics (baymard.com, accessed August 2026). For cosmetics the guidance goes further, asking that shades be viewable on multiple models of diverse skin tones rather than shown only as swatches.
The scale problem underneath it is measured. 42% of users attempt to gauge the overall scale and size of a product from its images, and 28% of the sites benchmarked provide no in-scale image at all (baymard.com, accessed August 2026). 56% of test subjects began exploring product images as their first action on arriving at a product page.
Everything outside those four categories needs an in-scale reference, and a person is only one way to supply it. A stand mixer reads correctly beside a kitchen counter. A candle reads correctly beside a book. Reaching for a generated model when a known object would do is how a simple brief turns expensive.
| Product type | What the shopper cannot judge | The cheapest fix |
|---|---|---|
| Apparel | Fit, drape, length | Full body, standing |
| Bags, watches, jewelry | Size against a body | Partial body: wrist, neck, shoulder |
| Skincare, supplements, drinkware | Bottle or jar size, how it is held | A hand |
| Cosmetics | Shade on real skin | Multiple hands or faces, varied tones |
| Homeware, tools, electronics | Size in a room | A known object, no person needed |
The four levels of person, and what each one replaces
Treat the person as a ladder rather than a switch. Each rung answers a different question and carries a different production cost in the physical world, which is the cost an AI version is actually competing against.
Level one is a hand. It answers “how big is this and how do I hold it”. It is the correct level for most non-apparel products, and it is the level almost every tutorial skips straight past. A booked hand model at one production studio starts at $79 for up to an hour, covering bare hands and arms to the elbow with clean unpolished nails, with press-on or custom manicures from $39 on top (soona.co/pro-services/hand-model, August 2026).
Level two is a partial body. A wrist, a neck, a shoulder, an ear. It answers scale for worn accessories without committing to a face, which keeps casting, likeness and disclosure simpler. Watches and jewelry live here. The watch product photography shot list and the jewelry product photography one both turn on this rung.
Level three is a full body without a held pose. The garment or bag is worn, the person is present, the product is still the subject. This is the level Amazon requires on adult clothing.
Level four is a full body in a scene. Now you are producing lifestyle and ad creative, not catalog. The person carries mood, and the product competes with the room for attention. Useful for paid social, wasteful for a listing image.
Pick the lowest rung that answers the shopper’s question. Every rung above it adds a casting decision, a consistency problem across the range, and a disclosure obligation you did not need.
Where the shot goes, and where the platform bans it
Producing the image is the easy half. The gallery slot it can legally occupy is set by the marketplace, and the rule inverts between categories on the same platform.
Amazon requires that main images for clothing in adult sizes are shot on a model, and that the model is standing rather than sitting, kneeling, leaning or lying down. Children’s clothing is the opposite: shown flat, with no model. Footwear main images show a single shoe facing left at a 45 degree angle, so no model appears at all. Across every category, main images must not show any part of a mannequin, whatever its finish (sellercentral.amazon.com, product image requirements, retrieved August 2026).
So the same generated person is mandatory in one category, banned in another, and irrelevant in a third. Decide the slot before the brief. A model shot produced for a listing that forbids models is a wasted generation, and the same trap runs through handbag product photography, where the platform wants the bag alone and the shopper wants a body.
The practical order is: name the slot, check the category rule, then pick the rung from the ladder above. Not the reverse.
How to brief an AI model so the product survives
The failure mode is not an unconvincing person. It is a convincing person holding a product that no longer matches what ships.
Start from the product photograph, never a description of it. The generation should treat your image as the fixed element and the person as the variable. If the tool asks you to describe the product in words, the output will be a lookalike.
Read the product first when the image comes back, not the face. Check the label text, the closure, the finish and the proportions against the real item. A hand shot fails on grip before it fails on skin: fingers wrapped around a label that should be readable, or a jar held at a size that contradicts the dimensions in your listing copy.
Fix the shot list rather than the image. When one output is wrong, the brief is usually wrong for the whole batch. Regenerating a single frame hides a problem that will reappear across the range.
Hold the person consistent when the range needs it. A skincare line photographed across six SKUs with six different hands reads as six different brands. The same anchoring problem in fashion is covered in how to create an AI fashion model, and the mechanics transfer directly to hands.
What this costs on DesignerBox
DesignerBox generates the person from your uploaded product photo, so the product in frame is the product you sell. Model Studio covers the worn and held shots, and the model catalog behind it is the same 13 image and video models on one bill.
Plans run Free at 112 credits, Basic at $15 a month for 500, Pro at $35 for 1,000, Premium at $75 for 2,500, and Ultra at $200 for 8,000. Credits per generation vary by model, so check the current rate on the pricing page before you budget a catalog.
Two gates matter for this work. Importing your own photos and using custom prompts start at Pro, and so does the commercial license, which is the tier most catalog work needs. Try-on for clothing sits higher, at Premium. For non-apparel products the try-on gate is not the one you hit, because a hand holding a bottle is an image generation job rather than a garment fitting job. That distinction saves $40 a month for brands that assumed otherwise.
If you are choosing between the models rather than the plans, the model comparison for product photography runs one brief through each of them.
Disclosure, and what changed on 2 August 2026
Two separate obligations apply, and they are often confused.
The EU AI Act’s Article 50 transparency rules apply from 2 August 2026. Providers of systems generating synthetic image content must mark outputs in a machine-readable format so they are detectable as artificially generated (artificialintelligenceact.eu, accessed August 2026). The AI Omnibus provisional agreement of May 2026 gives generative systems already on the market until 2 December 2026 to meet the machine-readable marking requirement. That obligation sits with the model provider. It does not discharge whatever your marketplace asks of you separately.
The advertising question is different, and narrower than most brands assume. The FTC Endorsement Guides define an endorsement as an advertising message that consumers are likely to believe reflects the opinions or beliefs of someone other than the sponsoring advertiser (ftc.gov, accessed August 2026). A generated hand holding your serum expresses no opinion, so it is product photography. The same generated person presented as a customer describing a result is an endorsement by a person who does not exist, and the guides are direct that content from non-existent people is deceptive.
The line is whether the synthetic person appears to be vouching. Staying on the photography side of it is a briefing choice, made before the generation, not a caption added afterwards. Platform-level labelling rules are covered separately in labeling AI-generated fashion images.
FAQ
Do I need an AI model for every product photo?
No. Baymard’s research names apparel, accessories such as bags, jewelry and watches, and cosmetics as the categories requiring a human model. Other products need an in-scale reference, which a known object can provide as well as a person. Adding a model to a homeware or electronics listing adds production cost without answering a question the shopper was asking.
Can I use an AI hand instead of a full model?
Yes, and for most non-apparel products a hand is the correct choice. It answers size and handling, which is what the shopper is checking, without a face, a casting decision or a likeness question. Hired hand models start around $79 an hour at one production studio, so the comparison is a modest line item rather than a studio day.
Will marketplaces accept product photos with an AI model?
It depends on the slot and the category. Amazon requires adult clothing main images to be on a standing model, shows children’s clothing flat with no model, and shows footwear as a single shoe with no model at all. Check the category rule for the specific image slot before producing the shot, because the requirement inverts between categories.
How do I keep the same model across a whole product range?
Anchor every image to one reference frame rather than regenerating the person each time. This applies to hands as much as faces. Six SKUs shot with six different hands reads as six different brands, which undoes the reason for using a consistent model in the first place.
Do AI models for product photos need a disclosure label?
The EU AI Act requires the model provider to mark outputs in a machine-readable format, applying from 2 August 2026. That is separate from any visible label your marketplace requires and separate again from advertising law. If the generated person appears to vouch for the product, US endorsement rules apply on top.
Does an AI model in a catalog photo count as an endorsement?
Not on its own. The FTC defines an endorsement as an advertising message consumers would believe reflects the opinions of someone other than the advertiser. A person simply wearing or holding a product conveys no opinion. Framing the same person as a customer sharing a result crosses into endorsement territory.
What does it cost to produce these on DesignerBox?
Plans start free with 112 credits and run to $200 a month for 8,000. Importing your own product photos, custom prompts and the commercial license start at Pro, which is $35 a month. Credit cost per generation varies by model, so check the current rate before budgeting a full catalog.
Sources
- Baymard Institute, product categories requiring a human model: https://baymard.com/blog/human-model
- Baymard Institute, in-scale product images research: https://baymard.com/blog/in-scale-product-images
- FTC, The FTC’s Endorsement Guides: What People Are Asking: https://www.ftc.gov/business-guidance/resources/ftcs-endorsement-guides-what-people-are-asking
- EU Artificial Intelligence Act, Article 50 transparency obligations: https://artificialintelligenceact.eu/article/50/
- European Commission, transparency obligations under Article 50: https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act
- Amazon Seller Central, product image requirements (retrieved August 2026)
- soona, hand model pro service pricing (retrieved August 2026)
Platform rules, model pricing and disclosure requirements verified from the sources above as of August 2026. Marketplace image requirements change by category and region, so confirm the rule for your specific listing slot before producing a batch. Individual results vary.