You create an AI fashion model in two stages, and only the second stage is written down anywhere. Stage one produces a single frame of a person who does not exist. Stage two anchors every later image to that frame. Stage one is a casting decision, stage two is a copying job, and almost every guide starts at stage two.
That gap is expensive because the copying stage has a hard limit. Google documents up to 5 character-consistency reference images for its Pro image model and up to 4 for the Flash tier (ai.google.dev, August 2026). Four or five frames of the wrong face is still the wrong face. You cannot fix a weak first frame by sending more of it.
This guide covers where frame one comes from when you have no photo of anyone and cannot use a real person, the three routes available, and which one leaves you owning the face at the end.
Key Takeaways
- The first frame is the asset. Every later image is a copy of it. Casting happens once per brand, not once per shoot.
- The anchor budget is small. Google documents 5 character-consistency images on its Pro image model, 4 on Flash (ai.google.dev, August 2026). Volume does not rescue a bad frame.
- Three routes exist: pick from parameters, describe once then throw the words away, or take a face from a vendor’s fixed set.
- “No prompts” is a casting claim. It says the face was not typed into existence. It says nothing about whether you can reproduce it next season.
- Words generate frame one. Words never reproduce it. A written description samples a new person on every pass.
- Owning the frame is what makes the model reusable. A face you can only reach inside one vendor’s tool leaves with the subscription.
How do you create an AI fashion model?
Generate one image of a person who does not exist, then use that image as the reference for every following generation. The first image can come from parametric selection, from a written description you use once, or from a vendor’s prebuilt set. After that frame exists, the description stops mattering. The file does the work.
The step every consistency guide skips
Every guide on holding a model steady says the same correct thing: anchor identity to a reference image, because a written description regenerates a different person each pass. That advice is right, and this blog repeats it, including in the walkthrough on how to lock a model across a whole drop.
None of it answers the question a brand actually starts with. You have a garment, a shoot date, and no photograph of any person you are allowed to use. There is no reference image yet. Something has to produce the first one.
That first generation is casting. A casting director does not audition a new face for every look in a campaign, and neither should you. You choose once, then you shoot that person repeatedly.
Three routes to the first frame
| Route | What you supply | Same face tomorrow? | Who owns the face |
|---|---|---|---|
| Parametric casting | Menu selections: build, age range, look, scene | Yes, once you save the frame | You |
| Describe, then discard | One written brief, used a single time | Only after you save a frame | You |
| Vendor’s fixed set | A pick from that vendor’s catalogue | Yes, inside that vendor | The vendor |
Route one: parametric casting
You pick attributes from a fixed menu rather than writing a sentence. DesignerBox’s AI Character Creator takes a look, gender, ethnicity, age range and scene, then assembles a character from those choices (designerbox.ai, August 2026).
The advantage is that a menu has no ambiguity. “Late twenties” as a slider position means one thing. “A woman in her late twenties” as a sentence means thousands of things, and the model picks a different one each time. The limit is the menu itself: you get the range the menu covers.
Route two: describe once, then discard the words
Write the brief, generate a batch, pick the single strongest frame, then delete the brief from your process. The words were a casting call, not a specification.
This route reaches looks a menu does not cover. It also fails in a specific way that catches people: they keep the prompt, reuse it next month expecting the same person, and get a stranger. The prompt was never the anchor. The saved frame is.
Route three: a vendor’s fixed set
Some tools built only for apparel state that their models are “100% AI-generated” with “no stock models, no real people, no photo references”, and that generated photos carry no usage-rights fees for commercial use (botika.com, August 2026). That claim does real work. A figure matching no living person carries no right-of-publicity exposure and no talent agreement to renew.
The trade is scope rather than quality. The face lives inside that vendor’s product. If you move tools, or that catalogue changes, the model does not come with you.
Which route fits which brand
Pick on how long the face has to survive, not on output quality. All three routes produce publishable frames.
- One campaign, one season, no recurrence: any route works. Take the fastest.
- A face that recurs across drops: route one or two, and save the frame the moment you like it.
- A house model that outlives your current tool stack: route one or two only. Route three ties the asset to a subscription.
- A named brand character with its own following: route two, then build a full reference set around the frame you chose.
The anchor budget is why casting matters
The copying stage runs on a smaller allowance than most people assume. Google’s documentation separates reference images by job: up to 5 character-consistency images on Gemini 3 Pro Image, up to 4 on Gemini 3.1 Flash Image, with style references counted separately at up to 3 (ai.google.dev, August 2026). Black Forest Labs describes FLUX.1 Kontext as built to “precisely preserve identity, of e.g. a reference character or object, across multiple scenes and environments” and publishes no reference count for it at all. Its FLUX.2 family does document one, up to 8 references through the API, and names the job outright: “fashion editorials where models stay consistent” (docs.bfl.ai, August 2026).
Four or five slots is enough to cover front, both three-quarter angles and a closer crop. It is not enough to average away a frame that was wrong to begin with. That asymmetry is the whole argument for treating frame one as a decision rather than a first attempt.
Where the same face has to hold across hundreds of SKUs, the character LoRA training workflow trades the small reference budget for a trained identity. That starts at Premium, $75 a month.
What you owe once the face exists
A synthetic model removes one obligation and creates another. There is no consent form to collect for a person who never existed, and there is no likeness to license. Disclosure is the separate duty, and it applies from 2 August 2026 under Article 50 of the EU AI Act.
Inventing the person does not exempt you. The Commission’s guidelines state it is enough for a simulated person to resemble someone who “can plausibly exist” to count, and name realistic AI-generated human personas directly (ec.europa.eu, C(2026) 5054 final, July 2026). A photorealistic house model is in scope.
One detail catches brands out. The same guidelines state that deployers cannot rely on the machine-readable marking the tool embeds in the file, because those markings are not clear and distinguishable to the people who see the image. A provenance tag inside the file does not discharge your duty. The label has to be visible. How this lands per market is covered in where AI fashion models actually come from.
Keep the record either way. Save the first frame, the route that produced it, and the date. A provenance claim you cannot document is worth nothing at the moment somebody asks you to prove it.
Building the first frame on DesignerBox
Cast once, then run the garment work against the saved frame.
- Generate the character from the parametric picker, or from a written brief if the menu does not reach the look.
- Save the single best frame to your library. That file is now the model.
- Generate three or four more angles conditioned on that frame. That set is your anchor budget.
- Run garments against the set. Which catalogue model suits the look is covered on the fashion and editorial model page.
- Save the pass as a workflow so drop two matches drop one without anyone retyping anything.
Two gates worth knowing before you commit a season. The commercial licence starts at Pro, $35 a month. Try-on clothes, AI video and LoRA start at Premium, $75 a month. Current tiers are on the pricing page.
The garment side of this runs on its own input rules, and the photo you shoot sets the ceiling for what any model can do with it. That is covered in how to create fashion visuals with AI. Fashion Factory is where the two halves meet: one cast face, one garment photo, a full set out.
The reason to keep both halves in one workspace is that a face rebuilt from a different file by a different person in a different tool three months later is a different face. The real cost is not the subscriptions. It is the seams.
FAQ
How do you create an AI fashion model without a real person?
Generate the first frame from parametric selections or from a written brief, then use that saved image as the reference for everything after. No photograph of a real individual enters the process at any point, so there is no likeness to license and no release form to collect.
Do you need a prompt to create an AI fashion model?
Not necessarily. Parametric tools assemble a character from menu selections instead of a sentence. Where a prompt is used, treat it as a casting call rather than a specification: it produces candidate frames once, and the frame you keep becomes the anchor from then on.
How many reference images do you need to hold the same face?
Fewer than most people send. Google documents up to 5 character-consistency images on its Pro image model and up to 4 on Flash (ai.google.dev, August 2026). Four covering front, both three-quarter angles and a closer crop is a working set.
Can you reuse the same AI fashion model across seasons?
Yes, if you saved the frame. Reusing the prompt does not work, because a written description samples a different person on each pass. Version the saved frame the way you would version a logo file, and keep it somewhere the whole team reaches.
What happens if you lose the first frame?
You recast. There is no way to recover a specific generated face from the settings that produced it, which is why the file matters more than the process that made it. Back it up on the day you choose it.
Does an AI fashion model need a style reference too?
Only if the look has to match across shots as well as the face. A style reference is counted separately from character references, so it does not consume the identity budget.
Sources
- Character-consistency and reference-image limits by model tier, including up to 5 character images on Gemini 3 Pro Image and up to 4 on Gemini 3.1 Flash Image: ai.google.dev, August 2026
- FLUX.1 Kontext identity preservation across scenes and environments, with no published reference-image count, and the FLUX.2 family’s documented limit of up to 8 API references: docs.bfl.ai and docs.bfl.ai/flux_2, August 2026
- EU AI Act Article 50 transparency duties, the 2 August 2026 application date, the treatment of persons who “can plausibly exist”, and the rule that machine-readable marking does not discharge a deployer’s disclosure duty: Commission Guidelines C(2026) 5054 final, 20 July 2026, digital-strategy.ec.europa.eu
- Model provenance stated as “100% AI-generated” with “no stock models, no real people, no photo references”, and commercial photos free from usage-rights fees: botika.com, August 2026
- DesignerBox character creation inputs, pricing, credit costs and feature gating verified against live product configuration, August 2026
Model documentation verified from Google and Black Forest Labs as of August 2026, and regulatory positions from the European Commission as of the same date. This is not legal advice. Confirm your obligations for your own markets and placements. Individual results vary.