Skip to main content
Get started free

AI Character Reference Images: How to Build One That Holds

A character reference locks a face, not an outfit. What each model accepts, why more angles makes drift worse, and the reference spec that survives a campaign.

AI Character Reference Images: How to Build One That Holds

An AI character reference image is a photo you hand a model so it renders the same person every time instead of inventing a new one from your description. It works by collapsing the space the model can wander in. The strongest references are single-subject, front or three-quarter facing, evenly lit, neutral in expression, and shot against a plain background.

Most guides on this topic give you a list of habits. Upload something sharp, iterate, stay organised. All reasonable, none of it tells you the thing that decides the outcome: every model accepts a different number of reference images, calls the feature something different, and two of the six video models in the DesignerBox catalog will not take a photo of a real human face at all. A reference spec copied from one tool’s help page quietly stops working on the next model.

This covers what a reference actually carries, why the common advice to upload more angles is documented to make identity drift worse for people, what each model accepts according to its own docs as of July 2026, and how to build a reference that survives a campaign.

Key Takeaways

  • The reference caps are published and they are not the same. Veo 3.1 takes up to three subject images (cloud.google.com, July 2026). Nano Banana Pro takes up to five character images, Nano Banana 2 takes four (ai.google.dev, July 2026). Seedream 5 Pro takes up to ten references, FLUX 2 Flex up to eight (volcengine.com and docs.bfl.ml, July 2026).
  • More angles is the wrong instinct for faces. ByteDance’s own prompt guide advises against multi-view character sheets for people, because the model reads different angles as different subjects and that worsens ID drift (volcengine.com, July 2026).
  • Two video models will not accept a real face. Sora 2’s image reference rejects input images containing human faces, and its character uploads block human likeness by default (developers.openai.com, July 2026). Seedance 2.0 does not accept reference images containing real human faces without a verification and consent step (volcengine.com, July 2026).
  • Identity travels, lighting does not have to. Runway documents its image references holding a character across different lighting, locations and treatments from a single reference, which is why it asks for even, neutral light in the source (help.runwayml.com, July 2026).
  • Wardrobe is a per-model answer, not a rule. Black Forest Labs states FLUX.2 maintains face, clothing, proportions and style across generations (docs.bfl.ml, July 2026). Runway shows wardrobe varying inside a single generation and saves an output as a new reference to lock it. Do not assume either behaviour.
  • The vendors admit the ceiling. OpenAI’s image docs list consistency as a limitation, noting the model may struggle to hold recurring characters across multiple generations (developers.openai.com, July 2026).

What is an AI character reference image?

A character reference image is an input photo that conditions a generation on a specific person’s appearance, rather than on a text description of them.

The difference matters because words are a wide net. “Dark shoulder-length hair, brown eyes, mid-thirties” describes millions of people, and the model picks a new one each run. An image narrows that to one face and holds the generation near it.

Providers give the feature different names, and the name tells you how it behaves. Google calls it a subject image, passed as a reference with the asset type. OpenAI splits it into an image reference that conditions the opening frame and a separate character asset that persists across requests. Kling calls its version elements. ByteDance calls it multimodal reference. Runway calls them references and lets you address them in a prompt with an @ tag.

Those are not synonyms for one feature. Some condition a single generation. Others create a reusable asset. Knowing which one you have changes whether your character survives past the shot you are working on.

What a character reference carries, and what it does not

A reference reliably carries identity: the face, the bone structure, the features that make someone recognisable. That is the job it was built for and the thing every vendor claims.

Lighting is where it gets useful. Runway documents its image references generating consistent characters across different lighting conditions, locations and treatments from a single reference image (help.runwayml.com, July 2026). That is why the same docs ask for even, natural lighting and a neutral expression in the source. The reference is meant to be a blank canvas that your prompt then lights, not a mood you are stuck with.

Wardrobe is genuinely contested, and this is where most advice overreaches. Black Forest Labs states that FLUX.2 maintains a character’s identity including face, clothing, proportions and style across multiple generations (docs.bfl.ml, July 2026). Runway’s own worked example runs the other way, showing several different wardrobe variations returned inside a single generation, then saving one output as a new reference to pin the outfit down. Both are accurate about their own model.

The practical rule: specify wardrobe in the prompt every time, whatever the model claims. If it was going to hold anyway you have lost nothing. If it was not, you have caught it before the batch renders.

Camera framing is not documented by anyone as something a reference carries. Google’s Veo prompting examples still write the shot (“medium shot of the detective behind his desk”) even when subject images are supplied. Treat framing as always your job.

Why more angles can make identity drift worse

The most repeated tip in character-reference guides is to give the model several angles of your character. For a product or a prop, that is sound. For a human face, ByteDance’s Seedance prompt guide advises the opposite.

The guide recommends against using multi-view character sheets for people, on the basis that multi-view material contains the same person from different angles and the model is prone to reading those as several different subjects, which worsens the ID drift problem it was meant to solve (volcengine.com, July 2026).

Its recommendation for a face is narrower than most people expect: a head-only close-up, face retained, neutral expression best, with shoulders, neck and background minimised as interfering elements.

The same guide adds two things worth stealing regardless of which model you run. Do not fill the reference slots just because they exist, because too many assets make it harder for the model to judge which features take priority. And put the most precision-critical asset earliest in the prompt.

There is an honest caveat here. ByteDance’s own documentation ships a worked example that uses a multi-view reference of a woman, contradicting its written advice. Both are official. Read the guidance as a strong default rather than a law, and test your own set before you commit a campaign to it.

The broader point holds across models. The ceiling on a reference set is not how many images you send. It is whether they read as one unambiguous person. Two clean frames that agree beat six that argue.

Which models accept a character reference, and how many

Every number below comes from the provider’s own documentation as of July 2026. Where a provider publishes nothing, the table says so rather than guessing.

ModelTakes a character referenceDocumented maxProvider’s name for it
Veo 3.1Yes3 subject imagesReference images, asset type
Veo 3.1 FastYes3 subject imagesReference images, asset type
Sora 2 ProNot for real faces1 image referenceImage reference, characters
Seedance 2.0Not for real faces without verification1 to 9 references totalMultimodal reference
Kling 2.6 ProNot documented for this versionNot publishedElements, on other versions
Runway Gen-4.5Text and image to video onlyNot applicablePrompt image, first frame
Nano Banana ProYes5 character imagesCharacter consistency images
Nano Banana 2Yes4 character imagesCharacter consistency images
Seedream 5 ProYes10 referencesReference images, subject consistency
GPT Image 2YesNo maximum publishedImage references
FLUX 2 FlexYes8 via the APIMulti-reference editing
Kontext MultiYesNo maximum publishedCharacter consistency
FLUX Pro 1.1NoNot applicableImage prompt, variation only

Five things in that table will change how you plan a shoot.

Sora 2 Pro will not take a photo of a real person. OpenAI’s docs state that input images with faces of humans are currently rejected, and that character uploads depicting human likeness are blocked by default. Its character feature is also built from a short video clip, not a still, and OpenAI notes that passing the character ID alone is not enough to reliably preserve the character in the shot. You have to name the character in the prompt too (developers.openai.com, July 2026).

Seedance 2.0 has the same restriction with a documented path around it. Real human faces cannot be uploaded directly. The sanctioned routes are your own recent Seedance or Seedream outputs, assets from the preset virtual-portrait library, or real-person material that has cleared identity verification and consent capture (volcengine.com, July 2026).

Runway Gen-4.5 is a text and image to video model. Its image input is a first frame, not an identity lock, and Runway’s spec page notes support for additional inputs is coming. Runway’s reference feature, capped at three active references with a documented minimum source size of 640 by 640 pixels, lives on its image model (help.runwayml.com, July 2026). Runway also publishes the most detailed reference-quality guidance of any vendor here, which is worth reading whichever model you end up on.

Veo 3.1 has a length side effect. Google’s console documentation notes that the Veo 3.1 preview model returns only eight-second videos when you use subject images (cloud.google.com, July 2026). Plan the edit around that rather than discovering it mid-batch.

Two caps are not published by the model’s maker. Black Forest Labs documents Kontext holding character consistency across multiple edits, but states no multi-reference maximum anywhere in its docs. OpenAI shows a four-image example for GPT Image 2 without stating a limit. Platforms that host these models may publish their own per-generation caps, and DesignerBox lists one for Kontext Multi on its model page. Treat a platform limit and a model limit as different claims, and check which one a number refers to before you plan around it.

For choosing between them on a specific brief rather than in the abstract, the main site runs one prompt through every model on its consistent characters comparison.

How to build a reference image that holds

Six steps. The first three do most of the work.

1. Pick or shoot one hero frame. Front-facing or three-quarter, sharp, evenly lit, plain background, neutral expression. Runway asks for a neutral expression and even lighting specifically so the reference behaves as a blank canvas. If your only source photo is heavily filtered, backlit or busy, generate a clean reference from it rather than trying to rescue the original.

2. Crop tight on the face. ByteDance’s guidance for identity is a head-only close-up with shoulders, neck and background minimised. A full-body source spends the model’s attention on clothing and setting when you wanted it on bone structure.

3. Stop at two or three. Send the hero frame plus at most one or two that clearly agree with it. Resist filling every slot. If you do add angles, check they read as the same person before you trust them.

4. Write the identity line once and freeze it. One or two sentences covering age band, build, hair and any fixed marker such as glasses. Keep it byte-identical across every prompt. Changing one adjective moves the face, which is the same discipline that keeps an AI brand character usable past its first campaign.

5. Never contradict the reference. Asking for blonde hair over a dark-haired reference forces the model to choose, and it will not choose consistently. Vary scene, wardrobe and framing. Hold identity fixed. The same rule governs writing prompts that produce realistic AI video.

6. Promote your best output to a reference. When a generation nails the character in the wardrobe and lighting you want, save it and use it as the reference for the next shot. This is how you lock the things the original reference could not carry.

Name the file, version it, and store it where the team can find it. A reference set that lives in one person’s downloads folder is not an asset.

Where character references break

Faces hold at the centre and fail at the edges. Front-facing headshots survive almost anything. Profile angles, hands, and full-body distance shots drift first, which is why shot lists built around what actually holds beat shot lists built around what you hoped for. The failure pattern is the same one behind hands and faces breaking in AI video.

Crowds are the other cliff. ByteDance documents output stability dropping when the number of reference people passes four, and names two specific failure modes: ID drift, where the face changes mid-video, and the twin problem, where one person renders twice as two different people (volcengine.com, July 2026).

Multi-shot sequences compound errors rather than resetting them. The mechanism researchers describe for diffusion sampling is recursive: errors in earlier steps produce iterates that drift away from the training distribution, and each step inherits the last one’s mistakes (arXiv:2302.09057). Note that this describes diffusion sampling in general, not any named commercial model, since no vendor publishes its architecture.

There is also a documented trade-off worth planning around. Recent work on subject-driven generation frames identity consistency and prompt diversity as fundamentally in tension (arXiv:2511.08061). The further you push a scene away from the reference, the more identity gives. If a shot demands an unusual pose, an extreme angle and a new location at once, expect to spend iterations on it, or split it into two shots.

Vendors say the quiet part themselves. OpenAI lists consistency under limitations for its image model, noting it may occasionally struggle to maintain visual consistency for recurring characters or brand elements across multiple generations. Plan for three to five attempts on a genuinely new scene, not one.

Locations drift on exactly the same mechanics as faces, and most guides stop before covering them. That is treated separately in keeping locations consistent across AI shots.

Where DesignerBox fits

The reference problem gets expensive when you solve it per tool. Each model has its own upload flow, its own cap, its own name for the feature, and its own asset library that does not talk to the others. Six subscriptions later, your references live in six places and the campaign drifts between them. The real cost is not the subscriptions. It is the seams.

DesignerBox puts all thirteen image and video models behind one subscription, one workspace and one searchable library, so the reference set you build is the reference set every model draws from. Swap Veo for Seedream on a single shot without re-uploading anything or rebuilding the character.

The Character Swap feature holds one model across a campaign’s assets. If you are building a recurring creator identity rather than a one-off, the UGC creator persona is the shape to start from. Both run inside Video Studio, which is also where saved workflows rerun the campaign that worked for the next product or drop.

Plans start free with 112 credits and no credit card. Basic is $15 a month for 500 credits, Pro $35 for 1,000, Premium $75 for 2,500, and Ultra $200 for 8,000. Video is by far the most expensive operation, priced per second of output, so budget it separately from stills.

For garment work specifically, where the face has to hold while the clothing changes, the constraints differ enough to be worth reading on their own in consistent AI fashion images. And for the video side of the same problem, keeping characters consistent across AI video clips covers the drift mechanics in depth.

FAQ

How many reference images should I use for an AI character?

Fewer than most guides suggest. Two or three that clearly show the same person beat six that disagree. ByteDance advises against filling the reference limit, on the basis that too many assets make it harder for the model to judge which features take priority (volcengine.com, July 2026). Start with one strong frame and add only when a specific shot fails without it.

What resolution should a character reference image be?

Runway is the only vendor publishing a figure, and recommends against reference images smaller than 640 by 640 pixels or larger than 4K (docs.dev.runwayml.com, July 2026). Veo, GPT Image 2, Nano Banana and Seedream publish no resolution guidance for references at all. Above 1024 pixels on the short edge is a safe working floor.

Can I use a photo of a real person as an AI character reference?

It depends on the model, and two will refuse. OpenAI states that input images with human faces are rejected for Sora 2 image references, and that character uploads depicting human likeness are blocked by default. Seedance 2.0 requires identity verification and consent capture before real-person material can be used. Veo 3.1, Nano Banana Pro, Seedream 5 and FLUX 2 Flex accept them (provider docs, July 2026). Separately from what the model allows, a real person’s face needs a signed release that names AI generation.

Why does my AI character look different in every generation?

Three causes stack. Sampling is stochastic unless the seed is pinned, and Google documents that specifying a seed without changing other parameters guides the model to produce the same output. Your prompt may be contradicting the reference. And identity consistency trades against prompt diversity, so the further a scene sits from the reference, the more the face gives (arXiv:2511.08061). Note that Seedance 2.0 does not expose a seed parameter, so that lever is not available on every model.

Does a character reference image keep the clothing consistent?

Only on some models, and you should not rely on it. Black Forest Labs states FLUX.2 maintains face, clothing, proportions and style across generations. Runway shows wardrobe varying inside a single generation and pins it by saving an output as a new reference (provider docs, July 2026). Specify wardrobe in every prompt regardless.

Should I use multiple angles of the same character?

For products, yes. For faces, ByteDance advises against multi-view character sheets, because the model can read different angles as different subjects and that increases ID drift (volcengine.com, July 2026). Its recommendation for people is a single head-only close-up with a neutral expression. Test your set on three shots before committing a campaign to it.

Can I keep two characters consistent in the same shot?

Yes, with declining reliability as you add people. ByteDance documents output stability dropping past four reference people, and names the twin problem, where one person renders twice as two visibly different people. OpenAI limits a single Sora video to two characters. Test each character separately before combining them, and see putting two characters in one AI image for the blending failure modes specifically.

Sources

  • Runway, Creating with Gen-4 Image References and Creating with Gen-4.5 (help.runwayml.com, accessed July 2026)
  • Runway developer documentation, model capability and image constraints (docs.dev.runwayml.com, accessed July 2026)
  • Google Cloud, Guide video generation using asset and style images (cloud.google.com, accessed July 2026)
  • Google, Gemini API image generation documentation (ai.google.dev, accessed July 2026)
  • OpenAI, video generation and image generation guides (developers.openai.com, accessed July 2026)
  • Volcengine, Seedance 2.0 prompt guide, video generation API and Seedream reference limits (volcengine.com, accessed July 2026)
  • Black Forest Labs, FLUX.2 overview, character consistency guide and Kontext documentation (docs.bfl.ml, accessed July 2026)
  • Consistent Diffusion Models, arXiv:2302.09057
  • Taming Identity Consistency and Prompt Diversity in Diffusion Models, arXiv:2511.08061

Model reference-image specifications verified against each provider’s own documentation as of July 2026. Model capabilities in this category change monthly. Individual results vary.

Cristian

Head of Content at DesignerBox

Cristian covers AI product photography, video ad tools and model comparisons. He runs the same prompt and the same product across models, then publishes the output side by side, so you pick on evidence instead of marketing copy.

Follow along on Instagram at @designerboxai for campaign breakdowns.

Reshoot the clip, not the campaign

Veo 3.1, Sora 2 Pro, Kling 2.6 Pro, Seedance 2.0 and Runway Gen-4.5 are built in. When one model warps a hand, rerun the same shot on another without a second subscription.

Start free

Upload one product photo. Ship the whole campaign, without a photoshoot.