Two characters blend in one AI image because most models bind identity to the whole frame rather than to a named person inside it. Give each character its own reference slot, separate them by position in the prompt, and keep them from touching. Only Google publishes a per-character number: Nano Banana Pro documents up to 5 character reference images.
You locked your presenter weeks ago. She holds across twelve shots. Then the brief asks for her handing the product to a customer, and both faces arrive wrong: her jaw on the customer, his jacket colour bleeding onto her sleeve, and in one output the same woman twice.
Nothing about your reference set changed. The second person did. Everything that makes one character hold works against you the moment a second one enters the frame, and almost no provider documents what happens next.
This covers why identity bleeds between two subjects, the number nearly everyone misreads when they check a model’s specs, what each model in the catalogue actually documents, the four ways a two-character shot fails, and the shot order that gets both faces through. It is written for marketing teams and agencies producing multi-person creative.
Key Takeaways
- Reference image counts are not people counts. Providers publish how many files you may upload. Almost none publish how many distinct people the model will hold. Reading the first number as the second is the most common mistake in multi-character work.
- Google is the exception, twice over. Nano Banana Pro documents up to 5 character reference images inside a 14-image total, and Google’s launch post describes maintaining “the consistency and resemblance of up to 5 people” (ai.google.dev and blog.google, July 2026).
- Video documentation points the other way. Veo 3.1 accepts up to three asset images of “a single person, character, or product” (ai.google.dev, July 2026). Sora publishes the only explicit ceiling: “A single video can include up to two characters” (developers.openai.com, July 2026).
- Identity bleeds at contact points. Side by side holds. Facing each other holds less. Touching, hugging, or handing something over is where traits cross, and it is also where hands fail.
- Build each character alone before they meet. Two locked single-character sets, generated and approved separately, then combined. Debugging a blended pair shot without them tells you nothing about which reference is at fault.
- Budget for retries, not for a rate. A two-character shot needs several times the generations of a single-character shot. That volume, not the per-image price, is what a multi-person campaign actually costs.
Why do two AI characters blend into one?
An image model does not hold a roster of people. It holds one description of a picture, and identity is one property of that description among many.
When you supply a single reference, the model has one identity to satisfy and the whole frame to spend it on. Add a second person and the model has to decide which features belong to which body, using a prompt that describes both in the same sentence and reference images that arrive as a flat set with no labels attached. Nothing in that input says which face goes where.
So the model does what it does with any ambiguity: it averages. The result is a plausible picture containing features from both references distributed across two bodies, which is precisely the failure you are looking at.
OpenAI names the general version of this in its own documentation, noting that the model “may occasionally struggle to maintain visual consistency for recurring characters or brand elements across multiple generations” (developers.openai.com, July 2026). Add a second character and the surface area for that struggle roughly doubles.
This is a different problem from single-character drift across a sequence. If your character changes between shots while alone in frame, the fix is the reference discipline in the guide to building an AI brand character. If two characters change into each other inside one frame, keep reading.
Reference images are not people
Here is the misreading that wastes the most time.
Every provider publishes a reference-image count. Almost none publish a people count. These are different numbers, and the gap between them is where multi-character projects get planned wrong.
“Up to 10 reference images” tells you how many files the endpoint accepts. It does not tell you whether those ten files may show ten different people, or whether they should all show one person from ten angles. For most models the documentation simply stops before answering, and the reader fills the silence with the more useful interpretation.
Google is the one provider that crosses the gap explicitly. Its API reference counts images, listing “Up to 5 images of characters to maintain character consistency” for Nano Banana Pro (ai.google.dev, July 2026). Its product blog counts people, describing “the consistency and resemblance of up to 5 people” and, in its prompting guidance, characters held “even when they appear together in a group” (blog.google, July 2026).
ByteDance asserts the capability in prose without a number, describing how Seedream 5, “given several separate photos of different people”, extracts each person’s features and combines them into a single scene (seed.bytedance.com, July 2026). That is a clear statement that distinct people work, attached to no ceiling.
Everyone else publishes a file cap and stops. Treat a reference count as a budget for your inputs, never as a guarantee about your cast.
What each model actually documents
Assembled from provider documentation, accessed July 2026. Where a provider is silent, that silence is reported rather than filled.
| Model | What the provider documents | Read it as |
|---|---|---|
| Nano Banana Pro | ”Up to 5 images of characters to maintain character consistency”, inside a 14-image total, plus up to 6 high-fidelity object images and up to 3 style references (ai.google.dev) | The clearest per-character allowance in the catalogue |
| Nano Banana 2 | ”Up to 4 images of characters to maintain character consistency”, up to 10 object images, no style references (ai.google.dev) | A smaller character budget than Pro |
| Seedream 5 | Up to 10 reference images on the Pro variant (docs.byteplus.com). ByteDance describes combining “several separate photos of different people” into one scene (seed.bytedance.com) | Distinct people stated in prose, no count published |
| GPT Image 2 | No character or subject count published. OpenAI notes the model “may occasionally struggle to maintain visual consistency for recurring characters” (developers.openai.com) | Silent on multi-character |
| Kontext Multi | Preserves the identity “of e.g. a reference character or object across multiple scenes and environments” (bfl.ai). Additional reference slots carry an experimental label | Singular wording throughout |
| Veo 3.1 | Up to three asset images “of a single person, character, or product” (ai.google.dev) | Documented for one subject |
| Sora 2 Pro | ”A single video can include up to two characters” (developers.openai.com) | The only published two-character ceiling |
| Seedance 2.0 | Accepts up to 9 input images. At launch ByteDance called multi-subject interaction industry-leading while noting “room for optimization regarding multi-subject consistency” (seed.bytedance.com, February 2026) | Both claims from the same post |
Three things worth pulling out of that table.
The 14 is a total, not an object budget. Nano Banana Pro’s object allowance is 6. The 14 is the ceiling across every input type combined, so a two-character scene with products in it spends from one pool.
Video is documented more conservatively than stills. Google’s video reference wording is singular, and Sora’s guidance recommends no more than two characters per generation (developers.openai.com, July 2026). Plan multi-person video around two people, and read the video-side character consistency guide before committing a sequence.
Element features are version-specific. Kuaishou’s capability map lists element control and multi-image to video as not supported on the 2.6 generation; the Elements feature that takes 1 to 4 subject images sits on 1.6 (app.klingai.com, July 2026). Check the generation, not the family name.
For a side-by-side of which model suits which brief, DesignerBox runs one prompt through every model at consistent characters, and the per-model detail sits at Nano Banana Pro and Seedream 5.
The four ways a two-character shot fails
Diagnose which one you have before touching the prompt. They have different fixes and they look similar at thumbnail size.
Trait bleed. Features cross between the two people. Her hair colour lands on him, his jacket appears on her sleeve, one person inherits the other’s jawline. The references are competing for the same slot, usually because they arrived as an unlabelled set.
The same-face collapse. Both characters render as the same person, sometimes twice in the same pose. The model resolved two identities into one average rather than holding both. This is the most common failure with text-only character descriptions, because “a woman in her thirties with dark hair” describes both of your characters equally well.
One dominates. Character A holds perfectly. Character B arrives as a hybrid, carrying maybe 60% of their own reference and the rest borrowed. Usually the stronger reference set wins, so the fix is levelling the two sets rather than reweighting the prompt.
Drop and duplicate. One character vanishes from the frame entirely, or a third person appears who was never referenced. Crowded prompts and complex staging both raise the odds.
Trait bleed and the same-face collapse concentrate where the two people meet. A handover, an embrace, an arm around a shoulder: every one of those puts two identities inside overlapping pixels, and it is also where hands go wrong. The hand and face failure patterns compound here rather than sitting alongside.
How to shoot two characters so both hold
Six steps, in order. The order carries most of the value.
1. Build each character alone first. One reference set per person, generated and approved separately, before either appears in a shared frame. Skipping this is why most pair shots cannot be debugged: when both faces are wrong you have no idea which reference failed.
2. Give each character its own reference slot. Supply separate images per person rather than one photo containing both. A combined reference hands the model a scene to copy, which is a different instruction from holding two identities.
3. Give the model a spatial handle. Name each character and state where they stand: one on the left, one on the right. Position is the only lever most prompts have for attaching a description to a body, and it does more work than any amount of added facial detail.
4. Stage them apart, then close the distance. Side by side generates most reliably. Facing each other in conversation is harder. Contact, hugging or handing something over, is hardest. Generate the easy staging first and confirm both identities hold before you ask for the difficult one.
5. Frame above the contact point where the shot allows. A two-shot cropped at the chest avoids the overlapping hands and touching sleeves that cause most bleed, and it costs you nothing a wider frame was going to deliver.
6. Promote the winner to a plate. Once a pair shot holds both faces, use that output as the reference for the rest of the set instead of regenerating from the two original sets each time. The approved pair carries the spatial relationship your separate references never had. The same logic applies to holding a set across shots, covered in keeping locations consistent.
Save that pass once it works. In DesignerBox the sequence lives in Video Studio and reruns as a workflow, which matters more here than in single-character work because the setup cost is roughly doubled and you do not want to pay it twice.
What a two-character set actually costs
The number that decides the budget is not the price of a generation. It is how many generations you burn to get a usable one.
A single-character shot with a locked reference set lands often. A two-character shot lands considerably less often, because every one of the four failure modes above is an additional way for an output to be unusable. Plan several times the attempts per approved asset, and plan them per staging: your side-by-side shots will clear quickly and your contact shots will not.
That changes what to schedule rather than what to spend. Front-load the difficult staging, so the shots most likely to need forty attempts are not the ones you are generating the night before the campaign ships.
Video is the line that moves a plan from one tier to another, because it is priced per second of output rather than per image, and a multi-person sequence multiplies both the retries and the seconds. Budget it separately from stills rather than folding it into the same allowance. Current rates and tiers are on the pricing page.
Where this still breaks
Honest limits, so the shot list is built around what holds.
- Three or more people compounds fast. Every additional character adds another identity to confuse and another contact point. Google documents up to five character reference images on its top model, and that is a documented input allowance, not a promise about a five-person group shot.
- Contact shots stay unreliable. Hugging, hand-holding, and handovers are the hardest staging in multi-character work, and no provider documents guidance for it. Composite from two clean single-character frames when the shot has to happen.
- Similar-looking characters bleed more. Two people of the same apparent age, build, and hair colour give the model less to separate them with. Design visible difference into the pair deliberately: hair length, wardrobe colour, height.
- Nobody publishes an identity-bleed spec. No provider documents what happens when two references compete, or how to bind one reference to one position. Everything in the method above is craft built on top of that silence, so test it against your own characters rather than assuming it transfers.
- Video multi-character is behind stills. ByteDance flagged multi-subject consistency as an optimisation target at Seedance 2.0’s launch (seed.bytedance.com, February 2026), and Runway’s Gen-4.5 documentation notes it “currently offers Text to Video and Image to Video control, with support for additional inputs coming soon” (help.runwayml.com, July 2026).
Build the check in. Put every pair shot beside both original reference sets at full size and look at the jaw and the mouth on each person separately. Bleed is obvious in that comparison and nearly invisible in a contact sheet.
FAQ
How many characters can one AI image hold?
No provider publishes a hard maximum for people in one frame. Google documents up to 5 character reference images on Nano Banana Pro and describes maintaining resemblance for up to 5 people (ai.google.dev and blog.google, July 2026), which is the highest documented figure in the catalogue. In practice two characters is reliable, three is workable with clean separation, and beyond that most teams composite.
Why do my two AI characters look like the same person?
The model collapsed two identities into one average. It happens most often when both characters are described in text rather than anchored to separate reference images, and when the two descriptions overlap. Give each character its own reference slot and design visible difference into the pair.
Can I use one reference image that already contains both characters?
It works differently from what you probably want. A combined reference reads as a scene to reproduce, which locks the pose and staging along with the faces. Separate references per person keep the identities portable across new staging.
Which AI model is best for two characters in one image?
Nano Banana Pro publishes the clearest per-character allowance, and Seedream 5 explicitly documents combining separate photos of different people into one scene. Both are worth testing against your own characters, since neither publishes an identity-bleed guarantee. Side-by-side outputs are at consistent characters.
Does this work in AI video?
Less reliably than in stills, and the documentation says so. Veo 3.1’s reference wording covers “a single person, character, or product”, while Sora publishes a two-character ceiling per video (ai.google.dev and developers.openai.com, July 2026). Generate an approved two-character still first, then animate from it.
Why do characters blend more when they touch?
Contact puts two identities inside overlapping pixels, so the model has to decide feature by feature which person each one belongs to. Hands in contact are the worst version, since hands are already the least reliable region. Frame above the contact point when the shot allows it.
Can I fix a blended image without regenerating everything?
Sometimes. If one character holds and the other is a hybrid, an edit pass conditioned on the failed character’s reference often recovers it without disturbing the one that worked. If both faces are wrong, regenerate, because an edit is working from a bad average.
Sources
- Gemini API image generation reference, character and object reference limits: ai.google.dev/gemini-api/docs/image-generation (accessed July 2026)
- Google, Nano Banana Pro announcement and prompting guidance: blog.google (accessed July 2026)
- Gemini API Veo reference images: ai.google.dev/gemini-api/docs/veo (accessed July 2026)
- OpenAI image generation guide and model limitations: developers.openai.com/api/docs/guides/image-generation (accessed July 2026)
- OpenAI video generation guide, character ceiling: developers.openai.com/api/docs/guides/video-generation (accessed July 2026)
- Black Forest Labs, FLUX.1 Kontext identity preservation: bfl.ai/models/flux-kontext (accessed July 2026)
- ByteDance, Seedream 5 multi-person composition: seed.bytedance.com (July 2026)
- ByteDance, Seedance 2.0 launch notes: seed.bytedance.com (February 2026)
- BytePlus ModelArk, Seedream reference image limits: docs.byteplus.com (accessed July 2026)
- Runway, Gen-4.5 input support: help.runwayml.com (accessed July 2026)
- Kuaishou, Kling capability map and Elements documentation: app.klingai.com (accessed July 2026)
Model capabilities and reference limits verified against provider documentation as of July 2026. This category changes monthly, so re-check before relying on a number. Individual results vary.