Skip to main content
Ship Name Lab

Product limits and failure analysis

Name Pair Limits Report

Eight deliberately different pairs show where character-based comparison stops being useful. The software result is preserved, but every section explains what it cannot establish and which human check should override it.

Last reviewed 2026-07-29

Why these cases belong in one report

The regression checks answer whether the engine behaves consistently. They cannot answer whether a person should use the result. These eight reviews share one method and are consolidated here so each failure mode adds evidence to a single substantial document instead of becoming a short, repeated template page.

Case 1

Taylor + Travis: When a Compact Blend Still Shows Both Sources

Can a short result preserve enough of Taylor and Travis to be recognizable without an explanation?

Inputs
Taylor + Travis
First structural result
Tayvis
Warnings
0

What the software result shows

The current engine ranks Tayvis first. The opening Tay is a visible Taylor fragment, while vis preserves the end of Travis. Neither contribution is only a token letter, so the result clears the lab's source-clarity and source-balance checks.

The join also removes the repeated tra sound that would make a direct concatenation long. In lowercase, tayvis keeps the same visible boundary and does not depend on capitalization to explain the construction.

A high structural score does not prove that Tayvis is the established label for a real pairing. Search results, community convention, and the preferences of the represented people remain outside the engine.

Human checks before use

  1. Ask a reader to point to the Taylor and Travis fragments before revealing the inputs.
  2. Search the exact candidate with the relevant couple, characters, or fandom.
  3. Compare the blend with Taylor/Travis when immediate clarity matters more than compactness.

Limit demonstrated

Tayvis is easy to trace back to both inputs and is a useful software test case. It remains a draft to review, not evidence of popularity, preference, or ownership.

Case 2

Chloe + Mason: Why a High Score Cannot Settle Pronunciation

What should a reviewer do when a candidate is balanced on paper but its spoken boundary is uncertain?

Inputs
Chloe + Mason
First structural result
Masloe
Warnings
0

What the software result shows

The production engine currently ranks Masloe first. Mas comes from Mason and loe preserves a substantial part of Chloe, so the candidate is compact and balanced by the lab's character-based measures.

The written result does not tell a reader whether loe should sound like the end of Chloe, like the English word low, or in another way. The engine deliberately has no phonetic dictionary, accent model, or language claim that could resolve that question.

An earlier engine favored a longer fragment beginning with a difficult written consonant sequence. The benchmark helped expose that failure and led to a narrow heuristic change, but the new result still requires speech testing.

Human checks before use

  1. Show Masloe to several readers without the source names and record their first pronunciation.
  2. Compare it with Chason, Chlason, and the full Chloe/Mason form instead of trusting rank one.
  3. Reject the blend if the intended audience repeatedly hesitates or cannot recover both sources.

Limit demonstrated

Masloe meets the software's structural invariant, but that is intentionally not a linguistic judgment. Human pronunciation evidence should decide whether it is usable.

Case 3

Bo + Jo: The Hard Limit of Two-Letter Name Blending

Does joining two complete two-letter names create a blend, or merely a compact concatenation?

Inputs
Bo + Jo
First structural result
Bojo
Warnings
0

What the software result shows

Bo and Jo each have only one internal cut. Removing more material would reduce one source to a single character or erase it, so the engine's candidate space is necessarily small.

Bojo preserves both inputs completely and is easy to trace, but that strength is also its limitation: the result performs almost no transformation. Its lower test invariant reflects the lack of alternatives rather than declaring Bojo especially creative.

Short names also make collision checks important. A four-letter form is more likely to match an existing word, name, handle, or abbreviation in an unrelated context.

Human checks before use

  1. Compare Bojo with Bo/Jo and Bo x Jo; the unblended forms may be equally compact and clearer.
  2. Search the exact four-letter candidate before using it as a public tag or handle.
  3. Do not manufacture extra variants by repeating letters unless the represented people prefer them.

Limit demonstrated

Bojo is a valid source-preserving output, but the case demonstrates when a generator has little useful work to do. Slash or x notation may be the better answer.

Case 4

Ann + Lee: When Letter Overlap Improves Score but Hides Meaning

Can overlap make a candidate look efficient while making one or both source names harder to recognize?

Inputs
Ann + Lee
First structural result
Anee
Warnings
0

What the software result shows

Anee is compact and receives a strong structural score because it uses material associated with both inputs without becoming long. The repeated letters around the join allow the engine to avoid a plain Annlee concatenation.

To a reader who does not know the inputs, Anee can look like an ordinary given name or a spelling variant. That ambiguity is not visible in the source-coverage score.

This case separates traceability from discoverability. A reviewer who already knows Ann and Lee can reconstruct the blend, while a new reader may not infer either source.

Context changes the cost of that ambiguity. A private nickname can work after one explanation, while a public tag, wedding sign, or shared account name must remain legible to people encountering it for the first time.

Human checks before use

  1. Ask readers to guess the two source names from Anee without prompts.
  2. Test Annlee as a clearer alternative even though it is less compressed.
  3. Use the full pair in profiles or captions if the label will be seen outside a familiar group.

Limit demonstrated

Anee is structurally efficient but semantically ambiguous. It should not win solely because its numeric score is high.

Case 5

Mary-Jane + O'Connor: What Normalization Removes

What information is lost when a generator removes hyphens and apostrophes before constructing candidates?

Inputs
Mary-Jane + O'Connor
First structural result
Marnor
Warnings
0

What the software result shows

The engine normalizes punctuation so it can apply the same deterministic cuts to every input. That produces Marnor as the current leader, using material from both normalized strings.

Normalization is mechanically consistent but not culturally neutral. Mary-Jane may be treated as one compound given name, while O'Connor is a surname whose apostrophe carries familiar written structure. Removing both marks makes computation simpler and interpretation poorer.

The full strings may also be the wrong source units. Mary + Connor, Jane + Connor, or another form people actually use can produce a more faithful label than blending legal or formal spellings.

A displayed candidate can restore punctuation after generation, but that would be an editorial presentation choice rather than the exact string the engine scored. The distinction should remain visible instead of silently changing the result.

Human checks before use

  1. Confirm which parts of each compound name the represented people use publicly.
  2. Compare normalized and punctuation-preserving display forms before publishing the result.
  3. Treat Marnor as a software observation, not a claim about correct handling of either name.

Limit demonstrated

Marnor shows that deterministic normalization works technically while still requiring a human decision about identity and source selection.

Case 6

Zoë + Chloé: Preserving Diacritics Without Claiming Pronunciation

Can software preserve accented characters while admitting that it cannot judge the resulting pronunciation?

Inputs
Zoë + Chloé
First structural result
Zoloé
Warnings
1

What the software result shows

The engine retains the source characters and currently ranks Zoloé first. It does not transliterate Zoë or Chloé into ASCII merely to satisfy an English-only implementation.

The candidate receives an explicit language-review warning. The lab's basic vowel and consonant heuristics cannot establish French, Dutch, or any individual speaker's pronunciation, and the same spelling can be used across languages.

Removing both diacritics might simplify typing but would alter the entered forms. That choice belongs to the people and context involved, not to an automatic score.

The visual join also deserves review on the target platform. Fonts, case conversion, and search normalization can treat accented characters differently even when the application stores the Unicode text correctly.

Human checks before use

  1. Ask a speaker familiar with the actual names to read the candidate aloud.
  2. Check whether the target platform preserves and searches the accented spelling reliably.
  3. Keep the source spelling unless the represented person prefers a different public form.

Limit demonstrated

Zoloé is a candidate with a visible limitation, not a validated pronunciation. The warning is part of the result rather than a defect to hide.

Case 7

小明 + 小红: Why Character Slicing Is Not Language Understanding

What can a general-purpose blending engine truthfully say about Chinese-character inputs?

Inputs
小明 + 小红
First structural result
小明小红
Warnings
1

What the software result shows

The engine can retain, count, and join Unicode characters. Its current leading output is 小明小红, a direct preservation of both source strings rather than a linguistically informed nickname.

Chinese names cannot be evaluated by the lab's Latin-script vowel, consonant, or syllable approximations. Character meaning, surname and given-name boundaries, dialect, sound, and social convention are not represented in the score.

Publishing the output without that limitation would turn technical Unicode acceptance into a false language claim. The regression check therefore expects a warning and labels the result for human review.

Even a fluent speaker would need context: the same characters may represent fictional characters, public nicknames, or real people with different preferences. The software receives none of that information from two text fields.

Human checks before use

  1. Ask a fluent speaker who understands the specific names and intended context.
  2. Do not infer pronunciation from character count or reuse English blending rules.
  3. Prefer an established community label or an intentionally chosen concept name when one exists.

Limit demonstrated

The output proves only that the software handles the characters without crashing. It does not prove that the result is a natural or appropriate Chinese pairing name.

Case 8

Elizabeth + Jonathan: Choosing Source Forms Before Scoring

Should a generator blend full formal names when the audience normally uses shorter forms?

Inputs
Elizabeth + Jonathan
First structural result
Jonabeth
Warnings
0

What the software result shows

The current full-name benchmark leader is Jonabeth. It is much shorter than ElizabethJonathan and preserves a substantial fragment from each input.

The result may still solve the wrong problem. If the represented people are known as Liz and Jon, a candidate built from the formal names can be structurally elegant but socially unrecognizable.

Long inputs also create many possible cuts, which increases the chance that a scoring rule finds a smooth-looking word by accident. More candidates do not provide more evidence of suitability.

The input decision should be recorded before comparing scores. Otherwise a reviewer can keep changing between Elizabeth, Liz, Beth, Jonathan, Jon, and Johnny until the software happens to produce a preferred-looking answer.

Human checks before use

  1. Start with the public names, nicknames, surnames, or character labels the audience already recognizes.
  2. Compare Jonabeth with candidates from Liz + Jon before selecting a winner.
  3. Reject any result that erases the preferred identity form merely to improve compactness.

Limit demonstrated

Jonabeth is a strong full-input construction, but source-form choice comes before scoring. A better input pair can matter more than a better algorithm.

Reproduce the software result

Every case points back to an input in the public engine regression checks. The limit demonstrated here is deliberately kept separate from the engine's self-maintained structural invariants.

Add a structured human judgment

The public human review study asks visitors to identify a fixed candidate's source pair and separately rate readability and draft retention. Its aggregate results are descriptive rather than linguistic validation.