Repository navigation
Use letter case to disambiguate suffix vs. surname: "Jack MA" vs "Jack Ma" #289
Description
Activity
A second position where case would decide it: the post-comma given slot.
"Smith, MA"putsMAingiven, and case is ignored there too:>>> parse("Smith, MA").given, parse("Smith, Ma").given, parse("Smith, ma").given ('MA', 'Ma', 'ma')
All three take the same path — the first post-comma piece becomes the given name before any suffix test runs.
"Smith, Ma"is plausibly a person;"Smith, MA"is more plausibly a surname followed by a stray credential, and"Smith, ma"is probably neither.This came up while deciding whether the parser should report an
Ambiguityfor the post-comma case. It doesn't, deliberately — a comma usually settles the structure, and flagging it would be noise on ordinary input. But if case ever becomes an input to the decision, this position is a second place it would apply, alongside the trailing-acronym case above.Not proposing a change to the parse today:
"Smith, MA"is unusual input, and any special handling here should be case-sensitive rather than firing on all three spellings.- added a commit that references this issue
on Aug 22, 2026 - added a commit that references this issue
on Sep 19, 2026 Shipped in 2.4 (PR #530). A bare ambiguous credential acronym is now read by the evidence the writing carries. In a name written in more than one case, a member of
suffix_acronyms_ambiguouswritten in capitals reads as the credential even with nothing to spare, soJack MAgives suffixMAwhere it gave family. One written in any other cased form that is not wholly lower reads as the surname even with words to spare, soJohn Smith Magives familyMa. A name written wholly in one case carries no contrast and keeps the positional rule:JOHN SMITH MA,ANH DOandjack maare unchanged.The same reading reaches the comma forms, which was the second position raised in the comment above.
Smith, MAgives familySmith, suffixMA, whileSmith, MaandSmith, mastay given names. At a comma the words-to-spare count is of name words and decides first, soSmith Jr., MAkeeps its family, and two name words before the comma read the credential whatever the case. That restores 1.4.0's reading ofJohn Smith, MA,John Smith, Ed,john smith, maandJOHN SMITH, MA.Every decision at these slots reports the existing
suffix-or-name, comma paths included. That is the first time the comma's own decision about such a word has been reported. One slot stays silent and is tracked in #531: the end of the given part after a family comma, as inDoe, John MA.Two costs are accepted rather than repaired, both recorded at
decisions.md#S2.Freiherr von Berg MAreads familyvon Berg, suffixMA.abdul Smith Jr MareadsJras a middle name, because the surname lean stops the peel before the suffix behind it is reached.- added 5 commits that reference this issue
on Sep 22, 2026 - added 6 commits that reference this issue
on Oct 8, 2026
Four post-nominals in
suffix_acronyms_ambiguousare also ordinary surnames:do,ed,jd,ma. The parser decides between the two readings by position — an ambiguous acronym is a credential when removing it still leaves a given and a family name, and a surname otherwise:That rule works well, but it throws away a signal the input already carries. Letter case is currently ignored entirely — these parse identically today:
Yet
MAandMaread very differently to a person. In"Jack MA", all-caps suggests a degree; in"John Smith Ma", title case suggests a surname. Both are cases where the positional rule and the case signal disagree.The constraint that makes this non-trivial: case is only informative when the input as a whole varies in case.
JOHN SMITH MAandjohn smith macome from all-caps and all-lower data sources, where case carries no information at all — and a lot of real name data looks like that. So this can't be a hard rule; it would have to be a tiebreaker consulted only when the string shows mixed case, which the parser would need to determine first.Sketch, if it's worth doing:
This stays deterministic — no model, no training data — so it doesn't conflict with the parser's guarantee that the same input always parses the same way.
Came out of the 2.0 work on the positional rule (8147ac6);
capitalized()already reasons about single-case vs mixed-case input, so some of the detection logic exists.