Skip to content

Use letter case to disambiguate suffix vs. surname: "Jack MA" vs "Jack Ma" #289

Description

@derek73

Four post-nominals in suffix_acronyms_ambiguous are also ordinary surnames: do, ed, jd, ma. The parser decides between the two readings by position — an ambiguous acronym is a credential when removing it still leaves a given and a family name, and a surname otherwise:

>>> parse("Jack MA").family
'MA'
>>> parse("John Smith MA").suffix
'MA'

That rule works well, but it throws away a signal the input already carries. Letter case is currently ignored entirely — these parse identically today:

>>> parse("Jack MA").family,       parse("Jack Ma").family
('MA', 'Ma')
>>> parse("John Smith MA").suffix, parse("John Smith Ma").suffix
('MA', 'Ma')

Yet MA and Ma read very differently to a person. In "Jack MA", all-caps suggests a degree; in "John Smith Ma", title case suggests a surname. Both are cases where the positional rule and the case signal disagree.

The constraint that makes this non-trivial: case is only informative when the input as a whole varies in case. JOHN SMITH MA and john smith ma come from all-caps and all-lower data sources, where case carries no information at all — and a lot of real name data looks like that. So this can't be a hard rule; it would have to be a tiebreaker consulted only when the string shows mixed case, which the parser would need to determine first.

Sketch, if it's worth doing:

  • decide whether the input is uniformly cased (all upper, all lower) or mixed
  • only for mixed-case input, let an all-caps ambiguous acronym lean credential and a title-case one lean surname
  • keep the positional rule as the fallback whenever case is uninformative

This stays deterministic — no model, no training data — so it doesn't conflict with the parser's guarantee that the same input always parses the same way.

Came out of the 2.0 work on the positional rule (8147ac6); capitalized() already reasons about single-case vs mixed-case input, so some of the detection logic exists.

Activity

  1. derek73 commented on Jul 20, 2026

    @derek73
    OwnerAuthor

    A second position where case would decide it: the post-comma given slot.

    "Smith, MA" puts MA in given, and case is ignored there too:

    >>> parse("Smith, MA").given, parse("Smith, Ma").given, parse("Smith, ma").given
    ('MA', 'Ma', 'ma')

    All three take the same path — the first post-comma piece becomes the given name before any suffix test runs. "Smith, Ma" is plausibly a person; "Smith, MA" is more plausibly a surname followed by a stray credential, and "Smith, ma" is probably neither.

    This came up while deciding whether the parser should report an Ambiguity for the post-comma case. It doesn't, deliberately — a comma usually settles the structure, and flagging it would be noise on ordinary input. But if case ever becomes an input to the decision, this position is a second place it would apply, alongside the trailing-acronym case above.

    Not proposing a change to the parse today: "Smith, MA" is unusual input, and any special handling here should be case-sensitive rather than firing on all three spellings.

  2. added this to the v2.2 milestone on Aug 1, 2026
  3. removed this from the v2.2 milestone on Aug 30, 2026
  4. added this to the 2.4 milestone on Sep 9, 2026
  5. added 3 commits that reference this issue on Sep 18, 2026
  6. added a commit that references this issue on Sep 19, 2026
  7. derek73 commented on Sep 19, 2026

    @derek73
    OwnerAuthor

    Shipped in 2.4 (PR #530). A bare ambiguous credential acronym is now read by the evidence the writing carries. In a name written in more than one case, a member of suffix_acronyms_ambiguous written in capitals reads as the credential even with nothing to spare, so Jack MA gives suffix MA where it gave family. One written in any other cased form that is not wholly lower reads as the surname even with words to spare, so John Smith Ma gives family Ma. A name written wholly in one case carries no contrast and keeps the positional rule: JOHN SMITH MA, ANH DO and jack ma are unchanged.

    The same reading reaches the comma forms, which was the second position raised in the comment above. Smith, MA gives family Smith, suffix MA, while Smith, Ma and Smith, ma stay given names. At a comma the words-to-spare count is of name words and decides first, so Smith Jr., MA keeps its family, and two name words before the comma read the credential whatever the case. That restores 1.4.0's reading of John Smith, MA, John Smith, Ed, john smith, ma and JOHN SMITH, MA.

    Every decision at these slots reports the existing suffix-or-name, comma paths included. That is the first time the comma's own decision about such a word has been reported. One slot stays silent and is tracked in #531: the end of the given part after a family comma, as in Doe, John MA.

    Two costs are accepted rather than repaired, both recorded at decisions.md#S2. Freiherr von Berg MA reads family von Berg, suffix MA. abdul Smith Jr Ma reads Jr as a middle name, because the surname lean stops the peel before the suffix behind it is reached.

  8. added 5 commits that reference this issue on Sep 22, 2026
  9. added 6 commits that reference this issue on Oct 8, 2026
  10. added 2 commits that reference this issue on Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions