Skip to content

docs: bring concepts.rst up to date with 2.1–2.4 behavior - #592

Merged
derek73 merged 3 commits into
masterfrom
docs/concepts-refresh
Oct 3, 2026
Merged

derek73 merged 3 commits into
masterfrom
docs/concepts-refresh

Conversation

@derek73

@derek73 derek73 commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

concepts.rst was last edited on 2026-09-08. I read it against the tree and ran every checkable claim. Most still hold: the 8 tokens, spans, hashability, parsing never raising, the ambiguity examples, and STABLE_TAGS. Six things had drifted:

  1. script_orders covered only "a name written wholly in one East Asian script". It also applies to kanji mixed with kana (高橋 みなみ, Japanese name segmentation via optional dependency (nameparser[ja]) #272).

  2. "…nothing else does" (change what "first" and "last" mean) was false:

    • a never-given particle opening the name opens the surname under every order;
    • the opt-in patronymic_rules reorder a name.

    Both are now named. I measured the particle claim under all three orders. de Mesnil Jean under FAMILY_FIRST gives family de Mesnil, given Jean: the particle opens the surname, it doesn't take the whole name.

  3. The two-layer model had no place for readings based on how a word is written. A word ending in a period at the front of the given part is a title with no list entry. Since 2.4, an unlisted dotted acronym is read as a credential, and so is an all-caps word after a comma by default. A new paragraph names each, says what position it needs, and links to the sections in usage.rst and customize.rst.

  4. replace() / revise(): replace() tokens do carry a tag (vocab:unclassified), and revise() doesn't preserve tags; it parses the new value and produces fresh ones.

  5. "A name of one name word… is reported" was broader than rules.md's "one name word nothing else decided". "Dr. Smith" reports nothing, and rules.md lists it as a boundary case.

  6. The token-tags paragraph sat at the end of "Honest ambiguity", which it isn't about. It moves next to the spans paragraph, and "the four stable tags" drops its count.

Following docs/AGENTS.md: every new claim was measured, including the variations that showed my first wording of item 2 was wrong. A diff of inline literals and cross-references shows none dropped. sphinx-build -b doctest and -b html both print nothing.

Review rounds. One reviewer checked the PR for claim accuracy and for whether the model still reads coherently. A second reviewer then checked the fix commit. All findings were verified against the parser before fixing (83795a4f, e00ea412):

  • Scripts: the family-first order covers any Japanese mix of kanji and kana, or of the two kanas, except katakana alone. That includes やまだ たろう, 山田 エミ and やまだ エミ. The list now has its own sentence instead of a nested parenthetical.
  • All-caps credential: the reading needs a name with a mixed-case word of its own, like Smith. JOHN SMITH, XYZ and john smith, XYZ both keep XYZ as the given name.
  • patronymic_rules reorder only under the default order (measured for both rules under all three orders).
  • The one-name-word silencers are now introduced with "such as" and include the ones rules.md lists. Only a script that settles its own order silences; マイケル and Иван still report.
  • The heading "Two layers decide the roles" became "How roles are decided" (nothing linked to it). The advice on which container to use now puts the unlisted-credential settings in the Policy, but not the period-title rule, which has no setting.
  • Precision: the period-title rule is for a word of two or more letters with a single period at its end, and revise() classifies a value "the way a parse of that value on its own would".

No behavior change.

🤖 Generated with Claude Code

concepts.rst was last touched 2026-09-08. Read against the tree:

- script_orders applies to kanji mixed with kana too (#272), not only
  to "a name written wholly in one East Asian script".
- "nothing else does" (change what first and last mean) was false: a
  never-given particle opening the name opens the surname under every
  order, and the opt-in patronymic_rules reorder a name. Both named,
  the particle claim measured under all three orders ("de Mesnil Jean"
  under FAMILY_FIRST is family de Mesnil, given Jean -- the particle
  opens the surname, it does not swallow the name).
- The two-layer model had no place for readings by written form: a
  period-final word opening the given part is a title with no list
  entry, and since 2.4 an unlisted dotted acronym, and by default an
  all-caps word after a comma, are credentials. A paragraph names
  them, each with the position it needs, and links usage.rst's and
  customize.rst's sections; "the whole parser in two sentences"
  becomes "in outline".
- replace() tokens do carry a tag (vocab:unclassified), and revise()
  does not preserve tags, it classifies the value as a parse would.
- "a name of one name word ... reported" was broader than rules.md's
  "one name word nothing else decided": "Dr. Smith" reports nothing,
  a listed boundary.
- The paragraph on token tags sat at the end of "Honest ambiguity",
  which it is not about; it moves beside the spans paragraph, and
  "the four stable tags" loses its standing count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added the docs Documentation fixes and updates label Oct 3, 2026
@derek73 derek73 self-assigned this Oct 3, 2026
derek73 and others added 2 commits October 2, 2026 22:27
- script_orders: the kana license needs no kanji (やまだ たろう and
  ヤマダ たろう read family-first; pure katakana stays positional), so
  "kanji mixed with kana" becomes "Japanese written with kana other
  than katakana alone".
- The capitals reading needs a name not itself in capitals ("JOHN
  SMITH, XYZ" keeps given XYZ); said.
- patronymic_rules stand down under the family-first orders (measured:
  Ivan Ivanovich Ivanov under FAMILY_FIRST stays positional); the
  sentence beside "under every order" now says "under the default
  order".
- The one-name-word silencers were listed as if complete; rules.md
  also has a nickname or maiden name beside it, the script, and the
  word's own shape ('Smitty' Jones, Smith née Jones, 김, J. all report
  nothing). Now "such as", with the fuller list.
- The heading said "Two layers" over a section now describing a third
  kind of reading; it becomes "How roles are decided" (no inbound
  links), the body keeping two layers as the main mechanism.
- The Lexicon/Policy advice now places the unlisted-credential switches
  in the Policy. Not "written-form readings" generally: the
  period-title reading has no switch at all.
- "A word ending in a period" is "an unlisted abbreviation of two or
  more letters ending in a period" (J. is an initial, not a title).
- revise() classifies the value "the way a parse of that value on its
  own would": its docstring's caveat, and roles are forced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- "Japanese written with kana other than katakana alone" read as
  hiragana and dropped kanji with katakana (山田 エミ is family-first).
  The script list moves to its own sentence -- the old one nested a
  colon inside a dash parenthetical inside a three-item subject -- and
  names the mixes: kanji with kana, or the two kanas with each other,
  anything but katakana alone (やまだ エミ, 山田 エミ family-first;
  ヤマダ タロウ positional).
- "A full name that is not itself written in capitals" was false for
  an all-lowercase name ("john smith, XYZ" keeps given XYZ); the
  condition is a word of its own in mixed case, as Smith is.
- "the script" as a silencer is only a script that settles its own
  order (マイケル and Иван still report).
- The period-title wording excluded interior periods too loosely; now
  "a single period at its end".
- unlisted_caps_suffixes is a placement choice, not only a switch;
  "choosing where the unlisted-credential readings apply".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73
derek73 merged commit 7a9788a into master Oct 3, 2026
9 checks passed
@derek73
derek73 deleted the docs/concepts-refresh branch October 3, 2026 05:33
@codecov

codecov Bot commented Oct 3, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.98%. Comparing base (495b2b3) to head (e00ea41).
⚠️ Report is 4 commits behind head on master.

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #592   +/-   ##
=======================================
  Coverage   98.98%   98.98%           
=======================================
  Files          45       45           
  Lines        4220     4220           
=======================================
  Hits         4177     4177           
  Misses         43       43           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

derek73 added a commit that referenced this pull request Oct 3, 2026
…kana license correctly (#593)

* docs: state the kana license as the code does, on every page

Five pages paraphrased the script_orders kana license as "kanji mixed
with kana" (concepts.rst's #592 wording, "mixing kanji with kana or the
two kanas", was narrower still). The predicate (_vocab's kana license,
plus HIRAGANA's own script_orders entry) admits any name within kanji
and kana that holds some kana and is not katakana alone: wholly
hiragana やまだ はなこ and unspaced さくらエミ read family-first along
with 山田 エミ and 高橋 みなみ, while ヤマダ タロウ stays positional
(measured). customize.rst (three sites), locales.rst, concepts.rst and
modules.rst now say "kanji and kana other than katakana alone".
usage.rst stated the condition right but omitted the katakana
exception from it, and gave the reason as "a transcription would have
been kana alone" -- it is katakana alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: correct locales.rst's drifted claims

Each measured on the tree:
- given_name_titles alone does not raise; the entry is inert (the
  check was dropped 2026-07-19 and AGENTS.md forbids its return).
- The validate-alone example used given_name_titles, which works fine
  in a fragment; the rule bites for particles_ambiguous,
  suffix_acronyms_ambiguous and honorific_tails, which raise
  ValueError when the word they mark is absent.
- Merge rules: two packs setting the same scalar warn even when the
  values agree, a pack overrides the base silently, and
  segment_scripts unions like the other set fields.
- The ru row described the shape backwards; it is family/given/
  patronymic, three words, the last patronymic-shaped and the middle
  not. The tr_az row now says four words ending in the marker, the
  first read as family (Aliyev Ilham Heydar oglu).
- A segmenter is not offered "every token": only an unspaced token the
  surname list could not divide, and not where a space, family comma
  or 间隔号 already divides the name; and a wrong-type answer or an
  out-of-token cut raises, not only the segmenter's own exceptions.
- ja without a segmenter now warns at construction.
- The stand-down example (Мицкевич Адам Юзеф) is a name the ru rule
  does not match at all; replaced with one it does, and the claim
  restated: on a matched name the declared order and the rotation
  agree, so they never compete.
- "seven scripts" was a standing count that also left out the CJK
  vocabulary; the list stays, the count goes.
- #146 (Vietnamese) closed 2026-08-07; locales.rst now cites the open
  pack issue #345, and customize.rst drops its "no vn pack yet
  (#146)" pointer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: give locales.rst subheadings and move each rule beside its topic

- "What works without a pack": "Adding a word the defaults hold back"
  and "East Asian names need no pack", which now also links the
  opt-out (east-asian-defaults).
- "Using a pack": "Stacking packs" (the merge rules move here from the
  end of "Creating your own Locale", as a three-item list),
  "Finding packs by code" (the --locale equivalence stated once, not
  twice) and "Shipped packs".
- The family-first stand-down note leaves the shape-not-language
  warning box, which it is not about.
- "Creating your own Locale": the Kapitan example now follows the
  PolicyPatch introduction it illustrates, and the validate-alone rule
  gets "The lexicon fragment is validated on its own".
- Contributing item 3's three pack kinds are sub-bullets.
- :doc: links become section refs: concepts' containers section
  (newly labeled config-containers) and ambiguous-words.
- Doctest variables get distinct names (shaikh, sayyid_lex, sayyid,
  script_pack, kapitan_lex, corp_pack, kapitan) instead of rebinding
  name/lex/mine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: correct migrate.rst's drifted claims and add the post-2.0 changes

- The capitalization_exceptions advice, dataclasses.replace(lexicon,
  capitalization_exceptions={...}), replaced every shipped mask
  ("john smith phd" then repairs to PHD) and failed mypy; both sites
  now point at customize.rst's extend-the-defaults recipe (newly
  labeled case-exceptions).
- The 1.x docs link pointed at readthedocs "stable", which now serves
  2.3.0; it targets /en/v1.4.0/.
- "Runs clean on 1.4 under -W error ... it will run on 2.0, with four
  exceptions" is scoped to 2.0.0 and loses its count, and a second
  list names what later 2.x releases broke without a 1.4 warning:
  TITLES.add raising (2.2), retired config names warning (2.2), a
  non-mask capitalization_exceptions value raising (2.4), an
  unmatchable Constants entry dropped with a UserWarning (2.4).
- "Behavior changes" pointed only at the 2.0.0 release-log section and
  stopped at 2.1. It now points at every 2.x section and lists the
  shapes 2.2-2.4 moved on HumanName, each 1.4.0 reading measured with
  PYTHONSAFEPATH=1 from outside the worktree (nameparser.__file__
  asserted) and each current reading pinned by a doctest.
- Names the Lexicon fields with no CONSTANTS attribute
  (conjunctions_ambiguous, maiden_markers, surnames, honorific_tails).
- The dean recipes become doctests, with the default reading shown so
  they are non-vacuous.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: give migrate.rst subheadings and put the flip warning by its row

- "Before you upgrade" splits into what breaks without a warning, the
  silent == change, the re-pickle step, and the warned-removal table.
- "Config map" splits into "Vocabulary sets → Lexicon", "Renamed word
  lists (2.2)", "Default word lists are frozen" and "Behavior and
  render settings → Policy". The particles flip warning moves up
  beside the table row whose "see the warning below" it answers -- it
  sat about 200 lines further down -- and the rename caveat's pointer
  to it now says "above".
- "Behavior changes" splits into "Changed in 2.0", "East Asian names
  (2.1)", "Turning the East Asian readings off" and "Changed in
  2.2–2.4".
- The suffix_delimiter cell shrinks to one line plus a ref; its
  scalar-raises detail moves into prose below the table rather than
  being lost.
- Table rows and prose gain section refs: rendering-arguments,
  suffix-delimiters, brackets, strip-flags, east-asian-names,
  east-asian-defaults.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: fix what the review of the locales/migrate review found

migrate.rst -- every 1.4.0 reading held; the one-line causes did not:
- Queen/Prince: only a given-name title makes the lone name given
  ("Dr Harry" still reads last); "a form of address" was too broad.
- md -> MD: md left the exceptions map; it was not a mask change.
- "a one-letter connective ... is an initial" holds for e only (y still
  joins); "i links two surnames" holds in a mixed-case name only.
- "a capitalized credential" and "an unlisted dotted credential" were
  broader than the rules (Jack Ma, JACK MA and Jack X.Y.Z. unchanged);
  each now states its condition and an unchanged counterexample.
- No listed shape changed in 2.2 (2.2.0 reads all as 1.4, measured);
  the subsection is "Changed in 2.3 and 2.4" and each bullet names its
  release (2.3.0 measured too).
- The unmatchable-entry warning also covers entries that fold to empty
  (full stops) and capitalization_exceptions keys.
- The frozen-sets bullet links its own subsection, and the moved flip
  warning points forward to the rename it cites.

locales.rst:
- The family-first stand-down is not "never disagree" under
  FAMILY_FIRST_GIVEN_LAST, where the patronymic reads as the given
  name; both orders are now stated, and the stand-down names both
  rotations.
- The segmenter is still asked beside a Latin word ("Dr. 高橋一郎"); only
  a second East Asian word, the nakaguro, a family comma or a 间隔号
  divides the name first.
- Only particles_ambiguous, suffix_acronyms_ambiguous and
  honorific_tails are checked against their parent set;
  given_name_titles and conjunctions_ambiguous are not.
- Katakana ships no honorific, and the East Asian scripts ship no
  conjunction or particle.
- Scalar-conflict warnings fire between packs of different codes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: fix what the second review round found in the locales/migrate fixes

- migrate.rst: "(Queen, Prince, Princess)" read as the list of
  given-name titles; sir, dame, king and swami are too, and the 2.3
  change was two mechanisms -- a title run matched by its last title
  (Her Majesty Queen, Rev Sir: measured 2.2.0 last -> 2.3.0 first) and
  newly added given-name titles (Prince, Princess, Swami, Guru). Both
  named; no list claims completeness.
- locales.rst: under FAMILY_FIRST_GIVEN_LAST the last word reads as the
  given name, and for tr_az that is often the separate oglu/qizi
  marker, not the patronymic ("Aliyev Ilham Heydar oglu" gives given
  oglu).
- locales.rst: any non-CJK-classified neighbour leaves the segmenter
  asked -- Cyrillic and halfwidth katakana as well as Latin.
- locales.rst: "the East Asian scripts honorifics only" took katakana
  back in; the scripts are named.
- Rewraps a line the previous fix left long.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docs Documentation fixes and updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant