Skip to content

docs: restructure customize.rst into sections and add a user-docs style guide - #588

Merged
derek73 merged 13 commits into
masterfrom
docs/policy-table-unlisted-suffixes
Oct 3, 2026
Merged

derek73 merged 13 commits into
masterfrom
docs/policy-table-unlisted-suffixes

Conversation

@derek73

@derek73 derek73 commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

The unlisted_dotted_suffixes and unlisted_caps_suffixes rows of the Policy table in customize.rst had grown to ~35 and ~25 lines each, which made the table hard to read as an index of fields.

Changes

  • Both rows now say what the field does, give its default, and link to the new section.

  • New subsection Credentials the vocabulary doesn't list (.. _unlisted-credentials:), after "Suffixes not separated by commas". It explains the shared positional reading and the two-letter initials exception once, then has one part per field. Every other row whose field has a section now links to it too (see the review round below).

  • The examples that sat in table cells as untested prose are now .. doctest:: blocks. Every claim was re-run against the tree first; none had gone stale.

  • One claim is narrowed, not just moved: "the fork is reported either way" holds only for the dotted field, and only where the dotted word ends a name, given part or clause. The caps field reports nothing for Smith, XYZ or García Márquez, MJ, and the dotted field reports nothing for John Smith, X.Y.Z..

  • usage.rst's pointer now targets the new section instead of the whole page.

  • The "Turning title detection off" example built its parser on Lexicon.empty(), which also drops the Korean surname list while hangul segmentation stays on, so the doctest build printed a segmenterless UserWarning. It now reuses lean, the title-only lexicon from the block above, with the same outputs, and the doctest build runs with no output.

  • maiden_delimiters: the row now says what the field does and links to "Nicknames, maiden names, and brackets", which gets a .. _brackets: label and names the field it used to call "this knob". The section also picks up the two pieces it lacked, both as doctests: the two kinds of clause the pair decides (markerless content and a lone marker word), and the marker-dropping rule with its token nuance ((Nee) and 山田花子(旧姓佐藤) keep the marker). "Whatever pair encloses it" is narrowed to any configured pair: Jane Smith [née Jones] is not a maiden name, because square brackets are not a default pair.

  • lenient_comma_suffixes: the CJK aside repeated usage.rst's paragraph on commas around CJK names, which already shows 田中さん, V. but didn't mention the switch. It moves there as one sentence, under that paragraph's caveat that the reading can change. The row is short enough now to need no section.

  • "Words that are also ordinary names": the section opened on a 45-line paragraph. It now has a short intro and one ^ subsection per field: Credentials that are also surnames (the reading rules as a bullet list in precedence order), Mark a word, or leave it out (the choice between unambiguous, marked and removed, as a list), Particles that are also given names (which now holds the stray Mc Donald and Ste Marie read the particle as the given name #360 sentence), and One-letter connectives that are also initials. Every example was re-run first. The section title is unchanged, so its one inbound link still resolves.

  • "Fixing the case of a particular word": the recipe and the replace-not-extend note stay at the top. The 28-line paragraph after them splits into How a key matches a word, Words that need no entry and Acronyms the vocabulary doesn't list (plain vs. dotted, as a list). Five claims that were prose are now doctests (ph.d., p.h.d., md iv, psyd, edd). The matching example now uses PHD instead of Phd: a mixed-case Phd does match the key, but it is left as written unless force=True, and the text now says so.

  • "East Asian defaults, and turning them off": split into Switching off order and splitting, Teaching the splitter a surname, Japanese names and Behaviors no policy field controls. The long closing paragraph had covered four topics; it is now spread across the last two subsections, with the two behaviors no field controls (the ・ separator and the glued honorific) as a list. New doctests pin that script_orders=() clears the kana-licensed order (山田 エミ), that ・ and さん still split with both fields cleared, and that honorific_tails is the peel's off-switch. The note about () / frozenset() spellings moves up beside the first example that uses them. I checked the peel-cost claim against the if not tails guard in _script_segment.

  • "Nicknames, maiden names, and brackets": split into How a bracketed clause is read, Routing a pair to maiden names and Adding a delimiter pair. The section said a clause "is settled in steps" but then described them in running prose. The steps are now a numbered list, and one doctest walks through all three. It includes the (née Jr.) example, which was prose before, and a quoted "née Jones" showing that step 2 holds inside any configured pair.

  • Maiden markers: customize.rst never said where marker words come from or how to add one. A new Maiden markers subsection opens the brackets section. It shows the bare née Jones form in one line and links to usage.rst for the full treatment. It links the shipped marker list (in the module docs) instead of copying it. And it gives an add() recipe showing an added marker (formerly) working both bare and in brackets. usage.rst gains a nicknames-and-maiden-names label, and its maiden_delimiters pointer now goes to the brackets section, not the top of customize.

  • All-caps words: each CapsSuffixes setting now has its own heading: AFTER_COMMA (the default), EVERYWHERE and OFF. They use a fifth heading level on the page (", rendered as <h5>). The contrast rule stays above them because it applies to all three.

  • docs/AGENTS.md (new): a style guide for the .rst user docs, written up from this PR. Its rules are: tables are indexes, use subheadings per reader task, match list shape to content shape, put background before behavior, place notes next to what they explain, follow the heading-level sequence, write recipe doctests that are non-vacuous and warning-free, measure every claim before moving it, check it end to end, and link instead of restating. Each rule names the failure from this PR that it prevents. It loads automatically when an agent edits files under docs/, the same way docs/design/AGENTS.md does, and the root AGENTS.md now points to it. Sphinx reads only .rst files, so it doesn't appear in the built site.

Review round (three parallel reviewers: claim accuracy, content preservation, the guide), fixed in 4bde4b1 and 127aa83:

  • The two-letter initials exception's give-way rule had landed under "Dotted acronyms" alone, but it covers capitals too (John Smith, PhD MJ). It moved into the shared intro.
  • The dotted field reads a fourth position silently (John Smith, X.Y.Z.). It's now listed, and the reporting claim is narrowed.
  • Restored what the move had dropped: the "since 2.4" and "since 2.2" notes and two EVERYWHERE examples.
  • The marker-dropping rule was described as a token count, which (z domu) contradicts. It's restated, with a doctest.
  • Five table rows (name_order, nickname_delimiters, extra_suffix_delimiters and both strip flags) had sections they didn't link to. They link now.
  • docs/AGENTS.md: fixed its own cited evidence (row counts, the chancellor incident predating this PR, the dotted reporting claim), settled the recipe-vs-pin rule, and added three lessons from the round.
  • Out of scope, filed as a separate task: a stale code comment in _script_segment.py claims jr is title vocabulary.

Second review round (on the fix commits, per AGENTS.md), fixed in the last commit:

  • The initials exception gives way to an unlisted word of either kind (MJ X.Y.Z.), not just the same kind.
  • "Since 2.4" is now said once for both fields; it had tagged only two positions, implying the others were older.
  • The dotted reading reports a fork only when bare; behind a title (Ms G.J.) nothing is reported.
  • Set-field examples (extra_suffix_delimiters, and two in migrate.rst) used bare set literals that fail mypy. They're now frozenset(...), checked with mypy --strict against the old spelling as a control.
  • The table omitted script_orders and segment_scripts; they now have short rows linking to the East Asian section.
  • docs/AGENTS.md: only the eight STAGES modules carry Reads: lines.

No behavior change.

Verified: sphinx-build -b doctest docs: 0 failures (126 tests in customize). sphinx-build -b html docs: exit 0.

🤖 Generated with Claude Code

The unlisted_dotted_suffixes and unlisted_caps_suffixes rows had grown
to ~35 and ~25 lines, one boundary case per fix, burying the table's
job as an index of fields. Each row now says what the field does and
its default, and links to a new "Credentials the vocabulary doesn't
list" subsection, matching how name_order, the delimiters and
extra_suffix_delimiters already work on this page.

The subsection states the shared positional reading and the two-letter
initials exception once, then covers each field. The examples that were
untested prose inside table cells are now doctests, so CI checks them.
One claim is narrowed rather than moved: the fork is reported either
way for the dotted field only -- "Smith, XYZ" reports nothing.
usage.rst's pointer now targets the section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added the docs Documentation fixes and updates label Oct 3, 2026
@derek73 derek73 self-assigned this Oct 3, 2026
@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.98%. Comparing base (ca28f04) to head (8cb9358).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #588   +/-   ##
=======================================
  Coverage   98.98%   98.98%           
=======================================
  Files          45       45           
  Lines        4220     4220           
=======================================
  Hits         4177     4177           
  Misses         43       43           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

derek73 and others added 3 commits October 2, 2026 20:57
The "Turning title detection off" example built Parser(lexicon=
Lexicon.empty()), which also drops the Korean surname list while the
default Policy keeps hangul segmentation active, so the doctest build
(and anyone pasting the example) got a segmenterless UserWarning about
something the example is not about. Reuse `lean` from the block above,
which empties exactly the title fields; both outputs are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
maiden_delimiters' row restated most of the "Nicknames, maiden names,
and brackets" section and carried the rest on its own: the two kinds of
clause the pair decides, and the marker-dropping rule with its token
nuance. The row now says what the field does and links to the section,
which gains a label, names the field it had called "this knob", and
takes those two pieces as doctests. "Whatever pair encloses it" is
narrowed to a configured pair: "Jane Smith [née Jones]" is no maiden
name, square brackets not being a default pair.

lenient_comma_suffixes' CJK aside repeated usage.rst's tolerated-input
paragraph on CJK commas, which already shows 田中さん, V. and lacked
only the switch; it moves there as one sentence, under that
paragraph's can-change caveat, and the row needs no section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The section opened on a 45-line paragraph holding four things: what the
three *_ambiguous fields are, how a bare marked acronym is read, how to
choose between marking a word and removing it, and a stray sentence
about particles_ambiguous. It now opens on a short intro and splits
into one subsection per field, with the acronym part in two: the
reading rules as a bullet list in precedence order (capitals, Title
case, no signal, comma, brackets) after the `ma` doctest, and the
three-way unambiguous / marked / left-out choice as its own list. The
#360 particle sentence moves into the particles subsection. Every
example was re-run against the tree; none had gone stale, and the
section title is unchanged, so the one inbound link still resolves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 changed the title docs: move the unlisted-credential detail out of the Policy table docs: break up the overgrown Policy rows and ambiguous-words section in customize.rst Oct 3, 2026
derek73 and others added 9 commits October 2, 2026 21:06
The section ran its intro and recipe into a 28-line paragraph covering
how a key matches, words that need no mask, the name-word exception and
unlisted acronyms. The recipe and the replace-not-extend note stay up
top; the rest splits into three subsections, the plain/dotted fallback
becomes a list, and five claims that were prose become doctests
(ph.d., p.h.d., md iv, psyd, edd). The matching example swaps "Phd"
for "PHD": a mixed-case Phd matches the key but is left as written
unless force=True, which the text now says.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The section's last paragraph ran four topics together: what the two
fields do to Japanese names, the segmenter's own off-switch, and the
two behaviors no policy field reaches (the nakaguro and the glued
honorific peel, with the peel's off-switch and cost). It now splits
into four subsections, the two policy-proof behaviors become a list,
and each claim that was prose gains a doctest: the kana-licensed order
clearing with script_orders, ・ and さん surviving both fields cleared,
and honorific_tails as the peel's off-switch. The type-spelling note
moves up beside the first example that uses () and frozenset(), since
under a subheading at the end it would read as the peel's alone. The
peel-cost claim was checked against _script_segment's entry guard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The section said a clause "is settled in steps" and then told the steps
as running prose. They are now a numbered list under "How a bracketed
clause is read", with one doctest walking all three (the née Jr.
example, previously prose, among them, and a quoted "née Jones" showing
that step 2 holds in any configured pair). The maiden_delimiters recipe
and the lone-marker rule go under "Routing a pair to maiden names", and
the DEFAULT_NICKNAME_DELIMITERS recipe under "Adding a delimiter pair".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
customize.rst never said where the maiden marker words come from or how
to add one; its only mention of maiden_markers was as an exception to
word-at-a-time matching. A "Maiden markers" subsection now opens the
brackets section, ahead of the steps list whose step 2 depends on it:
the bare "née Jones" form in one line with a link to usage.rst's fuller
treatment, the shipped list by module link rather than restated, and an
add() recipe showing an added marker ("formerly") working bare and in
brackets alike. usage.rst gains a label at its "Nicknames and maiden
names" heading, and its maiden_delimiters pointer now targets the
brackets section rather than the top of customize.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"All-caps words" walked the three settings in consecutive paragraphs.
Each now sits under its own heading (AFTER_COMMA marked as the
default), a fifth heading level on the page using '"'. The contrast
rule stays above them as the subsection's intro, since it applies to
all three.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Distills this PR's restructuring of customize.rst into rules for
docs/*.rst, each naming the failure it prevents: tables as indexes
with detail in linked sections, subheadings per reader task, list
shape following content shape, notes beside what they explain, the
heading-level sequence; recipe doctests that are non-vacuous and
warning-free; measuring every claim before moving it and checking it
end to end; linking rather than restating. Two rules promote existing
private session notes into the repo (background before behavior, from
PR #294's review; recipes versus pinning examples, from PR #216). The
doctest mechanics stay in the root AGENTS.md and are pointed to, not
copied. Nested discovery loads it when docs/ is edited, as with
docs/design/AGENTS.md; the root file gains a pointer for tools without
it. Sphinx reads only .rst, so the file stays out of the built site.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The two-letter initials exception and the evidence that overrides it
  now sit together in the shared intro. The give-way rule had landed
  under "Dotted acronyms" alone, so it read as not covering capitals,
  which it does ("John Smith, PhD MJ", "García Márquez, MJ XYZ"). The
  intro also says the dotted spelling reports that reading and the
  capitals do not, a difference the merged sentence hid.
- The dotted field reads a fourth position, right after a comma behind
  a full name, and reads it SILENTLY ("John Smith, X.Y.Z." reports
  nothing). The positions list gains it, and "either way the parse
  reports the fork" is narrowed to the positions where it holds.
- Restores the "since 2.4" and "since 2.2" notes and the two
  EVERYWHERE examples ("Doe, John XYZ", the maiden clause) the move
  had dropped.
- The marker-dropping rule was stated as a token count, which "(z
  domu)" contradicts (two tokens, kept whole). It is now "only where
  it stands as its own word with a name word after it", with a doctest.
- Every Policy row whose field has a section now links to it:
  name_order, nickname_delimiters, extra_suffix_delimiters and both
  strip flags were dead ends, with labels added at their sections.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Its own cited evidence was partly wrong: the table "model" claimed rows
linked that did not (now true, after the previous commit), the opening
row counts were off, the chancellor incident predates this PR (0ab19fd),
and "the fork is reported either way" was cited as true of the dotted
field when the review showed it is not. The recipe/pin rule now states
the line this PR actually drew (a claim the prose makes earns a doctest),
the row threshold is per description cell, and Verifying says what the
exit code hides: an unresolved :ref: or title link prints a WARNING and
exits 0, and CI's HTML step has no -W. Adds three lessons from the
review round: splitting breaks claims as merging does, check that
nothing was dropped, and spell examples so they type-check; and says
how to find the code behind a code claim (the stage Reads: lines).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The initials exception gives way to an unlisted word of EITHER kind
  beside the letters ("MJ X.Y.Z.", "G.J. XYZ"), not only the same kind,
  where that word's own field reads it.
- "Since 2.4" marked two of the dotted field's positions when the whole
  field is new in 2.4, implying the others were older; the section
  intro now says both fields are new in 2.4, once.
- The dotted spelling reports the BARE initials reading; behind a title
  ("Ms G.J.") nothing is reported.
- The extra_suffix_delimiters example, and two migrate.rst mappings,
  spelled set fields as bare set literals, which fail mypy against the
  frozenset annotations; now frozenset(...), checked with mypy --strict
  against the old spelling as a control.
- The Policy table claimed to list every field but omitted
  script_orders and segment_scripts; both get short rows linking to the
  East Asian section, which gains a label. The script order winning
  over every name_order was measured.
- docs/AGENTS.md: only the eight STAGES modules carry Reads: lines;
  _vocab.py and _pieces.py read the lexicon without one. And the first
  pass LEFT the five unlinked rows; the review found them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 changed the title docs: break up the overgrown Policy rows and ambiguous-words section in customize.rst docs: restructure customize.rst into sections and add a user-docs style guide Oct 3, 2026
@derek73
derek73 merged commit 7a56756 into master Oct 3, 2026
11 checks passed
@derek73
derek73 deleted the docs/policy-table-unlisted-suffixes branch October 3, 2026 04:50
derek73 added a commit that referenced this pull request Oct 3, 2026
Links and placement:
- The leading-particle "see name-order" landed on a section that never
  discusses leading particles; it now targets "Where the vocabulary
  answers first" (labeled order-vocabulary-first), and "Family-first
  name order"'s own "noted at the end of this section" -- wrong before
  this PR, the exceptions being in the next section -- links there too.
- The sentence about the whole attachment table had landed under the
  leading-particle subheading; it moves above it, and "covered next"
  becomes a link, the East Asian section no longer being next.
- "That default is built for display" opened the new "Surviving a
  reparse" with no antecedent, which is also where the maiden-roundtrip
  link lands; it names the str() rendering now.
- The Shakespeare example moves beside the paragraph it illustrates,
  and 王·Smith, the case the boundary paragraph is about, gets one.
- The ambiguous-words pointer's colon introduced a doctest showing no
  comma rules; the pointer moves after the doctest.

Claims narrowed or broadened by the move:
- "two or more words before it" drops "words to spare"'s meaning: a
  title does not count ("Dr. John ma" keeps the surname). Fixed here
  and in customize.rst's identical bullet from #588.
- "A word in Han, kana or hangul" read as wholly CJK; a word CARRYING
  such a character escapes the title reading ("田中x. Smith").
- "Every piece of each is optional" had widened from forms 1-3 to all
  seven (form 7's divider is not optional); it is scoped back.
- The Particles bullet read as a fourth view; it is capitalized()'s
  second exception and merges into that bullet.
- The facade parenthetical claimed initials() takes capitalized()'s
  whole fallback; it asks only the conjunction question (measured:
  J. S. for a spliced "y e"), and "only capitalized() is handed a
  vocabulary" is restored.

Also: "East Asian forms" renamed "Forms the script carries" to match
its condition-named siblings, a double blank line removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docs Documentation fixes and updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant