docs: restructure customize.rst into sections and add a user-docs style guide - #588
Merged
Merged
Conversation
The unlisted_dotted_suffixes and unlisted_caps_suffixes rows had grown to ~35 and ~25 lines, one boundary case per fix, burying the table's job as an index of fields. Each row now says what the field does and its default, and links to a new "Credentials the vocabulary doesn't list" subsection, matching how name_order, the delimiters and extra_suffix_delimiters already work on this page. The subsection states the shared positional reading and the two-letter initials exception once, then covers each field. The examples that were untested prose inside table cells are now doctests, so CI checks them. One claim is narrowed rather than moved: the fork is reported either way for the dotted field only -- "Smith, XYZ" reports nothing. usage.rst's pointer now targets the section. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #588 +/- ##
=======================================
Coverage 98.98% 98.98%
=======================================
Files 45 45
Lines 4220 4220
=======================================
Hits 4177 4177
Misses 43 43 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
The "Turning title detection off" example built Parser(lexicon= Lexicon.empty()), which also drops the Korean surname list while the default Policy keeps hangul segmentation active, so the doctest build (and anyone pasting the example) got a segmenterless UserWarning about something the example is not about. Reuse `lean` from the block above, which empties exactly the title fields; both outputs are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
maiden_delimiters' row restated most of the "Nicknames, maiden names, and brackets" section and carried the rest on its own: the two kinds of clause the pair decides, and the marker-dropping rule with its token nuance. The row now says what the field does and links to the section, which gains a label, names the field it had called "this knob", and takes those two pieces as doctests. "Whatever pair encloses it" is narrowed to a configured pair: "Jane Smith [née Jones]" is no maiden name, square brackets not being a default pair. lenient_comma_suffixes' CJK aside repeated usage.rst's tolerated-input paragraph on CJK commas, which already shows 田中さん, V. and lacked only the switch; it moves there as one sentence, under that paragraph's can-change caveat, and the row needs no section. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The section opened on a 45-line paragraph holding four things: what the three *_ambiguous fields are, how a bare marked acronym is read, how to choose between marking a word and removing it, and a stray sentence about particles_ambiguous. It now opens on a short intro and splits into one subsection per field, with the acronym part in two: the reading rules as a bullet list in precedence order (capitals, Title case, no signal, comma, brackets) after the `ma` doctest, and the three-way unambiguous / marked / left-out choice as its own list. The #360 particle sentence moves into the particles subsection. Every example was re-run against the tree; none had gone stale, and the section title is unchanged, so the one inbound link still resolves. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The section ran its intro and recipe into a 28-line paragraph covering how a key matches, words that need no mask, the name-word exception and unlisted acronyms. The recipe and the replace-not-extend note stay up top; the rest splits into three subsections, the plain/dotted fallback becomes a list, and five claims that were prose become doctests (ph.d., p.h.d., md iv, psyd, edd). The matching example swaps "Phd" for "PHD": a mixed-case Phd matches the key but is left as written unless force=True, which the text now says. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The section's last paragraph ran four topics together: what the two fields do to Japanese names, the segmenter's own off-switch, and the two behaviors no policy field reaches (the nakaguro and the glued honorific peel, with the peel's off-switch and cost). It now splits into four subsections, the two policy-proof behaviors become a list, and each claim that was prose gains a doctest: the kana-licensed order clearing with script_orders, ・ and さん surviving both fields cleared, and honorific_tails as the peel's off-switch. The type-spelling note moves up beside the first example that uses () and frozenset(), since under a subheading at the end it would read as the peel's alone. The peel-cost claim was checked against _script_segment's entry guard. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The section said a clause "is settled in steps" and then told the steps as running prose. They are now a numbered list under "How a bracketed clause is read", with one doctest walking all three (the née Jr. example, previously prose, among them, and a quoted "née Jones" showing that step 2 holds in any configured pair). The maiden_delimiters recipe and the lone-marker rule go under "Routing a pair to maiden names", and the DEFAULT_NICKNAME_DELIMITERS recipe under "Adding a delimiter pair". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
customize.rst never said where the maiden marker words come from or how
to add one; its only mention of maiden_markers was as an exception to
word-at-a-time matching. A "Maiden markers" subsection now opens the
brackets section, ahead of the steps list whose step 2 depends on it:
the bare "née Jones" form in one line with a link to usage.rst's fuller
treatment, the shipped list by module link rather than restated, and an
add() recipe showing an added marker ("formerly") working bare and in
brackets alike. usage.rst gains a label at its "Nicknames and maiden
names" heading, and its maiden_delimiters pointer now targets the
brackets section rather than the top of customize.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"All-caps words" walked the three settings in consecutive paragraphs. Each now sits under its own heading (AFTER_COMMA marked as the default), a fifth heading level on the page using '"'. The contrast rule stays above them as the subsection's intro, since it applies to all three. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Distills this PR's restructuring of customize.rst into rules for docs/*.rst, each naming the failure it prevents: tables as indexes with detail in linked sections, subheadings per reader task, list shape following content shape, notes beside what they explain, the heading-level sequence; recipe doctests that are non-vacuous and warning-free; measuring every claim before moving it and checking it end to end; linking rather than restating. Two rules promote existing private session notes into the repo (background before behavior, from PR #294's review; recipes versus pinning examples, from PR #216). The doctest mechanics stay in the root AGENTS.md and are pointed to, not copied. Nested discovery loads it when docs/ is edited, as with docs/design/AGENTS.md; the root file gains a pointer for tools without it. Sphinx reads only .rst, so the file stays out of the built site. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The two-letter initials exception and the evidence that overrides it
now sit together in the shared intro. The give-way rule had landed
under "Dotted acronyms" alone, so it read as not covering capitals,
which it does ("John Smith, PhD MJ", "García Márquez, MJ XYZ"). The
intro also says the dotted spelling reports that reading and the
capitals do not, a difference the merged sentence hid.
- The dotted field reads a fourth position, right after a comma behind
a full name, and reads it SILENTLY ("John Smith, X.Y.Z." reports
nothing). The positions list gains it, and "either way the parse
reports the fork" is narrowed to the positions where it holds.
- Restores the "since 2.4" and "since 2.2" notes and the two
EVERYWHERE examples ("Doe, John XYZ", the maiden clause) the move
had dropped.
- The marker-dropping rule was stated as a token count, which "(z
domu)" contradicts (two tokens, kept whole). It is now "only where
it stands as its own word with a name word after it", with a doctest.
- Every Policy row whose field has a section now links to it:
name_order, nickname_delimiters, extra_suffix_delimiters and both
strip flags were dead ends, with labels added at their sections.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Its own cited evidence was partly wrong: the table "model" claimed rows linked that did not (now true, after the previous commit), the opening row counts were off, the chancellor incident predates this PR (0ab19fd), and "the fork is reported either way" was cited as true of the dotted field when the review showed it is not. The recipe/pin rule now states the line this PR actually drew (a claim the prose makes earns a doctest), the row threshold is per description cell, and Verifying says what the exit code hides: an unresolved :ref: or title link prints a WARNING and exits 0, and CI's HTML step has no -W. Adds three lessons from the review round: splitting breaks claims as merging does, check that nothing was dropped, and spell examples so they type-check; and says how to find the code behind a code claim (the stage Reads: lines). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The initials exception gives way to an unlisted word of EITHER kind
beside the letters ("MJ X.Y.Z.", "G.J. XYZ"), not only the same kind,
where that word's own field reads it.
- "Since 2.4" marked two of the dotted field's positions when the whole
field is new in 2.4, implying the others were older; the section
intro now says both fields are new in 2.4, once.
- The dotted spelling reports the BARE initials reading; behind a title
("Ms G.J.") nothing is reported.
- The extra_suffix_delimiters example, and two migrate.rst mappings,
spelled set fields as bare set literals, which fail mypy against the
frozenset annotations; now frozenset(...), checked with mypy --strict
against the old spelling as a control.
- The Policy table claimed to list every field but omitted
script_orders and segment_scripts; both get short rows linking to the
East Asian section, which gains a label. The script order winning
over every name_order was measured.
- docs/AGENTS.md: only the eight STAGES modules carry Reads: lines;
_vocab.py and _pieces.py read the lexicon without one. And the first
pass LEFT the five unlinked rows; the review found them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 3, 2026
derek73
added a commit
that referenced
this pull request
Oct 3, 2026
Links and placement:
- The leading-particle "see name-order" landed on a section that never
discusses leading particles; it now targets "Where the vocabulary
answers first" (labeled order-vocabulary-first), and "Family-first
name order"'s own "noted at the end of this section" -- wrong before
this PR, the exceptions being in the next section -- links there too.
- The sentence about the whole attachment table had landed under the
leading-particle subheading; it moves above it, and "covered next"
becomes a link, the East Asian section no longer being next.
- "That default is built for display" opened the new "Surviving a
reparse" with no antecedent, which is also where the maiden-roundtrip
link lands; it names the str() rendering now.
- The Shakespeare example moves beside the paragraph it illustrates,
and 王·Smith, the case the boundary paragraph is about, gets one.
- The ambiguous-words pointer's colon introduced a doctest showing no
comma rules; the pointer moves after the doctest.
Claims narrowed or broadened by the move:
- "two or more words before it" drops "words to spare"'s meaning: a
title does not count ("Dr. John ma" keeps the surname). Fixed here
and in customize.rst's identical bullet from #588.
- "A word in Han, kana or hangul" read as wholly CJK; a word CARRYING
such a character escapes the title reading ("田中x. Smith").
- "Every piece of each is optional" had widened from forms 1-3 to all
seven (form 7's divider is not optional); it is scoped back.
- The Particles bullet read as a fourth view; it is capitalized()'s
second exception and merges into that bullet.
- The facade parenthetical claimed initials() takes capitalized()'s
whole fallback; it asks only the conjunction question (measured:
J. S. for a spliced "y e"), and "only capitalized() is handed a
vocabulary" is restored.
Also: "East Asian forms" renamed "Forms the script carries" to match
its condition-named siblings, a double blank line removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The
unlisted_dotted_suffixesandunlisted_caps_suffixesrows of thePolicytable incustomize.rsthad grown to ~35 and ~25 lines each, which made the table hard to read as an index of fields.Changes
Both rows now say what the field does, give its default, and link to the new section.
New subsection Credentials the vocabulary doesn't list (
.. _unlisted-credentials:), after "Suffixes not separated by commas". It explains the shared positional reading and the two-letter initials exception once, then has one part per field. Every other row whose field has a section now links to it too (see the review round below).The examples that sat in table cells as untested prose are now
.. doctest::blocks. Every claim was re-run against the tree first; none had gone stale.One claim is narrowed, not just moved: "the fork is reported either way" holds only for the dotted field, and only where the dotted word ends a name, given part or clause. The caps field reports nothing for
Smith, XYZorGarcía Márquez, MJ, and the dotted field reports nothing forJohn Smith, X.Y.Z..usage.rst's pointer now targets the new section instead of the whole page.The "Turning title detection off" example built its parser on
Lexicon.empty(), which also drops the Korean surname list while hangul segmentation stays on, so the doctest build printed a segmenterlessUserWarning. It now reuseslean, the title-only lexicon from the block above, with the same outputs, and the doctest build runs with no output.maiden_delimiters: the row now says what the field does and links to "Nicknames, maiden names, and brackets", which gets a.. _brackets:label and names the field it used to call "this knob". The section also picks up the two pieces it lacked, both as doctests: the two kinds of clause the pair decides (markerless content and a lone marker word), and the marker-dropping rule with its token nuance ((Nee)and山田花子(旧姓佐藤)keep the marker). "Whatever pair encloses it" is narrowed to any configured pair:Jane Smith [née Jones]is not a maiden name, because square brackets are not a default pair.lenient_comma_suffixes: the CJK aside repeatedusage.rst's paragraph on commas around CJK names, which already shows田中さん, V.but didn't mention the switch. It moves there as one sentence, under that paragraph's caveat that the reading can change. The row is short enough now to need no section."Words that are also ordinary names": the section opened on a 45-line paragraph. It now has a short intro and one
^subsection per field: Credentials that are also surnames (the reading rules as a bullet list in precedence order), Mark a word, or leave it out (the choice between unambiguous, marked and removed, as a list), Particles that are also given names (which now holds the strayMc DonaldandSte Marieread the particle as the given name #360 sentence), and One-letter connectives that are also initials. Every example was re-run first. The section title is unchanged, so its one inbound link still resolves."Fixing the case of a particular word": the recipe and the replace-not-extend note stay at the top. The 28-line paragraph after them splits into How a key matches a word, Words that need no entry and Acronyms the vocabulary doesn't list (plain vs. dotted, as a list). Five claims that were prose are now doctests (
ph.d.,p.h.d.,md iv,psyd,edd). The matching example now usesPHDinstead ofPhd: a mixed-casePhddoes match the key, but it is left as written unlessforce=True, and the text now says so."East Asian defaults, and turning them off": split into Switching off order and splitting, Teaching the splitter a surname, Japanese names and Behaviors no policy field controls. The long closing paragraph had covered four topics; it is now spread across the last two subsections, with the two behaviors no field controls (the ・ separator and the glued honorific) as a list. New doctests pin that
script_orders=()clears the kana-licensed order (山田 エミ), that ・ andさんstill split with both fields cleared, and thathonorific_tailsis the peel's off-switch. The note about()/frozenset()spellings moves up beside the first example that uses them. I checked the peel-cost claim against theif not tailsguard in_script_segment."Nicknames, maiden names, and brackets": split into How a bracketed clause is read, Routing a pair to maiden names and Adding a delimiter pair. The section said a clause "is settled in steps" but then described them in running prose. The steps are now a numbered list, and one doctest walks through all three. It includes the
(née Jr.)example, which was prose before, and a quoted"née Jones"showing that step 2 holds inside any configured pair.Maiden markers: customize.rst never said where marker words come from or how to add one. A new Maiden markers subsection opens the brackets section. It shows the bare
née Jonesform in one line and links to usage.rst for the full treatment. It links the shipped marker list (in the module docs) instead of copying it. And it gives anadd()recipe showing an added marker (formerly) working both bare and in brackets. usage.rst gains anicknames-and-maiden-nameslabel, and itsmaiden_delimiterspointer now goes to the brackets section, not the top of customize.All-caps words: each
CapsSuffixessetting now has its own heading:AFTER_COMMA(the default),EVERYWHEREandOFF. They use a fifth heading level on the page (", rendered as<h5>). The contrast rule stays above them because it applies to all three.docs/AGENTS.md(new): a style guide for the.rstuser docs, written up from this PR. Its rules are: tables are indexes, use subheadings per reader task, match list shape to content shape, put background before behavior, place notes next to what they explain, follow the heading-level sequence, write recipe doctests that are non-vacuous and warning-free, measure every claim before moving it, check it end to end, and link instead of restating. Each rule names the failure from this PR that it prevents. It loads automatically when an agent edits files underdocs/, the same waydocs/design/AGENTS.mddoes, and the root AGENTS.md now points to it. Sphinx reads only.rstfiles, so it doesn't appear in the built site.Review round (three parallel reviewers: claim accuracy, content preservation, the guide), fixed in 4bde4b1 and 127aa83:
John Smith, PhD MJ). It moved into the shared intro.John Smith, X.Y.Z.). It's now listed, and the reporting claim is narrowed.EVERYWHEREexamples.(z domu)contradicts. It's restated, with a doctest.name_order,nickname_delimiters,extra_suffix_delimitersand both strip flags) had sections they didn't link to. They link now.docs/AGENTS.md: fixed its own cited evidence (row counts, thechancellorincident predating this PR, the dotted reporting claim), settled the recipe-vs-pin rule, and added three lessons from the round._script_segment.pyclaimsjris title vocabulary.Second review round (on the fix commits, per AGENTS.md), fixed in the last commit:
MJ X.Y.Z.), not just the same kind.Ms G.J.) nothing is reported.extra_suffix_delimiters, and two in migrate.rst) used bare set literals that fail mypy. They're nowfrozenset(...), checked withmypy --strictagainst the old spelling as a control.script_ordersandsegment_scripts; they now have short rows linking to the East Asian section.docs/AGENTS.md: only the eightSTAGESmodules carryReads:lines.No behavior change.
Verified:
sphinx-build -b doctest docs: 0 failures (126 tests in customize).sphinx-build -b html docs: exit 0.🤖 Generated with Claude Code