CNLT-1: Canonical Normalized Lexical Transcription v1 22 Aug 2026 PURPOSE CNLT-1 is a structural proto-transcription, not plaintext. It rewrites each RF1b/CFST-1 token as: L= | LEMMA=[] | R= | STATE=<5-state grammar class>. EVIDENCE FOR THE LEMMA LAYER 1. Internal family-body sequences agree between directly aligned ZL and GC tokens slightly more often than complete family sequences. 2. Across ZL, GC and RF, different edge-marked surface forms sharing the same internal body have significantly more similar non-adjacent within-line body contexts than frequency- and ending-state-matched forms with different bodies. 3. This direction holds in 15/15 page-group folds. 4. This supports treating the body as a lemma-like invariant for structural analysis. It does NOT establish pronunciation, language, or meaning. CONTEXT INVARIANCE TEST - Tokens: simple STA family sequences of length >=3. - Body: token sequence after removing one physical family at each edge. - Surface variants: distinct complete family sequences sharing a body. - Context: bodies at token distance 2..5 within the same physical line; immediately adjacent tokens excluded to avoid simply re-measuring the known edge grammar. - Metric: Jensen-Shannon divergence of context distributions. - Control: different-body surface forms matched approximately on frequency and final-family/state class. - Full corpus: one-sided paired Wilcoxon across eligible body families. - Robustness: five page-group partitions in each ZL/GC/RF corpus. INTERPRETATION CNLT-1 is the closest supported normalized transcription currently available in this project. Anonymous lemma IDs are deliberately semantic-free.