Land pack-14 generality pass: detection patterns and pair data to langpacks and embedded target data, book-scoped overlay, pair-agnostic checkers, generic palladius parser, prompts under zh-ru

This commit is contained in:
Claude (backend session) 2026-07-24 23:01:30 +03:00
parent 77dd9b8156
commit 952469b278
53 changed files with 2019 additions and 651 deletions

View file

@ -39,6 +39,6 @@ Go-бэкенд издательского художественного пер
2. `docs/architecture/05-decisions-log.md` целиком (контракт).
3. По роли: Бэкенд — `backend/README.md` + `03-implementation-notes.md`; Полигон — `eval/README.md` + `experiments/00` + `09-pilot-protocol.md`; всем — `07-strategic-review.md` §69 (вердикт/риски/курс).
## Текущее состояние (2026-07-23, пост-D39.20 — эра «механизация выпускного качества»)
## Текущее состояние (2026-07-24, пост-D39.22/пак-14 — эра «механизация выпускного качества»)
Фаза 0 ✅; **Ф1-инфра ✅** (D20D28; golden = инвариант детерминизма, №8 README = экспорт-контракт). Арка «качество-первым» D26→D38.5 ЗАКРЫТА; **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ**: 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1D39.10) → план (D39.12) → **стройка пака-11 завершена** (D39.13D39.19: волновой исполнитель + чанкер output-бюджет + src→dst-редактор + Go-майнер паритет-EXACT + банкнота + `internal/lang`+langpacks + reject-set + echo-сплит; хроника — D-лог). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07, `books/gu-zhenren/rerun2/`):** вся операционка подтверждена живьём — ОДИН resnapshot, **$0-резюм драфта на всех 3 армах**, echo 0, $6.63<$15; майнинг-стоп→подпись→reject-set отработали. **Вердикт качества:** планка ≤2 претензии НЕ пройдена (пилот не открыт); три независимых сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232 при floor-шуме 0 · два слепых чтения) → **анти-корреляция гладкость↔верность доказана трижды** (чисто-художественные выборы обоих читателей идентичны = mistral 4/5 — арм с критическими смысловыми); **mistral исключён как несущий редактор** (голос = north-star промпта), **dspro-vs-glm решает итерация №2 после механизации** (драфт $0 → ~$0.08/арм + судья-lite). **Эксп ЗАКРЫТ решением владельца (D39.22), итерация №2 отложена** (её вопросы — довесок следующего платного прогона); **редактор = deepseek-v4-pro ИНТЕРИМ** (судья №1 · проза/глоссарий · дефекты чинибельны · ×2 дешевле; glm-5 резерв, отмена = один арм-конфиг); стайл-решения D39.21: меры → единицы читателя (时辰=2ч) · generic-«гу» средний род, именные по-термово · стихи смыслом · 资质=«талант». research/21: транспорт наш опережает/вровень, рефакторинг ОТВЕРГНУТ. **Курс — разработка бэкенда:** пак-12 транспорт-гигиена (выдан, гейт открыт, лендится первым) → пак-13 «выпускной QA» (выдан: заголовок-шаблон · юнит/число-чекеры · broken-word/latin · mined_delta-крэш-фикс · аудит банк-инъекции · стайл-канон · дефолт-редактор dspro; wire-двигающий, один resnapshot) → сид-дельта (元-корень · 赤城 · 学堂家老 · 春秋蝉) → **масштаб целой книги** (волны/майнинг-каденс/потолки на сотни глав) · канал B (18+) вживую · Ф2-механизмы (голос/состояние D21 · native-Gemini судья) · ja→ru §B5 → пилот Ф2.5. Стек: draft flash thinking-ON → **editor deepseek-v4-pro** → судья gemini-preview + grok-фолбэк на 18+ (~$0.85/ранобэ D30.4). **Висит на владельце:** мини-каноны сид-дельты (元-корень · 春秋蝉) · ja→ru-реплика (§B5) · старые (пилот Ф2.5: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Ключ xAI: единый, data-sharing, off перед продом (D27). Книга на стенде `/home/ubuntu/books/gu-zhenren/` (GB18030; текст и производные ВНЕ git). Ключи: DEEPSEEK, ZAI, KIMI, OPENAI, GEMINI, XAI, MISTRAL. Стенд: WSL2, GTX 1070 8GB (localhost из-под прокси = 403 — no-proxy транспорт для local).
Фаза 0 ✅; **Ф1-инфра ✅** (D20D28; golden = инвариант детерминизма, №8 README = экспорт-контракт). Арка «качество-первым» D26→D38.5 ЗАКРЫТА; **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ**: 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1D39.10) → план (D39.12) → **стройка пака-11 завершена** (D39.13D39.19: волновой исполнитель + чанкер output-бюджет + src→dst-редактор + Go-майнер паритет-EXACT + банкнота + `internal/lang`+langpacks + reject-set + echo-сплит; хроника — D-лог). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07, `books/gu-zhenren/rerun2/`):** вся операционка подтверждена живьём — ОДИН resnapshot, **$0-резюм драфта на всех 3 армах**, echo 0, $6.63<$15; майнинг-стоп→подпись→reject-set отработали. **Вердикт качества:** планка ≤2 претензии НЕ пройдена (пилот не открыт); три независимых сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232 при floor-шуме 0 · два слепых чтения) → **анти-корреляция гладкость↔верность доказана трижды** (чисто-художественные выборы обоих читателей идентичны = mistral 4/5 — арм с критическими смысловыми); **mistral исключён как несущий редактор** (голос = north-star промпта), **dspro-vs-glm решает итерация №2 после механизации** (драфт $0 → ~$0.08/арм + судья-lite). **Эксп ЗАКРЫТ решением владельца (D39.22), итерация №2 отложена** (её вопросы — довесок следующего платного прогона); **редактор = deepseek-v4-pro ИНТЕРИМ** (судья №1 · проза/глоссарий · дефекты чинибельны · ×2 дешевле; glm-5 резерв, отмена = один арм-конфиг); стайл-решения D39.21: меры → единицы читателя (时辰=2ч) · generic-«гу» средний род, именные по-термово · стихи смыслом · 资质=«талант». research/21: транспорт наш опережает/вровень, рефакторинг ОТВЕРГНУТ. **Курс — разработка бэкенда:** паки 12/13/14 ЗАЛЕНДЕНЫ 24.07 (транспорт-гигиена `7fe0a5b` · выпускной QA `2f91b04` · **генеральность-пасс**: пара/книга-данные → langpack+`internal/lang/data` embedded, книго-overlay `langpack_extend` (古月 вон из общего zh-пака), пар-агностичный `checkers.go`, generic-Палладий; приёмка воркфлоу 11+4 линз, байт-сверка против HEAD, отчёт с ревью-шапкой в `docs/archive/reports/`) → **пак-15 «структура+форма»** (хойст `internal/text`+value-типов → сплит miner/checks/membank/chunk по карте связности отчёта пака-14 · Export/RequestHash/resolveChunkState · конфиг-слой пары · fail-loud-хвосты оверлея) → сид-дельта (元-корень · 赤城 · 学堂家老 · 春秋蝉) → канал B (18+) вживую · слой-3 diff-редактор (прецедент AiNiee/LinguaGacha, research/22) · Ф2-механизмы (голос/состояние D21 · native-Gemini судья) · ja→ru §B5 → пилот Ф2.5. Стек: draft flash thinking-ON → **editor deepseek-v4-pro** → судья gemini-preview + grok-фолбэк на 18+ (~$0.85/ранобэ D30.4). **Висит на владельце:** мини-каноны сид-дельты (元-корень · 春秋蝉) · ja→ru-реплика (§B5) · старые (пилот Ф2.5: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Ключ xAI: единый, data-sharing, off перед продом (D27). Книга на стенде `/home/ubuntu/books/gu-zhenren/` (GB18030; текст и производные ВНЕ git). Ключи: DEEPSEEK, ZAI, KIMI, OPENAI, GEMINI, XAI, MISTRAL. Стенд: WSL2, GTX 1070 8GB (localhost из-под прокси = 403 — no-proxy транспорт для local).

View file

@ -56,5 +56,5 @@ set -a; . ./.env; set +a; TM_LIVE=1 go test -tags live -run TestLive -v ./intern
- **D15.2:** READ-половина `tmctl export` ПОСТРОЕНА (D39.5, минимальная форма); annotation/override-половина и content-addressed resume — отложены (v3.1-спека = дизайн-оф-рекорд, `docs/D15.2-…spec.md`).
- **Retries.RegenerateBeforeEscalate НЕ в снапшоте** (pre-existing, D39.4-LOW): в крэш-окне пониженный ретрай-бюджет молча бросает оплаченные OK-чекпоинты. Дизайн-фикс при следующем заходе в stagerun.
- F3 at-most-once (после D15.2); Escal.Chains — валидируемый мёртвый конфиг (раннер читает только `escalate_to`); флоры min_max_tokens = схема (D24.3).
- `prompts/editor-mono.md` — референс-вариант D13.1 (не боевой, тестом закреплён); боевой = `editor.md` v3-discourse (БИЛИНГВ). Глосс-капы санитайзера (6 CJK / 24 лат-токен) — tunable-эвристики (D39.5), ревизия по чтению пере-прогона.
- `prompts/zh-ru/editor-mono.md` — референс-вариант D13.1 (не боевой, тестом закреплён); боевой = `prompts/zh-ru/editor.md` v3-discourse (БИЛИНГВ) (промпты pair-keyed под `prompts/zh-ru/`, pair-14 §9). Глосс-капы санитайзера (6 CJK / 24 лат-токен) — tunable-эвристики (D39.5), ревизия по чтению пере-прогона.
- Хвост LOW/NOTE-находок адверсариала D39.4 — в ledger-очереди (golden-ячейка глосса, escalated+stripped тест, trustGateEvent первый-отказ и пр.).

View file

@ -0,0 +1,51 @@
# WS5 defect-class checker DATA for zh→ru (checkers.go). Pair-14 data-out: the DETECTION PATTERNS (regexes +
# literal probes) are DATA too, not engine code — a new pair ships its own. `category<TAB>key[<TAB>value]`.
# pattern values are VERBATIM (a regex or a literal probe); the checker ALGORITHM (compare counts, ×2 hours,
# suppress-if-ok) stays generic Go. The `*_re` keys are compiled; the others are literal strings.Contains probes.
#
# DC1 时辰 (double-hour): small CJK count → value (1 时辰 = 2 modern hours).
cjk_numeral 一 1
cjk_numeral 二 2
cjk_numeral 两 2
cjk_numeral 三 3
cjk_numeral 四 4
cjk_numeral 五 5
cjk_numeral 六 6
cjk_numeral 七 7
cjk_numeral 八 8
cjk_numeral 九 9
cjk_numeral 十 10
# DC1: Russian hours-count word → value (the alternatives ru_hours_re captures).
ru_hour один 1
ru_hour два 2
ru_hour двух 2
ru_hour три 3
ru_hour трёх 3
ru_hour трех 3
ru_hour четыре 4
ru_hour пять 5
ru_hour шесть 6
# DC6 register negative-list: fairy-tale/chancery lexemes out of the xianxia register (терем + case forms).
register_neg терем
register_neg терема
register_neg тереме
register_neg теремом
register_neg терему
register_neg теремах
# --- DETECTION patterns (pair-14 data-out). `pattern<TAB>key<TAB>value`; value VERBATIM. ---
# DC1: count before 时辰 (src); Russian «<count> час…» rendering (final).
pattern shichen_re ([0-9一二两三四五六七八九十])\s*个?\s*时辰
pattern ru_hours_re (\d+|один|два|двух|три|трёх|трех|четыре|пять|шесть)\s+час
# DC2 千万=10^7: src probe, ok-suppressor (case-insensitive), fire/veto words (case-sensitive Contains).
pattern qianwan_src 千万
pattern qianwan_ok_re (?i)десят\p{L}* миллион|10\s*000\s*000|10000000
pattern qianwan_fire_word тысяч
pattern qianwan_veto_word миллион
# DC2 数十万: src probe, ok-suppressor, under-render fire (case-sensitive, per reference).
pattern shushiwan_src 数十万
pattern shushiwan_ok_re (?i)сотн\p{L}* тысяч|нескольк\p{L}* сот\p{L}* тысяч|[1-9]00\s*000
pattern shushiwan_fire_re десятк\p{L}* тысяч
# percent 成=tenths: src expression; error-cue fraction; percent-form suppressor word.
pattern cheng_re ([0-9一二三四五六七八九十])成([0-9一二三四五六七八九]?)
pattern decimal_fraction_re десят(?:ая|ых|ой|ые)|сот(?:ая|ых|ой|ые)|\d+[.,]\d
pattern percent_word процент

View file

@ -0,0 +1,20 @@
# Palladius syllable-generator phonotactic constraints (miner_palladius.buildPalladiusCyrSyllables). `category<TAB>member`, pinyin tokens.
# retroflex: retroflex/sibilant initials that take SPECIAL_I, not a bare "i" final.
retroflex zh
retroflex ch
retroflex sh
retroflex r
retroflex z
retroflex c
retroflex s
# vfinal: the ü-finals (pinyin writes ü as plain u); legal only after a vfinal_initial.
vfinal v
vfinal ve
vfinal van
vfinal vn
# vfinal_initial: the initials after which a ü-final is legal (j/q/x + l/n).
vfinal_initial j
vfinal_initial q
vfinal_initial x
vfinal_initial l
vfinal_initial n

View file

@ -0,0 +1,2 @@
# Source sentence-ending marks the alias miner splits on (alias.cooccur_same_sentence 。!?). A rune SET; stored sorted. The miner adds "\n" itself.
。!?

View file

@ -3,7 +3,6 @@
东方
公孙
南宫
古月
司徒
司马
夏侯

View file

@ -0,0 +1,2 @@
# Rank/measure formant chars a candidate treats as a TITLE (patterns.formant_type 等/转/阶 → title). A rune SET; stored sorted.
等转阶

View file

@ -37,7 +37,7 @@ stages:
role: translator
model: deepseek-v4-flash
prompts:
zh-ru: ../prompts/translator.md
zh-ru: ../prompts/zh-ru/translator.md
prompt_version: v1-reflow
temperature: 0.3
reasoning: "off"
@ -48,7 +48,7 @@ stages:
# через ReasoningNone no-op). Пиннут (editor не эскалирует — D12).
model: deepseek-v4-pro
prompts:
zh-ru: ../prompts/editor.md
zh-ru: ../prompts/zh-ru/editor.md
prompt_version: v3-discourse-reflow
temperature: 0.4
reasoning: "off" # NO-OP на deepseek-v4-pro (ReasoningNone) → thinking ON; НЕ вооружает эхо-мину

View file

@ -36,7 +36,7 @@ stages:
role: translator
model: deepseek-v4-flash
prompts:
zh-ru: ../prompts/translator.md
zh-ru: ../prompts/zh-ru/translator.md
prompt_version: v1-reflow
temperature: 0.3
reasoning: "off"
@ -47,7 +47,7 @@ stages:
# P1a-дискурс few-shot держится (glm не reasoning-модель, CoT-конфликта нет, в отличие от dspro-арма).
model: glm-5
prompts:
zh-ru: ../prompts/editor.md
zh-ru: ../prompts/zh-ru/editor.md
prompt_version: v3-discourse-reflow
temperature: 0.4
reasoning: "off" # glm thinking:disabled — таймауты ×3 (полигон)

View file

@ -35,7 +35,7 @@ stages:
role: translator
model: deepseek-v4-flash # черновик — тот же, что базлайн (свап только редактора)
prompts:
zh-ru: ../prompts/translator.md
zh-ru: ../prompts/zh-ru/translator.md
prompt_version: v1-reflow
temperature: 0.3
reasoning: "off"
@ -46,7 +46,7 @@ stages:
# WS6(г); не reasoning-модель → эхо-мины нет). Идёт ТОЛЬКО за rate-guard (models.yaml rate_limit).
model: mistral-large-2512
prompts:
zh-ru: ../prompts/editor.md
zh-ru: ../prompts/zh-ru/editor.md
prompt_version: v3-discourse-reflow
temperature: 0.4
reasoning: "off"

View file

@ -51,7 +51,7 @@ stages:
# loadTemplates (закрывает латентный баг «ja едет через китайский паратаксис»). Язык промпта —
# свойство пакета. Прод zh→ru резолвит тот же translator.md → тот же SHA (поведение-нейтрально).
prompts:
zh-ru: ../prompts/translator.md
zh-ru: ../prompts/zh-ru/translator.md
# D30.2 reflow: из translator.md убрано «Сохраняй разбивку на абзацы» (построчная
# вёрстка исходника деструктивна для ru — корень нечитаемости exp12 §2.5-S). Bump
# v0-draft→v1-reflow: сменил SHA промпта, осознанный --resnapshot.
@ -76,11 +76,11 @@ stages:
model: deepseek-v4-pro
# Слой 2 (D39): pair-keyed пакет конвенций (см. draft-стадию) — наполнена zh→ru.
prompts:
zh-ru: ../prompts/editor.md
zh-ru: ../prompts/zh-ru/editor.md
# Bump v2-bilingual-reflow→v3-discourse-reflow: в editor.md вшито валидированное P1a-ядро
# ДИСКУРС-ПЕРЕВЁРСТКИ + few-shot (exp14/D37 §2а — претензия-1 «рубленые абзацы» = рычаг ПРОМПТА,
# не модель; P1a реверстает дефект у всех семей). Меняет SHA промпта → осознанный --resnapshot.
# Билингв-каркас (D30.1) и omission-осторожность (D34.3) сохранены. Моно-вариант — prompts/editor-mono.md.
# Билингв-каркас (D30.1) и omission-осторожность (D34.3) сохранены. Моно-вариант — prompts/zh-ru/editor-mono.md.
prompt_version: v3-discourse-reflow
temperature: 0.4
reasoning: "off" # NO-OP на deepseek-v4-pro (ReasoningNone) → thinking ON; НЕ вооружает эхо-мину

View file

@ -15,6 +15,15 @@ context:
cache_ttl: "5m"
# STMDepth/OverlapTokens сняты (WS2 §2а — carryover не строим).
# Явный zh-ru пар-калибровочный блок (pair-14 §7 — все шиппинг-конфиги ставят его ЯВНО, пар-конфиг =
# источник истины; значения = дореформенный Go-fallback → байт-нейтрально, но не полагаемся на fallback).
segmentation:
draft_budget_out: 1797
edit_ceiling_out: 3200
fertility:
cjk: 1.1978
other: 0.3852
retries:
regenerate_before_escalate: 1
@ -23,7 +32,7 @@ stages:
role: translator
model: deepseek-v4-flash
prompts:
zh-ru: ../prompts/translator.md
zh-ru: ../prompts/zh-ru/translator.md
prompt_version: v1-reflow # D30.2 reflow (тот же translator.md, что в C1) — label-SHA дисциплина
temperature: 0.8 # C2: кандидаты сэмплируются горячее (Р2: T=0.61.0)
reasoning: "off"
@ -35,7 +44,7 @@ stages:
role: judge
model: glm-5
prompts:
zh-ru: ../prompts/judge-selector.md
zh-ru: ../prompts/zh-ru/judge-selector.md
prompt_version: v0-skeleton
temperature: 0
reasoning: "off"
@ -47,7 +56,7 @@ stages:
# select выше — judge/glm-5, Gemini-слот Ф2 — НЕ трогаем: это судья, не редактор.)
model: deepseek-v4-pro
prompts:
zh-ru: ../prompts/editor.md
zh-ru: ../prompts/zh-ru/editor.md
# Бамп v2-bilingual-reflow→v3-discourse-reflow (label-SHA дисциплина: тот же editor.md, что в C1 —
# вшито P1a-ядро ДИСКУРС-ПЕРЕВЁРСТКИ, D37 §2а; лейбл обязан следовать за новым SHA файла).
prompt_version: v3-discourse-reflow

View file

@ -65,6 +65,13 @@ type Book struct {
// set ⇒ the runner loads the pack in W0, failing LOUD when the pair's catalog dir exists but a file is
// missing/corrupt, and folds pack.Version() into the snapshot so a pack edit is a loud --resnapshot (R1).
LangpackRoot string `yaml:"langpack_root"`
// LangpackExtend is the optional root of a BOOK-SCOPED langpack overlay (pair-14 §1): a book's PRIVATE
// canon — a clan surname (古月 in 蛊真人 is a book clan, not a 百家姓 family), a sect term — that must reach
// the miner FOR THIS BOOK without polluting the shared pair langpack. Same layout as langpack_root; only
// the files a book ships are present, each UNIONED onto the shared pack (additive) with its bytes folded
// into pack.Version() (a book-canon edit is a loud --resnapshot for this book only). Resolved relative to
// book.yaml. Empty ⇒ no overlay (the shared pack is used as-is).
LangpackExtend string `yaml:"langpack_extend"`
// MinedDelta is the optional path to the owner-curated mined-term delta YAML (the W1.5 sign artifact,
// R1). Its terms are loaded as Source:"mined" (NOT Source:"seed"), so they fold into the ENRICHED bank
// version but NOT the base — adding them moves only snapshot_W2 (the pay-once invariant). Same seedTerm
@ -108,6 +115,7 @@ func LoadBook(path string) (*Book, error) {
b.SourceFile = resolve(b.SourceFile)
b.GlossarySeed = resolve(b.GlossarySeed)
b.LangpackRoot = resolve(b.LangpackRoot)
b.LangpackExtend = resolve(b.LangpackExtend)
b.MinedDelta = resolve(b.MinedDelta)
b.MinedRejects = resolve(b.MinedRejects)
if b.ProjectDB == "" {

View file

@ -361,8 +361,16 @@ func LoadPipeline(path string, models *Models) (*Pipeline, error) {
if p.Waves.Workers <= 0 {
p.Waves.Workers = 1
}
// WS2 segmentation defaults (ratified zh-ru, §2в): a config that omits the block gets the
// output-token budget + independently-re-derived fertility. Zero/negative → default.
// Segmentation calibration (pair-14 §7). These numbers are the zh-ru PAIR CALIBRATION — the budgets
// tuned for zh→ru and the fertility (est_out per source char-class) independently re-derived on the zh-ru
// rerun (R²=0.96). They are NOT a language-neutral engine constant: a new pair must set its OWN
// calibration in its pair-config, so every SHIPPING pipeline (pipeline-c1 / the arm yamls) sets the whole
// block explicitly — the pair-config is the source of truth ("brать из пар-конфига"). The literals below
// are only the last-resort GENERIC FALLBACK for a config that omits the block. The values are held EXACT
// on purpose: a book with no langpack chunks entirely on these (the ja→ru golden fixture relies on this
// fallback — a re-derived number would shift its chunk boundaries → the wire). Relocating the canonical
// zh-ru calibration into the langpack was considered and NOT done: it would be dead for every live path
// (shipping configs set the block; the golden has no langpack to read it from) — least-mechanism §12.1.
if p.Segmentation.DraftBudgetOut <= 0 {
p.Segmentation.DraftBudgetOut = 1797
}

View file

@ -132,8 +132,8 @@ func TestBoevoyPipelineC1ResolvesZhRu(t *testing.T) {
for _, st := range p.Stages {
if st.Role == "translator" {
zh, _ := st.PromptPathFor("zh-ru")
if !strings.HasSuffix(zh, filepath.Join("prompts", "translator.md")) {
t.Errorf("draft zh-ru must be prompts/translator.md, got %q", zh)
if !strings.HasSuffix(zh, filepath.Join("prompts", "zh-ru", "translator.md")) {
t.Errorf("draft zh-ru must be prompts/zh-ru/translator.md, got %q", zh)
}
}
}

View file

@ -0,0 +1,34 @@
# Universal CJK section-numeral data (pair-14 §4 + addendum-A): the ONE source of truth for the chapter/
# section numeral inventory shared by the ingest chapter-splitter (章节節回 headers) and the chunker heading
# rule. A linguistic CONSTANT (the CJK numeral system does not vary by book/pair), so it is embedded in the
# language-data package and available even to a book with NO langpack (the ja→ru golden splits 第X章 here).
# `category<TAB>rune[<TAB>value]`. digit/unit carry a value; zero/chapter_unit are membership only.
digit 一 1
digit 二 2
digit 两 2
digit 兩 2
digit 三 3
digit 四 4
digit 五 5
digit 六 6
digit 七 7
digit 八 8
digit 九 9
unit 十 10
unit 百 100
unit 千 1000
zero
zero 零
# chapter_unit: the section-level markers auto-detected in a source txt (第N章/节/節/回). 卷 (volume) is
# deliberately excluded (coarser than a chapter).
chapter_unit 章
chapter_unit 节
chapter_unit 節
chapter_unit 回
# magnitude: the myriad-scale BIG units (万/億/兆) → value, for the number-magnitude gate (cheapgates). ONE
# source with digit/unit above — the gate's CJK numeral parser reads this section instead of its own switch.
magnitude 万 10000
magnitude 萬 10000
magnitude 億 100000000
magnitude 亿 100000000
magnitude 兆 1000000000000

View file

@ -0,0 +1,11 @@
# Target-language glossary/injection wire-text (pair-14 §2). Rendered into the model request, so it is
# TARGET-language data (a Russian header for a →ru book); gated on the target so a non-ru book never
# gets a Russian block. The ja→ru golden injects these with NO langpack, so they are engine/target data
# here, not a book pack. `target<TAB>key<TAB>value`; VALUE IS VERBATIM (leading spaces are significant).
# Add a target = add its rows; remove one = delete its rows.
ru glossary_header ГЛОССАРИЙ (используй эти утверждённые переводы имён и терминов последовательно; строки с пометкой ⟨проверить⟩ — неподтверждённые кандидаты):
ru editor_header КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имён и терминов (в черновике термин исходника слева ДОЛЖЕН быть передан именно указанной формой справа — приводи к ней любые расхождения, склоняя по контексту; не вводи иных вариантов и не меняй ничего другого):
ru unverified_marker ⟨проверить⟩
ru gender_male (муж. — мужские родовые формы)
ru gender_female (жен. — женские родовые формы)
ru gender_hidden (пол СКРЫТ до раскрытия — избегай родовых форм; при неизбежности — мужские)

View file

@ -0,0 +1,25 @@
# Universal refusal blacklist (pair-14 §3, ported 1:1 from eval/refusal_bench.py). Engine safety data:
# a model can refuse in ANY language regardless of the book's pair, so EVERY pattern here fires for EVERY
# book (the ja→ru golden flags a Russian refusal with NO langpack). ONE regex per non-comment line, joined
# with "|" under (?im); order is immaterial to the boolean match but preserved.
# NOTE: the `# --- xx ---` markers are ORGANIZATIONAL comments, NOT functional scoping — RefusalPatterns()
# reads every line flat into the one universal set. Deleting a language's lines TRIMS that universal set
# (drops those phrases for ALL books); it does NOT "turn off" detection for one pair. Add a language's phrases
# = add lines under (or beside) its marker.
# --- en ---
I (?:can(?:'|no)t|cannot|won'?t|am (?:not able|unable)) (?:to )?(?:assist|help|translate|continue|provide|fulfill)
I'?m sorry,? but
against (?:my|our) (?:content )?polic
content polic(?:y|ies)
# --- ru ---
не могу (?:помочь|перевести|продолжить|выполнить)
я не буду переводить
противоречит (?:политике|правилам)
# --- ja ---
としてお答えできません
申し訳ありませんが
# --- zh ---
我(?:不能|无法)(?:协助|翻译|继续)
抱歉[,]
# --- universal explicit marker ---
^\s*\[?(?:REFUSED|ОТКАЗ)\]?\s*$

View file

@ -0,0 +1,35 @@
# Source-language sentence-splitter abbreviations (pair-14 §4): trailing tokens after which a lone "." is an
# abbreviation, not a sentence end (case-insensitive). Per SOURCE language — a CJK source uses 。 and needs
# none; `en` is the one that guards an ASCII period. `lang<TAB>token`. Add a source = add its section.
# NOTE: the chunker currently consumes the `en` section universally (its historical en-only design); making
# the splitter consume the BOOK's source section is a shallow follow-up (thread sourceLang into SplitChunks).
# --- en ---
en mr
en mrs
en ms
en dr
en prof
en st
en jr
en sr
en vs
en no
en vol
en ch
en fig
en col
en gen
en sgt
en capt
en lt
en rev
en gov
en sen
en rep
en etc
en inc
en ltd
en co
en mt
en ave
en rd

View file

@ -0,0 +1,65 @@
# Target-language (ru) readability checker DATA (pair-14 data-out): the wordlists and DETECTION patterns the
# TARGET-GENERAL checkers/sanitizer run on ANY →ru output (regardless of source pair). Engine/target data —
# a no-langpack →ru book (the ja→ru golden) still runs these — so it is embedded here, not a book pack; the
# checker/sanitizer ALGORITHMS stay generic Go. `category<TAB>value`; value VERBATIM (a regex or a literal);
# lines are grouped by category IN ORDER (order matters for the sanitizer preamble preview). Add a target =
# add data/target-<tgt>.txt + its accessor entry.
#
# broken-word: a Russian word ending in this suffix is a mangled infinitive (no valid word ends in «-йть»).
broken_suffix йть
# yofikator homograph whitelist: е-spellings that are DISTINCT words from their ё-counterpart (все≠всё,
# небо≠нёбо, берет≠берёт …), so the inconsistency rule must NOT flag «все»+«всё» as one word two ways.
yo_homograph все
yo_homograph всех
yo_homograph всем
yo_homograph всеми
yo_homograph небо
yo_homograph небом
yo_homograph узнаем
yo_homograph узнаете
yo_homograph узнает
yo_homograph падеж
yo_homograph падежа
yo_homograph совершенный
yo_homograph совершенное
yo_homograph совершенная
yo_homograph совершенно
yo_homograph совершенны
yo_homograph чем
yo_homograph тем
yo_homograph тема
yo_homograph теме
yo_homograph темы
yo_homograph берет
yo_homograph берете
yo_homograph берета
yo_homograph осел
yo_homograph осла
yo_homograph ослы
yo_homograph ослов
yo_homograph мел
yo_homograph мела
yo_homograph лен
yo_homograph лена
yo_homograph нес
yo_homograph несем
yo_homograph несет
yo_homograph вел
yo_homograph ведет
yo_homograph ведем
# magnitude_stem: Russian magnitude WORD stem → base-10 exponent, for the number-magnitude gate (a stem
# covers [base, base+2] since an unseen multiplier can lift it two orders). `magnitude_stem<TAB>stem<TAB>exp`.
magnitude_stem тысяч 3
magnitude_stem миллион 6
magnitude_stem миллиард 9
magnitude_stem триллион 12
# sanitizer DETECTION patterns (pair-14 data-out). The sanitizer is isRuTarget-gated, so these are
# ru-target data; the strip/detect ALGORITHM stays in sanitizer.go. `sanitizer_preamble` is ORDERED
# (detectLeadingPreamble returns the FIRST match's preview). Values VERBATIM (regexes).
sanitizer_preamble (?is)^\s*(?:вот|представляю|привожу|держите)\s+[^\n:]{0,40}?(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:
sanitizer_preamble (?is)^\s*(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:
sanitizer_preamble (?is)^\s*(?:вот\s+)?перевод\s+(?:фрагмента|текста|отрывка|главы|черновика)\s*:
sanitizer_preamble (?is)^\s*ниже\s+(?:привед[её]н|представлен|дан|след)[а-яё]*\s+[^\n:]{0,40}?(?:перевод|отредактированн|текст)[а-яё]*[^\n:]{0,20}?:
sanitizer_trailing_note (?i)^[*_>\s-]{0,4}(?:(?:примечани[ея]|заметк[аи]|комментари[йяю]|пояснени[ея]|сноск[аи])\s*:|(?:примечани[ея]|заметк[аи]|комментари[йяю])\s+(?:переводчик|редактор)[а-яё]*|прим\.\s*(?:перев|ред)\.?)
sanitizer_edit_meta (?im)^[*_>\s-]{0,4}(?:(?:внесённ|внесен)[а-яё]+\s+правк|основн[а-яё]+\s+правк[а-яё]*(?:\s+для\s+справк[а-яё]+|\s*:)|что\s+(?:было\s+)?(?:изменен|исправлен)|список\s+(?:правок|изменений)|список\s+внесённых)
sanitizer_invalid_sign (?i)(?:^|\P{L})[ъь][а-яё]|[ъь][ъь]

View file

@ -0,0 +1,392 @@
package lang
import (
"embed"
"fmt"
"sort"
"strconv"
"strings"
"sync"
)
// embedded.go: language data the engine needs even for a book with NO langpack pack — engine-universal
// (a CJK numeral is a linguistic constant; a refusal phrase is engine safety behaviour) or target-generic
// (the Russian glossary-injection wire-text). It is versioned with the engine BINARY (like code), not with
// a book's pack, because a no-pack book still splits 第X章 chapters, still flags a refusal, and still injects
// a Russian glossary block (the ja→ru golden proves all three). Kept OUT of internal/pipeline so the
// data/algorithm boundary holds; embedded so it is available without a book's langpack_root. Each file is
// SECTIONED per language/target, so adding a language = add a section, removing one = delete its section.
//go:embed data/cjk-section.txt data/refusal.txt data/injection.txt data/sentence-abbrev.txt data/target-ru.txt
var embeddedFS embed.FS
// targetCheckFiles maps a TARGET language to its embedded readability-checker data file (pair-14 data-out).
// Add a target = add data/target-<tgt>.txt to the go:embed line above and a row here.
var targetCheckFiles = map[string]string{"ru": "data/target-ru.txt"}
// TargetChecks is the TARGET-language readability wordlists + DETECTION patterns (pair-14 data-out) the
// target-general checkers/sanitizer run on ANY →target output. A target with no file yields an empty value
// (HasData()==false → the consumers stay inert). Values are grouped per category in AUTHORED ORDER.
type TargetChecks struct{ byKey map[string][]string }
// List returns the authored-order values for a category ("" categories → nil).
func (t TargetChecks) List(key string) []string { return t.byKey[key] }
// HasData reports whether this target ships any checker data (used to gate the target-general checkers so a
// target without data flags nothing rather than mis-firing).
func (t TargetChecks) HasData() bool { return len(t.byKey) > 0 }
var (
targetChecksOnce sync.Once
targetChecksBy map[string]TargetChecks
)
// TargetChecksFor returns the readability-checker data for a TARGET language, parsed once from the embedded
// per-target file. Unknown/absent target → empty TargetChecks. Panics on a corrupt embed.
func TargetChecksFor(tgt string) TargetChecks {
targetChecksOnce.Do(func() {
targetChecksBy = map[string]TargetChecks{}
for lang, file := range targetCheckFiles {
byKey, err := parseOrderedByCategory(mustEmbed(file))
if err != nil {
panic(fmt.Sprintf("lang: embedded %s is corrupt: %v", file, err))
}
targetChecksBy[lang] = TargetChecks{byKey: byKey}
}
})
return targetChecksBy[tgt]
}
// parseOrderedByCategory reads `category<TAB>value` lines into category → ordered values. Value is VERBATIM
// (a regex may carry a trailing metachar), so only the whole line's \r is stripped; blank/`#`-comment lines
// (by their trimmed form) are dropped.
func parseOrderedByCategory(b []byte) (map[string][]string, error) {
out := map[string][]string{}
for i, raw := range strings.Split(string(b), "\n") {
line := strings.TrimRight(raw, "\r")
if strings.TrimSpace(line) == "" || strings.HasPrefix(strings.TrimSpace(line), "#") {
continue
}
f := strings.SplitN(line, "\t", 2)
if len(f) != 2 || strings.TrimSpace(f[0]) == "" || f[1] == "" {
return nil, fmt.Errorf("line %d: want `category<TAB>value`, got %q", i+1, line)
}
out[f[0]] = append(out[f[0]], f[1])
}
return out, nil
}
// CJKSection is the universal chapter/section-numeral inventory shared by the ingest chapter-splitter and
// the chunker heading rule (pair-14 §4 + addendum-A: ONE source, no ingest↔chunker byte-duplication). The
// heading-numeral membership test derives from Digit Unit Zero — nothing is stored twice.
type CJKSection struct {
Digit map[rune]int // 一→1 … 九→9 (incl. 两/兩→2)
Unit map[rune]int // 十→10 百→100 千→1000
Zero map[rune]bool // 零 (positional zero: value*10 in a numeral run)
// ChapterUnit membership + ChapterUnitOrdered (authored order) — ingest's detectChapterUnit iterates in
// order and breaks a tie by first-seen, so the ORDER is behaviour, not just a set.
ChapterUnit map[rune]bool
ChapterUnitOrdered []rune
// Magnitude: the myriad-scale big units (万/億/兆 → value) for the number-magnitude gate. int64 (10^12
// overflows int32). ONE source with Digit/Unit — the gate no longer keeps its own CJK numeral switch.
Magnitude map[rune]int64
headingRunes string // sorted ZeroDigitUnit, cached for the ingest chapter-numeral regex class
}
// IsNumeralRune reports whether r can be part of a CJK numeral EXPRESSION (ZeroDigitUnitMagnitude) — the
// alphabet of a magnitude candidate run (the gate adds ASCII digits itself). BigUnit returns a big unit's
// value (万/億/兆). Both read the single CJKSection source (pair-14 data-out dedup).
func (c *CJKSection) IsNumeralRune(r rune) bool {
if c.IsHeadingNumeral(r) {
return true
}
_, ok := c.Magnitude[r]
return ok
}
func (c *CJKSection) BigUnit(r rune) (int64, bool) { v, ok := c.Magnitude[r]; return v, ok }
// IsHeadingNumeral reports whether r can be part of a CJK chapter-number run (Zero Digit Unit). Arabic
// digits are handled by the caller (they are not language data).
func (c *CJKSection) IsHeadingNumeral(r rune) bool {
if c.Zero[r] {
return true
}
if _, ok := c.Digit[r]; ok {
return true
}
_, ok := c.Unit[r]
return ok
}
// HeadingNumeralClass returns the CJK heading-numeral runes (ZeroDigitUnit) as a sorted string, for use
// INSIDE a regex character class (the ingest chapter-numeral pattern). Sorted for a stable pattern; order
// inside a class is immaterial to matching. None of the runes are class metacharacters, so no escaping.
func (c *CJKSection) HeadingNumeralClass() string { return c.headingRunes }
var (
cjkOnce sync.Once
cjkVal *CJKSection
cjkErr error
)
// DefaultCJKSection returns the process-wide CJK section-numeral data, parsed once from the embedded file.
// It panics on a corrupt embed (a build-time asset the tests exercise, so a bad edit fails the suite, never
// production silently) — the parse is deterministic and input-free.
func DefaultCJKSection() *CJKSection {
cjkOnce.Do(func() { cjkVal, cjkErr = parseCJKSection(mustEmbed("data/cjk-section.txt")) })
if cjkErr != nil {
panic(fmt.Sprintf("lang: embedded cjk-section.txt is corrupt: %v", cjkErr))
}
return cjkVal
}
func parseCJKSection(b []byte) (*CJKSection, error) {
c := &CJKSection{Digit: map[rune]int{}, Unit: map[rune]int{}, Zero: map[rune]bool{}, ChapterUnit: map[rune]bool{}, Magnitude: map[rune]int64{}}
valueRune := func(i int, f []string) (rune, int, error) {
if len(f) != 3 {
return 0, 0, fmt.Errorf("line %d: %s wants `%s<TAB>rune<TAB>value`, got %q", i+1, f[0], f[0], strings.Join(f, "\t"))
}
r := []rune(f[1])
if len(r) != 1 {
return 0, 0, fmt.Errorf("line %d: key %q must be a single rune", i+1, f[1])
}
v, err := strconv.Atoi(strings.TrimSpace(f[2]))
if err != nil {
return 0, 0, fmt.Errorf("line %d: value %q: %w", i+1, f[2], err)
}
return r[0], v, nil
}
setRune := func(i int, f []string) (rune, error) {
if len(f) != 2 {
return 0, fmt.Errorf("line %d: %s wants `%s<TAB>rune`, got %q", i+1, f[0], f[0], strings.Join(f, "\t"))
}
r := []rune(f[1])
if len(r) != 1 {
return 0, fmt.Errorf("line %d: key %q must be a single rune", i+1, f[1])
}
return r[0], nil
}
for i, raw := range strings.Split(string(b), "\n") {
t := strings.TrimRight(raw, "\r")
if strings.TrimSpace(t) == "" || strings.HasPrefix(strings.TrimSpace(t), "#") {
continue
}
f := strings.Split(t, "\t")
switch f[0] {
case "digit":
r, v, err := valueRune(i, f)
if err != nil {
return nil, err
}
c.Digit[r] = v
case "unit":
r, v, err := valueRune(i, f)
if err != nil {
return nil, err
}
c.Unit[r] = v
case "zero":
r, err := setRune(i, f)
if err != nil {
return nil, err
}
c.Zero[r] = true
case "chapter_unit":
r, err := setRune(i, f)
if err != nil {
return nil, err
}
if !c.ChapterUnit[r] {
c.ChapterUnit[r] = true
c.ChapterUnitOrdered = append(c.ChapterUnitOrdered, r)
}
case "magnitude":
if len(f) != 3 {
return nil, fmt.Errorf("line %d: magnitude wants `magnitude<TAB>rune<TAB>value`, got %q", i+1, t)
}
r := []rune(f[1])
if len(r) != 1 {
return nil, fmt.Errorf("line %d: magnitude key %q must be a single rune", i+1, f[1])
}
v, err := strconv.ParseInt(strings.TrimSpace(f[2]), 10, 64)
if err != nil {
return nil, fmt.Errorf("line %d: magnitude value %q: %w", i+1, f[2], err)
}
c.Magnitude[r[0]] = v
default:
return nil, fmt.Errorf("line %d: unknown category %q (want digit|unit|zero|chapter_unit|magnitude)", i+1, f[0])
}
}
if len(c.Digit) == 0 || len(c.Unit) == 0 || len(c.Zero) == 0 || len(c.ChapterUnit) == 0 || len(c.Magnitude) == 0 {
return nil, fmt.Errorf("cjk-section needs non-empty digit, unit, zero, chapter_unit and magnitude sections")
}
// Cache the sorted heading-numeral class (ZeroDigitUnit) for the ingest regex.
var runes []rune
for r := range c.Zero {
runes = append(runes, r)
}
for r := range c.Digit {
runes = append(runes, r)
}
for r := range c.Unit {
runes = append(runes, r)
}
sort.Slice(runes, func(i, j int) bool { return runes[i] < runes[j] })
c.headingRunes = string(runes)
return c, nil
}
// InjectionTexts is the TARGET-language glossary/injection wire-text (pair-14 §2): the Russian headers and
// annotations the memory renderers emit into a request. TARGET data (a Russian header for a →ru book), so
// the renderers gate on HasData — a target with no injection texts (e.g. →en) injects NOTHING rather than a
// stray Russian block. The ja→ru golden renders these with NO book pack, so they are engine/target data.
type InjectionTexts struct {
GlossaryHeader string // "ГЛОССАРИЙ (…):" — introduces the draft glossary block
EditorHeader string // "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ …:" — introduces the editor constraint block
UnverifiedMarker string // " ⟨проверить⟩" — tags an unverified draft candidate (leading spaces significant)
GenderMale string // " (муж. — …)" DC3 gender directive (leading space significant)
GenderFemale string // " (жен. — …)"
GenderHidden string // " (пол СКРЫТ …)"
}
// HasData reports whether the target has injection texts (its header is present). The renderers use it to
// gate: a target without texts injects nothing, so a non-ru book never gets a Russian glossary block.
func (t InjectionTexts) HasData() bool { return t.GlossaryHeader != "" }
var (
injectionOnce sync.Once
injectionBy map[string]*InjectionTexts
injectionErr error
)
// InjectionTextsFor returns the injection wire-text for a TARGET language (pair-14 §2), parsed once from the
// embedded file. A target with no rows returns a zero value (HasData()==false → no injection). Panics on a
// corrupt embed. VALUE bytes are verbatim (leading spaces in gender_*/unverified_marker are significant).
func InjectionTextsFor(targetLang string) InjectionTexts {
injectionOnce.Do(func() { injectionBy, injectionErr = parseInjection(mustEmbed("data/injection.txt")) })
if injectionErr != nil {
panic(fmt.Sprintf("lang: embedded injection.txt is corrupt: %v", injectionErr))
}
if t, ok := injectionBy[targetLang]; ok {
return *t
}
return InjectionTexts{}
}
func parseInjection(b []byte) (map[string]*InjectionTexts, error) {
out := map[string]*InjectionTexts{}
for i, raw := range strings.Split(string(b), "\n") {
line := strings.TrimRight(raw, "\r")
if strings.TrimSpace(line) == "" || strings.HasPrefix(strings.TrimSpace(line), "#") {
continue
}
f := strings.SplitN(line, "\t", 3) // value (f[2]) VERBATIM — leading spaces are significant
if len(f) != 3 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" {
return nil, fmt.Errorf("line %d: want `target<TAB>key<TAB>value`, got %q", i+1, line)
}
tgt, key, val := f[0], f[1], f[2]
if out[tgt] == nil {
out[tgt] = &InjectionTexts{}
}
t := out[tgt]
switch key {
case "glossary_header":
t.GlossaryHeader = val
case "editor_header":
t.EditorHeader = val
case "unverified_marker":
t.UnverifiedMarker = val
case "gender_male":
t.GenderMale = val
case "gender_female":
t.GenderFemale = val
case "gender_hidden":
t.GenderHidden = val
default:
return nil, fmt.Errorf("line %d: unknown injection key %q", i+1, key)
}
}
return out, nil
}
var (
refusalOnce sync.Once
refusalVal []string
)
// RefusalPatterns returns the universal refusal-blacklist regex patterns (pair-14 §3), parsed once from the
// embedded sectioned file in authored order. Engine safety data: a model can refuse in any language for any
// pair, so this is not book-pack data — the caller joins the patterns into one case-insensitive regex. Each
// non-comment line is ONE pattern, taken verbatim (a regex may contain leading/trailing metacharacters, so
// it is NOT trimmed beyond the line's \r). Panics on a corrupt embed.
func RefusalPatterns() []string {
refusalOnce.Do(func() { refusalVal = embeddedLines(mustEmbed("data/refusal.txt")) })
return refusalVal
}
// embeddedLines returns the non-comment, non-blank lines of an embedded file VERBATIM (only the trailing \r
// is stripped; the content is NOT trimmed — significant for a regex pattern). A blank/`#`-comment line is
// dropped by its TRIMMED form, but a kept line keeps its own leading/trailing bytes.
func embeddedLines(b []byte) []string {
var out []string
for _, raw := range strings.Split(string(b), "\n") {
s := strings.TrimRight(raw, "\r")
if strings.TrimSpace(s) == "" || strings.HasPrefix(strings.TrimSpace(s), "#") {
continue
}
out = append(out, s)
}
return out
}
var (
abbrevOnce sync.Once
abbrevBy map[string]map[string]bool
)
// SentenceAbbrev returns the lower-cased sentence-splitter abbreviation SET for a SOURCE language (pair-14
// §4), parsed once from the embedded sectioned file. An unknown/absent language returns an empty set (a
// source with no ASCII-period abbreviations — e.g. a CJK source using 。). Panics on a corrupt embed.
func SentenceAbbrev(srcLang string) map[string]bool {
abbrevOnce.Do(func() {
var err error
abbrevBy, err = parseSectionedSet(mustEmbed("data/sentence-abbrev.txt"))
if err != nil {
panic(fmt.Sprintf("lang: embedded sentence-abbrev.txt is corrupt: %v", err))
}
})
if m, ok := abbrevBy[srcLang]; ok {
return m
}
return map[string]bool{}
}
// parseSectionedSet reads `key<TAB>member` lines into per-key sets (keys lower-cased on read is NOT done
// here — the key is the language tag; members are stored verbatim). Used for the sentence-abbrev table.
func parseSectionedSet(b []byte) (map[string]map[string]bool, error) {
out := map[string]map[string]bool{}
for i, raw := range strings.Split(string(b), "\n") {
t := strings.TrimRight(raw, "\r")
if strings.TrimSpace(t) == "" || strings.HasPrefix(strings.TrimSpace(t), "#") {
continue
}
f := strings.Split(t, "\t")
if len(f) != 2 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" {
return nil, fmt.Errorf("line %d: want `key<TAB>member`, got %q", i+1, t)
}
key := strings.TrimSpace(f[0])
if out[key] == nil {
out[key] = map[string]bool{}
}
out[key][strings.TrimSpace(f[1])] = true
}
return out, nil
}
func mustEmbed(name string) []byte {
b, err := embeddedFS.ReadFile(name)
if err != nil {
panic(fmt.Sprintf("lang: missing embedded asset %q: %v", name, err))
}
return b
}

View file

@ -21,6 +21,7 @@ import (
"fmt"
"os"
"path/filepath"
"strconv"
"strings"
)
@ -28,7 +29,9 @@ import (
// a schema change (a new file, a format change). The DATA content is versioned separately by hashing the
// files into Version(), so editing a table also invalidates — you cannot forget to bump a version when
// you change the data, because the data's bytes ARE the version (the memnorm.go drift-proofing).
const packAlgoVersion = "langpack-v1"
// v2 (pair-14): new pair/source files (title-formant, sentence-terminator, palladius-phonotactics, dc-checkers)
// + new formats (the pattern/category rows, the generic Palladius parser) — a schema change, so the tag bumps.
const packAlgoVersion = "langpack-v2"
// Pack is a loaded, versioned language-data pack for one source→target pair. Fields are the DATA the
// pipeline algorithms read; the zero value is unusable (load via Load). Maps are membership sets / lookup
@ -46,12 +49,25 @@ type Pack struct {
GradePrefix map[rune]bool // grade/stem prefix chars 甲乙丙…
Numeral map[rune]bool // CJK numeral chars
AliasParticle map[rune]bool // trailing-particle set marking a boundary fragment
// TitleFormant: the rank/measure formant chars a miner classifies as a TITLE (patterns.formant_type
// 等/转/阶 → title). Pair-14: moved out of the pipeline formant switch, sitting beside TopoSuffix (its
// place-side neighbour). A rune SET.
TitleFormant map[rune]bool
// SentenceTerminator: the source sentence-ending marks the alias miner splits on (。!?, alias.
// cooccur_same_sentence). Pair-14: moved out of the pipeline literal. A rune SET; the miner adds "\n"
// (a structural newline, not a language mark) itself.
SentenceTerminator map[rune]bool
// <src>-<tgt> pair transliteration (configs/langpacks/<pair>/) — the Palladius (Палладий) table.
PalladiusInitials map[string]string
PalladiusFinals map[string]string
PalladiusYW map[string]string
PalladiusSpecialI map[string]string
// Palladius is the <src>-<tgt> transliteration table (configs/langpacks/<pair>/palladius.txt +
// palladius-phonotactics.txt) — the pinyin→Cyrillic maps and the syllable-generator's phonotactic
// constraints, as ONE typed value (owner addendum 24.07: a struct, not four+three parallel Pack fields).
Palladius Palladius
// DCCheckers is the OPTIONAL pair data for the WS5 defect-class checkers (configs/langpacks/<pair>/
// dc-checkers.txt, checkers_zh_ru.go). nil when the pair ships no file — the checker ALGORITHMS then read
// empty tables and fire 0 (a pair that does not opt in is never flagged). DATA only (pair-14 §6: the
// LOOKUP TABLES move out of the pipeline; the regex DETECTION patterns stay as the checker algorithm).
DCCheckers *DCCheckerData
// Heading is the OPTIONAL chapter-heading rule (configs/langpacks/<pair>/heading.txt). nil when the pair
// carries no heading.txt — the chapter-title feature is then inert (the chunker keeps the source header
@ -63,6 +79,37 @@ type Pack struct {
version string
}
// Palladius is the pair transliteration table: the pinyin→Cyrillic maps and the syllable-generator's
// phonotactic constraint sets (owner addendum 24.07 — one typed struct instead of four+three parallel Pack
// fields). Populated from palladius.txt (Initials/Finals/YW/SpecialI) + palladius-phonotactics.txt
// (Retroflex/VFinal/VFinalInitial) by the generic category parser; which categories are REQUIRED is enforced
// by validate(), not the parser (an unknown category is collected, never a parse error).
type Palladius struct {
Initials, Finals, YW, SpecialI map[string]string // pinyin→Cyrillic
Retroflex, VFinal, VFinalInitial map[string]bool // phonotactic constraint sets (pinyin tokens)
}
// newPalladius returns a Palladius with all tables allocated (so merge/union is nil-safe).
func newPalladius() Palladius {
return Palladius{
Initials: map[string]string{}, Finals: map[string]string{}, YW: map[string]string{}, SpecialI: map[string]string{},
Retroflex: map[string]bool{}, VFinal: map[string]bool{}, VFinalInitial: map[string]bool{},
}
}
// DCCheckerData is the per-pair lookup data the WS5 defect-class checkers read (checkers_zh_ru.go). The
// checker DETECTION regexes stay in the pipeline as the pair-scoped algorithm (§12.2); only the LOOKUP
// TABLES live here as data. Parsed from dc-checkers.txt.
type DCCheckerData struct {
Numeral map[rune]int // DC1: small CJK count → value (一→1 … 十→10, incl. 两→2)
RuHours map[string]int // DC1: Russian hours-count word → value (один→1 … шесть→6)
RegisterNeg []string // DC6: register negative-list lexemes (терем + case forms), authored order
// Patterns are the DETECTION patterns as DATA (pair-14 data-out): key → verbatim regex or literal probe.
// The `*_re` keys are compiled by the consumer; the rest are literal strings.Contains probes. A pair that
// ships no pattern for a key runs that sub-checker inert.
Patterns map[string]string
}
// HeadingRule is the per-pair data for the chapter-title policy: instead of letting the model render a
// chapter heading (which drifted to «Раздел 2» / «Первая глава» / an orphaned « :» across models), the
// chunker detects a source header (Marker + a numeral + a Unit rune), strips it from the model input, and
@ -78,13 +125,18 @@ type HeadingRule struct {
func (p *Pack) Version() string { return p.version }
// srcFiles are the source-morphology files (under configs/langpacks/<src>/), read in this fixed order.
// SCOPE (honest): this manifest is the zh-family NAME-MINER's morphology schema (百家姓 surnames, 甲乙丙 grade
// prefixes, 等/转/阶 title-formants, CJK numerals, …), and validate() forbids an empty table. So "add a pair =
// drop a directory, no recompile" holds for a source in THIS family (a second CJK source drops its files); a
// different source family that wants mining needs new Pack fields + parse + miner channels, not just a dir.
var srcFiles = []string{
"surnames-single.txt", "surnames-compound.txt", "title-suffix.txt", "ordinal-title.txt",
"rank-word.txt", "topo-suffix.txt", "grade-prefix.txt", "numeral.txt", "alias-particle.txt",
"title-formant.txt", "sentence-terminator.txt",
}
// pairFiles are the pair-transliteration files (under configs/langpacks/<pair>/).
var pairFiles = []string{"palladius.txt"}
var pairFiles = []string{"palladius.txt", "palladius-phonotactics.txt"}
// Load resolves and reads the pack for (sourceLang, targetLang) from root (e.g. "configs/langpacks"): the
// source-morphology files under root/<src>/ and the pair-transliteration files under root/<src>-<tgt>/.
@ -93,7 +145,7 @@ var pairFiles = []string{"palladius.txt"}
// files in a fixed order, so it is deterministic and drift-proof.
func Load(root, sourceLang, targetLang string) (*Pack, error) {
pair := sourceLang + "-" + targetLang
p := &Pack{Pair: pair}
p := &Pack{Pair: pair, Palladius: newPalladius()}
h := sha256.New()
h.Write([]byte(packAlgoVersion))
@ -144,6 +196,21 @@ func Load(root, sourceLang, targetLang string) (*Pack, error) {
p.Heading = hr
}
// Optional per-pair DC-checker tables (pair-14 §6). ABSENT → nil, the checkers read empty tables and fire
// 0 (a pair that does not opt in is never re-billed / flagged); PRESENT → bytes fold into the hash and it
// is parsed; CORRUPT → fail loud. Same optional contract as heading.txt.
if db, ok, derr := readOptional(root, pair, "dc-checkers.txt"); derr != nil {
return nil, fmt.Errorf("langpack %q dc-checkers.txt: %w", pair, derr)
} else if ok {
h.Write([]byte("\x00" + pair + "/dc-checkers.txt\x00"))
h.Write(db)
dc, perr := parseDCCheckers(db)
if perr != nil {
return nil, fmt.Errorf("langpack %q dc-checkers.txt: %w", pair, perr)
}
p.DCCheckers = dc
}
if err := p.validate(); err != nil {
return nil, fmt.Errorf("langpack %q: %w", pair, err)
}
@ -151,6 +218,138 @@ func Load(root, sourceLang, targetLang string) (*Pack, error) {
return p, nil
}
// LoadWithOverlay loads the shared pair pack (Load) and then UNIONS a book-scoped OVERLAY on top: a book's
// PRIVATE canon (a clan surname 古月, a sect term) that belongs to ONE book, not the shared pair langpack
// (pair-14 §1 — a book term in the shared pair layer is a leak). overlayRoot has the same layout as root
// (<src>/… + <pair>/…); ONLY the files a book chooses to ship are present, each read OPTIONALLY and UNIONED
// (additive — sets gain members, ordered slices append; an overlay never removes a base entry). The overlay
// bytes fold into Version(), so a book-canon edit is a loud --resnapshot for THAT book while the shared pack
// stays byte-stable. overlayRoot == "" ⇒ identical to Load (same *Pack, same Version).
func LoadWithOverlay(root, src, tgt, overlayRoot string) (*Pack, error) {
p, err := Load(root, src, tgt)
if err != nil {
return nil, err
}
if overlayRoot == "" {
return p, nil
}
pair := src + "-" + tgt
// FAIL LOUD on a misnamed overlay file (review finding, pair-14 scale lens): the fold loop below reads
// ONLY the manifest names, so a typo (surname-compound.txt) or a non-mergeable file (heading.txt) shipped
// in an overlay would be SILENTLY ignored — the book's private canon never reaches the miner, recall
// degrades with no load-time signal. Enumerate the overlay dirs and refuse any unexpected file, keeping
// the "never silently empty / drift-proof" guarantee the shared loader makes.
mergeable := map[string]bool{}
for _, n := range srcFiles {
mergeable[src+"/"+n] = true
}
for _, n := range pairFiles {
mergeable[pair+"/"+n] = true
}
for _, dir := range []string{src, pair} {
names, derr := overlayDirFiles(filepath.Join(overlayRoot, dir))
if derr != nil {
return nil, fmt.Errorf("langpack %q overlay %s: %w", pair, dir, derr)
}
for _, n := range names {
if !mergeable[dir+"/"+n] {
return nil, fmt.Errorf("langpack %q overlay: unexpected file %s/%s — an overlay merges only the source/pair manifest files (a typo, or a non-mergeable file like heading.txt/dc-checkers.txt, would be silently ignored)", pair, dir, n)
}
}
}
// Seed a fresh hash with the base version (which already uniquely encodes every base byte), then fold the
// overlay files in a fixed order — deterministic and drift-proof (edit the overlay → Version() moves).
h := sha256.New()
h.Write([]byte(p.version))
merged := false
fold := func(dir, name string, pairFile bool) error {
b, ok, rerr := readOptional(overlayRoot, dir, name)
if rerr != nil {
return fmt.Errorf("langpack %q overlay %s/%s: %w", pair, dir, name, rerr)
}
if !ok {
return nil
}
h.Write([]byte("\x00" + dir + "/" + name + "\x00"))
h.Write(b)
merged = true
if pairFile {
return p.mergePair(name, b)
}
return p.mergeSrc(name, b)
}
for _, name := range srcFiles {
if err := fold(src, name, false); err != nil {
return nil, err
}
}
for _, name := range pairFiles {
if err := fold(pair, name, true); err != nil {
return nil, err
}
}
if merged {
p.version = packAlgoVersion + "-x" + hex.EncodeToString(h.Sum(nil))[:12]
}
return p, nil
}
// mergeSrc unions an overlay source file into the loaded pack (additive; see LoadWithOverlay). Rune/string
// SETS gain members; ordered slices append (a book's extra title/rank tokens follow the shared ones).
func (p *Pack) mergeSrc(name string, b []byte) error {
switch name {
case "surnames-single.txt":
unionRuneSet(p.SurnamesSingle, runeSet(b))
case "surnames-compound.txt":
unionStringSet(p.SurnamesCompound, stringSet(b))
case "title-suffix.txt":
p.TitleSuffix = append(p.TitleSuffix, lines(b)...)
case "ordinal-title.txt":
p.OrdinalTitle = append(p.OrdinalTitle, lines(b)...)
case "rank-word.txt":
p.RankWord = append(p.RankWord, lines(b)...)
case "topo-suffix.txt":
unionRuneSet(p.TopoSuffix, runeSet(b))
case "grade-prefix.txt":
unionRuneSet(p.GradePrefix, runeSet(b))
case "numeral.txt":
unionRuneSet(p.Numeral, runeSet(b))
case "alias-particle.txt":
unionRuneSet(p.AliasParticle, runeSet(b))
case "title-formant.txt":
unionRuneSet(p.TitleFormant, runeSet(b))
case "sentence-terminator.txt":
unionRuneSet(p.SentenceTerminator, runeSet(b))
default:
return fmt.Errorf("unknown source file")
}
return nil
}
// mergePair unions an overlay pair file into the loaded pack (additive). Both pair files carry Palladius
// categories, parsed generically and unioned into p.Palladius (same path as assignPair).
func (p *Pack) mergePair(name string, b []byte) error {
return p.assignPair(name, b)
}
func unionRuneSet(dst, src map[rune]bool) {
for k := range src {
dst[k] = true
}
}
func unionStringSet(dst, src map[string]bool) {
for k := range src {
dst[k] = true
}
}
func unionStringMap(dst, src map[string]string) {
for k, v := range src {
dst[k] = v
}
}
// validate makes the "never silently empty" contract real: a present-but-empty or comment-only data file
// (a fat-fingered edit once packs are hand-authored at R1) parses to an empty table with no error and would
// silently disable a miner channel — recall degradation with no load-time signal. Every required table must
@ -172,10 +371,16 @@ func (p *Pack) validate() error {
req("grade-prefix", len(p.GradePrefix))
req("numeral", len(p.Numeral))
req("alias-particle", len(p.AliasParticle))
req("palladius/initials", len(p.PalladiusInitials))
req("palladius/finals", len(p.PalladiusFinals))
req("palladius/yw", len(p.PalladiusYW))
req("palladius/special_i", len(p.PalladiusSpecialI))
req("title-formant", len(p.TitleFormant))
req("sentence-terminator", len(p.SentenceTerminator))
// The Palladius REQUIRED-category list lives HERE (the consumer), not in the generic parser (addendum).
req("palladius/initials", len(p.Palladius.Initials))
req("palladius/finals", len(p.Palladius.Finals))
req("palladius/yw", len(p.Palladius.YW))
req("palladius/special_i", len(p.Palladius.SpecialI))
req("palladius-phonotactics/retroflex", len(p.Palladius.Retroflex))
req("palladius-phonotactics/vfinal", len(p.Palladius.VFinal))
req("palladius-phonotactics/vfinal_initial", len(p.Palladius.VFinalInitial))
if len(empty) > 0 {
return fmt.Errorf("empty required table(s) %s — a present-but-empty/comment-only data file is a corrupt pack, not a valid one", strings.Join(empty, ", "))
}
@ -202,26 +407,49 @@ func (p *Pack) assignSrc(name string, b []byte) error {
p.Numeral = runeSet(b)
case "alias-particle.txt":
p.AliasParticle = runeSet(b)
case "title-formant.txt":
p.TitleFormant = runeSet(b)
case "sentence-terminator.txt":
p.SentenceTerminator = runeSet(b)
default:
return fmt.Errorf("unknown source file")
}
return nil
}
// assignPair reads a pair file into p.Palladius. Both pair files (palladius.txt, palladius-phonotactics.txt)
// carry `category<TAB>…` rows, parsed by ONE generic category parser (owner addendum 24.07); the known
// Palladius categories are UNIONED into p.Palladius and an unknown category is simply ignored, NEVER a parse
// error — which categories are REQUIRED is enforced downstream by validate(), not here.
func (p *Pack) assignPair(name string, b []byte) error {
switch name {
case "palladius.txt":
ini, fin, yw, si, err := parsePalladius(b)
if err != nil {
return err
}
p.PalladiusInitials, p.PalladiusFinals, p.PalladiusYW, p.PalladiusSpecialI = ini, fin, yw, si
default:
return fmt.Errorf("unknown pair file")
cats, err := parseCategoryRows(b)
if err != nil {
return err
}
p.Palladius.merge(cats)
return nil
}
// overlayDirFiles lists the regular-file names directly in dir (an overlay's <src> or <pair> subdir). A
// MISSING dir is fine (returns nil — a book may overlay only sources or only the pair). Any other read
// error is loud. Nested dirs are ignored (only top-level manifest files are mergeable).
func overlayDirFiles(dir string) ([]string, error) {
ents, err := os.ReadDir(dir)
if err != nil {
if os.IsNotExist(err) {
return nil, nil
}
return nil, err
}
var out []string
for _, e := range ents {
if !e.IsDir() {
out = append(out, e.Name())
}
}
return out, nil
}
// readOptional reads an OPTIONAL pack file. A missing file returns (nil, false, nil) — the feature it
// backs is simply inert — while any OTHER read error (permission, a directory) is a loud failure; a
// present file returns (bytes, true, nil). Used for the pack-13 heading rule, which a pair opts into.
@ -314,34 +542,116 @@ func lines(b []byte) []string {
return out
}
// parsePalladius reads the `category<TAB>pinyin<TAB>cyrillic` table into the four maps. It scans the raw
// lines directly (not contentLines) so an error names the PHYSICAL file line — the point of the diagnostic
// is to send a human editing the table to the right line.
func parsePalladius(b []byte) (ini, fin, yw, si map[string]string, err error) {
ini, fin, yw, si = map[string]string{}, map[string]string{}, map[string]string{}, map[string]string{}
// parseCategoryRows is the GENERIC Palladius category parser (owner addendum 24.07): it reads a pair file's
// `category<TAB>key[<TAB>value]` rows into category → key → value, WITHOUT a per-category switch and WITHOUT
// treating an unknown category as an error (which categories are required is the CONSUMER's call — validate()).
// A 3-field row (initials/finals/yw/special_i) stores key→cyrillic; a 2-field row (retroflex/vfinal/…) stores
// key→"" (a set member). Scans raw lines so an error names the PHYSICAL file line; the only parse errors are a
// bad field count / an empty key.
func parseCategoryRows(b []byte) (map[string]map[string]string, error) {
cats := map[string]map[string]string{}
for i, raw := range strings.Split(string(b), "\n") {
t := strings.TrimSpace(strings.TrimRight(raw, "\r"))
if t == "" || strings.HasPrefix(t, "#") {
continue
}
f := strings.Split(t, "\t")
if len(f) != 3 {
return nil, nil, nil, nil, fmt.Errorf("line %d: want 3 tab-separated fields, got %d (%q)", i+1, len(f), t)
if len(f) < 2 || len(f) > 3 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" {
return nil, fmt.Errorf("line %d: want `category<TAB>key[<TAB>value]`, got %q", i+1, t)
}
if cats[f[0]] == nil {
cats[f[0]] = map[string]string{}
}
val := "" // a 2-field row is a set member (value "")
if len(f) == 3 {
val = f[2]
}
cats[f[0]][f[1]] = val
}
return cats, nil
}
// merge unions parsed categories into the Palladius table. KNOWN categories populate the typed fields (a map
// category takes key→cyrillic, a set category takes its keys); an UNKNOWN category is simply not consumed —
// validate() enforces that every REQUIRED category ended up non-empty.
func (pal *Palladius) merge(cats map[string]map[string]string) {
unionStringMap(pal.Initials, cats["initials"])
unionStringMap(pal.Finals, cats["finals"])
unionStringMap(pal.YW, cats["yw"])
unionStringMap(pal.SpecialI, cats["special_i"])
unionKeysAsSet(pal.Retroflex, cats["retroflex"])
unionKeysAsSet(pal.VFinal, cats["vfinal"])
unionKeysAsSet(pal.VFinalInitial, cats["vfinal_initial"])
}
// unionKeysAsSet adds the KEYS of a parsed category (a 2-field set) to dst.
func unionKeysAsSet(dst map[string]bool, src map[string]string) {
for k := range src {
dst[k] = true
}
}
// parseDCCheckers reads the DC-checker pair tables from category-keyed lines (checkers_zh_ru.go data):
//
// cjk_numeral<TAB>rune<TAB>value | ru_hour<TAB>word<TAB>value | register_neg<TAB>lexeme
//
// Scans raw lines so an error names the PHYSICAL file line. Fail-loud on a bad field count / non-integer
// value / unknown category (a malformed table is a corrupt pack). RegisterNeg keeps its authored order.
func parseDCCheckers(b []byte) (*DCCheckerData, error) {
d := &DCCheckerData{Numeral: map[rune]int{}, RuHours: map[string]int{}, Patterns: map[string]string{}}
for i, raw := range strings.Split(string(b), "\n") {
// A `pattern` line's VALUE is verbatim (a regex may carry trailing metachars), so only \r is stripped
// from the whole line, not the value; other categories tolerate the trimmed form.
line := strings.TrimRight(raw, "\r")
t := strings.TrimSpace(line)
if t == "" || strings.HasPrefix(t, "#") {
continue
}
if strings.HasPrefix(line, "pattern\t") {
f := strings.SplitN(line, "\t", 3) // value (f[2]) VERBATIM
if len(f) != 3 || strings.TrimSpace(f[1]) == "" || f[2] == "" {
return nil, fmt.Errorf("line %d: pattern wants `pattern<TAB>key<TAB>value`, got %q", i+1, line)
}
d.Patterns[f[1]] = f[2]
continue
}
f := strings.Split(t, "\t")
switch f[0] {
case "initials":
ini[f[1]] = f[2]
case "finals":
fin[f[1]] = f[2]
case "yw":
yw[f[1]] = f[2]
case "special_i":
si[f[1]] = f[2]
case "cjk_numeral":
if len(f) != 3 {
return nil, fmt.Errorf("line %d: cjk_numeral wants `cjk_numeral<TAB>rune<TAB>value`, got %q", i+1, t)
}
r := []rune(f[1])
if len(r) != 1 {
return nil, fmt.Errorf("line %d: cjk_numeral key %q must be a single rune", i+1, f[1])
}
v, err := strconv.Atoi(strings.TrimSpace(f[2]))
if err != nil {
return nil, fmt.Errorf("line %d: cjk_numeral value %q: %w", i+1, f[2], err)
}
d.Numeral[r[0]] = v
case "ru_hour":
if len(f) != 3 {
return nil, fmt.Errorf("line %d: ru_hour wants `ru_hour<TAB>word<TAB>value`, got %q", i+1, t)
}
v, err := strconv.Atoi(strings.TrimSpace(f[2]))
if err != nil {
return nil, fmt.Errorf("line %d: ru_hour value %q: %w", i+1, f[2], err)
}
d.RuHours[f[1]] = v
case "register_neg":
if len(f) != 2 || strings.TrimSpace(f[1]) == "" {
return nil, fmt.Errorf("line %d: register_neg wants `register_neg<TAB>lexeme`, got %q", i+1, t)
}
d.RegisterNeg = append(d.RegisterNeg, f[1])
default:
return nil, nil, nil, nil, fmt.Errorf("line %d: unknown category %q", i+1, f[0])
return nil, fmt.Errorf("line %d: unknown category %q (want cjk_numeral|ru_hour|register_neg)", i+1, f[0])
}
}
return ini, fin, yw, si, nil
if len(d.Numeral) == 0 || len(d.RuHours) == 0 || len(d.RegisterNeg) == 0 {
return nil, fmt.Errorf("dc-checkers needs non-empty cjk_numeral, ru_hour and register_neg sections")
}
return d, nil
}
// contentLines splits into lines, dropping '#'-comment and blank lines.

View file

@ -29,8 +29,11 @@ func TestLoadResolvesRealZhRu(t *testing.T) {
if p.SurnamesSingle['凝'] {
t.Error("surnames-single must NOT contain 凝 (discarded — not a surname)")
}
if !p.SurnamesCompound["古月"] || !p.SurnamesCompound["欧阳"] {
t.Error("surnames-compound missing 古月/欧阳")
if !p.SurnamesCompound["欧阳"] {
t.Error("surnames-compound missing 欧阳")
}
if p.SurnamesCompound["古月"] {
t.Error("surnames-compound must NOT contain 古月 (pair-14 §1: a 蛊真人 clan, moved to the book overlay)")
}
if len(p.TitleSuffix) == 0 || p.TitleSuffix[0] != "公子" {
t.Errorf("title-suffix order not preserved: %v", p.TitleSuffix)
@ -38,12 +41,22 @@ func TestLoadResolvesRealZhRu(t *testing.T) {
if !p.TopoSuffix['山'] || !p.Numeral['三'] || !p.GradePrefix['甲'] || !p.AliasParticle['的'] {
t.Error("rune-set membership missing an expected char (山/三/甲/的)")
}
// Pair transliteration spot-check.
if p.PalladiusInitials["b"] != "б" || p.PalladiusInitials["zh"] != "чж" {
t.Errorf("palladius initials wrong: b=%q zh=%q", p.PalladiusInitials["b"], p.PalladiusInitials["zh"])
// pair-14 moves: title-formant + sentence-terminator rune sets, and the Palladius phonotactic sets.
if !p.TitleFormant['等'] || !p.TitleFormant['转'] || !p.TitleFormant['阶'] {
t.Error("title-formant missing 等/转/阶")
}
if p.PalladiusSpecialI["zhi"] != "чжи" {
t.Errorf("palladius special_i zhi = %q, want чжи", p.PalladiusSpecialI["zhi"])
if !p.SentenceTerminator['。'] || !p.SentenceTerminator[''] || !p.SentenceTerminator[''] {
t.Error("sentence-terminator missing 。//")
}
if !p.Palladius.Retroflex["zh"] || !p.Palladius.VFinal["v"] || !p.Palladius.VFinalInitial["j"] {
t.Error("palladius phonotactics missing retroflex zh / vfinal v / vfinal_initial j")
}
// Pair transliteration spot-check.
if p.Palladius.Initials["b"] != "б" || p.Palladius.Initials["zh"] != "чж" {
t.Errorf("palladius initials wrong: b=%q zh=%q", p.Palladius.Initials["b"], p.Palladius.Initials["zh"])
}
if p.Palladius.SpecialI["zhi"] != "чжи" {
t.Errorf("palladius special_i zhi = %q, want чжи", p.Palladius.SpecialI["zhi"])
}
if !strings.HasPrefix(p.Version(), packAlgoVersion+"-") {
t.Errorf("Version() = %q, want %s-<hash>", p.Version(), packAlgoVersion)
@ -82,6 +95,63 @@ func TestRoutesByPairToDifferentBytes(t *testing.T) {
}
}
// TestBookOverlayUnionsAndReVersions pins the pair-14 §1 book-overlay contract: an overlay UNIONS its
// private canon onto the shared pack (additive — the base entries survive, the overlay entry is added) and
// SHIFTS Version() (a book-canon edit is a loud --resnapshot for that book), while a no-overlay Load is
// byte-identical to before. Guards the exact mechanism the miner's {古月:22} parity rides.
func TestBookOverlayUnionsAndReVersions(t *testing.T) {
base, err := Load(realRoot, "zh", "ru")
if err != nil {
t.Fatalf("load base: %v", err)
}
if base.SurnamesCompound["古月"] {
t.Fatal("shared pack must not carry 古月 (it is a book clan)")
}
overlay := t.TempDir()
if err := os.MkdirAll(filepath.Join(overlay, "zh"), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(overlay, "zh", "surnames-compound.txt"), []byte("# book canon\n古月\n"), 0o644); err != nil {
t.Fatal(err)
}
ext, err := LoadWithOverlay(realRoot, "zh", "ru", overlay)
if err != nil {
t.Fatalf("load with overlay: %v", err)
}
if !ext.SurnamesCompound["古月"] {
t.Error("overlay must add 古月 to the effective pack")
}
if !ext.SurnamesCompound["欧阳"] {
t.Error("overlay must UNION (keep the shared 欧阳), not replace")
}
if ext.Version() == base.Version() {
t.Errorf("overlay must shift Version() (base=%s overlay=%s)", base.Version(), ext.Version())
}
// An empty overlayRoot is exactly Load — same version, no re-hash.
same, err := LoadWithOverlay(realRoot, "zh", "ru", "")
if err != nil {
t.Fatal(err)
}
if same.Version() != base.Version() {
t.Errorf("empty overlay must equal Load (%s != %s)", same.Version(), base.Version())
}
// A MISNAMED overlay file must FAIL LOUD, not be silently ignored (review finding, scale lens): a typo'd
// canon file that never reaches the miner would degrade recall with no signal.
bad := t.TempDir()
if err := os.MkdirAll(filepath.Join(bad, "zh"), 0o755); err != nil {
t.Fatal(err)
}
if err := os.WriteFile(filepath.Join(bad, "zh", "surname-compound.txt"), []byte("古月\n"), 0o644); err != nil { // typo: missing 's'
t.Fatal(err)
}
if _, err := LoadWithOverlay(realRoot, "zh", "ru", bad); err == nil {
t.Error("a misnamed overlay file must fail loud (silently-ignored canon would degrade recall)")
} else if !strings.Contains(err.Error(), "surname-compound.txt") {
t.Errorf("error must name the unexpected file, got: %v", err)
}
}
// TestFailsLoudOnMissingPair pins the fail-loud contract: a pair with no pack directory errors, naming the
// missing file — never a silent empty pack (mirrors the prompt seam's PromptPathFor fail-loud). This is
// the load-time invariant a live consumer relies on (fail before billing).
@ -120,16 +190,19 @@ func writeSyntheticPack(t *testing.T, root, src, tgt string) {
t.Helper()
pair := src + "-" + tgt
files := map[string]string{
filepath.Join(src, "surnames-single.txt"): "# synthetic\n甴甶甹\n",
filepath.Join(src, "surnames-compound.txt"): "# synthetic\n甲乙\n",
filepath.Join(src, "title-suffix.txt"): "# synthetic\n阁下\n",
filepath.Join(src, "ordinal-title.txt"): "# synthetic\n第甲\n",
filepath.Join(src, "rank-word.txt"): "# synthetic\n級\n",
filepath.Join(src, "topo-suffix.txt"): "# synthetic\n峰\n",
filepath.Join(src, "grade-prefix.txt"): "# synthetic\n子丑\n",
filepath.Join(src, "numeral.txt"): "# synthetic\n壹貳\n",
filepath.Join(src, "alias-particle.txt"): "# synthetic\n之乎\n",
filepath.Join(pair, "palladius.txt"): "# synthetic\ninitials\tb\tб\nfinals\ta\tа\nyw\tyi\tи\nspecial_i\tzhi\tчжи\n",
filepath.Join(src, "surnames-single.txt"): "# synthetic\n甴甶甹\n",
filepath.Join(src, "surnames-compound.txt"): "# synthetic\n甲乙\n",
filepath.Join(src, "title-suffix.txt"): "# synthetic\n阁下\n",
filepath.Join(src, "ordinal-title.txt"): "# synthetic\n第甲\n",
filepath.Join(src, "rank-word.txt"): "# synthetic\n級\n",
filepath.Join(src, "topo-suffix.txt"): "# synthetic\n峰\n",
filepath.Join(src, "grade-prefix.txt"): "# synthetic\n子丑\n",
filepath.Join(src, "numeral.txt"): "# synthetic\n壹貳\n",
filepath.Join(src, "alias-particle.txt"): "# synthetic\n之乎\n",
filepath.Join(src, "title-formant.txt"): "# synthetic\n甼\n",
filepath.Join(src, "sentence-terminator.txt"): "# synthetic\n。\n",
filepath.Join(pair, "palladius.txt"): "# synthetic\ninitials\tb\tб\nfinals\ta\tа\nyw\tyi\tи\nspecial_i\tzhi\tчжи\n",
filepath.Join(pair, "palladius-phonotactics.txt"): "# synthetic\nretroflex\tzh\nvfinal\tv\nvfinal_initial\tj\n",
}
for rel, body := range files {
p := filepath.Join(root, rel)

View file

@ -5,6 +5,8 @@ import (
"sort"
"strings"
"unicode"
"textmachine/backend/internal/lang"
)
// cheapgates.go: four cheap, deterministic post-check flaggers on the FINAL chunk text (04-unhappy
@ -64,6 +66,11 @@ type cheapGateConfig struct {
// OPT-IN observability flagger (draft→final length collapse + number drift) folded into this
// result. Off → the two regression fields stay 0 and the output is byte-identical to before.
regressionEnabled bool
// checkers is the compiled WS5/pack-13 checker spec (pair-14 data-out): the pair's DETECTION patterns +
// lookup tables (from the pair langpack) plus the target-general lists (from embedded target data),
// resolved once per run. nil for a bare config or a book with no data → the language-specific checkers
// run inert (fire 0, the no-pack golden path); the general Latin-residue check needs no spec.
checkers *dcCheckers
}
// cheapGateResult is the per-chunk outcome: a count per flagger plus human-readable detail lines
@ -103,33 +110,35 @@ func runCheapGates(source, draft, final string, cfg cheapGateConfig) cheapGateRe
var r cheapGateResult
n, det := lintDialogueDash(final)
r.DialogueDash, r.Detail = n, append(r.Detail, det...)
n, det = lintYofikation(final, cfg.yoPolicy)
n, det = cfg.checkers.lintYofikation(final, cfg.yoPolicy)
r.YoInconsistent = n
r.Detail = append(r.Detail, det...)
n, det = lintTranslitInterjections(final, cfg.allowlist)
r.TranslitInterj = n
r.Detail = append(r.Detail, det...)
n, det = lintNumberMagnitude(source, final)
n, det = cfg.checkers.lintNumberMagnitude(source, final)
r.NumberMagnitude = n
r.Detail = append(r.Detail, det...)
// WS5 defect-class checkers (DC1/DC2/DC6, checkers_zh_ru.go) — src↔target observability flaggers.
n, det = lintTimeUnits(source, final)
// WS5 defect-class checkers (DC1/DC2/DC6, checkers.go) — src↔target observability flaggers. All patterns
// + tables are pair langpack DATA (pair-14 data-out); the spec is nil/inert for a no-pack book → fire 0.
n, det = cfg.checkers.lintTimeUnits(source, final)
r.DC1TimeUnits = n
r.Detail = append(r.Detail, det...)
n, det = lintMagnitudeScale(source, final)
n, det = cfg.checkers.lintMagnitudeScale(source, final)
r.DC2Magnitude = n
r.Detail = append(r.Detail, det...)
n, det = lintRegisterLexicon(final)
n, det = cfg.checkers.lintRegisterLexicon(final)
r.DC6Register = n
r.Detail = append(r.Detail, det...)
// pack-13 general checkers (checkers_zh_ru.go) — src↔target / Russian-side observability flaggers.
n, det = lintPercentScale(source, final)
// pack-13 general checkers (checkers.go): percent scale (pair data), Latin residue (language-general),
// broken word (target data).
n, det = cfg.checkers.lintPercentScale(source, final)
r.PercentScale = n
r.Detail = append(r.Detail, det...)
n, det = lintLatinResidue(final, cfg.allowlist)
r.LatinResidue = n
r.Detail = append(r.Detail, det...)
n, det = lintBrokenWord(final)
n, det = cfg.checkers.lintBrokenWord(final)
r.BrokenWord = n
r.Detail = append(r.Detail, det...)
if cfg.regressionEnabled {
@ -248,33 +257,16 @@ func quoteShape(after []rune) bool {
// --- 2. yofikator ---------------------------------------------------------------
// yoHomographEForms are е-spellings that are DISTINCT words from their ё-counterpart (все≠всё,
// небо≠нёбо, берет≠берёт, осел≠осёл, …). The inconsistency rule below would otherwise false-flag
// «все»+«всё» as one word spelled two ways. Expanded (self-review major) to cover the frequent
// distinct-word pairs and their common inflections; still not exhaustive — full disambiguation
// needs a ё-dictionary (deferred, B-tier). Names — the primary target (Пётр/Петр) — are never
// homographs, so they are always caught regardless.
var yoHomographEForms = map[string]bool{
"все": true, "всех": true, "всем": true, "всеми": true,
"небо": true, "небом": true,
"узнаем": true, "узнаете": true, "узнает": true,
"падеж": true, "падежа": true,
"совершенный": true, "совершенное": true, "совершенная": true, "совершенно": true, "совершенны": true,
"чем": true, "тем": true, "тема": true, "теме": true, "темы": true,
"берет": true, "берете": true, "берета": true,
"осел": true, "осла": true, "ослы": true, "ослов": true,
"мел": true, "мела": true,
"лен": true, "лена": true,
"нес": true, "несем": true, "несет": true,
"вел": true, "ведет": true, "ведем": true,
}
// lintYofikation flags inconsistent ё. Default ("auto"): the same word appears BOTH with ё and, as
// a separate token, with its exact ё→е form (Пётр/Петр) — excluding the homograph traps above.
// Policy "all-e": any ё present is a violation (the project wants no ё). Policy "all-yo": an е-form
// of a word that ALSO appears somewhere with ё is flagged (the partial-yofikation case), same signal
// as auto — full "every word that SHOULD have ё" enforcement needs a ё-dictionary (deferred, B-tier).
func lintYofikation(text, policy string) (int, []string) {
// a separate token, with its exact ё→е form (Пётр/Петр) — excluding the target's homograph whitelist
// (е-spellings that are DISTINCT words from their ё-counterpart, все≠всё …; TARGET data, lang.TargetChecks
// "yo_homograph" — pair-14 data-out, so the ё↔е FOLD stays as generic orthography and only the wordlist is
// data). Policy "all-e": any ё present is a violation. Policy "all-yo": same signal as auto. Full "every
// word that SHOULD have ё" enforcement needs a ё-dictionary (deferred, B-tier). Inert if no target data.
func (c *dcCheckers) lintYofikation(text, policy string) (int, []string) {
if c == nil {
return 0, nil
}
words := tokenizeCyrillic(text)
if policy == "all-e" {
seen := map[string]bool{}
@ -303,7 +295,7 @@ func lintYofikation(text, policy string) (int, []string) {
continue
}
eForm := strings.ReplaceAll(w, "ё", "е")
if eForm == w || yoHomographEForms[eForm] {
if eForm == w || c.yoHomograph[eForm] {
continue // no ё, or a distinct-word homograph (все/всё) — not an inconsistency
}
if present[eForm] && !seenPair[w] {
@ -404,7 +396,10 @@ func isWordRune(r rune) bool {
// range for their possible multiplier, plus Arabic-number orders). It fires only when the source's
// TOP order is outside every output range — a conservative, multiplier-tolerant signal that leaves
// 三億→«триста миллионов» (8 within миллион's [6,8]) silent while catching 三万→«три миллиона».
func lintNumberMagnitude(source, final string) (int, []string) {
func (c *dcCheckers) lintNumberMagnitude(source, final string) (int, []string) {
if c == nil || len(c.magnitudeStem) == 0 {
return 0, nil // no target magnitude-word data → the gate can never confirm coverage; stay inert
}
srcOrders := cjkMagnitudeOrders(source)
if len(srcOrders) == 0 {
return 0, nil
@ -427,7 +422,7 @@ func lintNumberMagnitude(source, final string) (int, []string) {
return 0, nil // an Arabic figure of the source order confirms coverage (30000 for 三万)
}
}
wordRanges := magnitudeWordRanges(final)
wordRanges := c.magnitudeWordRanges(final)
for _, rg := range wordRanges {
if maxSrc >= rg[0] && maxSrc <= rg[1] {
return 0, nil // a magnitude WORD covers the source order (三万 → «тридцать тысяч»)
@ -444,14 +439,11 @@ func lintNumberMagnitude(source, final string) (int, []string) {
return 1, []string{fmt.Sprintf("the source magnitude 10^%d (万/億) is not reflected in the translation's orders of magnitude — a possible magnitude error (e.g. 三万→«три миллиона»)", maxSrc)}
}
// cjkNumeralRunes are the characters that can form a CJK numeral expression (digits, small units,
// big markers). A maximal run of these is one candidate number.
var cjkNumeralRunes = map[rune]bool{}
func init() {
for _, r := range "0123456789零一二三四五六七八九十百千两兩万萬億亿兆" {
cjkNumeralRunes[r] = true
}
// isCJKNumeralRune reports whether r can form a CJK numeral expression: an ASCII digit or a shared
// lang.CJKSection numeral (ZeroDigitUnitMagnitude, pair-14 data-out — the gate no longer keeps its own
// rune set). A maximal run of these is one candidate number.
func isCJKNumeralRune(r rune) bool {
return (r >= '0' && r <= '9') || lang.DefaultCJKSection().IsNumeralRune(r)
}
// cjkMagnitudeOrders returns the base-10 order of every CJK numeral run in text that contains a big
@ -460,12 +452,12 @@ func cjkMagnitudeOrders(text string) []int {
var orders []int
rs := []rune(text)
for i := 0; i < len(rs); {
if !cjkNumeralRunes[rs[i]] {
if !isCJKNumeralRune(rs[i]) {
i++
continue
}
j := i
for j < len(rs) && cjkNumeralRunes[rs[j]] {
for j < len(rs) && isCJKNumeralRune(rs[j]) {
j++
}
run := string(rs[i:j])
@ -474,7 +466,7 @@ func cjkMagnitudeOrders(text string) []int {
// 万一 (in case), 万分 (extremely), 万物 (all things), 千万 (by all means), 亿万 (myriads) — where 万/億
// is not preceded by a digit. It also drops bare-unit magnitudes (十万/百万) — an accepted recall
// trade for not false-flagging the far more frequent idioms.
if strings.ContainsAny(run, "万萬億亿兆") && hasDigitBeforeBigMarker(run) {
if containsBigMarker(run) && hasDigitBeforeBigMarker(run) {
if v, ok := parseCJKNumber(run); ok && v > 0 {
orders = append(orders, orderOf(v))
}
@ -484,14 +476,27 @@ func cjkMagnitudeOrders(text string) []int {
return orders
}
// isBigMarker / containsBigMarker read the shared lang.CJKSection magnitude set (万/萬/億/亿/兆, pair-14
// data-out) instead of a hard-coded rune string.
func isBigMarker(r rune) bool { _, ok := lang.DefaultCJKSection().BigUnit(r); return ok }
func containsBigMarker(run string) bool {
for _, r := range run {
if isBigMarker(r) {
return true
}
}
return false
}
// hasDigitBeforeBigMarker reports whether a digit (一-九 / 两 / 0-9) appears before the FIRST big
// marker (万/億/兆) in the run — the signature of a real magnitude expression (三万) vs an idiom (万一).
func hasDigitBeforeBigMarker(run string) bool {
sec := lang.DefaultCJKSection()
for _, r := range run {
if strings.ContainsRune("万萬億亿兆", r) {
if isBigMarker(r) {
return false // hit a big marker with no digit before it → idiom / bare unit
}
if (r >= '0' && r <= '9') || strings.ContainsRune("一二三四五六七八九两兩", r) {
if _, isDigit := sec.Digit[r]; (r >= '0' && r <= '9') || isDigit {
return true
}
}
@ -503,28 +508,29 @@ func hasDigitBeforeBigMarker(run string) bool {
// <10^4 section; big units (万億兆) flush the section times the big unit into the total. Returns
// ok=false on a shape it cannot parse (conservative — an unparseable run does not flag).
func parseCJKNumber(s string) (int64, bool) {
sec := lang.DefaultCJKSection()
var total, section, cur int64
sawBig := false
for _, r := range s {
switch {
case r >= '0' && r <= '9':
cur = cur*10 + int64(r-'0')
case r == '' || r == '零':
case sec.Zero[r]:
cur = cur * 10
default:
if d, ok := cjkDigit(r); ok {
cur = cur*10 + d // positional accumulation (一二→12), matching the Arabic-digit branch (self-review)
if d, ok := sec.Digit[r]; ok {
cur = cur*10 + int64(d) // positional accumulation (一二→12), matching the Arabic-digit branch (self-review)
continue
}
if u, ok := cjkSmallUnit(r); ok {
if u, ok := sec.Unit[r]; ok {
if cur == 0 {
cur = 1
}
section += cur * u
section += cur * int64(u)
cur = 0
continue
}
if u, ok := cjkBigUnit(r); ok {
if u, ok := sec.BigUnit(r); ok {
sawBig = true
section += cur
if section == 0 {
@ -544,54 +550,6 @@ func parseCJKNumber(s string) (int64, bool) {
return total + section + cur, true
}
func cjkDigit(r rune) (int64, bool) {
switch r {
case '一':
return 1, true
case '二', '两', '兩':
return 2, true
case '三':
return 3, true
case '四':
return 4, true
case '五':
return 5, true
case '六':
return 6, true
case '七':
return 7, true
case '八':
return 8, true
case '九':
return 9, true
}
return 0, false
}
func cjkSmallUnit(r rune) (int64, bool) {
switch r {
case '十':
return 10, true
case '百':
return 100, true
case '千':
return 1000, true
}
return 0, false
}
func cjkBigUnit(r rune) (int64, bool) {
switch r {
case '万', '萬':
return 10000, true
case '億', '亿':
return 100000000, true
case '兆':
return 1000000000000, true
}
return 0, false
}
func orderOf(v int64) int {
o := 0
for v >= 10 {
@ -606,10 +564,10 @@ func orderOf(v int64) int {
// two orders (триста миллионов = 3·10^8, base 6 → order 8). Arabic integers are handled separately
// by the caller (they may only CONFIRM coverage, never trigger a mismatch — D20.4), so they are NOT
// folded in here.
func magnitudeWordRanges(text string) [][2]int {
func (c *dcCheckers) magnitudeWordRanges(text string) [][2]int {
low := strings.ToLower(text)
var ranges [][2]int
for stem, base := range map[string]int{"тысяч": 3, "миллион": 6, "миллиард": 9, "триллион": 12} {
for stem, base := range c.magnitudeStem { // target-general data (pair-14 data-out)
if strings.Contains(low, stem) {
ranges = append(ranges, [2]int{base, base + 2})
}

View file

@ -56,9 +56,10 @@ func TestLintYofikation(t *testing.T) {
{"all-e clean", "Петр ел мед.", "all-e", 0},
{"artem both spellings", "Артём и Артем — один человек.", "auto", 1},
}
dcc := testCheckers(t)
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
n, _ := lintYofikation(c.text, c.policy)
n, _ := dcc.lintYofikation(c.text, c.policy)
if n != c.want {
t.Errorf("lintYofikation(%q, %q) = %d, want %d", c.text, c.policy, n, c.want)
}
@ -112,8 +113,8 @@ func TestParseCJKNumber(t *testing.T) {
{"30万", 300000, true},
{"五億", 500000000, true},
{"一兆", 1000000000000, true},
{"三", 0, false}, // no big marker
{"三千", 0, false}, // no big marker (千 is a small unit)
{"三", 0, false}, // no big marker
{"三千", 0, false}, // no big marker (千 is a small unit)
{"", 0, false},
}
for _, c := range cases {
@ -152,9 +153,10 @@ func TestLintNumberMagnitude(t *testing.T) {
// output number must NOT flag.
{"wan in a name plus stray year not flagged", "他叫赵三万生于1985年。", "Его звали Чжао Саньвань, родился в 1985 году.", 0},
}
dcc := testCheckers(t)
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
n, _ := lintNumberMagnitude(c.source, c.final)
n, _ := dcc.lintNumberMagnitude(c.source, c.final)
if n != c.want {
t.Errorf("lintNumberMagnitude(%q → %q) = %d, want %d", c.source, c.final, n, c.want)
}
@ -167,7 +169,7 @@ func TestLintNumberMagnitude(t *testing.T) {
func TestRunCheapGatesCombined(t *testing.T) {
src := "他有三万石粮食。"
final := "- Ара-ара, — у Пётр было три миллиона мешков. Потом Петр ушёл."
cfg := cheapGateConfig{yoPolicy: "auto", allowlist: map[string]bool{}}
cfg := cheapGateConfig{yoPolicy: "auto", allowlist: map[string]bool{}, checkers: testCheckers(t)}
r := runCheapGates(src, final, final, cfg)
if r.DialogueDash == 0 {
t.Error("expected a dialogue-dash flag (hyphen-led line)")

View file

@ -0,0 +1,355 @@
package pipeline
import (
"fmt"
"regexp"
"sort"
"strconv"
"strings"
"unicode"
"textmachine/backend/internal/lang"
)
// checkers.go: the WS5 defect-class checkers (DC1 时辰 double-hour units, DC2 千万/数十万 magnitude scale, DC6
// register negative-list) + the pack-13 general checkers (percent scale, Latin residue, broken word) —
// deterministic, $0 OBSERVABILITY flaggers on the source↔FINAL text, ported from ws5_checkers_verify.py.
// Like the four cheap style gates they are NEVER a disposition (a hit is recorded, never drops a chunk),
// tuned PRECISION over recall.
//
// PAIR-AGNOSTIC (pair-14 data-out): this file no longer holds any language literal. Every DETECTION pattern,
// lookup table and wordlist is DATA — the SOURCE-gated pair checkers (DC1/DC2/percent: they need a zh source
// token and compare src↔tgt) read the pair pack configs/langpacks/<pair>/dc-checkers.txt (lang.DCCheckerData);
// the TARGET-general ones (register, broken word: they run on ANY →target output) read the embedded target
// data (lang.TargetChecks). The ALGORITHM (compare counts, ×2 hours, suppress-if-ok, whole-word match) stays
// here. A pair/target that ships no data runs the relevant sub-checker inert (empty → 0, the no-pack golden
// path). Version rides the langpack Version() (data) + cheapGateVersion (algorithm) — a data or rule edit is
// a loud --resnapshot.
// dcCheckers is the compiled, per-run checker spec: the pair's DETECTION patterns compiled ONCE + its lookup
// tables + the target-general lists, resolved from the langpack. A nil receiver, a nil pattern or an empty
// table leaves that sub-checker inert. Built once per run (compileCheckers), carried in cheapGateConfig.
type dcCheckers struct {
numeral map[rune]int
ruHours map[string]int
registerNeg []string
brokenSuffix []string // target-general (any →ru output)
yoHomograph map[string]bool // target-general: ё↔е homograph whitelist (cheapgates yofikator)
magnitudeStem map[string]int // target-general: ru magnitude word stem → base-10 exponent (cheapgates)
shichenRE, ruHoursRE, chengRE, decimalFractionRE *regexp.Regexp
qianwanOKRE, shushiwanOKRE, shushiwanFireRE *regexp.Regexp
qianwanSrc, qianwanFireWord, qianwanVetoWord string
shushiwanSrc, percentWord string
}
// compileCheckers resolves the checker spec from the pair pack (dc) and the target data (tc). A malformed
// regex in the pack is a corrupt pack → panic (deterministic, caught by the checker/golden tests). dc==nil
// (a no-pack book) → the pair sub-checkers are inert; tc still supplies the target-general lists.
func compileCheckers(dc *lang.DCCheckerData, tc lang.TargetChecks) *dcCheckers {
c := &dcCheckers{
brokenSuffix: tc.List("broken_suffix"),
yoHomograph: listToSet(tc.List("yo_homograph")),
magnitudeStem: listToStemExp(tc.List("magnitude_stem")),
}
if dc != nil {
c.numeral, c.ruHours, c.registerNeg = dc.Numeral, dc.RuHours, dc.RegisterNeg
p := dc.Patterns
c.shichenRE = mustPairRE(p, "shichen_re")
c.ruHoursRE = mustPairRE(p, "ru_hours_re")
c.chengRE = mustPairRE(p, "cheng_re")
c.decimalFractionRE = mustPairRE(p, "decimal_fraction_re")
c.qianwanOKRE = mustPairRE(p, "qianwan_ok_re")
c.shushiwanOKRE = mustPairRE(p, "shushiwan_ok_re")
c.shushiwanFireRE = mustPairRE(p, "shushiwan_fire_re")
c.qianwanSrc, c.qianwanFireWord, c.qianwanVetoWord = p["qianwan_src"], p["qianwan_fire_word"], p["qianwan_veto_word"]
c.shushiwanSrc, c.percentWord = p["shushiwan_src"], p["percent_word"]
}
return c
}
// listToSet turns an ordered value list into a membership set (target wordlists).
func listToSet(xs []string) map[string]bool {
m := make(map[string]bool, len(xs))
for _, x := range xs {
m[x] = true
}
return m
}
// listToStemExp parses `stem<TAB>exp` values (the magnitude_stem list carries a second tab-field) into a
// stem→exponent map. A malformed value is a corrupt embed → panic (deterministic, caught by tests).
func listToStemExp(xs []string) map[string]int {
m := make(map[string]int, len(xs))
for _, x := range xs {
f := strings.SplitN(x, "\t", 2)
if len(f) != 2 {
panic(fmt.Sprintf("pipeline: magnitude_stem wants `stem<TAB>exp`, got %q", x))
}
v, err := strconv.Atoi(strings.TrimSpace(f[1]))
if err != nil {
panic(fmt.Sprintf("pipeline: magnitude_stem exponent %q: %v", f[1], err))
}
m[f[0]] = v
}
return m
}
// dcCheckerData returns the pack's checker data, or nil when the book has no pack (nil-safe helper for the
// compile step, which runs even for a no-langpack book).
func dcCheckerData(p *lang.Pack) *lang.DCCheckerData {
if p != nil {
return p.DCCheckers
}
return nil
}
// mustPairRE compiles a pair detection pattern by key; a missing key → nil (inert sub-checker), a malformed
// regex → panic (a corrupt pack, not a silent no-op — the same fail-loud discipline the loader keeps).
func mustPairRE(p map[string]string, key string) *regexp.Regexp {
s := p[key]
if s == "" {
return nil
}
re, err := regexp.Compile(s)
if err != nil {
panic(fmt.Sprintf("pipeline: langpack checker pattern %q is not a valid regex: %v", key, err))
}
return re
}
// --- DC-1: 时辰 (double-hour) unit checker (ws5.shichen_checker) -----------------------------------
// lintTimeUnits flags a 时辰 (=2h) unit error: N个时辰 rendered as N часов (the count copied as hours)
// instead of ~2N hours (三个时辰 → «три часа» should be ~6h). It fires ONLY on an explicit mismatch — a
// paraphrase with no hours count is a valid rendering, not a defect. Pure and deterministic. Inert when the
// pair ships no DC1 pattern (the shichen_re / ru_hours_re detection patterns are pair langpack DATA).
func (c *dcCheckers) lintTimeUnits(source, final string) (int, []string) {
if c == nil || c.shichenRE == nil || c.ruHoursRE == nil {
return 0, nil
}
m := c.shichenRE.FindStringSubmatch(source)
if m == nil {
return 0, nil
}
n, ok := dcParseCount(m[1], c.numeral)
if !ok {
return 0, nil
}
hm := c.ruHoursRE.FindStringSubmatch(final)
if hm == nil {
return 0, nil // no explicit hours rendering → a valid paraphrase, not a defect
}
ruNum, ok := dcParseRuHours(hm[1], c.ruHours)
if !ok {
return 0, nil
}
expectedHours := n * 2
if ruNum == n && ruNum != expectedHours {
return 1, []string{fmt.Sprintf("DC1 时辰: %d个时辰 rendered as «%d час…» (counted as hours) instead of ~%d h (1 时辰 = 2 h)", n, ruNum, expectedHours)}
}
return 0, nil
}
// dcParseCount parses the DC1 count group: an Arabic digit string or a single small CJK numeral (looked up
// in the pair's DCCheckerData.Numeral, empty for a no-pack book → CJK counts don't resolve, the check is inert).
func dcParseCount(s string, dcNum map[rune]int) (int, bool) {
if v, err := strconv.Atoi(s); err == nil {
return v, true
}
r := []rune(s)
if len(r) == 1 {
if v, ok := dcNum[r[0]]; ok {
return v, true
}
}
return 0, false
}
// dcParseRuHours parses the DC1 hours group: an Arabic digit string or a Russian count word (pair data).
func dcParseRuHours(s string, dcRu map[string]int) (int, bool) {
if v, err := strconv.Atoi(s); err == nil {
return v, true
}
if v, ok := dcRu[s]; ok {
return v, true
}
return 0, false
}
// --- DC-2: number-scale magnitude checker (ws5.magnitude_checker, 千万 / 数十万) --------------------
// lintMagnitudeScale flags a 千万 (10^7) / 数十万 (~several×10^5) magnitude rendered at a WRONG smaller
// scale. A CORRECT rendering anywhere in the chunk (the ok-suppressor) suppresses the flag (ws5 reference
// parity). The ok-suppressors are case-INsensitive (their data carries the (?i)); the FIRE predicates are
// case-SENSITIVE literal Contains (the reference does not pass re.I to the inner searches). ⚠ 千万 is also
// stock HYPERBOLE whose «тысячи» rendering is in-register (§5 A4). Observability only. All probes/patterns
// are pair langpack DATA — inert when the pair ships no DC2.
func (c *dcCheckers) lintMagnitudeScale(source, final string) (int, []string) {
if c == nil {
return 0, nil
}
var flags []string
if c.qianwanOKRE != nil && c.qianwanSrc != "" && strings.Contains(source, c.qianwanSrc) && !c.qianwanOKRE.MatchString(final) {
if strings.Contains(final, c.qianwanFireWord) && !strings.Contains(final, c.qianwanVetoWord) { // case-sensitive, per reference
flags = append(flags, "DC2 千万=10^7 rendered as «тысячи» (≈10000× under) — a possible magnitude error (hyperbole risk, §5-A4)")
}
}
if c.shushiwanOKRE != nil && c.shushiwanFireRE != nil && c.shushiwanSrc != "" && strings.Contains(source, c.shushiwanSrc) && !c.shushiwanOKRE.MatchString(final) {
if c.shushiwanFireRE.MatchString(final) {
flags = append(flags, "DC2 数十万≈several×10^5 rendered as «десятки тысяч» (≈10× under)")
}
}
return len(flags), flags
}
// --- DC-6: register-lexicon negative-list (ws5.register_checker) ----------------------------------
// lintRegisterLexicon flags whole-word occurrences of a register negative-list lexeme in the FINAL text.
// The negative-list is pair langpack DATA (DCCheckerData.RegisterNeg): fairy-tale-Russian / chancery lexemes
// that break the xianxia register. Lower-cased; whole-word matched. Empty (a no-pack book) → nothing to flag.
func (c *dcCheckers) lintRegisterLexicon(final string) (int, []string) {
if c == nil || len(c.registerNeg) == 0 {
return 0, nil
}
low := []rune(strings.ToLower(final))
hitSet := map[string]bool{}
for _, w := range c.registerNeg {
wr := []rune(w)
for i := 0; i+len(wr) <= len(low); i++ {
if !runesEqual(low[i:i+len(wr)], wr) {
continue
}
if (i == 0 || !isCyrLetter(low[i-1])) && (i+len(wr) == len(low) || !isCyrLetter(low[i+len(wr)])) {
hitSet[w] = true
}
}
}
if len(hitSet) == 0 {
return 0, nil
}
hits := make([]string, 0, len(hitSet))
for w := range hitSet {
hits = append(hits, w)
}
sort.Strings(hits)
return len(hits), []string{"DC6 register: fairy-tale Russian lexis out of the xianxia genre: " + strings.Join(hits, ", ")}
}
// isCyrLetter reports whether r is a Cyrillic letter (the word boundary for the register match).
func isCyrLetter(r rune) bool { return unicode.IsLetter(r) && unicode.Is(unicode.Cyrillic, r) }
// --- percent-scale checker (成 = tenths) -----------------------------------------------------------
//
// In Chinese, 成 is one tenth: 六成 = 60%, 六成六 = 66%. A common error renders this as a decimal FRACTION
// instead of a percentage (a ~100× scale error). Precision over recall: it fires only when the source has a
// «<count>成[<count>]» (cheng_re, pair data), the output has NO percent form (percent_word / «%»), AND the
// output carries a decimal-fraction cue (decimal_fraction_re). Inert when the pair ships no percent patterns.
func (c *dcCheckers) lintPercentScale(source, final string) (int, []string) {
if c == nil || c.chengRE == nil || c.decimalFractionRE == nil {
return 0, nil
}
m := c.chengRE.FindStringSubmatch(source)
if m == nil {
return 0, nil
}
low := strings.ToLower(final)
if (c.percentWord != "" && strings.Contains(low, c.percentWord)) || strings.Contains(final, "%") {
return 0, nil // the output uses a percent form — the scale is handled correctly
}
if !c.decimalFractionRE.MatchString(low) {
return 0, nil // no fraction cue — the magnitude was paraphrased, not mis-scaled
}
tens, _ := dcParseCount(m[1], c.numeral)
pct := tens * 10
if m[2] != "" {
if ones, ok := dcParseCount(m[2], c.numeral); ok {
pct += ones
}
}
return 1, []string{fmt.Sprintf("成-percent: %s成%s = %d%% rendered as a decimal fraction instead of a percentage (~%d%%)", m[1], m[2], pct, pct)}
}
// --- Latin residue in the Russian output -----------------------------------------------------------
//
// A whole Latin WORD left untranslated in the output (e.g. «открыл их again, …»). Language-general (Latin is
// not pair data): it splits the output into maximal alphanumeric tokens and flags an all-lowercase Latin
// token (leaked prose is lowercase; a capital signals a proper noun/brand) with no digit, ≥ minLatinResidueLen
// letters, not a Roman numeral, not on the per-project allowlist. Precision over recall.
const minLatinResidueLen = 3
func lintLatinResidue(final string, allow map[string]bool) (int, []string) {
hits := map[string]bool{}
rs := []rune(final)
for i := 0; i < len(rs); {
if !isLatinLetterOrDigit(rs[i]) {
i++
continue
}
j := i
reject := false // set on any digit or uppercase letter — not a lowercase leaked word
for j < len(rs) && isLatinLetterOrDigit(rs[j]) {
if (rs[j] >= '0' && rs[j] <= '9') || (rs[j] >= 'A' && rs[j] <= 'Z') {
reject = true
}
j++
}
tok := string(rs[i:j])
i = j
if !reject && len([]rune(tok)) >= minLatinResidueLen && !isRomanNumeral(tok) && !allow[tok] {
hits[tok] = true
}
}
if len(hits) == 0 {
return 0, nil
}
surfaces := make([]string, 0, len(hits))
for s := range hits {
surfaces = append(surfaces, s)
}
sort.Strings(surfaces)
return len(surfaces), []string{"Latin word left untranslated in the Russian output: " + strings.Join(surfaces, ", ")}
}
// isLatinLetterOrDigit reports whether r is an ASCII Latin letter or digit (the alphanumeric-token alphabet).
func isLatinLetterOrDigit(r rune) bool {
return (r >= 'A' && r <= 'Z') || (r >= 'a' && r <= 'z') || (r >= '0' && r <= '9')
}
// isRomanNumeral reports whether a lowercase token is a Roman numeral (all chars in ivxlcdm) — «iii» reads
// as a numeral, not a leaked word; excluded to hold precision. (Uppercase «II» is already skipped as a cap.)
func isRomanNumeral(tok string) bool {
for _, r := range tok {
switch r {
case 'i', 'v', 'x', 'l', 'c', 'd', 'm':
default:
return false
}
}
return true
}
// --- broken word forms (target-general) ------------------------------------------------------------
//
// Flags a target word ending in a structurally-impossible suffix — for ru, «-йть», which no well-formed
// Russian word does (the shape of a mangled infinitive, «войть» for «войти»). The suffix set is TARGET data
// (lang.TargetChecks "broken_suffix"), so the ALGORITHM is language-general; a target with no suffix data
// flags nothing. Zero false-positive by construction (only a structural signature, no dictionary).
func (c *dcCheckers) lintBrokenWord(final string) (int, []string) {
if c == nil || len(c.brokenSuffix) == 0 {
return 0, nil
}
seen := map[string]bool{}
var det []string
for _, w := range tokenizeCyrillic(final) {
wl := len([]rune(w))
for _, suf := range c.brokenSuffix {
if wl >= 4 && strings.HasSuffix(w, suf) && !seen[w] {
seen[w] = true
det = append(det, "malformed word ending in «-"+suf+"» (no valid Russian word does): "+w)
}
}
}
sort.Strings(det)
return len(det), det
}

View file

@ -6,17 +6,18 @@ import "testing"
// word lists): each defect shape fires and a clean counterpart / golden-shape Russian output stays 0
// (precision over recall). The 时辰/千万 cases confirm the pre-existing DC1/DC2 still catch the corpus shapes.
func TestPack13Checkers(t *testing.T) {
dcc := testCheckers(t)
// Percent scale (成 = tenths): a decimal-fraction rendering fires; a correct percent is suppressed.
if n, _ := lintPercentScale("море истинной ци 六成六 …", "море истинной ци — шесть десятых и шесть сотых"); n != 1 {
if n, _ := dcc.lintPercentScale("море истинной ци 六成六 …", "море истинной ци — шесть десятых и шесть сотых"); n != 1 {
t.Errorf("percent: «шесть десятых и шесть сотых» should fire, got %d", n)
}
if n, _ := lintPercentScale("六成六", "море истинной ци — шесть и шесть десятых"); n != 1 {
if n, _ := dcc.lintPercentScale("六成六", "море истинной ци — шесть и шесть десятых"); n != 1 {
t.Errorf("percent: «шесть и шесть десятых» should fire, got %d", n)
}
if n, _ := lintPercentScale("六成六", "заполнено на шестьдесят шесть процентов"); n != 0 {
if n, _ := dcc.lintPercentScale("六成六", "заполнено на шестьдесят шесть процентов"); n != 0 {
t.Errorf("percent: a correct «процентов» rendering must be suppressed, got %d", n)
}
if n, _ := lintPercentScale("нет числа", "шесть десятых чего-то"); n != 0 {
if n, _ := dcc.lintPercentScale("нет числа", "шесть десятых чего-то"); n != 0 {
t.Errorf("percent: no 成 in source must not fire, got %d", n)
}
@ -28,9 +29,9 @@ func TestPack13Checkers(t *testing.T) {
for _, clean := range []string{
"ОТРЕДАКТИРОВАННЫЙ ПЕРЕВОД 5abc35ddfb65. Судзуки шёл по коридорам.", // a body-hash id (has digits)
"Классы таланта А, Б, В и Г — от высшей к низшей.", // Cyrillic class letters
"Глава II начинается.", // a Roman numeral (uppercase)
"Он держал в руках iPhone и MacBook.", // brands — any capital → skipped
"Судзуки шёл в Академию.", // Cyrillic proper noun
"Глава II начинается.", // a Roman numeral (uppercase)
"Он держал в руках iPhone и MacBook.", // brands — any capital → skipped
"Судзуки шёл в Академию.", // Cyrillic proper noun
} {
if n, det := lintLatinResidue(clean, nil); n != 0 {
t.Errorf("latin: clean text must not fire (%q): %d %v", clean, n, det)
@ -42,7 +43,7 @@ func TestPack13Checkers(t *testing.T) {
}
// Broken word — the general «-йть» rule only (no book-specific lists): «войть» fires, valid words do not.
if n, _ := lintBrokenWord("хотел тихонько войть и закрыть"); n != 1 {
if n, _ := dcc.lintBrokenWord("хотел тихонько войть и закрыть"); n != 1 {
t.Errorf("broken: «войть» (-йть) should fire")
}
for _, clean := range []string{
@ -50,16 +51,16 @@ func TestPack13Checkers(t *testing.T) {
"глава клана Гуюэ поклонился", // valid prose (no -йть)
"впереди идёт Фан Юань", // valid впереди
} {
if n, det := lintBrokenWord(clean); n != 0 {
if n, det := dcc.lintBrokenWord(clean); n != 0 {
t.Errorf("broken: clean text must not fire (%q): %d %v", clean, n, det)
}
}
// The pre-existing DC1/DC2 still catch the corpus shapes (general Chinese units/idioms).
if n, _ := lintTimeUnits("僵持了三个时辰", "прошло три часа"); n != 1 {
if n, _ := dcc.lintTimeUnits("僵持了三个时辰", "прошло три часа"); n != 1 {
t.Errorf("time-unit: 三个时辰→«три часа» should fire")
}
if n, _ := lintMagnitudeScale("千万生灵", "погубил тысячи жизней"); n != 1 {
if n, _ := dcc.lintMagnitudeScale("千万生灵", "погубил тысячи жизней"); n != 1 {
t.Errorf("magnitude: 千万→«тысячи» should fire")
}
}

View file

@ -1,295 +0,0 @@
package pipeline
import (
"fmt"
"regexp"
"sort"
"strconv"
"strings"
"unicode"
)
// checkers_zh_ru.go: the WS5 defect-class checkers DC1 (时辰 double-hour units), DC2 (千万/数十万 magnitude
// scale) and DC6 (register negative-list) — deterministic, $0 flaggers on the source↔FINAL text, ported
// from the frozen ws5_checkers_verify.py. Like the four cheap style gates they are OBSERVABILITY, NEVER a
// disposition (research/20: judges too noisy → deterministic gates; a hit is recorded in the retrieval-
// state / report, never dropping a chunk). They are tuned PRECISION over recall — DC1/DC2 fire only on an
// EXPLICIT src↔target mismatch; the LANDING (whether the flag is trusted) is gated by a §5(д) false-
// positive measure on a fresh rerun, not by this build. Their rule VERSION rides cheapGateVersion (folded
// into the snapshot as StyleCheckVersion), so editing a rule / the pack is a loud --resnapshot.
//
// PER-PAIR PACK (layer-2 data, §5(в)): the unit table + negative-list below are the zh-ru convention pack.
// They are code consts versioned by cheapGateVersion (the same discipline the existing cheap-gate data
// uses), not a config-loaded per-pair file — a config-loaded versioned pack is a cleaner future form
// (noted residual). The checkers self-gate on source content (时辰/千万 are zh-specific → no fire on a
// non-zh source), and DC6 is target-side (gated behind ru like the other readability flaggers).
// --- DC-1: 时辰 (double-hour) unit checker (ws5.shichen_checker) -----------------------------------
// dcCNNum maps the small CJK numerals the 时辰 pattern accepts to their value (1 时辰 = 2 modern hours).
var dcCNNum = map[rune]int{
'一': 1, '二': 2, '两': 2, '三': 3, '四': 4, '五': 5, '六': 6, '七': 7, '八': 8, '九': 9, '十': 10,
}
// dcShichenRE captures the count before 时辰 (an Arabic digit or a small CJK numeral), tolerating 个.
var dcShichenRE = regexp.MustCompile(`([0-9一二两三四五六七八九十])\s*个?\s*时辰`)
// dcRuHoursRE captures a Russian "<count> час…" rendering; the alternatives mirror the reference map.
var dcRuHoursRE = regexp.MustCompile(`(\d+|один|два|двух|три|трёх|трех|четыре|пять|шесть)\s+час`)
// dcRuHourWord maps the Russian count words dcRuHoursRE captures to their value.
var dcRuHourWord = map[string]int{
"один": 1, "два": 2, "двух": 2, "три": 3, "трёх": 3, "трех": 3, "четыре": 4, "пять": 5, "шесть": 6,
}
// lintTimeUnits flags a 时辰 (=2h) unit error: N个时辰 rendered as N часов (the count copied as hours)
// instead of ~2N hours (三个时辰 → «три часа» should be ~6h). It fires ONLY on an explicit mismatch — a
// paraphrase with no hours count is a valid rendering, not a defect (the reference dropped the "no hours
// found" branch that false-flagged 1.8%). Pure and deterministic.
func lintTimeUnits(source, final string) (int, []string) {
m := dcShichenRE.FindStringSubmatch(source)
if m == nil {
return 0, nil
}
n, ok := dcParseCount(m[1])
if !ok {
return 0, nil
}
hm := dcRuHoursRE.FindStringSubmatch(final)
if hm == nil {
return 0, nil // no explicit hours rendering → a valid paraphrase, not a defect
}
ruNum, ok := dcParseRuHours(hm[1])
if !ok {
return 0, nil
}
expectedHours := n * 2
if ruNum == n && ruNum != expectedHours {
return 1, []string{fmt.Sprintf("DC1 时辰: %d个时辰 rendered as «%d час…» (counted as hours) instead of ~%d h (1 时辰 = 2 h)", n, ruNum, expectedHours)}
}
return 0, nil
}
// dcParseCount parses the DC1 count group: an Arabic digit string or a single small CJK numeral.
func dcParseCount(s string) (int, bool) {
if v, err := strconv.Atoi(s); err == nil {
return v, true
}
r := []rune(s)
if len(r) == 1 {
if v, ok := dcCNNum[r[0]]; ok {
return v, true
}
}
return 0, false
}
// dcParseRuHours parses the DC1 hours group: an Arabic digit string or a Russian count word.
func dcParseRuHours(s string) (int, bool) {
if v, err := strconv.Atoi(s); err == nil {
return v, true
}
if v, ok := dcRuHourWord[s]; ok {
return v, true
}
return 0, false
}
// --- DC-2: number-scale magnitude checker (ws5.magnitude_checker, 千万 / 数十万) --------------------
// DC2 regexes — a BYTE-FAITHFUL port of ws5.magnitude_checker (_MAG). The ok_re SUPPRESSION guards are
// case-INsensitive (the reference passes re.I); the FIRE predicates are case-SENSITIVE (the reference
// does NOT pass re.I to the inner тысяч/десятки-тысяч searches). \w is ported as \p{L}* (RE2's \w is
// ASCII; the Cyrillic word-continuation the reference intends is letters).
var (
dcQianwanOkRE = regexp.MustCompile(`(?i)десят\p{L}* миллион|10\s*000\s*000|10000000`) // 千万 correct render → suppress
dcShushiwanOkRE = regexp.MustCompile(`(?i)сотн\p{L}* тысяч|нескольк\p{L}* сот\p{L}* тысяч|[1-9]00\s*000`) // 数十万 correct → suppress
dcDesyatkiTysRE = regexp.MustCompile(`десятк\p{L}* тысяч`) // 数十万 under-render (case-SENSITIVE, per reference)
)
// lintMagnitudeScale flags a 千万 (10^7) / 数十万 (~several×10^5) magnitude rendered at a WRONG smaller
// scale (an order-of-magnitude error DC2 targets). 千万 → «тысячи» without «миллион» = 10000× under; 数十万 → «десятки
// тысяч» = 10× under. A CORRECT rendering anywhere in the chunk (the ok_re guard) SUPPRESSES the flag, so
// a chunk that renders 数十万 as «сотни тысяч» does not fire even if «десятки тысяч» appears elsewhere as
// an unrelated quantity (the ws5 reference parity — required so the §5(д) FP-measure taken against the
// frozen reference matches what ships). ⚠ 千万 is also a stock HYPERBOLE («несметно») whose «тысячи»
// rendering is in-register literary, NOT a hard error (§5 A4) — hyperbole-exposed, landing FP-gated §5(д).
// Observability only. (Word↔word fraction inversion — 四成四=44% via q4a rule_b1/b2 — is a residual
// sub-checker, not built here.)
func lintMagnitudeScale(source, final string) (int, []string) {
var flags []string
if strings.Contains(source, "千万") && !dcQianwanOkRE.MatchString(final) {
if strings.Contains(final, "тысяч") && !strings.Contains(final, "миллион") { // case-sensitive, per reference
flags = append(flags, "DC2 千万=10^7 rendered as «тысячи» (≈10000× under) — a possible magnitude error (hyperbole risk, §5-A4)")
}
}
if strings.Contains(source, "数十万") && !dcShushiwanOkRE.MatchString(final) {
if dcDesyatkiTysRE.MatchString(final) {
flags = append(flags, "DC2 数十万≈several×10^5 rendered as «десятки тысяч» (≈10× under)")
}
}
return len(flags), flags
}
// --- DC-6: register-lexicon negative-list (ws5.register_checker) ----------------------------------
// dcRegisterNegList is the zh-ru register negative-list (layer-2 pack): fairy-tale-Russian / chancery
// lexemes that break the xianxia register («терем» in a cultivation novel is a hard register error). The
// owner extends this per corpus finding. Lower-cased; whole-word matched.
var dcRegisterNegList = []string{"терем", "терема", "тереме", "теремом", "терему", "теремах"}
// lintRegisterLexicon flags whole-word occurrences of a register negative-list lexeme in the FINAL text.
func lintRegisterLexicon(final string) (int, []string) {
low := []rune(strings.ToLower(final))
hitSet := map[string]bool{}
for _, w := range dcRegisterNegList {
wr := []rune(w)
for i := 0; i+len(wr) <= len(low); i++ {
if !runesEqual(low[i:i+len(wr)], wr) {
continue
}
if (i == 0 || !isCyrLetter(low[i-1])) && (i+len(wr) == len(low) || !isCyrLetter(low[i+len(wr)])) {
hitSet[w] = true
}
}
}
if len(hitSet) == 0 {
return 0, nil
}
hits := make([]string, 0, len(hitSet))
for w := range hitSet {
hits = append(hits, w)
}
sort.Strings(hits)
return len(hits), []string{"DC6 register: fairy-tale Russian lexis out of the xianxia genre: " + strings.Join(hits, ", ")}
}
// isCyrLetter reports whether r is a Cyrillic letter (the word boundary for the register match).
func isCyrLetter(r rune) bool { return unicode.IsLetter(r) && unicode.Is(unicode.Cyrillic, r) }
// --- percent-scale checker (成 = tenths) -----------------------------------------------------------
//
// In Chinese, 成 is one tenth: 六成 = 60%, 六成六 = 66%. A common translation error renders this as a
// decimal FRACTION instead of a percentage — «шесть десятых и шесть сотых» (0.66) or «шесть и шесть
// десятых» (6.6) for 六成六 — a ~100× scale error. This flags that. Precision over recall: it fires only
// when the source has «<count>成[<count>]», the output has NO percent form («процент»/«%»), AND the output
// carries a decimal-fraction cue (a «десятых»/«сотых» ordinal or a «N,N» number). So an output that renders
// the percent correctly is suppressed, and an output without a fraction cue (a paraphrase) stays silent.
// A general zh→ru unit convention, not tied to any book.
// chengPercentRE matches a 成-percent expression: a count (CJK or Arabic), 成, and an optional second count.
var chengPercentRE = regexp.MustCompile(`([0-9一二三四五六七八九十])成([0-9一二三四五六七八九]?)`)
// decimalFractionRE is the error cue: the tenths rendered as a Russian fraction ordinal or a decimal number.
var decimalFractionRE = regexp.MustCompile(`десят(?:ая|ых|ой|ые)|сот(?:ая|ых|ой|ые)|\d+[.,]\d`)
// lintPercentScale flags a 成-percent rendered as a decimal fraction instead of a percentage.
func lintPercentScale(source, final string) (int, []string) {
m := chengPercentRE.FindStringSubmatch(source)
if m == nil {
return 0, nil
}
low := strings.ToLower(final)
if strings.Contains(low, "процент") || strings.Contains(final, "%") {
return 0, nil // the output uses a percent form — the scale is handled correctly
}
if !decimalFractionRE.MatchString(low) {
return 0, nil // no fraction cue — the magnitude was paraphrased, not mis-scaled
}
tens, _ := dcParseCount(m[1])
pct := tens * 10
if m[2] != "" {
if ones, ok := dcParseCount(m[2]); ok {
pct += ones
}
}
return 1, []string{fmt.Sprintf("成-percent: %s成%s = %d%% rendered as a decimal fraction instead of a percentage (~%d%%)", m[1], m[2], pct, pct)}
}
// --- Latin residue in the Russian output -----------------------------------------------------------
//
// A whole Latin WORD left untranslated in the Russian output (e.g. «открыл их again, …»). It splits the
// output into maximal alphanumeric tokens and flags a token that is worth reporting as leaked prose. The
// target is a lowercase mid-sentence English word, so the guards keep precision high:
// - the token must be ALL LOWERCASE Latin letters. Leaked prose is lowercase; a token with any capital
// is a proper noun / brand / acronym (iPhone, Google, Suzuki, «II») — legitimate in Russian text, and
// the sanitizer already treats capitals as a brand signal, so we skip them here too.
// - it must have NO digit: an id / hash like «5abc35ddfb65» is an alphanumeric token, not a word.
// - at least minLatinResidueLen letters, not a Roman numeral, and not on the per-project allowlist.
// A foreignizing brief that keeps intentional Latin (a motto, a scientific name) puts those on the
// allowlist. Tuned for precision over recall: a leaked word that is Capitalized, or a bare URL host
// («example.com» → «example»/«com»), is not caught — accepted for a low-noise observability signal.
const minLatinResidueLen = 3
func lintLatinResidue(final string, allow map[string]bool) (int, []string) {
hits := map[string]bool{}
rs := []rune(final)
for i := 0; i < len(rs); {
if !isLatinLetterOrDigit(rs[i]) {
i++
continue
}
j := i
reject := false // set on any digit or uppercase letter — not a lowercase leaked word
for j < len(rs) && isLatinLetterOrDigit(rs[j]) {
if (rs[j] >= '0' && rs[j] <= '9') || (rs[j] >= 'A' && rs[j] <= 'Z') {
reject = true
}
j++
}
tok := string(rs[i:j])
i = j
if !reject && len([]rune(tok)) >= minLatinResidueLen && !isRomanNumeral(tok) && !allow[tok] {
hits[tok] = true
}
}
if len(hits) == 0 {
return 0, nil
}
surfaces := make([]string, 0, len(hits))
for s := range hits {
surfaces = append(surfaces, s)
}
sort.Strings(surfaces)
return len(surfaces), []string{"Latin word left untranslated in the Russian output: " + strings.Join(surfaces, ", ")}
}
// isLatinLetterOrDigit reports whether r is an ASCII Latin letter or digit (the alphanumeric-token alphabet).
func isLatinLetterOrDigit(r rune) bool {
return (r >= 'A' && r <= 'Z') || (r >= 'a' && r <= 'z') || (r >= '0' && r <= '9')
}
// isRomanNumeral reports whether a lowercase token is a Roman numeral (all chars in ivxlcdm) — «iii» reads
// as a numeral, not a leaked word; excluded to hold precision. (Uppercase «II» is already skipped as a cap.)
func isRomanNumeral(tok string) bool {
for _, r := range tok {
switch r {
case 'i', 'v', 'x', 'l', 'c', 'd', 'm':
default:
return false
}
}
return true
}
// --- broken Russian word forms ---------------------------------------------------------------------
//
// Deliberately GENERAL, with NO dictionary and NO book-specific word lists: it flags a Russian word ending
// in «-йть», which no well-formed Russian word does — the shape of a mangled infinitive (e.g. «войть» for
// «войти»). It is zero false-positive and applies to any book. Malformations WITHOUT such a structural
// signature — a plausible misspelling («Вперди» for «Впереди») or a case-agreement error («глава клан»
// for «главы клана») — are NOT detectable deterministically without a morphology/dictionary pass, which is
// out of scope here; the report states this recall limit honestly. The existing sanitizer broken-word class
// (invalid soft/hard-sign bigrams, script-mixed homoglyph tokens) is orthogonal and still runs.
func lintBrokenWord(final string) (int, []string) {
seen := map[string]bool{}
var det []string
for _, w := range tokenizeCyrillic(final) {
if len([]rune(w)) >= 4 && strings.HasSuffix(w, "йть") && !seen[w] {
seen[w] = true
det = append(det, "malformed word ending in «-йть» (no valid Russian word does): "+w)
}
}
sort.Strings(det)
return len(det), det
}

View file

@ -3,8 +3,18 @@ package pipeline
import (
"strings"
"testing"
"textmachine/backend/internal/lang"
)
// testCheckers builds the compiled checker spec from the REAL zh-ru pack + ru target data (pair-14 data-out):
// the fixtures exercise the checker ALGORITHM over the pack's own DETECTION patterns / tables / wordlists, so
// no pair literal is duplicated in the test — a pattern edit in the langpack flows straight into these cases.
func testCheckers(t *testing.T) *dcCheckers {
t.Helper()
return compileCheckers(testLangPack(t).DCCheckers, lang.TargetChecksFor("ru"))
}
// checkers_zh_ru_test.go: WS5 (г) — the DC1/DC2/DC6 checkers must CATCH the empirical trap positives
// (ws5_checkers_verify.py §a) and stay SILENT on clean/in-register text (precision over recall). These
// are the regression fixtures the plan requires (positives from exp15 §7.9 / exp14b).
@ -20,9 +30,10 @@ func TestDC1TimeUnits(t *testing.T) {
{"no-shichen", "他走了三里。", "Он прошёл три ли.", false},
{"arabic-count", "过了2个时辰。", "Прошло 2 часа.", true}, // 2时辰=4h rendered «2 часа»
}
dcc := testCheckers(t)
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
n, _ := lintTimeUnits(c.src, c.tgt)
n, _ := dcc.lintTimeUnits(c.src, c.tgt)
if (n > 0) != c.wantFlag {
t.Fatalf("lintTimeUnits(%q,%q) fired=%v, want %v", c.src, c.tgt, n > 0, c.wantFlag)
}
@ -47,9 +58,10 @@ func TestDC2MagnitudeScale(t *testing.T) {
{"case-sensitive-capitalized", "聚集了数十万人。", "Десятки тысяч человек собрались.", false},
{"no-magnitude", "他有三个朋友。", "У него три друга.", false},
}
dcc := testCheckers(t)
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
n, _ := lintMagnitudeScale(c.src, c.tgt)
n, _ := dcc.lintMagnitudeScale(c.src, c.tgt)
if (n > 0) != c.wantFlag {
t.Fatalf("lintMagnitudeScale(%q,%q) fired=%v, want %v", c.src, c.tgt, n > 0, c.wantFlag)
}
@ -67,9 +79,10 @@ func TestDC6RegisterLexicon(t *testing.T) {
{"clean", "Он вошёл в высокий зал павильона.", 0},
{"substring-not-word", "Термин был странным.", 0}, // «терм» inside «термин» must NOT fire (whole-word)
}
dcc := testCheckers(t)
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
n, det := lintRegisterLexicon(c.tgt)
n, det := dcc.lintRegisterLexicon(c.tgt)
if n != c.wantN {
t.Fatalf("lintRegisterLexicon(%q) = %d (%v), want %d", c.tgt, n, det, c.wantN)
}
@ -92,15 +105,15 @@ func TestDC3GenderInjection(t *testing.T) {
mk := func(gender string) []pickedEntry {
return []pickedEntry{{entry: &memoryEntry{src: "方源", dst: "Фан Юань", status: "approved", gender: gender}, via: "方源", disp: memConfirmed}}
}
male := renderEditorConstraintBlock(mk("male"))
male := renderEditorConstraintBlock(mk("male"), ruTX())
if !strings.Contains(male, "方源 → «Фан Юань» (муж.") {
t.Fatalf("male term must carry a masculine directive, got:\n%s", male)
}
hidden := renderEditorConstraintBlock(mk("hidden"))
hidden := renderEditorConstraintBlock(mk("hidden"), ruTX())
if !strings.Contains(hidden, "пол СКРЫТ") {
t.Fatalf("hidden term must carry the gender-avoidance mandate, got:\n%s", hidden)
}
none := renderEditorConstraintBlock(mk(""))
none := renderEditorConstraintBlock(mk(""), ruTX())
if !strings.HasSuffix(none, "«Фан Юань»") {
t.Fatalf("a genderless term must have NO gender note (line ends at the dst), got:\n%s", none)
}

View file

@ -198,13 +198,14 @@ func matchHeaderLine(line string, hr *lang.HeadingRule) (n int, subtitle string,
}
// isHeadingNumeral reports whether a rune can be part of a chapter-number run: an Arabic digit
// (half/fullwidth) or a CJK numeral character.
// (half/fullwidth) or a CJK numeral character (the shared lang.CJKSection — pair-14 §4, the SAME source
// ingest's chapter-numeral regex reads, so the two can never byte-drift apart).
func isHeadingNumeral(r rune) bool {
switch {
case r >= '0' && r <= '9', r >= '' && r <= '':
return true
}
return strings.ContainsRune("〇零一二三四五六七八九十百千两兩", r)
return lang.DefaultCJKSection().IsHeadingNumeral(r)
}
// isHeaderContentRune reports whether a rune is CONTENT (a letter/digit/ideograph/kana) rather than a
@ -239,7 +240,7 @@ func parseSectionNumeral(s string) (int, bool) {
case r >= '' && r <= '':
num = num*10 + int(r-'')
any = true
case r == '' || r == '零':
case lang.DefaultCJKSection().Zero[r]:
num = num * 10
any = true
default:
@ -267,40 +268,16 @@ func parseSectionNumeral(s string) (int, bool) {
return v, true
}
// cjkSectionDigit / cjkSectionUnit read the shared lang.CJKSection (pair-14 §4): the SAME digit/unit value
// tables the ingest chapter-numeral inventory derives from, so there is ONE source of truth.
func cjkSectionDigit(r rune) (int, bool) {
switch r {
case '一':
return 1, true
case '二', '两', '兩':
return 2, true
case '三':
return 3, true
case '四':
return 4, true
case '五':
return 5, true
case '六':
return 6, true
case '七':
return 7, true
case '八':
return 8, true
case '九':
return 9, true
}
return 0, false
v, ok := lang.DefaultCJKSection().Digit[r]
return v, ok
}
func cjkSectionUnit(r rune) (int, bool) {
switch r {
case '十':
return 10, true
case '百':
return 100, true
case '千':
return 1000, true
}
return 0, false
v, ok := lang.DefaultCJKSection().Unit[r]
return v, ok
}
// chapterDraftChunks packs one chapter's paragraphs into DRAFT chunks (the fine tiling). Rule:
@ -621,17 +598,13 @@ func isASCIISpace(r rune) bool {
return false
}
// sourceAbbrevs are trailing tokens after which a lone "." is treated as an
// abbreviation, not a sentence end (case-insensitive). A pragmatic set for prose;
// the acceptance path is ja→ru where CJK terminators dominate, so this only guards
// the en source. Single-letter initials are handled separately (isAbbrevBefore).
var sourceAbbrevs = map[string]bool{
"mr": true, "mrs": true, "ms": true, "dr": true, "prof": true, "st": true,
"jr": true, "sr": true, "vs": true, "no": true, "vol": true, "ch": true,
"fig": true, "col": true, "gen": true, "sgt": true, "capt": true, "lt": true,
"rev": true, "gov": true, "sen": true, "rep": true, "etc": true, "inc": true,
"ltd": true, "co": true, "mt": true, "ave": true, "rd": true,
}
// sourceAbbrevs are trailing tokens after which a lone "." is treated as an abbreviation, not a sentence
// end (case-insensitive). DATA lives in internal/lang (embedded, sectioned per SOURCE language, pair-14 §4);
// the splitter consumes the "en" section — the one source whose ASCII period needs the guard (the acceptance
// path is zh/ja→ru where 。 terminators dominate, so those sources ship no abbreviations). Making the splitter
// consume the BOOK's source section is a shallow follow-up (thread sourceLang into SplitChunks). Single-
// letter initials are handled separately (isAbbrevBefore).
var sourceAbbrevs = lang.SentenceAbbrev("en")
// isAbbrevBefore reports whether the text immediately before a lone "." ends in a
// known abbreviation or a single-letter initial (so the "." is not a boundary).

View file

@ -5,6 +5,8 @@ import (
"maps"
"slices"
"strings"
"textmachine/backend/internal/lang"
"textmachine/backend/internal/store"
)
@ -77,8 +79,9 @@ func (r *Runner) classifyOutput(role, source, output, finish string, isFinal boo
return cls, ""
}
// injectionRenderer serializes a chunk's selected memory records into a role's injection message.
type injectionRenderer func(injected []pickedEntry) string
// injectionRenderer serializes a chunk's selected memory records into a role's injection message, using the
// target-language wire-text (lang.InjectionTexts, pair-14 §2) — a target with no texts renders nothing.
type injectionRenderer func(injected []pickedEntry, tx lang.InjectionTexts) string
// roleInjectionRenderers maps a stage ROLE to the memory-injection renderer it consumes — a registry
// instead of a hand-edited switch (D39 layer 7, L8-role-injection-hardcoded-switch), so a new Phase-2
@ -171,5 +174,6 @@ func (r *Runner) cheapGateConfig() cheapGateConfig {
yoPolicy: r.Book.YoPolicy,
allowlist: allow,
regressionEnabled: r.Pipeline.Gates.RegressionGuard.Enabled,
checkers: r.checkers, // compiled once in openRunner (pair-14 data-out); nil-inert for a no-pack book
}
}

View file

@ -7,6 +7,7 @@ import (
"unicode"
"unicode/utf8"
"textmachine/backend/internal/lang"
"textmachine/backend/internal/llm"
)
@ -231,23 +232,12 @@ func classify(in classifyInput) classification {
}
// --- refusal blacklist (en/ru/zh/ja) — ported 1:1 from eval/refusal_bench.py ---
var refusalPatterns = []string{
`I (?:can(?:'|no)t|cannot|won'?t|am (?:not able|unable)) (?:to )?(?:assist|help|translate|continue|provide|fulfill)`,
`I'?m sorry,? but`,
`against (?:my|our) (?:content )?polic`,
`content polic(?:y|ies)`,
`не могу (?:помочь|перевести|продолжить|выполнить)`,
`я не буду переводить`,
`противоречит (?:политике|правилам)`,
`としてお答えできません`,
`申し訳ありませんが`,
`我(?:不能|无法)(?:协助|翻译|继续)`,
`抱歉[,]`,
`^\s*\[?(?:REFUSED|ОТКАЗ)\]?\s*$`,
}
var refusalRE = regexp.MustCompile("(?im)" + strings.Join(refusalPatterns, "|"))
//
// The patterns are UNIVERSAL engine safety data (a model can refuse in any language regardless of the
// book's pair), so they live in internal/lang as an embedded, per-language-sectioned file (pair-14 §3), NOT
// a book pack — a no-langpack book (the ja→ru golden) still flags a refusal. Joined here into one
// case-insensitive regex, byte-identically to the ported reference.
var refusalRE = regexp.MustCompile("(?im)" + strings.Join(lang.RefusalPatterns(), "|"))
var thinkRE = regexp.MustCompile(`(?s)<think>.*?</think>\s*`)

View file

@ -64,7 +64,7 @@ func TestFewShotEnabledDefault(t *testing.T) {
// the 物是人非 chengyu-atom), and few_shot OFF keeps the discourse core (incl. the chengyu-atom)
// but drops the examples — never losing the meaning-preservation instructions.
func TestEditorFewShotToggleRealFile(t *testing.T) {
tpl, err := LoadPromptTemplate(filepath.Join("..", "..", "prompts", "editor.md"))
tpl, err := LoadPromptTemplate(filepath.Join("..", "..", "prompts", "zh-ru", "editor.md"))
if err != nil {
t.Fatalf("load editor.md: %v", err)
}

View file

@ -11,9 +11,12 @@ import (
"path"
"regexp"
"strings"
"sync"
"unicode"
"unicode/utf8"
"textmachine/backend/internal/lang"
"golang.org/x/text/encoding/simplifiedchinese"
xunicode "golang.org/x/text/encoding/unicode"
"golang.org/x/text/transform"
@ -112,12 +115,20 @@ func ingestTXT(p, encoding, sourceLang string) (*Document, error) {
// --- CJK chapter-header splitting (D18: real zh/ja txt mark chapters as «第N章/节/回») ---------
// chapterUnitRunes are the section-level chapter markers auto-detected in a txt. 卷 (volume) is
// deliberately EXCLUDED — it is coarser than a chapter and would carve a tiny title-only "chapter".
var chapterUnitRunes = []rune{'章', '节', '節', '回'}
// chapterNumeralRE matches the numeral run of a chapter header: Arabic (half/fullwidth) or CJK. Built ONCE
// from the shared lang.CJKSection (pair-14 §4 + addendum-A): the section markers (章节節回) and the CJK
// numeral class are the SAME data the chunker heading rule reads — ONE source, no ingest↔chunker byte-drift.
// A CONSTANT (the CJK numeral system is language-invariant), so it is available even to a no-langpack book
// (the ja→ru golden splits 第X章 here). Arabic ranges stay in code (not language data).
var chapterNumeralReOnce sync.Once
var chapterNumeralReVal *regexp.Regexp
// chapterNumeralRE matches the numeral run of a chapter header: Arabic (half/fullwidth) or CJK.
var chapterNumeralRE = regexp.MustCompile(`^\s*第[0-9-9〇零一二三四五六七八九十百千两兩]+`)
func chapterNumeralRE() *regexp.Regexp {
chapterNumeralReOnce.Do(func() {
chapterNumeralReVal = regexp.MustCompile(`^\s*第[0-9-` + lang.DefaultCJKSection().HeadingNumeralClass() + `]+`)
})
return chapterNumeralReVal
}
// chapterHeaderMaxRunes bounds a header line so a prose sentence that merely opens with «第三节…»
// (a longer line) is not mistaken for a header. The 蛊真人 headers are ≤23 runes; 60 leaves room for
@ -138,7 +149,7 @@ func isCJKChapterHeader(line string, unit rune) bool {
if utf8.RuneCountInString(t) == 0 || utf8.RuneCountInString(t) > chapterHeaderMaxRunes {
return false
}
loc := chapterNumeralRE.FindStringIndex(t)
loc := chapterNumeralRE().FindStringIndex(t)
if loc == nil {
return false
}
@ -170,7 +181,9 @@ func isHeaderSeparator(r rune) bool {
// single stray match). Returns 0 when no unit qualifies (→ the part stays a single chapter).
func detectChapterUnit(lines []string) rune {
best, bestN := rune(0), 0
for _, unit := range chapterUnitRunes {
// The section markers (章节節回) are shared lang.CJKSection data (pair-14 addendum-A); iterated in AUTHORED
// order so a tie between two units breaks deterministically (first-seen), as the fixed slice once did.
for _, unit := range lang.DefaultCJKSection().ChapterUnitOrdered {
n := 0
for _, ln := range lines {
if isCJKChapterHeader(ln, unit) {

View file

@ -11,6 +11,7 @@ import (
"strings"
"unicode"
"textmachine/backend/internal/lang"
"textmachine/backend/internal/store"
)
@ -498,17 +499,17 @@ func priorityRank(p pickedEntry) int {
return rank
}
// glossaryBlockHeader introduces the injected glossary block. The layout mirrors the
// SakuraLLM/GalTransl convention "src → dst", with unverified (ambiguous) records
// tagged so the model — and a human reviewer — see they are not authoritative (A2).
const glossaryBlockHeader = "ГЛОССАРИЙ (используй эти утверждённые переводы имён и терминов последовательно; строки с пометкой ⟨проверить⟩ — неподтверждённые кандидаты):"
// renderGlossaryBlock serializes the selected records into the injection message
// (§C). Records with no dst yet (ruby candidates) are skipped — a "src → " line
// carries nothing. Returns "" when nothing renders, so an empty selection injects NO
// message at all ("better nothing than garbage"). Deterministic: the injected order is the
// budget priority order fixed by Select.
func renderGlossaryBlock(injected []pickedEntry) string {
// budget priority order fixed by Select. tx is the TARGET-language wire-text (pair-14 §2): a
// target with no injection texts (HasData()==false) injects nothing (a non-ru book gets no
// Russian block), so the whole render is gated on it.
func renderGlossaryBlock(injected []pickedEntry, tx lang.InjectionTexts) string {
if !tx.HasData() {
return ""
}
var lines []string
for _, p := range injected {
if strings.TrimSpace(p.entry.dst) == "" {
@ -516,21 +517,21 @@ func renderGlossaryBlock(injected []pickedEntry) string {
}
line := p.entry.src + " → " + p.entry.dst
if p.disp != memConfirmed {
line += " ⟨проверить⟩"
line += tx.UnverifiedMarker
} else {
// pack-13 injection-completeness fix (D39.21 owner directive: «род должен доезжать»): the
// gender of a CONFIRMED named term (蛊 Надежды = ж.р., a cicada gu, …) now reaches the DRAFT
// wire too, not only the editor — so the TRANSLATOR renders the right родовые формы FIRST,
// instead of leaving the editor to repair a wrong gender. "" for a genderless term (the common
// case → byte-identical). Mirrors the editor block's confirmed-only gender (renderEditorConstraintBlock).
line += genderConstraintNote(p.entry.gender)
line += genderConstraintNote(p.entry.gender, tx)
}
lines = append(lines, line)
}
if len(lines) == 0 {
return ""
}
return glossaryBlockHeader + "\n" + strings.Join(lines, "\n")
return tx.GlossaryHeader + "\n" + strings.Join(lines, "\n")
}
// renderFormatVersion versions the FORMAT of the role-injection renderers whose output is NOT
@ -547,11 +548,10 @@ func renderGlossaryBlock(injected []pickedEntry) string {
// draft injected bytes → the wire, so a loud --resnapshot; a genderless bank is byte-identical to v2.
const renderFormatVersion = "renderfmt-v3-draft-gender+editor-src2dst+dc3-gender"
// editorConstraintHeader introduces the editor's canonical-constraint block. It gives the BILINGUAL
// editor (D30.1) the approved src→dst bindings as consistency constraints — it must render each
// listed source term with EXACTLY the paired canonical form (inflecting for context) and touch
// nothing else.
const editorConstraintHeader = "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имён и терминов (в черновике термин исходника слева ДОЛЖЕН быть передан именно указанной формой справа — приводи к ней любые расхождения, склоняя по контексту; не вводи иных вариантов и не меняй ничего другого):"
// The editor's canonical-constraint block header and the glossary header are TARGET-language wire-text
// (lang.InjectionTexts, pair-14 §2) — they gave the BILINGUAL editor (D30.1) the approved src→dst bindings
// as consistency constraints. Relocated out of the pipeline (a Russian header rendered for a →en book was a
// leak); now gated on the target and sourced from embedded per-target data.
// renderEditorConstraintBlock serializes the selected records' CONFIRMED renderings into the
// editor's injection as a src→dst MAPPING (WS2 §2в resolves the D30.1-open question): "源термин →
@ -568,7 +568,10 @@ const editorConstraintHeader = "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имё
// separate token budget. The src→dst FORMAT is snapshot-folded via renderFormatVersion (a format
// edit is a loud --resnapshot); the injection is a message, so it also enters request_hash
// directly (no silent false-hit).
func renderEditorConstraintBlock(injected []pickedEntry) string {
func renderEditorConstraintBlock(injected []pickedEntry, tx lang.InjectionTexts) string {
if !tx.HasData() {
return ""
}
var lines []string
seen := map[[2]string]bool{}
for _, p := range injected {
@ -585,12 +588,12 @@ func renderEditorConstraintBlock(injected []pickedEntry) string {
continue
}
seen[key] = true
lines = append(lines, "- "+src+" → «"+dst+"»"+genderConstraintNote(p.entry.gender))
lines = append(lines, "- "+src+" → «"+dst+"»"+genderConstraintNote(p.entry.gender, tx))
}
if len(lines) == 0 {
return ""
}
return editorConstraintHeader + "\n" + strings.Join(lines, "\n")
return tx.EditorHeader + "\n" + strings.Join(lines, "\n")
}
// genderConstraintNote is the DC3 gender directive appended to a CONFIRMED editor-constraint line (WS5
@ -598,14 +601,14 @@ func renderEditorConstraintBlock(injected []pickedEntry) string {
// the Bai Ninbing class — no coreference needed. male/female ⇒ hard gender forms; hidden ⇒ a mandate to
// AVOID gender-marking constructions until the reveal (a masculine default when unavoidable, D19.3). ""
// for a term with no gender datum (the common case → the line is unchanged, byte-identical to before).
func genderConstraintNote(gender string) string {
func genderConstraintNote(gender string, tx lang.InjectionTexts) string {
switch gender {
case "male", "m":
return " (муж. — мужские родовые формы)"
return tx.GenderMale
case "female", "f":
return " (жен. — женские родовые формы)"
return tx.GenderFemale
case "hidden":
return " (пол СКРЫТ до раскрытия — избегай родовых форм; при неизбежности — мужские)"
return tx.GenderHidden
}
return ""
}

View file

@ -23,7 +23,9 @@ func declJSON(invariant bool, forms ...string) string {
return string(b)
}
func alias(a string) store.GlossaryAlias { return store.GlossaryAlias{Alias: a, AliasType: "прозвище"} }
func alias(a string) store.GlossaryAlias {
return store.GlossaryAlias{Alias: a, AliasType: "прозвище"}
}
// specGlossary mirrors eval/memory_hotpath.py's G (with decl forms replacing the
// toy accept-regexps).
@ -237,7 +239,7 @@ func TestSingleKeyBanAndAllowShort(t *testing.T) {
func TestBudgetEviction(t *testing.T) {
entries := []store.GlossaryEntry{
gl("甲甲", "Альфа", "", "approved"),
gl("乙乙", "Бета", "", "auto"), // ambiguous → lower priority
gl("乙乙", "Бета", "", "auto"), // ambiguous → lower priority
gl("丙丙", "Гамма", "", "approved"),
}
b := bankFrom(entries)
@ -295,7 +297,7 @@ func TestEmptyDstNotMatchable(t *testing.T) {
if len(sel.injected) != 1 || sel.injected[0].entry.src != "乙乙" {
t.Fatalf("empty-dst phantom polluted selection: injected=%v", injMap(sel))
}
if renderGlossaryBlock(sel.injected) == "" {
if renderGlossaryBlock(sel.injected, ruTX()) == "" {
t.Error("a renderable line was evicted by an empty-dst phantom")
}
// Unbounded: the phantom never appears as an exact hit.
@ -368,8 +370,8 @@ func TestPerLanguageMinKeyAndCollisionDisposition(t *testing.T) {
gl("鈴木", "Судзуки", "", "approved"), // 2 Han → fires, CONFIRMED (ideographic anchor)
gl("すずき", "Судзуки-х", "s2", "approved"), // 3 kana → fires but AMBIGUOUS (collision-prone)
gl("ながいなまえ", "Длинное имя", "", "approved"), // 6 kana → fires, CONFIRMED (long enough)
gl("リン", "Рин", "", "approved"), // 2 kana → BANNED (phonetic min 3)
gl("ai", "ИИ", "", "approved"), // 2 latin → BANNED (phonetic min 3)
gl("リン", "Рин", "", "approved"), // 2 kana → BANNED (phonetic min 3)
gl("ai", "ИИ", "", "approved"), // 2 latin → BANNED (phonetic min 3)
}
b := bankFrom(entries)
sel := b.Select("鈴木とすずきとながいなまえが会った。", 1, nil, 0)
@ -548,7 +550,7 @@ func TestRenderEditorConstraintBlock(t *testing.T) {
{entry: tanakaDup, via: "田中", disp: memConfirmed},
{entry: homonym, via: "済", disp: memConfirmed},
}
block := renderEditorConstraintBlock(sel)
block := renderEditorConstraintBlock(sel, ruTX())
// src→dst mapping present under the editor's own header.
if !strings.Contains(block, "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ") {
@ -570,7 +572,7 @@ func TestRenderEditorConstraintBlock(t *testing.T) {
t.Errorf("homonymic dst must stay distinguishable by source term: %q", block)
}
// An all-AMBIGUOUS (or empty) selection yields no block at all.
if got := renderEditorConstraintBlock([]pickedEntry{{entry: cand, disp: memAmbiguous}}); got != "" {
if got := renderEditorConstraintBlock([]pickedEntry{{entry: cand, disp: memAmbiguous}}, ruTX()); got != "" {
t.Errorf("an all-AMBIGUOUS selection must yield no editor block, got %q", got)
}
}

View file

@ -43,7 +43,7 @@ func TestSuppressorTrustGateFires(t *testing.T) {
t.Errorf("draft 四代族长 must still inject AMBIGUOUS alongside, got %q", m["四代族长"])
}
// Editor constraint block: the approved term is now present (was empty before the fix).
if blk := renderEditorConstraintBlock(sel.injected); !strings.Contains(blk, "глава клана") {
if blk := renderEditorConstraintBlock(sel.injected, ruTX()); !strings.Contains(blk, "глава клана") {
t.Errorf("editor constraint block must carry the approved «глава клана», got %q", blk)
}
// Loud record of the refused suppression.

View file

@ -47,10 +47,10 @@ func isFragment(src string, seedSurfaces map[string]bool, p *lang.Pack) bool {
}
// cooccurSameSentence reports whether a and b appear in one source sentence anywhere (alias.
// cooccur_same_sentence: split on 。!?\n).
func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string) bool {
// cooccur_same_sentence: split on the pack's sentence terminators + \n).
func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string, p *lang.Pack) bool {
for _, c := range chunks {
for _, sent := range splitMinerSentences(c.NSource) {
for _, sent := range splitMinerSentences(c.NSource, p) {
if strings.Contains(sent, aNorm) && strings.Contains(sent, bNorm) {
return true
}
@ -59,9 +59,11 @@ func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string) bool {
return false
}
func splitMinerSentences(s string) []string {
// splitMinerSentences splits on the pair's source sentence terminators (langpack DATA, pair-14) plus the
// structural newline (a layout mark, not a language one, so it stays in code).
func splitMinerSentences(s string, p *lang.Pack) []string {
return strings.FieldsFunc(s, func(r rune) bool {
return r == '。' || r == '' || r == '' || r == '\n'
return r == '\n' || p.SentenceTerminator[r]
})
}
@ -112,7 +114,7 @@ func proposeAliasEdges(surfaces map[string]aliasSurface, chunks []MinerChunk, se
givenA := trimRunePrefix(a.src, surA)
givenB := trimRunePrefix(b.src, surB)
if givenA != givenB && givenA != "" && givenB != "" { // R4-i different given names → family only
if cooccurSameSentence(chunks, a.src, b.src) {
if cooccurSameSentence(chunks, a.src, b.src, p) {
family = append(family, aliasEdge{a.src, b.src, "R2+R4iv", "family_copresent"})
} else {
family = append(family, aliasEdge{a.src, b.src, "R2", "family"})

View file

@ -21,32 +21,28 @@ import (
// name-shape signal a future faithful-A variant (with a Go morphology backend) would consume, and the
// pattern P4 transliteration-by-pair seam (§B5).
var palladiusJQX = map[string]bool{"j": true, "q": true, "x": true}
// buildPalladiusCyrSyllables derives the distinct Cyrillic syllable set, longest-first, for greedy
// segmentation (palladius._CYR_SYL) from the pack's pinyin→Cyrillic table. Ties in length never both match
// a position (distinct strings), so the order among equal-length forms is immaterial — sorted (len desc,
// then value) for a stable artifact. The table (initials/finals/Y_W/SPECIAL_I) is langpack DATA.
// then value) for a stable artifact. The table (initials/finals/Y_W/SPECIAL_I) AND the phonotactic
// constraints (retroflex / ü-finals / their legal initials) are langpack DATA (pair-14).
func buildPalladiusCyrSyllables(p *lang.Pack) []string {
set := map[string]bool{}
for _, v := range p.PalladiusYW {
pal := p.Palladius
for _, v := range pal.YW {
set[v] = true
}
for _, v := range p.PalladiusSpecialI {
for _, v := range pal.SpecialI {
set[v] = true
}
retroflex := map[string]bool{"zh": true, "ch": true, "sh": true, "r": true, "z": true, "c": true, "s": true}
for pi, ci := range p.PalladiusInitials {
for pf, cf := range p.PalladiusFinals {
// ü finals (written v) are valid only after j/q/x (pinyin writes ü as plain u) or l/n.
switch pf {
case "v", "ve", "van", "vn":
if !palladiusJQX[pi] && pi != "l" && pi != "n" {
continue
}
for pi, ci := range pal.Initials {
for pf, cf := range pal.Finals {
// ü finals (written v) are valid only after a vfinal_initial (pinyin writes ü as plain u).
if pal.VFinal[pf] && !pal.VFinalInitial[pi] {
continue
}
// skip the retroflex/sibilant + bare i (handled by SPECIAL_I).
if pf == "i" && retroflex[pi] {
if pf == "i" && pal.Retroflex[pi] {
continue
}
set[ci+cf] = true

View file

@ -56,13 +56,14 @@ func runeHasPrefixAt(run []rune, i int, p []rune) bool {
return true
}
// formantType maps a formant char to a candidate type (patterns.formant_type): topo→place, 等/转/阶→
// title, else term.
// formantType maps a formant char to a candidate type (patterns.formant_type): topo→place, title-formant
// (等/转/阶)→title, else term. Both the place-side (TopoSuffix) and title-side (TitleFormant) inventories are
// langpack DATA (pair-14: the title-formant set moved out of this switch, beside its TopoSuffix neighbour).
func formantType(c rune, p *lang.Pack) string {
if p.TopoSuffix[c] {
return "place"
}
if c == '等' || c == '转' || c == '阶' {
if p.TitleFormant[c] {
return "title"
}
return "term"

View file

@ -14,7 +14,10 @@ import (
// exercised over the exact same data.
func testLangPack(t *testing.T) *lang.Pack {
t.Helper()
p, err := lang.Load("../../configs/langpacks", "zh", "ru")
// Load the shared zh-ru pack PLUS the 蛊真人 book-scoped overlay (pair-14 §1): 古月 is the book's clan,
// carried in the overlay, not the shared pair langpack. Every miner fixture (and the stand parity test)
// runs on this book's effective pack, so the surname-anchor cases (古月方源) and {古月:22} stay EXACT.
p, err := lang.LoadWithOverlay("../../configs/langpacks", "zh", "ru", "testdata/langpack-overlay-guzhenren")
if err != nil {
t.Fatalf("load langpack: %v", err)
}

View file

@ -3,8 +3,14 @@ package pipeline
import (
"strings"
"testing"
"textmachine/backend/internal/lang"
)
// ruTX is the ru target injection wire-text (pair-14 §2), for the renderer fixtures that assert the
// Russian glossary/editor block bytes (the values that moved to lang/data/injection.txt).
func ruTX() lang.InjectionTexts { return lang.InjectionTextsFor("ru") }
func TestMessagesInjectionLayout(t *testing.T) {
tpl := writeTemplate(t, "Стабильный системный префикс.\n---USER---\nПереведи: {{text}}")
v := RenderVars{Book: testBook(), Text: "исходный чанк"}
@ -73,7 +79,7 @@ func TestRenderGlossaryBlock(t *testing.T) {
ambiguous := pickedEntry{entry: &memoryEntry{src: "小D", dst: "Малыш Дэ", status: "auto"}, disp: memAmbiguous}
emptyDst := pickedEntry{entry: &memoryEntry{src: "鈴木", dst: "", status: "auto"}, disp: memAmbiguous}
block := renderGlossaryBlock([]pickedEntry{confirmed, ambiguous, emptyDst})
block := renderGlossaryBlock([]pickedEntry{confirmed, ambiguous, emptyDst}, ruTX())
if !strings.Contains(block, "阿Q → А-кью") {
t.Errorf("confirmed line missing: %q", block)
}
@ -84,10 +90,10 @@ func TestRenderGlossaryBlock(t *testing.T) {
t.Errorf("empty-dst candidate should be skipped (nothing to inject): %q", block)
}
// Nothing to inject → empty block (better nothing).
if renderGlossaryBlock(nil) != "" {
if renderGlossaryBlock(nil, ruTX()) != "" {
t.Error("empty selection must render an empty block")
}
if renderGlossaryBlock([]pickedEntry{emptyDst}) != "" {
if renderGlossaryBlock([]pickedEntry{emptyDst}, ruTX()) != "" {
t.Error("only-empty-dst selection must render an empty block")
}
}

View file

@ -69,6 +69,11 @@ type Runner struct {
// enriched `memory`. nil ⇔ memory is nil (materialized together in seedGlossary).
baseMemory *MemoryBank
// checkers is the compiled WS5/pack-13 observability checker spec (pair-14 data-out): the pair's
// DETECTION patterns + tables (from pack) + the target-general lists (from embedded target data),
// compiled once in loadLangPack. nil-safe: a no-pack / no-target-data book runs the checkers inert.
checkers *dcCheckers
// pack is the book's language-data pack (internal/lang), loaded once in openRunner from
// book.LangpackRoot (D39.15/16). The the bank-mining stop bank-miner reads its tables; pack.Version() is folded into
// the snapshot (a pack edit is a loud --resnapshot). nil when the book declares no langpack_root, or
@ -191,6 +196,10 @@ func openRunner(bookPath string, logger *slog.Logger, forWrite bool) (*Runner, e
// catalog — or a book with no langpack_root — runs with a nil pack (the miner is inert, the bank-mining stop auto-continues,
// and the snapshot fold is omitted). Presence of the pair directory is the "catalog exists" signal.
func (r *Runner) loadLangPack() error {
// The observability checkers compile from the pair pack (DETECTION patterns + tables, may be nil) AND the
// embedded target data (target-general lists, keyed by target lang). Built here so it exists even for a
// no-langpack book (a nil pack → inert pair checkers; the target lists still load). r.pack is set below.
defer func() { r.checkers = compileCheckers(dcCheckerData(r.pack), lang.TargetChecksFor(r.Book.TargetLang)) }()
if r.Book.LangpackRoot == "" {
return nil // no langpack declared → nil pack, miner inert
}
@ -199,7 +208,9 @@ func (r *Runner) loadLangPack() error {
// No catalog for this pair → nil-and-run (a ja book against a zh-only root just runs without mining).
return nil
}
pack, err := lang.Load(r.Book.LangpackRoot, r.Book.SourceLang, r.Book.TargetLang)
// A book-scoped overlay (LangpackExtend) unions the book's PRIVATE canon onto the shared pair pack
// (pair-14 §1: 古月 is a 蛊真人 clan, not a shared 百家姓 surname). "" ⇒ plain Load. Folds into Version().
pack, err := lang.LoadWithOverlay(r.Book.LangpackRoot, r.Book.SourceLang, r.Book.TargetLang, r.Book.LangpackExtend)
if err != nil {
return fmt.Errorf("pipeline: load langpack for %s: %w", r.Book.LangPair(), err)
}

View file

@ -316,7 +316,7 @@ func TestRunnerMemoryResnapshotOnApprovedChange(t *testing.T) {
// ({{draft}}). The monolingual variant is PRESERVED as prompts/editor-mono.md (the D13.1
// confirming pilot arm) — it must NOT reference {{text}}.
func TestEditorPromptIsBilingual(t *testing.T) {
raw, err := os.ReadFile(filepath.Join("..", "..", "prompts", "editor.md"))
raw, err := os.ReadFile(filepath.Join("..", "..", "prompts", "zh-ru", "editor.md"))
if err != nil {
t.Fatalf("read editor.md: %v", err)
}
@ -326,7 +326,7 @@ func TestEditorPromptIsBilingual(t *testing.T) {
if !strings.Contains(string(raw), "{{draft}}") {
t.Error("editor.md must reference {{draft}} (it edits the draft)")
}
mono, err := os.ReadFile(filepath.Join("..", "..", "prompts", "editor-mono.md"))
mono, err := os.ReadFile(filepath.Join("..", "..", "prompts", "zh-ru", "editor-mono.md"))
if err != nil {
t.Fatalf("read editor-mono.md: %v (D30.1: the monolingual variant must be preserved for the D13.1 pilot arm)", err)
}

View file

@ -7,10 +7,39 @@ import (
"strings"
"unicode"
"textmachine/backend/internal/lang"
"golang.org/x/text/unicode/norm"
"golang.org/x/text/width"
)
// ruSanitizer holds the ru-target sanitizer DETECTION patterns as DATA (pair-14 data-out): the sanitizer is
// isRuTarget-gated (chunkrun/waverun), so its preamble / trailing-note / edit-meta / invalid-sign patterns
// are ru-target data (lang.TargetChecks), loaded once at package init. The strip/detect ALGORITHM stays in
// this file. The "ru" key mirrors the isRuTarget call-site gate; threading the target lang so a future non-ru
// sanitizer loads its own patterns is a follow-up (like sourceAbbrevs).
var ruSanitizer = lang.TargetChecksFor("ru")
// compileSanitizerREs compiles an ORDERED list of sanitizer patterns by key (e.g. the 4 preamble shapes). A
// malformed pattern panics (a corrupt embed, caught by the sanitizer/golden tests).
func compileSanitizerREs(key string) []*regexp.Regexp {
vals := ruSanitizer.List(key)
out := make([]*regexp.Regexp, len(vals))
for i, v := range vals {
out[i] = regexp.MustCompile(v)
}
return out
}
// compileSanitizerRE compiles a single sanitizer pattern by key; the key must carry exactly one value.
func compileSanitizerRE(key string) *regexp.Regexp {
vals := ruSanitizer.List(key)
if len(vals) != 1 {
panic(fmt.Sprintf("pipeline: sanitizer pattern %q wants exactly 1 value, got %d", key, len(vals)))
}
return regexp.MustCompile(vals[0])
}
// sanitizer.go: the output-sanitizer (D30.3) — a deterministic verdict-axis gate on the
// FINAL chunk text that catches the "instant unreadability" defect classes NO existing
// gate detects (exp12 root-cause + flagman §5). It is OPT-IN (Gates.Sanitizer.Enabled),
@ -206,18 +235,11 @@ func sanitizeOutput(text string) sanitizerResult {
// вариант noun, or «перевод фрагмента/текста/…», or «ниже приведён … перевод/текст».
//
// Go RE2 \w/\b are ASCII-only, so Cyrillic uses [а-яё].
var leadingPreamblePatterns = []*regexp.Regexp{
// «Вот/Представляю/Привожу … отредактированный/исправленный перевод/текст/вариант …:»
// (the edit-adjective front-gate holds precision, so the tail-to-colon may be long — the
// gemini leak carries «… с соблюдением всех терминов из глоссария:»).
regexp.MustCompile(`(?is)^\s*(?:вот|представляю|привожу|держите)\s+[^\n:]{0,40}?(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:`),
// «Отредактированный/Исправленный перевод/текст/вариант …:» (label at the very start)
regexp.MustCompile(`(?is)^\s*(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:`),
// «(Вот) перевод фрагмента/текста/отрывка/главы/черновика:»
regexp.MustCompile(`(?is)^\s*(?:вот\s+)?перевод\s+(?:фрагмента|текста|отрывка|главы|черновика)\s*:`),
// «Ниже приведён/представлен … перевод/отредактированный текст:»
regexp.MustCompile(`(?is)^\s*ниже\s+(?:привед[её]н|представлен|дан|след)[а-яё]*\s+[^\n:]{0,40}?(?:перевод|отредактированн|текст)[а-яё]*[^\n:]{0,20}?:`),
}
// Pair-14 data-out: the actual ru patterns are DATA (lang.TargetChecks "sanitizer_preamble", ORDERED — the
// preview returns the FIRST match). The 4 shapes: (1) «Вот… отредактированный перевод…:» edit-adjective front-
// gate; (2) «Отредактированный перевод…:» label at the start; (3) «(Вот) перевод фрагмента/…:»; (4) «Ниже
// приведён … перевод/текст:».
var leadingPreamblePatterns = compileSanitizerREs("sanitizer_preamble")
func detectLeadingPreamble(text string) string {
t := strings.TrimSpace(text)
@ -237,7 +259,7 @@ func detectLeadingPreamble(text string) string {
// «Сноска объясняла обычай.» are ordinary narrative nouns/gerunds — NOT headers). So the header
// word MUST be followed by a COLON, or be «примечание/заметки/комментарий + переводчика/редактора»,
// or be the «(прим. перев./ред.)» marker. Only counted inside the trailing region (below).
var trailingNoteRE = regexp.MustCompile(`(?i)^[*_>\s-]{0,4}(?:(?:примечани[ея]|заметк[аи]|комментари[йяю]|пояснени[ея]|сноск[аи])\s*:|(?:примечани[ея]|заметк[аи]|комментари[йяю])\s+(?:переводчик|редактор)[а-яё]*|прим\.\s*(?:перев|ред)\.?)`)
var trailingNoteRE = compileSanitizerRE("sanitizer_trailing_note")
// editMetaAnywhereRE matches UNAMBIGUOUS editor meta-commentary that is a defect wherever it
// appears (not only trailing): «Внесённые правки:», «Что изменено:», «Список правок:», and the
@ -248,7 +270,7 @@ var trailingNoteRE = regexp.MustCompile(`(?i)^[*_>\s-]{0,4}(?:(?:примеча
// are common narrative openers), keeping only editor-specific compounds that never occur in prose.
// «основн…правк» is gated to «… для справки» OR a colon so a bare narrative «Основные правки внёс
// редактор» (no summary label) does not fire.
var editMetaAnywhereRE = regexp.MustCompile(`(?im)^[*_>\s-]{0,4}(?:(?:внесённ|внесен)[а-яё]+\s+правк|основн[а-яё]+\s+правк[а-яё]*(?:\s+для\s+справк[а-яё]+|\s*:)|что\s+(?:было\s+)?(?:изменен|исправлен)|список\s+(?:правок|изменений)|список\s+внесённых)`)
var editMetaAnywhereRE = compileSanitizerRE("sanitizer_edit_meta")
func detectTrailingNote(text string) string {
if loc := editMetaAnywhereRE.FindStringIndex(text); loc != nil {
@ -359,7 +381,7 @@ func hasUpperLatin(s string) bool {
// metalinguistic mention of the letter itself, «Буква ъ называлась ером.» (adversarial
// review nit). High precision otherwise: word-initial-sign-then-letter and doubled signs are
// impossible in well-formed text. \P{L} (non-letter), not RE2's ASCII-only \b.
var invalidSignRE = regexp.MustCompile(`(?i)(?:^|\P{L})[ъь][а-яё]|[ъь][ъь]`)
var invalidSignRE = compileSanitizerRE("sanitizer_invalid_sign")
func detectBrokenWords(text string) (int, []string) {
var det []string

View file

@ -0,0 +1,4 @@
# 蛊真人 (Reverend Insanity) BOOK-SCOPED langpack overlay (pair-14 §1). 古月 is the protagonist's CLAN,
# not a 百家姓 compound surname — it lives here, unioned onto the shared pack FOR THIS BOOK, so the
# shared zh langpack stays a clean pair layer while the miner still anchors 古月方源 for 蛊真人.
古月

View file

@ -8,6 +8,7 @@ import (
"sync"
"textmachine/backend/internal/config"
"textmachine/backend/internal/lang"
"textmachine/backend/internal/store"
)
@ -331,8 +332,9 @@ func (r *Runner) runDraftChunk(ctx context.Context, draftSnapshot string, ch Chu
// Serialize the per-role injection blocks via the role→renderer registry (D39 layer 7) from the
// precomputed base-bank selection: the translator gets its src→dst glossary block; any other
// role's block is rendered too but consumed only if a stage of that role runs in this wave.
tx := lang.InjectionTextsFor(r.Book.TargetLang)
for role, render := range roleInjectionRenderers {
injectionByRole[role] = render(memSel.injected) // pure → map-order-independent
injectionByRole[role] = render(memSel.injected, tx) // pure → map-order-independent
}
if len(memSel.trustGated) > 0 {
r.Log.WarnContext(ctx, "memory: lower-trust longer key refused from suppressing a higher-trust nested key (approved term preserved; reconcile the seed)",
@ -442,8 +444,9 @@ func (r *Runner) runEditUnit(ctx context.Context, editSnapshot string, unit edit
// Render the per-role injection from the ENRICHED unit selection via the registry (D39 layer 7): the
// editor gets its CONFIRMED-dst constraint block; other roles' blocks are rendered but consumed
// only if a stage of that role runs in the edit wave.
tx := lang.InjectionTextsFor(r.Book.TargetLang)
for role, render := range roleInjectionRenderers {
injectionByRole[role] = render(editSel.injected)
injectionByRole[role] = render(editSel.injected, tx)
}
}
seq, err := r.runStageSequence(ctx, editStages, editSnapshot, leader, unitDraft, injectionByRole)

View file

@ -9,6 +9,10 @@
> - **Ждём от владельца:** тачпойнты exp16 (карта подписи/пол · мини-голд алиасов · precision@30, `books/gu-zhenren/exp16/`, ~3040 мин — для сид-дельты к пере-прогону) · развилки плана §10 (W1.5-UX · DC5-стих-политика · Edit-ceiling · и др.) · реплика ja→ru (тест общности §B5) · **ре-чек прайса DeepSeek у катовера слагов 24.07** (до него платных deepseek-прогонов нет) · чтение пере-прогона (после R1+resnapshot) · старые висящие: планка запуска (лучше-фана/гибрид/издательский) · билингв-якорь пилота (D25.9-Q1) · контаминация пилот-корпуса (D27.4) · publishable/waiver (D25.1) · FN-bound L3 (D25.4) · юр-пакет · провенанс 12-*-доков.
> - Архивы хроники: `archive/PROGRESS-2026-07-04-10.md` (D31) · `archive/PROGRESS-2026-07-10-13.md` (D39.6-гигиена). Записи ниже — живая эра D39.
## Оркестратор №7 — ПАК-14 «генеральность-пасс» ПРИНЯТ (воркфлоу 11 линз + доприёмка 4) и залендён; пак-15 «структура+форма» готовится, 24.07
Сессия исполнила 9 пунктов + оба аддендума + **расширение по директиве владельца «full data-out»** (детект-регексы чекеров/гейтов/санитайзера → данные, не только lookup-таблицы; `checkers_zh_ru.go` → пар-агностичный `checkers.go`). Два дома данных: `configs/langpacks/`+книго-overlay (`book.yaml: langpack_extend`, аддитивный union, fail-loud на чужой файл) для пар/книго-данных · `internal/lang/data/` go:embed для того, что нужно движку БЕЗ пака (golden ja→ru без langpack это доказал: refusal/CJK-нарезка/→ru-инъекция). **Приёмка исполнением:** 0 блокеров; байт-сверка каждого переноса против HEAD; парити EXACT (`{方源:0 蛊:1 蛊师:2 古月:22}`, 古月 теперь из baseoverlay — общий zh-пак чист); golden байт-нетронут; тесты грузят фикстуры из реального пака. Блокер доприёмки (аддендум parsePalladius не был исполнен) закрыт дельтой: generic-парсер категорий + один типизированный `Pack.Palladius`, required валидирует потребитель; + c2-segmentation явно, packAlgoVersion v2. **Поправка отчёту (ревью-шапка):** Приложение A «distinct/matches 8» — греп-артефакт (0 кодовых ссылок; 35→32/14), хойст-вывод не задет. **Гигиена-инцидент мой:** `77dd9b8` утащил staged `git mv` промптов сессии (git commit коммитит весь индекс) — одна неконсистентная bisect-точка, вылечена этим лендингом; правило: перед лендинг-коммитом `git diff --cached`. Прод-overlay 古月 подключён в оба стендовых book.yaml. **Пак-15 (готовится):** хойст `internal/text`+value-типов → сплит miner/checks/membank/chunk (карта связности+экспорт-поверхности = приложения A/B отчёта пака-14, с моей поправкой) + форм-фиксы Export/RequestHash/resolveChunkState + конфиг-слой пары (пар-калибровки из ран-конфигов, промпт-резолв конвенцией) + перенесённый остаток (тесты Палладия · fail-loud оверлея · unionStringMap-семантика · empty-probe гард · терминаторы chunker/coverage · `第`).
## Оркестратор №7 — research/22 (доменные харнессы) ПРИНЯТ и залендён; вердикт: калибрующий, не переворачивающий, 24.07
Сессия сдала ревизию-2 (после СОБСТВЕННОГО 9-агентного адверсариала синтеза, поймавшего её факт-ошибки — прецедент мандата 12.07 в лучшей форме). Мой спот-чек 5/5 (detectCJKLeak/cheapgate-v4/ledger наши · ProofreadTask/GenDic их). **Оценка:** ядровые ставки ПОДТВЕРЖДЕНЫ рынком+академией (fused-майнинг>bolt-on доказан смертью KeywordGacha; DelTA-роадмап = наш пройденный путь; TransAgents назвали нашу D39.20-проблему и не решили; north-star «измеренная редактура» НЕ роадмапится никем — окно открыто). **Мы безоговорочно впереди:** COGS-телеметрия · 18+ · общность · измеренность (recall 0.932, вклад ролей). **Позади:** Q4 выпуск/epub (разрыв РАСТЁТ — bbm/Immersive активно куют) · gate+repair-петля QA (они чинят typed-error→адресный ре-перевод, мы флагаем). **В планы:** слой-3 diff-редактор получает прецедент AiNiee/LinguaGacha (приоритет↑, пак-15/16) · epub-tag-rewrite не держать до Ф3 бесконечно (кандидат после пилота) · Q7-леджер дополнить (глагольный вид/addition/voice-flattening) · Problem.py-классы reflow-aware (char-freq, control-code) · LinguaGacha = бенчмарк-кандидат №1 · SakuraLLM мониторить (канал B/ja→ru). «D39.7-коллизия» разрешена: не коллизия, ссылка на стандарт. Промт → архив.

View file

@ -11,17 +11,17 @@
## Структура
- `architecture/` — синтез. **Источник истины по решениям — [`05-decisions-log.md`](architecture/05-decisions-log.md) (D1D39.16); при конфликте с любым доком он выше.**
- `architecture/` — синтез. **Источник истины по решениям — [`05-decisions-log.md`](architecture/05-decisions-log.md) (D1D39.22+); при конфликте с любым доком он выше.**
- `01-decisions.md` — принципы Р1Р10; `02-mvp-plan.md` — фазы и приёмка (v3, 09.07); `03-implementation-notes.md` — контракты Фазы 0; `04-unhappy-paths.md` — ~70 режимов отказа → механизм; `06-memory-risk-registry.md` — реестр рисков банка памяти; **[`07-strategic-review.md`](architecture/07-strategic-review.md) — стратегический аудит (09.07): вердикт end-to-end, топ-риски, курс-коррекции**; **[`08-sync-audit-ledger.md`](architecture/08-sync-audit-ledger.md) — верифицированный ледджер синк-аудита «сломанного телефона» (65 находок, D39)** · **[`09-target-architecture.md`](architecture/09-target-architecture.md) — целевая 7-слойная архитектура + фазовый план арх-ресета (D39; инвариант общности §0.1)** · [`10-prompt-architecture.md`](architecture/10-prompt-architecture.md) — консолидированная промпт-заметка (концерн 4); **[`11-implementation-plan.md`](architecture/11-implementation-plan.md) — РАТИФИЦИРОВАННЫЙ план стройки пере-прогонного стека (D39.12; дизайн-оф-рекорд бэкенд-пака, три рубежа верификации)**; `components.puml`/`pipeline.puml` — диаграммы v3 (владелец смотрит PlantUML-расширением VS Code; вручную НЕ рендерить).
- `experiments/` — эмпирика «Полигона»: `00-provider-quirks` (читать перед любым вызовом провайдера), `01-token-calibration`, `02-refusal-benchmark`, `03-local-stand`, `04-editor-quality`, `06-local-extraction`, `07-coverage-precision`, `08-cost-model-v2` (актуальная денежная модель), `09-pilot-protocol` (пилот Ф2.5 + поправки D13), `10-explicit-benchmark` (18+ violence-рука канала B), `11-erotica-benchmark` (erotica по трём парам/регистрам — закрытие D14.4, D22).
- `research/` — фактура исследований 0405.07: `0110` базовые, `11-gap-*` добор критиком, `12-*` режимы отказа/отзывы/таксономии (+ два внешних материала с провенанс-шапками), `13` валидация памяти, `14` адаптивная память, `15` голос и состояние (принят, D21), **`16` ридер-IDE (принят с ревью-шапкой, D29)**, **`17` внешняя критика GPT-5.6 (принят с ревью-шапкой, D25)** — у 16/17 читать шапку прежде тела. **`18` рычаги качества (два отчёта, D36)** · **`19` нарезка+когезия+контракт t/e (D39.1)** · **`20` банк-майнинг W1.5 (D39.6)** · **`21` обзор LLM-транспорта чужих харнессов (23.07: наш транспорт опережает/вровень со всеми 11)** · **`22` доменные харнессы перевода (24.07: калибрующий — ядро подтверждено, впереди COGS/18+/общность/измеренность, позади Q4-выпуск и gate+repair-петля; сиквел `02`/`05`)** — у всех ревью-шапки. ⚠ Часть под superseded-баннерами (01/02/03/04/05/09 и gap-1/2/5) — **читай баннер прежде содержимого**.
- `PROGRESS.md`**журнал** (CURRENT-STATE сверху, ниже хронология; НЕ источник решений).
- **Активные хендофф-промты (пост-паки-12/13, 24.07):** [`BACKEND_GENERALITY_PASS_SESSION_PROMPT.md`](BACKEND_GENERALITY_PASS_SESSION_PROMPT.md) (**пак-14 генеральность-пасс — ТЕКУЩИЙ**: подтверждённые утечки пара/книга-данных из Go в langpack/сид; байт-точность+парити EXACT; норматив — [`architecture/12-go-style-notes.md`](architecture/12-go-style-notes.md)) · [`ORCHESTRATOR_SESSION_PROMPT.md`](ORCHESTRATOR_SESSION_PROMPT.md) (хендофф №5) · [`POLYGON_PACKAGE4_SESSION_PROMPT.md`](POLYGON_PACKAGE4_SESSION_PROMPT.md) (residual пилот/18+/echo). **Паки 12/13 ЗАЛЕНДЕНЫ 24.07** (`7fe0a5b`/`2f91b04`, приёмка 25-агентным воркфлоу; их промты + research/22-промт → архив после ревью research/22). Дальше: пак-15 = канал B вживую + хвосты слоя 7 → пак-16 = слой 3 diff-based редактор → голос/состояние D21 + C2/C3. Закрытые — в `archive/prompts/`.
- **Активные хендофф-промты (пост-пак-14, 24.07):** [`ORCHESTRATOR_SESSION_PROMPT.md`](ORCHESTRATOR_SESSION_PROMPT.md) (хендофф №5) · [`POLYGON_PACKAGE4_SESSION_PROMPT.md`](POLYGON_PACKAGE4_SESSION_PROMPT.md) (residual пилот/18+/echo). **Паки 12/13/14 ЗАЛЕНДЕНЫ 24.07** (`7fe0a5b`/`2f91b04`/пак-14 — приёмка воркфлоу-линзами; отчёт пака-14 с ревью-шапкой — [`archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md`](archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md); норматив общности — [`architecture/12-go-style-notes.md`](architecture/12-go-style-notes.md)). Дальше: **пак-15 «структура+форма»** (хойст text/model-базы → сплит miner/checks/membank/chunk + Export/RequestHash/resolveChunkState + конфиг-слой пары; промт готовится) → пак-16 = канал B вживую / слой 3 diff-based редактор (прецедент research/22) → голос/состояние D21 + C2/C3. Закрытые — в `archive/prompts/`.
- `archive/` — закрытые сессионные промты (только история, инструкции оттуда не исполнять).
## Статус (2026-07-23, пост-D39.20 — механизация выпускного качества)
Фаза 0 ✅; Ф1-инфра ✅ (D20D28). **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ:** 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1D39.10) → план (D39.12, `11-*`) → стройка пака-11 (D39.13D39.19: волновой исполнитель · чанкер output-бюджет · src→dst-редактор · Go-майнер паритет-EXACT · банкнота · `internal/lang`+langpacks · reject-set · echo-сплит). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07):** операционка вся подтверждена живьём (ОДИН resnapshot · **$0-резюм драфта ×3 арма** · echo 0 · $6.63<$15 · майнинг-стоп→подпись→reject-set); **планка ≤2 НЕ пройдена** — но дефект-классы полностью механизируемы. Три сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232, floor-шум 0 · два слепых чтения): **анти-корреляция гладкость↔верность доказана трижды**; **mistral исключён как несущий редактор** (голос = north-star промпта); dspro-vs-glm → итерация №2 после механизации (драфт $0). research/21: транспорт наш опережает/вровень со всеми 11, рефакторинг отвергнут. **Эксп ЗАКРЫТ (D39.22), итерация №2 отложена; редактор = deepseek-v4-pro ИНТЕРИМ (glm-резерв); стайл-каноны D39.21 ратифицированы.** **Очередь (курс — разработка бэкенда):** пак-12 транспорт-гигиена (лендится первым) → пак-13 «выпускной QA» (wire-двигающий; вкл. дефолт-редактор dspro) → сид-дельта (元/赤城/学堂家老/春秋蝉) → масштаб целой книги + канал B вживую + Ф2-механизмы (D21 голос/состояние, native-Gemini судья) + ja→ru §B5 → пилот Ф2.5 (блокеры прежние: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Судья 18+ = grok-фолбэк (квирк D22.6 подтверждён живьём). Ключ xAI: единый, data-sharing, off перед продом (D27).
Фаза 0 ✅; Ф1-инфра ✅ (D20D28). **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ:** 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1D39.10) → план (D39.12, `11-*`) → стройка пака-11 (D39.13D39.19: волновой исполнитель · чанкер output-бюджет · src→dst-редактор · Go-майнер паритет-EXACT · банкнота · `internal/lang`+langpacks · reject-set · echo-сплит). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07):** операционка вся подтверждена живьём (ОДИН resnapshot · **$0-резюм драфта ×3 арма** · echo 0 · $6.63<$15 · майнинг-стоп→подпись→reject-set); **планка ≤2 НЕ пройдена** — но дефект-классы полностью механизируемы. Три сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232, floor-шум 0 · два слепых чтения): **анти-корреляция гладкость↔верность доказана трижды**; **mistral исключён как несущий редактор** (голос = north-star промпта); dspro-vs-glm → итерация №2 после механизации (драфт $0). research/21: транспорт наш опережает/вровень со всеми 11, рефакторинг отвергнут. **Эксп ЗАКРЫТ (D39.22), итерация №2 отложена; редактор = deepseek-v4-pro ИНТЕРИМ (glm-резерв); стайл-каноны D39.21 ратифицированы.** **Очередь (курс — разработка бэкенда):** паки 12/13/14 залендены (транспорт-гигиена · выпускной QA · генеральность: данные→langpack/embedded, книго-overlay 古月, пар-агностичные чекеры) → пак-15 «структура+форма» → сид-дельта (元/赤城/学堂家老/春秋蝉) → канал B вживую + слой-3 diff-редактор + Ф2-механизмы (D21 голос/состояние, native-Gemini судья) + ja→ru §B5 → пилот Ф2.5 (блокеры прежние: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Судья 18+ = grok-фолбэк (квирк D22.6 подтверждён живьём). Ключ xAI: единый, data-sharing, off перед продом (D27).
## Доступные ключи от моделей
DEEPSEEK_API_KEY, ZAI_API_KEY, KIMI_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, XAI_API_KEY, MISTRAL_API_KEY

View file

@ -0,0 +1,253 @@
# PACK-14 «генеральность-пасс» — приёмочный отчёт (2026-07-24)
> **Ревью-шапка оркестратора №7 (24.07, приёмка исполнением, два воркфлоу: 11 линз ~970k ток. + доприёмка дельты 4 линзы; 0 блокеров).** Байт-точность КАЖДОГО переноса независимо сверена против git HEAD (refusal 12/12 · injection 6/6 с ведущими пробелами · dc-регексы · санитайзер · CJK-числительные · фонотактика · форманты — ноль расхождений); парити EXACT свежими `-count=1` прогонами; golden байт-нетронут; тесты не ослаблены (2 удалённых ассерта — законные замены). **Поправки к телу отчёта:** (а) Приложение A — ведро «distinct/matches (8 refs)» = греп-артефакт: все 8 в комментариях, кодовых ссылок 0; заголовок «miner→membank 35» реально 32 полных вхождений / 14 кодовых; остальные вёдра точны, вывод «хойст первым» НЕ задет; (б) ёфикатор = 38 слов (строка «40 ru-слов» — обсчёт в прозе); (в) свип-47 воспроизводится как 41 (точный класс) + 6 строк ё/кана (класс сессии был шире заявленного регекса), скрытых пар-данных нет; класс «пунктуация-сепараторы» в 47 не представлен (перенос из классификации-100). **Дельта доприёмки принята** (generic-Палладий по аддендуму, c2-segmentation, README-пути, want-list magnitude, packAlgoVersion v2). **Остаток в пак-15 (зафиксирован):** тесты unknown-category-толерантности и required-fail Палладия · fail-loud оверлея уровнем выше (пропавший overlayRoot/чужой подкаталог/категория-опечатка в overlay-файле — тихие) · семантика override unionStringMap · isRuTarget/'ru'-швы · empty-probe гард lintMagnitudeScale · дуп терминаторов chunker/coverage · `第`. Прод-оверлей 古月 подключён оркестратором в оба стендовых book.yaml (rerun2).
Бэкенд-сессия. Источник: `docs/BACKEND_GENERALITY_PASS_SESSION_PROMPT.md` (9 пунктов) + два аддендума
оркестратора (A: ingest `章节節回`+CJK-числительные → та же ед. точки истины, что у чанкера; B: форманты
`等/转/阶`). Норматив: `docs/architecture/12-go-style-notes.md` §0. **Перенос ДАННЫХ, не редизайн алгоритмов.**
**Сессия НЕ коммитит.**
## Вердикт: ✅ жёсткий инвариант выполнен
- **golden БАЙТ-ИДЕНТИЧЕН**`testdata/golden/capture.golden` не тронут (`git diff` пуст). Сильнее допущенного
«version-only»: у golden-книги нет langpack (ja→ru, `langpack_root` не задан → `pack==nil`), поэтому
`packVersion()==""` и её снапшот НЕ двигается ни от одного НОВОГО langpack-файла; `TestGolden` сверяет
capture побайтно и зелёный.
- **майнер-парити EXACT**`TM_MINER_PARITY=1`: `n=13618 catastrophe{方源:0 蛊:1 蛊师:2 古月:22}
recall@proposed=0.9655` — тождественно дореформенному (вкл. 古月-механику через книго-расширение).
- **`go build/vet/test -race` зелёные** (весь модуль, вкл. новый `internal/lang`).
## Ключевой факт, определивший архитектуру
golden — ja→ru книга БЕЗ langpack (`pack==nil`), но она живьём: (1) флагает `hard_refusal` (RU «Не могу
помочь»), (2) инъектит русские заголовки глоссария в wire (16×), (3) режет главы `第X章`, (4) style_flags=0.
Значит рефузал-паттерны, CJK-числительные глав и →ru-инъекцию НЕЛЬЗЯ держать в книжном паке (nil для golden).
`request_hash` фолдит `snapshotID` (render.go) — любое движение снапшота golden-книги сдвинуло бы wire. →
## Два дома данных (правило, выдержавшее «100 языков» линзу владельца)
1. **`configs/langpacks/<src>|<pair>/` + книжный overlay** — пер-исходник / пер-пара / пер-книга данные.
Добавить язык = **положить каталог, без перекомпиляции**; убрать = удалить каталог; загрузчик **падает
ГРОМКО, называя файл** (никогда тихо-пусто). Сюда: item 1 (surnames/古月), 5 (терминаторы, Палладий-
фонотактика), 6 (DC-таблицы), addB (форманты). Overlay: добавить/убрать КНИГУ = ±один каталог;
DC-таблицы OPT-IN (убрать файл → чекеры инертны, не болтаются).
2. **`internal/lang/data/` (go:embed, секционировано по языку/таргету)** — данные, нужные движку БЕЗ книжного
пака (golden доказал: рефузал, CJK-нарезка, →ru-инъекция — все без langpack). Добавить язык = добавить
секцию в один data-файл; убрать = удалить секцию. Перекомпиляция — приемлемо: это инженерные КОНСТАНТЫ
(версионируются с бинарём, как код), а байт-точность no-pack golden запрещает фолдить их в книжный пак.
Сюда: item 3 (refusal.txt), 4+addA (cjk-section.txt), 2 (injection.txt), sourceAbbrevs (sentence-abbrev.txt).
## По пунктам (что сделано, чем доказана нейтральность)
**1. 古月 ВОН из общего zh-langpack.** Удалён из `configs/langpacks/zh/surnames-compound.txt`. Механизм —
книго-скоупный overlay: новое поле `book.yaml: langpack_extend` (каталог той же раскладки), `lang.LoadWithOverlay`
делает АДДИТИВНЫЙ union (множества дополняются, слайсы append; overlay ничего не удаляет), байты overlay
фолдятся в `Version()` (правка книжного канона = громкий --resnapshot ТОЛЬКО этой книги). Парити: `testLangPack`
грузит `testdata/langpack-overlay-guzhenren` → эффективный пак = старый (base {古月}) → {古月:22} EXACT. Стенд-
overlay создан `/home/ubuntu/books/gu-zhenren/langpack-extend/zh/surnames-compound.txt`. Прод-book.yaml —
см. «Хвосты владельцу».
**2. Инъекция-тексты → таргет-данные, с гейтом.** `glossaryBlockHeader`/`editorConstraintHeader`/маркер
`⟨проверить⟩`/3 gender-ноты вынесены в `internal/lang/data/injection.txt` (target-keyed, `lang.InjectionTextsFor`),
рендереры гейтятся на `tx.HasData()` — →en книга получит ПУСТО, не русский блок (это и есть фикс утечки).
Значимые ведущие пробелы (` (муж.…`, ` ⟨проверить⟩`) сохранены дословно (парсер НЕ триммит value). `renderFormatVersion`
НЕ тронут → снапшот golden не двигается → wire байт-идентичен (проверено golden). Langpack-ФАЙЛ (а не embedded)
отложен: golden без пака потребовал бы target-loader вне book-langpack_root, что сдвинуло бы снапшот golden —
запрещено инвариантом.
**3. Рефузал-блэклист → универсальные embedded-данные.** `internal/lang/data/refusal.txt` (секции en/ru/ja/zh
+ маркер), `disposition.go` строит `refusalRE` из `lang.RefusalPatterns()`. Байт-в-байт 12 паттернов в том же
порядке (сверено скриптом). golden флагает `hard_refusal` тождественно.
**4 + addendum-A. CJK-числительные — ОДНА точка истины.** `internal/lang/data/cjk-section.txt` (digit/unit
значения, zero, chapter_unit) — читают И чанкер (`isHeadingNumeral`/`cjkSectionDigit`/`cjkSectionUnit`, путь
заголовков, nil для golden), И ingest (`chapterNumeralRE` строится из `HeadingNumeralClass()`, `detectChapterUnit`
итерит `ChapterUnitOrdered` — авторский порядок сохранён для детерминированного тай-брейка). Устранён
ingest↔chunker байт-дубль. golden режет `第X章` тождественно (ingest живёт для golden — верифицировано TestGolden).
`sourceAbbrevs``internal/lang/data/sentence-abbrev.txt` (секции по ИСХОДНИКУ, en=29 токенов); чанкер потребляет
секцию `en` (как и раньше — универсально; исходник-параметризация сплиттера = мелкий отложенный шаг, чтобы не
трогать 18 колл-сайтов SplitChunks на не-приёмочном пути).
**5. Терминаторы `。!?` + Палладий-фонотактика.** `zh/sentence-terminator.txt` (`p.SentenceTerminator`,
`splitMinerSentences`); фонотактика jqx/ретрофлексы/ü-финалы → `zh-ru/palladius-phonotactics.txt`
(`PalladiusRetroflex/VFinal/VFinalInitial`), l/n-исключение стало данными `vfinal_initial`. Майнер-only,
парити EXACT.
**6. DC-таблицы → пар-данные, регексы = алгоритм.** `dcCNNum`+`dcRuHourWord`+`dcRegisterNegList` → OPTIONAL
`zh-ru/dc-checkers.txt` (`lang.DCCheckerData`), прокинуто через `cheapGateConfig` (nil пак → пусто → чекеры
0 → golden style_flags=0). ДЕТЕКТ-регексы (`时辰`/`千万`/`成`/`терем`-паттерны) остались как пар-скоупный
алгоритм чекера (§12.2 — «данные наружу, алгоритм остаётся»).
**7. Калибровка (config/pipeline.go).** Только пере-документирование (числа НЕ тронуты): дефолты 1797/3200/
1.1978/0.3852 помечены как «zh-ru пар-калибровка = последний generic-fallback», шиппинг-конфиги (c1/армы) ставят
блок ЯВНО (пар-конфиг = источник истины). Релокация в langpack ОТВЕРГНУТА (была бы мёртвой для всех живых путей:
конфиги ставят явно, у golden нет langpack; least-mechanism §12.1). Числа держатся EXACT: no-langpack книга
(golden) чанкается на этом fallback — пере-выведенное число сдвинуло бы её границы→wire.
**addendum-B. Форманты `等/转/阶`.** `zh/title-formant.txt` (`p.TitleFormant`, `formantType`). Рядом с TopoSuffix
(сосед уже data-driven). Майнер-only, парити EXACT.
**9. Промпты → `prompts/zh-ru/`.** `backend/prompts/{translator,editor,editor-mono,judge-selector}.md`
`prompts/zh-ru/` (git mv, БАЙТЫ не тронуты — снапшот фолдит `PromptSHA256` от СОДЕРЖИМОГО, не пути). Обновлены 5
шиппинг-конфигов (`../prompts/zh-ru/…`) + 4 тест-ссылки на реальные файлы. Тесты, грузящие реальные c1/арм-конфиги
(prompt_pack_test, echo_mine_test), зелёные.
## Item 8 — свип-добор (полный grep non-test `.go`, построчно)
Метод: изолирован **non-comment** grep `[\x{4e00}-\x{9fff}]|[а-яА-Я]` (строчные комментарии срезаны). Было 133
дата-литерала → **перенесено 33** (items 16+addA/B) → **осталось 100**, классификация:
- **SQL-DDL комментарии (migrate.go, 26):** русская документация ВНУТРИ raw-string схемы. Не пар-данные, 0
поведения. OK / вне скоупа (это перевод комментариев, отдельная забота).
- **Диагностика/сообщения-флагов (~11):** `book.go:188`(Р7), `models.go:184`(Р4)/`231`, `pipeline.go:411`(Р2),
`checkers_zh_ru.go:63/119/124/200`, `cheapgates.go:451/294/319`, `checkers:285` — operator-диагностика /
observability-строки (англ.-первичные, цитируют матч для человека). НЕ wire, НЕ пар-данные. OK.
- **Пар-скоупный ДЕТЕКТ-алгоритм (регексы), §12.2 (данные уже вынесены, регекс остаётся):**
`checkers_zh_ru.go:34/38/101/102/103/117/118/122/175/178/187/283` (DC1/DC2/percent/broken-word детект);
`sanitizer.go:213/215/217/219/240/251/362` (ru-таргет editorial-preamble/note/invalid-sign детект). OK.
- **Алгоритм-инвариантные литералы (OK-generic):** `chunker.go:222` (класс СЕПАРАТОРОВ пунктуации :、,。…);
`memnorm.go:180-181` (ё→е фолд ключа памяти, 1-символьное правило); `memseed.go:570` (ー/・ в кана-ридингах,
ja-норм); `miner_palladius.go:71` (ъ-drop перед сегментацией). `ingest.go:128` — мой новый регекс, `第`
остался литералом (кандидат — см. ниже).
- **Ратифицированный DEFER (НЕ тронут):** `cheapgates.go:361` translit-interj блоклист (ара-ара/маа/…) —
D39.16 явно (ja-филлеры, мис-кей-ловушка). ✅ оставлен.
### Кандидаты СЛЕДУЮЩЕГО пака (тот же класс, НЕ в подтверждённом фикс-листе — решение оркестратора)
Одного класса с перенесёнными, но не названы в 9 пунктах/аддендумах; байт-нейтральны (golden фаерит 0), но
переносить их сейчас = расширять скоуп сверх санкционированного + повтор D39.16-концерна «живой пара-агностик».
Флагаю для след. пака:
- **`cheapgates.go:265-276` yofikator homograph-whitelist (40 ru-слов)** — ru-таргет readability-данные (класс
DC6). Нужен ru-TARGET langpack-слот (сейчас нет).
- **`cheapgates.go:459-619` 万/億 magnitude-gate** — zh-исходник числ-значения (万萬億亿兆 — ДРУГОЙ инвентарь, не
section) + ru-стемы (тысяч/миллион/…). Кандидат: расширить `lang.CJKSection` категорией `magnitude` + ru-слот.
- **`sanitizer.go` ru-фрагменты** — data-экстракция лексем-детекторов в ru-таргет-слот (регексы = алгоритм).
- **`ingest.go:128` `第`** — маркер главы; можно добавить в `cjk-section.txt` рядом с chapter_unit (1 руна).
## «100 языков» линза (ответ на вопрос владельца)
- **Добавить пару:** положить `configs/langpacks/<src>/`(11 required src-файлов)+`<src>-<tgt>/`(2 pair + опц.
heading/dc-checkers) + book.yaml. НИ строки под `internal/pipeline`. Загрузчик fail-loud называет отсутствующий
файл. Смелл: `title-formant`/`surnames-*` осмысленны только для CJK-исходника (весь минер CJK-центричен —
ПРЕ-существующее свойство, не ухудшено; для не-CJK исходника миннер сам подлежит переосмыслению).
- **Добавить таргет:** секция в `injection.txt`/`refusal.txt` — правка ДАННЫХ + перекомпиляция (go:embed).
Оправдано: инженерные универсалии/таргет-генерики.
- **Убрать язык:** удалить каталог пака / секцию embedded-файла — ничего не болтается (DC opt-in инертен;
overlay опционален; загрузчик громкий). Тесты не ломаются на удалении пары (fail-loud по требованию).
## Хвосты владельцу (операционка, вне зоны backend/)
- **Прод-`book.yaml` gu-zhenren** (стенд `rerun2/book.yaml`, `book-mistral.yaml`): добавить
`langpack_extend: /home/ubuntu/books/gu-zhenren/langpack-extend` — иначе прод-майнинг потеряет 古月 (overlay-
каталог уже создан на стенде). Не тронул сам: зона книжных данных владельца + закрытый эксперимент rerun2.
Любой ре-ран и так --resnapshot (langpack Version() сдвинут новыми файлами).
- **Стенд pipeline-конфиги** (`rerun2/pipeline-*.yaml`) и **docs**, ссылающиеся на старый путь `backend/prompts/…`
(истор. записи в archive/experiments) — обновить путь на `prompts/zh-ru/…` при следующем касании (зона
оркестратора/владельца).
## Само-проверка (мандат 12.07) + адверсариальный 4-линзовый ревью
Ревью ИСПОЛНЕНИЕМ: каждый item верифицирован golden+parity сразу после кода (не в конце). Линзы: байт-точность /
парити / общность-ja-ru / **масштаб-100-языков (доп. линза владельца)**.
### Итог адверсариального 4-линзового ревью (воркфлоу, 4 независимых агента, 289k токенов)
- **БАЙТ-ТОЧНОСТЬ: CLEAN** — не опровергнуто. Агент независимо байт-сверил КАЖДЫЙ перенесённый литерал против
git HEAD: injection (6 значений вкл. ведущие пробелы: unverified_marker len=13, gender len=31), refusal
(12 паттернов, joined_equal=True), cjk-section (digit/unit/zero MATCH, chapter_unit ПОРЯДОК [章 节 節 回],
heading-класс symmetric-diff пуст), sentence-abbrev (en=29 exact), dc-checkers (11+9+6 rows match), Палладий-
фонотактика/форманты/терминаторы MATCH, промпты = чистые rename (0 байт). golden capture пуст в git, TestGolden
байт-сверка PASS.
- **МАЙНЕР-ПАРИТИ: CLEAN** — не опровергнуто (свежий -count=1). Overlay-union доказан аддитивным (base古月
{古月} = старый 16-сет); l/n-исключение фонотактики логически тождественно старому switch.
- **ОБЩНОСТЬ: CONCERNS** — 1 MINOR + 3 NOTE (ниже, все либо пофикшены, либо документированы).
- **МАСШТАБ-100-ЯЗЫКОВ: CONCERNS** — 1 MAJOR (пофикшен) + 3 MINOR/NOTE (документированы).
**Пофикшено по ревью (пост-ревью, верифицировано golden+parity):**
- **MAJOR (масштаб):** `LoadWithOverlay` ТИХО игнорировал overlay-файл вне манифеста (опечатка `surname-compound.txt`
или книжный `heading.txt` в overlay → приватный канон не доезжает до майнера, recall падает БЕЗ сигнала). Фикс:
`overlayDirFiles` сканирует overlay-каталоги и **падает громко**, называя неожиданный файл. + тест
(misnamed-overlay fails loud). Восстанавливает fail-loud-гарантию на масштабе десятков книг-overlay.
- **NOTE (масштаб):** `refusal.txt` — комментарий-заголовки `# --- xx ---` вводили в заблуждение («удали секцию =
отключи детект для пары»). На деле `RefusalPatterns()` читает плоско, ВСЕ паттерны фаерят для ВСЕХ книг. Комментарий
переписан: секции ОРГАНИЗАЦИОННЫЕ; удаление триммит УНИВЕРСАЛЬНЫЙ сет (только `#`-строки — байты паттернов, golden не сдвинут).
- **MINOR (масштаб):** header `srcFiles` — уточнён скоуп: «drop a directory» верно ВНУТРИ zh-family name-miner
морфо-схемы; другое семейство исходников требует новых Pack-полей (не просто каталог).
**Документированные остатки (не фикшу — конфликт с инвариантом/мандатом/пре-существующее):**
- **MINOR (общность): калибровка = ТИХИЙ zh-ru дефолт** при пропуске блока (ломает fail-loud). НЕ фиксится:
item 7 предписал «оставить механизм», а golden (ja→ru) ПРОПУСКАЕТ блок и опирается на этот fallback — падение
громко на пропуске сломало бы golden. Тайтен возможен лишь когда golden сам явно проставит сегментацию (будущий
пак, который и так пере-снимет golden).
- **NOTE (общность+масштаб): `sourceAbbrevs` захардкожен на `en`** — исходник-параметризация сплиттера отложена (18
колл-сайтов SplitChunks, не-приёмочный путь). Секция не-en = мёртвые данные до треда `sourceLang`.
- **NOTE (общность+масштаб): DC2 `lintMagnitudeScale` держит НЕЗАВИСИМЫЙ CJK-числ-парсер** (万/億/兆), дублирующий
`lang.CJKSection` — пре-существующий, self-gating (инертен на не-zh), не тронут этим диффом. Кандидат след. пака
(в списке item-8).
- **NOTE (общность): DC ДЕТЕКТ-регексы пар-скоупны в Go** (§12.2 by-design: данные вынесены, детект-алгоритм остаётся;
self-gating → добавление ПАРЫ не требует Go, только СВОИ чекеры пары — да).
- **NOTE (масштаб): двух-домовая когезия — шов.** →ru target-wire-text в ДВУХ домах: injection (embedded, recompile)
vs heading.txt «Глава {n}» (pack, data-drop). Причина — байт-точность no-pack golden: то, что golden НУЖНО без пака
→ embedded; пар-гейтнутое, что golden НЕ трогает → pack. Честный шов, задокументирован; унификация потребовала бы
target-loader вне book-langpack_root (сдвиг снапшота golden — запрещён). + сиблинг target/source-данные ещё в Go
(cheapgates ё-политика/оному/magnitude) — граница data/algorithm проведена неравномерно (кандидаты item-8).
**Финальная верификация (пост-фиксы):** `go build/vet` clean; `TM_MINER_PARITY=1 go test -race ./...` весь зелёный;
парити `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}` EXACT; `git status` golden capture пуст (байт-идентичен).
---
# ДОБАВКА: полный data-out чекеров/гейтов/санитайзера (директива владельца + оркестратора)
**Триггер:** владелец (после core-пака) — «магические названия остались в `checkers_zh_ru.go`/`checkers_pack13_test.go`; что при +100 языках? это мелкие примеры, ты всё отловил?». Честный ответ: НЕТ — core-пак вынес LOOKUP-таблицы, но оставил DETECTION-регексы + сиблинг-гейты как «алгоритм/DEFER» (слишком мягко для планки 100 языков). Решение оркестратора: **доводи data-out, структуру НЕ трогай** (сплит — след. пак с форм-фиксами). Выбор владельца: «Full: все паттерны → данные, Go generic» + «тесты грузят фикстуры из пака».
## Внешняя сверка (харнесы `/home/ubuntu/projects/tmp`, 5-агентный воркфлоу)
- **crush** (Charm, Go AI-агент, ~300 .go) — **~71 фокус-пакет**, один каталог = один концерн; провайдеры = ДАННЫЕ (внешний SDK + JSON-каталог, 0 per-provider Go); промпты = embedded `.md`; пост-обработка РАЗНЕСЕНА по фокус-пакетам, НЕ «checkers-куча». Сильное свидетельство против нашей 38-файловой кучи И за «volatile-ось = данные».
- **anthropic/openai-go** — плоский root = кодоген-артефакт (не образец); но ручные части = МАЛЫЕ single-purpose пакеты (`apijson`/`apiquery`/…).
- **GalTransl/AiNiee** (домен, что мы зеркалим) — валидируют data-мандат (глоссарий TSV, промпты per-язык, WHICH-чеки = config-список, control-codes `regex.json` = ДАННЫЕ). НО обе ТЕКУТ языком в ТЕЛА чекеров (残留日文/独白男他 gated on target_lang; per-lang Unicode-регексы) — ровно наша утечка; AiNiee `regex.json` = чистый контр-пример = наш фикс.
## Что вынесено (Full data-out)
**`checkers_zh_ru.go``checkers.go`, ПОЛНОСТЬЮ пар-агностичен.** DC1/DC2/percent DETECTION-паттерны (`shichen_re`/`ru_hours_re`/`qianwan_*`/`shushiwan_*`/`cheng_re`/`decimal_fraction_re`/`percent_word`) → пар-пак `dc-checkers.txt` (`pattern<TAB>key<TAB>value`, verbatim); компилируются ОДИН раз в `compileCheckers``*dcCheckers` спек, живёт в `r.checkers` (openRunner) → `cheapGateConfig`. Target-general (broken-word `-йть`) → embedded `internal/lang/data/target-ru.txt`. dc*-символы стали ключами данных; чекеры — методы generic-спека, nil-инертны для no-pack.
**`cheapgates.go`:** yofikator homograph-whitelist (38 слов) → target-ru `yo_homograph`; 万億 magnitude — CJK-парсер (digit/unit/**magnitude 万萬億亿兆**) **ДЕДУП в `lang.CJKSection`** (закрыл reviewer-NOTE о дубле), ru-стемы (тысяч/миллион/…) → target-ru `magnitude_stem`. `lintYofikation`/`lintNumberMagnitude` — методы спека. **Оставлено (документировано):** ё↔е фолд (1-символьная орфография = алгоритм), dialogue-dash (пунктуация), **translit-interj `ара-ара`/… = РАТИФИЦИРОВАННЫЙ DEFER** (D39.16, prompt NOT-do — единственный ru-output-блоклист, придержан).
**`sanitizer.go`:** 4 preamble + trailing-note + edit-meta + invalid-sign регексы → target-ru (извлечены из исходника СКРИПТОМ, round-trip byte-match — не перепечатка). Санитайзер уже `isRuTarget`-гейтнут → паттерны = ru-target данные, загружаются `ruSanitizer = lang.TargetChecksFor("ru")`; strip/detect-АЛГОРИТМ остаётся. «ru»-ключ зеркалит гейт (тред таргета = follow-up, как sourceAbbrevs).
**Тесты грузят из пака:** `testCheckers(t) = compileCheckers(pack.DCCheckers, TargetChecksFor("ru"))`; checkers/cheapgates-фикстуры прогоняют АЛГОРИТМ над паттернами пака — 0 дублирования пар-значений (входные примеры-сценарии остаются, §0.3 легитимно).
**Дома по ГЕЙТИНГУ:** source-gated пара-чекеры (DC1/DC2/percent, нужен zh-исходник) → пар-пак (nil для golden→инертны); target-general (register/broken/yofikator/magnitude-стемы/sanitizer, бегут на любом →ru) → embedded target-ru (нужны no-pack golden'у); source-CJK (magnitude-руны) → embedded cjk-section.
**Байт-точность (пост-каждый-инкремент):** golden БАЙТ-ИДЕНТИЧЕН (`capture.golden` не тронут — вкл. ch6, где санитайзер СРЕЗАЕТ ru-преамбулу → доказано, что паттерны из данных = те же байты); парити EXACT `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}`; `-race` весь зелёный. **Свип: 100→47** остаток (26 SQL-коммент migrate.go · flag-МESSAGE-диагностики checkers/cheapgates · 1-символьные фолды ё→е/ъ/kana · пунктуация-сепараторы · translit-DEFER · `第`-маркер ingest — кандидат). Все классифицированы/легитимны.
## Приложение A — карта связности подсистем (измерено grep'ом, вкл. тесты) — ДЛЯ ПРОМТА СЛЕД. ПАКА
Кандидаты-пакеты и их файлы: **miner** (miner*.go, mining.go — 8) · **membank** (memory/memnorm/mempostcheck/memseed/seeding — 5) · **chunk** (chunker/ingest — 2) · **checks** (cheapgates/checkers/sanitizer/coverage/quality/regressionguard — 6). Драйвер (runner/bookrun/waverun/wave/stagerun/chunkrun/resume) + инфра (snapshot/render/export/status/banknote/disposition/escalation) = композит-корень.
**Кто ссылается ВНУТРЬ каждого (non-test · TEST):**
- **miner** ← non-test: runner:2, waverun:2, seeding:2 (узко!) · test: miner_test:63, miner_parity:7, seedlint:4, waverun:2.
- **membank** ← non-test: waverun:18, wave:14, runner:13, snapshot:13, +многие (горячий путь) · test: memory:78, memseed:54, memnorm:38.
- **chunk** ← non-test: waverun:12, wave:8, runner:5, status:5, stagerun:5, memseed:4 · test: chunker:35, ingest_encoding:28, ingest:23.
- **checks** ← non-test: chunkrun:17, waverun:16, status:16, export:9, disposition:9 · test: sanitizer:90, coverage:29, regressionguard:19, cheapgates:15.
**Латеральная связность — НЕ «ноль» (уточнение к survey «star topology»):** есть ОБЩИЙ СУБСТРАТ, который надо ХОЙСТНУТЬ ПЕРВЫМ, иначе сплит потечёт:
- **miner → membank: 35 refs**, но это ШАРЕД-утилиты, не membank-API: `normalizeSourceKey`@memnorm (12 — текст-нормализация, общая для майнера И банка), `loadGlossarySeed`/`seedTerm`/`seedAlias`/`seedFile`@memseed (9 — майнер читает СИД как GT/entity-guard), `runesEqual`@mempostcheck (3), `distinct`/`matches`@memory (8 — мелкие дженерик-хелперы).
- **membank → chunk: 7** (`Chunk`-тип), **checks → membank: 19** (`memoryEntry`/`pickedEntry`/`tokenizeCyrillic`), **chunk/miner → checks: 6+6**.
**Вывод для сплита (порядок):** (1) ХОЙСТ шаред-базы ПЕРЕД подсистемами — `Chunk`/`MinerChunk` value-типы + текст-нормализация (`normalizeSourceKey`/`runesEqual`/`tokenizeCyrillic`) + сид-типы (`SeedTerm`/`loadGlossarySeed`) в базовый пакет (напр. `internal/pipeline/model` + `internal/text`); (2) затем miner (самый узкий: 6 внешних не-тест ссылок, импортит только lang+store+base) → checks (чистые ф-ии над строками, импортят config+lang+base) → membank → chunk. Драйвер импортит всё (композит-корень, как crush `internal/agent`). Форм-фиксы оркестратора (Export/RequestHash/resolveChunkState) едут в тот же пак — golden ревьюится ОДНИМ слоем.
## Приложение B — предлагаемая экспорт-поверхность (что капитализируется / что прячется)
- **`internal/text` (база):** `NormalizeSourceKey`, `RunesEqual`, `TokenizeCyrillic`. (сейчас memnorm/mempostcheck; общие для miner+membank+checks).
- **`internal/pipeline/model` (база):** `Chunk`, `MinerChunk`, `SegBudget`, `Document` (value-типы, вокабуляр).
- **`miner`:** ЭКСПОРТ `MineBank`, `MinerConfig`, `MinedTerm`, `Contrast`, `LoadContrast`; ПРЯЧЕТ mineDetect/proposeAliasEdges/buildPalladiusCyrSyllables/patternCandidates/emissionEligible. Импорт: lang, store, text, model.
- **`membank`:** ЭКСПОРТ `MemoryBank`, `Select`, `Version`, `BaseVersion`, `MemoryEntry`, инъекц-рендереры; ПРЯЧЕТ pickedEntry/Aho-Corasick/suppressContained. Импорт: lang, store, text.
- **`chunk`:** ЭКСПОРТ `SplitChunks`, `Ingest`, `IngestEncoded`; ПРЯЧЕТ chapterDraftChunks/packSentences/splitSourceSentences. Импорт: lang, model.
- **`checks`:** ЭКСПОРТ `RunCheapGates`(+`CheapGateConfig`/`CheapGateResult`), `SanitizeOutput`(+`SanitizerResult`), `CompileCheckers`(+`Checkers`); ПРЯЧЕТ lint*-методы/dcCheckers-внутренности. Импорт: config, lang, text.
---
# ДОПРИЁМОЧНАЯ ДЕЛЬТА (лендинг-блокер оркестратора, 24.07)
**(1) ОБЯЗАТЕЛЬНО — аддендум `parsePalladius` ВЫПОЛНЕН.** (а) generic-парсер `parseCategoryRows``map[категория]map[key]value` (3-field key→cyr, 2-field key→"" set-member); НЕТ per-category switch, незнакомая категория ≠ ошибка парсера. (б) ОДИН типизированный `Palladius`-struct (Initials/Finals/YW/SpecialI + Retroflex/VFinal/VFinalInitial) вместо 4-картежа + 7 параллельных Pack-полей → `Pack.Palladius Palladius`. (в) required-категории валидирует ПОТРЕБИТЕЛЬ (`validate()`), не парсер. Удалены `parsePalladius`+`parsePalladiusPhonotactics`; `assignPair`/`mergePair` схлопнуты в parse+merge (оба пар-файла через generic-путь, union нил-сейф через `newPalladius()`); потребители (miner_palladius.go, langpack_test.go) обновлены. **Hash-нейтрально** (байты .txt не тронуты); `packAlgoVersion` v1→v2 (схемо-изменение по политике константы — выбрал БАМП; сдвиг Version() = version-only, golden без пака не тронут).
**(2) Хвосты:** `pipeline-c2.yaml` — явный segmentation-блок (1797/3200/1.1978/0.3852, байт-нейтрально, закрыл обещание коммента pipeline.go «все шиппинг ставят явно») · `backend/README.md:59` пути промптов → `prompts/zh-ru/` · `embedded.go` parseCJKSection unknown-category want-list → `+magnitude` · `packAlgoVersion` бамп (см. выше).
**(3) НЕ тронуто (кандидаты пака-15, по указанию оркестратора):** fail-loud оверлея уровнем выше · override-семантика unionStringMap · isRuTarget/'ru'-швы · empty-probe гард lintMagnitudeScale · дуп терминаторов chunker/coverage · 第.
**Верификация дельты:** `go build/vet/test -race` зелёные · golden БАЙТ-ИДЕНТИЧЕН (`capture.golden` не тронут) · парити EXACT `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}`. Сессия НЕ коммитит.