From 952469b2786c615bba6df6247b07c8c818ceac82 Mon Sep 17 00:00:00 2001 From: "Claude (backend session)" Date: Fri, 24 Jul 2026 23:01:30 +0300 Subject: [PATCH] Land pack-14 generality pass: detection patterns and pair data to langpacks and embedded target data, book-scoped overlay, pair-agnostic checkers, generic palladius parser, prompts under zh-ru --- CLAUDE.md | 4 +- backend/README.md | 2 +- .../configs/langpacks/zh-ru/dc-checkers.txt | 51 +++ .../zh-ru/palladius-phonotactics.txt | 20 + .../langpacks/zh/sentence-terminator.txt | 2 + .../langpacks/zh/surnames-compound.txt | 1 - .../configs/langpacks/zh/title-formant.txt | 2 + .../configs/pipeline-arm-deepseek-pro.yaml | 4 +- backend/configs/pipeline-arm-glm.yaml | 4 +- backend/configs/pipeline-arm-mistral.yaml | 4 +- backend/configs/pipeline-c1.yaml | 6 +- backend/configs/pipeline-c2.yaml | 15 +- backend/internal/config/book.go | 8 + backend/internal/config/pipeline.go | 12 +- backend/internal/config/prompt_pack_test.go | 4 +- backend/internal/lang/data/cjk-section.txt | 34 ++ backend/internal/lang/data/injection.txt | 11 + backend/internal/lang/data/refusal.txt | 25 ++ .../internal/lang/data/sentence-abbrev.txt | 35 ++ backend/internal/lang/data/target-ru.txt | 65 +++ backend/internal/lang/embedded.go | 392 ++++++++++++++++++ backend/internal/lang/langpack.go | 386 +++++++++++++++-- backend/internal/lang/langpack_test.go | 107 ++++- backend/internal/pipeline/cheapgates.go | 172 +++----- backend/internal/pipeline/cheapgates_test.go | 12 +- backend/internal/pipeline/checkers.go | 355 ++++++++++++++++ .../internal/pipeline/checkers_pack13_test.go | 23 +- backend/internal/pipeline/checkers_zh_ru.go | 295 ------------- .../internal/pipeline/checkers_zh_ru_test.go | 25 +- backend/internal/pipeline/chunker.go | 61 +-- backend/internal/pipeline/chunkrun.go | 8 +- backend/internal/pipeline/disposition.go | 24 +- backend/internal/pipeline/fewshot_test.go | 2 +- backend/internal/pipeline/ingest.go | 27 +- backend/internal/pipeline/memory.go | 47 ++- backend/internal/pipeline/memory_test.go | 16 +- .../pipeline/memory_trustgate_test.go | 2 +- backend/internal/pipeline/miner_alias.go | 14 +- backend/internal/pipeline/miner_palladius.go | 26 +- backend/internal/pipeline/miner_patterns.go | 7 +- backend/internal/pipeline/miner_test.go | 5 +- .../internal/pipeline/render_memory_test.go | 12 +- backend/internal/pipeline/runner.go | 13 +- .../internal/pipeline/runner_memory_test.go | 4 +- backend/internal/pipeline/sanitizer.go | 52 ++- .../zh/surnames-compound.txt | 4 + backend/internal/pipeline/waverun.go | 7 +- docs/PROGRESS.md | 4 + docs/README.md | 6 +- ...ERALITY_PASS_SESSION_PROMPT_2026-07-24.md} | 0 ...ND_PACK13_QA_SESSION_PROMPT_2026-07-24.md} | 0 ...ANSPORT_PACK_SESSION_PROMPT_2026-07-24.md} | 0 .../PACK14_GENERALITY_REPORT_2026-07-24.md | 253 +++++++++++ 53 files changed, 2019 insertions(+), 651 deletions(-) create mode 100644 backend/configs/langpacks/zh-ru/dc-checkers.txt create mode 100644 backend/configs/langpacks/zh-ru/palladius-phonotactics.txt create mode 100644 backend/configs/langpacks/zh/sentence-terminator.txt create mode 100644 backend/configs/langpacks/zh/title-formant.txt create mode 100644 backend/internal/lang/data/cjk-section.txt create mode 100644 backend/internal/lang/data/injection.txt create mode 100644 backend/internal/lang/data/refusal.txt create mode 100644 backend/internal/lang/data/sentence-abbrev.txt create mode 100644 backend/internal/lang/data/target-ru.txt create mode 100644 backend/internal/lang/embedded.go create mode 100644 backend/internal/pipeline/checkers.go delete mode 100644 backend/internal/pipeline/checkers_zh_ru.go create mode 100644 backend/internal/pipeline/testdata/langpack-overlay-guzhenren/zh/surnames-compound.txt rename docs/{BACKEND_GENERALITY_PASS_SESSION_PROMPT.md => archive/prompts/BACKEND_GENERALITY_PASS_SESSION_PROMPT_2026-07-24.md} (100%) rename docs/{BACKEND_PACK13_QA_SESSION_PROMPT.md => archive/prompts/BACKEND_PACK13_QA_SESSION_PROMPT_2026-07-24.md} (100%) rename docs/{BACKEND_TRANSPORT_PACK_SESSION_PROMPT.md => archive/prompts/BACKEND_TRANSPORT_PACK_SESSION_PROMPT_2026-07-24.md} (100%) create mode 100644 docs/archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md diff --git a/CLAUDE.md b/CLAUDE.md index 60f3f06b..c7765569 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -39,6 +39,6 @@ Go-бэкенд издательского художественного пер 2. `docs/architecture/05-decisions-log.md` целиком (контракт). 3. По роли: Бэкенд — `backend/README.md` + `03-implementation-notes.md`; Полигон — `eval/README.md` + `experiments/00` + `09-pilot-protocol.md`; всем — `07-strategic-review.md` §6–9 (вердикт/риски/курс). -## Текущее состояние (2026-07-23, пост-D39.20 — эра «механизация выпускного качества») +## Текущее состояние (2026-07-24, пост-D39.22/пак-14 — эра «механизация выпускного качества») -Фаза 0 ✅; **Ф1-инфра ✅** (D20–D28; golden = инвариант детерминизма, №8 README = экспорт-контракт). Арка «качество-первым» D26→D38.5 ЗАКРЫТА; **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ**: 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1–D39.10) → план (D39.12) → **стройка пака-11 завершена** (D39.13–D39.19: волновой исполнитель + чанкер output-бюджет + src→dst-редактор + Go-майнер паритет-EXACT + банкнота + `internal/lang`+langpacks + reject-set + echo-сплит; хроника — D-лог). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07, `books/gu-zhenren/rerun2/`):** вся операционка подтверждена живьём — ОДИН resnapshot, **$0-резюм драфта на всех 3 армах**, echo 0, $6.63<$15; майнинг-стоп→подпись→reject-set отработали. **Вердикт качества:** планка ≤2 претензии НЕ пройдена (пилот не открыт); три независимых сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232 при floor-шуме 0 · два слепых чтения) → **анти-корреляция гладкость↔верность доказана трижды** (чисто-художественные выборы обоих читателей идентичны = mistral 4/5 — арм с критическими смысловыми); **mistral исключён как несущий редактор** (голос = north-star промпта), **dspro-vs-glm решает итерация №2 после механизации** (драфт $0 → ~$0.08/арм + судья-lite). **Эксп ЗАКРЫТ решением владельца (D39.22), итерация №2 отложена** (её вопросы — довесок следующего платного прогона); **редактор = deepseek-v4-pro ИНТЕРИМ** (судья №1 · проза/глоссарий · дефекты чинибельны · ×2 дешевле; glm-5 резерв, отмена = один арм-конфиг); стайл-решения D39.21: меры → единицы читателя (时辰=2ч) · generic-«гу» средний род, именные по-термово · стихи смыслом · 资质=«талант». research/21: транспорт наш опережает/вровень, рефакторинг ОТВЕРГНУТ. **Курс — разработка бэкенда:** пак-12 транспорт-гигиена (выдан, гейт открыт, лендится первым) → пак-13 «выпускной QA» (выдан: заголовок-шаблон · юнит/число-чекеры · broken-word/latin · mined_delta-крэш-фикс · аудит банк-инъекции · стайл-канон · дефолт-редактор dspro; wire-двигающий, один resnapshot) → сид-дельта (元-корень · 赤城 · 学堂家老 · 春秋蝉) → **масштаб целой книги** (волны/майнинг-каденс/потолки на сотни глав) · канал B (18+) вживую · Ф2-механизмы (голос/состояние D21 · native-Gemini судья) · ja→ru §B5 → пилот Ф2.5. Стек: draft flash thinking-ON → **editor deepseek-v4-pro** → судья gemini-preview + grok-фолбэк на 18+ (~$0.85/ранобэ D30.4). **Висит на владельце:** мини-каноны сид-дельты (元-корень · 春秋蝉) · ja→ru-реплика (§B5) · старые (пилот Ф2.5: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Ключ xAI: единый, data-sharing, off перед продом (D27). Книга на стенде `/home/ubuntu/books/gu-zhenren/` (GB18030; текст и производные ВНЕ git). Ключи: DEEPSEEK, ZAI, KIMI, OPENAI, GEMINI, XAI, MISTRAL. Стенд: WSL2, GTX 1070 8GB (localhost из-под прокси = 403 — no-proxy транспорт для local). +Фаза 0 ✅; **Ф1-инфра ✅** (D20–D28; golden = инвариант детерминизма, №8 README = экспорт-контракт). Арка «качество-первым» D26→D38.5 ЗАКРЫТА; **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ**: 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1–D39.10) → план (D39.12) → **стройка пака-11 завершена** (D39.13–D39.19: волновой исполнитель + чанкер output-бюджет + src→dst-редактор + Go-майнер паритет-EXACT + банкнота + `internal/lang`+langpacks + reject-set + echo-сплит; хроника — D-лог). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07, `books/gu-zhenren/rerun2/`):** вся операционка подтверждена живьём — ОДИН resnapshot, **$0-резюм драфта на всех 3 армах**, echo 0, $6.63<$15; майнинг-стоп→подпись→reject-set отработали. **Вердикт качества:** планка ≤2 претензии НЕ пройдена (пилот не открыт); три независимых сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232 при floor-шуме 0 · два слепых чтения) → **анти-корреляция гладкость↔верность доказана трижды** (чисто-художественные выборы обоих читателей идентичны = mistral 4/5 — арм с критическими смысловыми); **mistral исключён как несущий редактор** (голос = north-star промпта), **dspro-vs-glm решает итерация №2 после механизации** (драфт $0 → ~$0.08/арм + судья-lite). **Эксп ЗАКРЫТ решением владельца (D39.22), итерация №2 отложена** (её вопросы — довесок следующего платного прогона); **редактор = deepseek-v4-pro ИНТЕРИМ** (судья №1 · проза/глоссарий · дефекты чинибельны · ×2 дешевле; glm-5 резерв, отмена = один арм-конфиг); стайл-решения D39.21: меры → единицы читателя (时辰=2ч) · generic-«гу» средний род, именные по-термово · стихи смыслом · 资质=«талант». research/21: транспорт наш опережает/вровень, рефакторинг ОТВЕРГНУТ. **Курс — разработка бэкенда:** паки 12/13/14 ЗАЛЕНДЕНЫ 24.07 (транспорт-гигиена `7fe0a5b` · выпускной QA `2f91b04` · **генеральность-пасс**: пара/книга-данные → langpack+`internal/lang/data` embedded, книго-overlay `langpack_extend` (古月 вон из общего zh-пака), пар-агностичный `checkers.go`, generic-Палладий; приёмка воркфлоу 11+4 линз, байт-сверка против HEAD, отчёт с ревью-шапкой в `docs/archive/reports/`) → **пак-15 «структура+форма»** (хойст `internal/text`+value-типов → сплит miner/checks/membank/chunk по карте связности отчёта пака-14 · Export/RequestHash/resolveChunkState · конфиг-слой пары · fail-loud-хвосты оверлея) → сид-дельта (元-корень · 赤城 · 学堂家老 · 春秋蝉) → канал B (18+) вживую · слой-3 diff-редактор (прецедент AiNiee/LinguaGacha, research/22) · Ф2-механизмы (голос/состояние D21 · native-Gemini судья) · ja→ru §B5 → пилот Ф2.5. Стек: draft flash thinking-ON → **editor deepseek-v4-pro** → судья gemini-preview + grok-фолбэк на 18+ (~$0.85/ранобэ D30.4). **Висит на владельце:** мини-каноны сид-дельты (元-корень · 春秋蝉) · ja→ru-реплика (§B5) · старые (пилот Ф2.5: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Ключ xAI: единый, data-sharing, off перед продом (D27). Книга на стенде `/home/ubuntu/books/gu-zhenren/` (GB18030; текст и производные ВНЕ git). Ключи: DEEPSEEK, ZAI, KIMI, OPENAI, GEMINI, XAI, MISTRAL. Стенд: WSL2, GTX 1070 8GB (localhost из-под прокси = 403 — no-proxy транспорт для local). diff --git a/backend/README.md b/backend/README.md index 56c718c0..da368540 100644 --- a/backend/README.md +++ b/backend/README.md @@ -56,5 +56,5 @@ set -a; . ./.env; set +a; TM_LIVE=1 go test -tags live -run TestLive -v ./intern - **D15.2:** READ-половина `tmctl export` ПОСТРОЕНА (D39.5, минимальная форма); annotation/override-половина и content-addressed resume — отложены (v3.1-спека = дизайн-оф-рекорд, `docs/D15.2-…spec.md`). - **Retries.RegenerateBeforeEscalate НЕ в снапшоте** (pre-existing, D39.4-LOW): в крэш-окне пониженный ретрай-бюджет молча бросает оплаченные OK-чекпоинты. Дизайн-фикс при следующем заходе в stagerun. - F3 at-most-once (после D15.2); Escal.Chains — валидируемый мёртвый конфиг (раннер читает только `escalate_to`); флоры min_max_tokens = схема (D24.3). -- `prompts/editor-mono.md` — референс-вариант D13.1 (не боевой, тестом закреплён); боевой = `editor.md` v3-discourse (БИЛИНГВ). Глосс-капы санитайзера (6 CJK / 24 лат-токен) — tunable-эвристики (D39.5), ревизия по чтению пере-прогона. +- `prompts/zh-ru/editor-mono.md` — референс-вариант D13.1 (не боевой, тестом закреплён); боевой = `prompts/zh-ru/editor.md` v3-discourse (БИЛИНГВ) (промпты pair-keyed под `prompts/zh-ru/`, pair-14 §9). Глосс-капы санитайзера (6 CJK / 24 лат-токен) — tunable-эвристики (D39.5), ревизия по чтению пере-прогона. - Хвост LOW/NOTE-находок адверсариала D39.4 — в ledger-очереди (golden-ячейка глосса, escalated+stripped тест, trustGateEvent первый-отказ и пр.). diff --git a/backend/configs/langpacks/zh-ru/dc-checkers.txt b/backend/configs/langpacks/zh-ru/dc-checkers.txt new file mode 100644 index 00000000..50f40bd5 --- /dev/null +++ b/backend/configs/langpacks/zh-ru/dc-checkers.txt @@ -0,0 +1,51 @@ +# WS5 defect-class checker DATA for zh→ru (checkers.go). Pair-14 data-out: the DETECTION PATTERNS (regexes + +# literal probes) are DATA too, not engine code — a new pair ships its own. `categorykey[value]`. +# pattern values are VERBATIM (a regex or a literal probe); the checker ALGORITHM (compare counts, ×2 hours, +# suppress-if-ok) stays generic Go. The `*_re` keys are compiled; the others are literal strings.Contains probes. +# +# DC1 时辰 (double-hour): small CJK count → value (1 时辰 = 2 modern hours). +cjk_numeral 一 1 +cjk_numeral 二 2 +cjk_numeral 两 2 +cjk_numeral 三 3 +cjk_numeral 四 4 +cjk_numeral 五 5 +cjk_numeral 六 6 +cjk_numeral 七 7 +cjk_numeral 八 8 +cjk_numeral 九 9 +cjk_numeral 十 10 +# DC1: Russian hours-count word → value (the alternatives ru_hours_re captures). +ru_hour один 1 +ru_hour два 2 +ru_hour двух 2 +ru_hour три 3 +ru_hour трёх 3 +ru_hour трех 3 +ru_hour четыре 4 +ru_hour пять 5 +ru_hour шесть 6 +# DC6 register negative-list: fairy-tale/chancery lexemes out of the xianxia register (терем + case forms). +register_neg терем +register_neg терема +register_neg тереме +register_neg теремом +register_neg терему +register_neg теремах +# --- DETECTION patterns (pair-14 data-out). `patternkeyvalue`; value VERBATIM. --- +# DC1: count before 时辰 (src); Russian « час…» rendering (final). +pattern shichen_re ([0-9一二两三四五六七八九十])\s*个?\s*时辰 +pattern ru_hours_re (\d+|один|два|двух|три|трёх|трех|четыре|пять|шесть)\s+час +# DC2 千万=10^7: src probe, ok-suppressor (case-insensitive), fire/veto words (case-sensitive Contains). +pattern qianwan_src 千万 +pattern qianwan_ok_re (?i)десят\p{L}* миллион|10\s*000\s*000|10000000 +pattern qianwan_fire_word тысяч +pattern qianwan_veto_word миллион +# DC2 数十万: src probe, ok-suppressor, under-render fire (case-sensitive, per reference). +pattern shushiwan_src 数十万 +pattern shushiwan_ok_re (?i)сотн\p{L}* тысяч|нескольк\p{L}* сот\p{L}* тысяч|[1-9]00\s*000 +pattern shushiwan_fire_re десятк\p{L}* тысяч +# percent 成=tenths: src expression; error-cue fraction; percent-form suppressor word. +pattern cheng_re ([0-9一二三四五六七八九十])成([0-9一二三四五六七八九]?) +pattern decimal_fraction_re десят(?:ая|ых|ой|ые)|сот(?:ая|ых|ой|ые)|\d+[.,]\d +pattern percent_word процент diff --git a/backend/configs/langpacks/zh-ru/palladius-phonotactics.txt b/backend/configs/langpacks/zh-ru/palladius-phonotactics.txt new file mode 100644 index 00000000..bd762a1b --- /dev/null +++ b/backend/configs/langpacks/zh-ru/palladius-phonotactics.txt @@ -0,0 +1,20 @@ +# Palladius syllable-generator phonotactic constraints (miner_palladius.buildPalladiusCyrSyllables). `categorymember`, pinyin tokens. +# retroflex: retroflex/sibilant initials that take SPECIAL_I, not a bare "i" final. +retroflex zh +retroflex ch +retroflex sh +retroflex r +retroflex z +retroflex c +retroflex s +# vfinal: the ü-finals (pinyin writes ü as plain u); legal only after a vfinal_initial. +vfinal v +vfinal ve +vfinal van +vfinal vn +# vfinal_initial: the initials after which a ü-final is legal (j/q/x + l/n). +vfinal_initial j +vfinal_initial q +vfinal_initial x +vfinal_initial l +vfinal_initial n diff --git a/backend/configs/langpacks/zh/sentence-terminator.txt b/backend/configs/langpacks/zh/sentence-terminator.txt new file mode 100644 index 00000000..80e5b69f --- /dev/null +++ b/backend/configs/langpacks/zh/sentence-terminator.txt @@ -0,0 +1,2 @@ +# Source sentence-ending marks the alias miner splits on (alias.cooccur_same_sentence 。!?). A rune SET; stored sorted. The miner adds "\n" itself. +。!? diff --git a/backend/configs/langpacks/zh/surnames-compound.txt b/backend/configs/langpacks/zh/surnames-compound.txt index efe3bd63..d2754afe 100644 --- a/backend/configs/langpacks/zh/surnames-compound.txt +++ b/backend/configs/langpacks/zh/surnames-compound.txt @@ -3,7 +3,6 @@ 东方 公孙 南宫 -古月 司徒 司马 夏侯 diff --git a/backend/configs/langpacks/zh/title-formant.txt b/backend/configs/langpacks/zh/title-formant.txt new file mode 100644 index 00000000..b1d6e3e0 --- /dev/null +++ b/backend/configs/langpacks/zh/title-formant.txt @@ -0,0 +1,2 @@ +# Rank/measure formant chars a candidate treats as a TITLE (patterns.formant_type 等/转/阶 → title). A rune SET; stored sorted. +等转阶 diff --git a/backend/configs/pipeline-arm-deepseek-pro.yaml b/backend/configs/pipeline-arm-deepseek-pro.yaml index fdac0417..548a0559 100644 --- a/backend/configs/pipeline-arm-deepseek-pro.yaml +++ b/backend/configs/pipeline-arm-deepseek-pro.yaml @@ -37,7 +37,7 @@ stages: role: translator model: deepseek-v4-flash prompts: - zh-ru: ../prompts/translator.md + zh-ru: ../prompts/zh-ru/translator.md prompt_version: v1-reflow temperature: 0.3 reasoning: "off" @@ -48,7 +48,7 @@ stages: # через ReasoningNone no-op). Пиннут (editor не эскалирует — D12). model: deepseek-v4-pro prompts: - zh-ru: ../prompts/editor.md + zh-ru: ../prompts/zh-ru/editor.md prompt_version: v3-discourse-reflow temperature: 0.4 reasoning: "off" # NO-OP на deepseek-v4-pro (ReasoningNone) → thinking ON; НЕ вооружает эхо-мину diff --git a/backend/configs/pipeline-arm-glm.yaml b/backend/configs/pipeline-arm-glm.yaml index 20b5d855..24b33fa5 100644 --- a/backend/configs/pipeline-arm-glm.yaml +++ b/backend/configs/pipeline-arm-glm.yaml @@ -36,7 +36,7 @@ stages: role: translator model: deepseek-v4-flash prompts: - zh-ru: ../prompts/translator.md + zh-ru: ../prompts/zh-ru/translator.md prompt_version: v1-reflow temperature: 0.3 reasoning: "off" @@ -47,7 +47,7 @@ stages: # P1a-дискурс few-shot держится (glm не reasoning-модель, CoT-конфликта нет, в отличие от dspro-арма). model: glm-5 prompts: - zh-ru: ../prompts/editor.md + zh-ru: ../prompts/zh-ru/editor.md prompt_version: v3-discourse-reflow temperature: 0.4 reasoning: "off" # glm thinking:disabled — таймауты ×3 (полигон) diff --git a/backend/configs/pipeline-arm-mistral.yaml b/backend/configs/pipeline-arm-mistral.yaml index b177d338..02ed1ca6 100644 --- a/backend/configs/pipeline-arm-mistral.yaml +++ b/backend/configs/pipeline-arm-mistral.yaml @@ -35,7 +35,7 @@ stages: role: translator model: deepseek-v4-flash # черновик — тот же, что базлайн (свап только редактора) prompts: - zh-ru: ../prompts/translator.md + zh-ru: ../prompts/zh-ru/translator.md prompt_version: v1-reflow temperature: 0.3 reasoning: "off" @@ -46,7 +46,7 @@ stages: # WS6(г); не reasoning-модель → эхо-мины нет). Идёт ТОЛЬКО за rate-guard (models.yaml rate_limit). model: mistral-large-2512 prompts: - zh-ru: ../prompts/editor.md + zh-ru: ../prompts/zh-ru/editor.md prompt_version: v3-discourse-reflow temperature: 0.4 reasoning: "off" diff --git a/backend/configs/pipeline-c1.yaml b/backend/configs/pipeline-c1.yaml index ccb303a8..7b1475d6 100644 --- a/backend/configs/pipeline-c1.yaml +++ b/backend/configs/pipeline-c1.yaml @@ -51,7 +51,7 @@ stages: # loadTemplates (закрывает латентный баг «ja едет через китайский паратаксис»). Язык промпта — # свойство пакета. Прод zh→ru резолвит тот же translator.md → тот же SHA (поведение-нейтрально). prompts: - zh-ru: ../prompts/translator.md + zh-ru: ../prompts/zh-ru/translator.md # D30.2 reflow: из translator.md убрано «Сохраняй разбивку на абзацы» (построчная # вёрстка исходника деструктивна для ru — корень нечитаемости exp12 §2.5-S). Bump # v0-draft→v1-reflow: сменил SHA промпта, осознанный --resnapshot. @@ -76,11 +76,11 @@ stages: model: deepseek-v4-pro # Слой 2 (D39): pair-keyed пакет конвенций (см. draft-стадию) — наполнена zh→ru. prompts: - zh-ru: ../prompts/editor.md + zh-ru: ../prompts/zh-ru/editor.md # Bump v2-bilingual-reflow→v3-discourse-reflow: в editor.md вшито валидированное P1a-ядро # ДИСКУРС-ПЕРЕВЁРСТКИ + few-shot (exp14/D37 §2а — претензия-1 «рубленые абзацы» = рычаг ПРОМПТА, # не модель; P1a реверстает дефект у всех семей). Меняет SHA промпта → осознанный --resnapshot. - # Билингв-каркас (D30.1) и omission-осторожность (D34.3) сохранены. Моно-вариант — prompts/editor-mono.md. + # Билингв-каркас (D30.1) и omission-осторожность (D34.3) сохранены. Моно-вариант — prompts/zh-ru/editor-mono.md. prompt_version: v3-discourse-reflow temperature: 0.4 reasoning: "off" # NO-OP на deepseek-v4-pro (ReasoningNone) → thinking ON; НЕ вооружает эхо-мину diff --git a/backend/configs/pipeline-c2.yaml b/backend/configs/pipeline-c2.yaml index 6963ba66..3841fe04 100644 --- a/backend/configs/pipeline-c2.yaml +++ b/backend/configs/pipeline-c2.yaml @@ -15,6 +15,15 @@ context: cache_ttl: "5m" # STMDepth/OverlapTokens сняты (WS2 §2а — carryover не строим). +# Явный zh-ru пар-калибровочный блок (pair-14 §7 — все шиппинг-конфиги ставят его ЯВНО, пар-конфиг = +# источник истины; значения = дореформенный Go-fallback → байт-нейтрально, но не полагаемся на fallback). +segmentation: + draft_budget_out: 1797 + edit_ceiling_out: 3200 + fertility: + cjk: 1.1978 + other: 0.3852 + retries: regenerate_before_escalate: 1 @@ -23,7 +32,7 @@ stages: role: translator model: deepseek-v4-flash prompts: - zh-ru: ../prompts/translator.md + zh-ru: ../prompts/zh-ru/translator.md prompt_version: v1-reflow # D30.2 reflow (тот же translator.md, что в C1) — label-SHA дисциплина temperature: 0.8 # C2: кандидаты сэмплируются горячее (Р2: T=0.6–1.0) reasoning: "off" @@ -35,7 +44,7 @@ stages: role: judge model: glm-5 prompts: - zh-ru: ../prompts/judge-selector.md + zh-ru: ../prompts/zh-ru/judge-selector.md prompt_version: v0-skeleton temperature: 0 reasoning: "off" @@ -47,7 +56,7 @@ stages: # select выше — judge/glm-5, Gemini-слот Ф2 — НЕ трогаем: это судья, не редактор.) model: deepseek-v4-pro prompts: - zh-ru: ../prompts/editor.md + zh-ru: ../prompts/zh-ru/editor.md # Бамп v2-bilingual-reflow→v3-discourse-reflow (label-SHA дисциплина: тот же editor.md, что в C1 — # вшито P1a-ядро ДИСКУРС-ПЕРЕВЁРСТКИ, D37 §2а; лейбл обязан следовать за новым SHA файла). prompt_version: v3-discourse-reflow diff --git a/backend/internal/config/book.go b/backend/internal/config/book.go index faa29dfd..34707ae7 100644 --- a/backend/internal/config/book.go +++ b/backend/internal/config/book.go @@ -65,6 +65,13 @@ type Book struct { // set ⇒ the runner loads the pack in W0, failing LOUD when the pair's catalog dir exists but a file is // missing/corrupt, and folds pack.Version() into the snapshot so a pack edit is a loud --resnapshot (R1). LangpackRoot string `yaml:"langpack_root"` + // LangpackExtend is the optional root of a BOOK-SCOPED langpack overlay (pair-14 §1): a book's PRIVATE + // canon — a clan surname (古月 in 蛊真人 is a book clan, not a 百家姓 family), a sect term — that must reach + // the miner FOR THIS BOOK without polluting the shared pair langpack. Same layout as langpack_root; only + // the files a book ships are present, each UNIONED onto the shared pack (additive) with its bytes folded + // into pack.Version() (a book-canon edit is a loud --resnapshot for this book only). Resolved relative to + // book.yaml. Empty ⇒ no overlay (the shared pack is used as-is). + LangpackExtend string `yaml:"langpack_extend"` // MinedDelta is the optional path to the owner-curated mined-term delta YAML (the W1.5 sign artifact, // R1). Its terms are loaded as Source:"mined" (NOT Source:"seed"), so they fold into the ENRICHED bank // version but NOT the base — adding them moves only snapshot_W2 (the pay-once invariant). Same seedTerm @@ -108,6 +115,7 @@ func LoadBook(path string) (*Book, error) { b.SourceFile = resolve(b.SourceFile) b.GlossarySeed = resolve(b.GlossarySeed) b.LangpackRoot = resolve(b.LangpackRoot) + b.LangpackExtend = resolve(b.LangpackExtend) b.MinedDelta = resolve(b.MinedDelta) b.MinedRejects = resolve(b.MinedRejects) if b.ProjectDB == "" { diff --git a/backend/internal/config/pipeline.go b/backend/internal/config/pipeline.go index 48d08a42..359f4af2 100644 --- a/backend/internal/config/pipeline.go +++ b/backend/internal/config/pipeline.go @@ -361,8 +361,16 @@ func LoadPipeline(path string, models *Models) (*Pipeline, error) { if p.Waves.Workers <= 0 { p.Waves.Workers = 1 } - // WS2 segmentation defaults (ratified zh-ru, §2в): a config that omits the block gets the - // output-token budget + independently-re-derived fertility. Zero/negative → default. + // Segmentation calibration (pair-14 §7). These numbers are the zh-ru PAIR CALIBRATION — the budgets + // tuned for zh→ru and the fertility (est_out per source char-class) independently re-derived on the zh-ru + // rerun (R²=0.96). They are NOT a language-neutral engine constant: a new pair must set its OWN + // calibration in its pair-config, so every SHIPPING pipeline (pipeline-c1 / the arm yamls) sets the whole + // block explicitly — the pair-config is the source of truth ("brать из пар-конфига"). The literals below + // are only the last-resort GENERIC FALLBACK for a config that omits the block. The values are held EXACT + // on purpose: a book with no langpack chunks entirely on these (the ja→ru golden fixture relies on this + // fallback — a re-derived number would shift its chunk boundaries → the wire). Relocating the canonical + // zh-ru calibration into the langpack was considered and NOT done: it would be dead for every live path + // (shipping configs set the block; the golden has no langpack to read it from) — least-mechanism §12.1. if p.Segmentation.DraftBudgetOut <= 0 { p.Segmentation.DraftBudgetOut = 1797 } diff --git a/backend/internal/config/prompt_pack_test.go b/backend/internal/config/prompt_pack_test.go index d35b35fd..ffc79462 100644 --- a/backend/internal/config/prompt_pack_test.go +++ b/backend/internal/config/prompt_pack_test.go @@ -132,8 +132,8 @@ func TestBoevoyPipelineC1ResolvesZhRu(t *testing.T) { for _, st := range p.Stages { if st.Role == "translator" { zh, _ := st.PromptPathFor("zh-ru") - if !strings.HasSuffix(zh, filepath.Join("prompts", "translator.md")) { - t.Errorf("draft zh-ru must be prompts/translator.md, got %q", zh) + if !strings.HasSuffix(zh, filepath.Join("prompts", "zh-ru", "translator.md")) { + t.Errorf("draft zh-ru must be prompts/zh-ru/translator.md, got %q", zh) } } } diff --git a/backend/internal/lang/data/cjk-section.txt b/backend/internal/lang/data/cjk-section.txt new file mode 100644 index 00000000..d4f0fced --- /dev/null +++ b/backend/internal/lang/data/cjk-section.txt @@ -0,0 +1,34 @@ +# Universal CJK section-numeral data (pair-14 §4 + addendum-A): the ONE source of truth for the chapter/ +# section numeral inventory shared by the ingest chapter-splitter (章节節回 headers) and the chunker heading +# rule. A linguistic CONSTANT (the CJK numeral system does not vary by book/pair), so it is embedded in the +# language-data package and available even to a book with NO langpack (the ja→ru golden splits 第X章 here). +# `categoryrune[value]`. digit/unit carry a value; zero/chapter_unit are membership only. +digit 一 1 +digit 二 2 +digit 两 2 +digit 兩 2 +digit 三 3 +digit 四 4 +digit 五 5 +digit 六 6 +digit 七 7 +digit 八 8 +digit 九 9 +unit 十 10 +unit 百 100 +unit 千 1000 +zero 〇 +zero 零 +# chapter_unit: the section-level markers auto-detected in a source txt (第N章/节/節/回). 卷 (volume) is +# deliberately excluded (coarser than a chapter). +chapter_unit 章 +chapter_unit 节 +chapter_unit 節 +chapter_unit 回 +# magnitude: the myriad-scale BIG units (万/億/兆) → value, for the number-magnitude gate (cheapgates). ONE +# source with digit/unit above — the gate's CJK numeral parser reads this section instead of its own switch. +magnitude 万 10000 +magnitude 萬 10000 +magnitude 億 100000000 +magnitude 亿 100000000 +magnitude 兆 1000000000000 diff --git a/backend/internal/lang/data/injection.txt b/backend/internal/lang/data/injection.txt new file mode 100644 index 00000000..f0d9e912 --- /dev/null +++ b/backend/internal/lang/data/injection.txt @@ -0,0 +1,11 @@ +# Target-language glossary/injection wire-text (pair-14 §2). Rendered into the model request, so it is +# TARGET-language data (a Russian header for a →ru book); gated on the target so a non-ru book never +# gets a Russian block. The ja→ru golden injects these with NO langpack, so they are engine/target data +# here, not a book pack. `targetkeyvalue`; VALUE IS VERBATIM (leading spaces are significant). +# Add a target = add its rows; remove one = delete its rows. +ru glossary_header ГЛОССАРИЙ (используй эти утверждённые переводы имён и терминов последовательно; строки с пометкой ⟨проверить⟩ — неподтверждённые кандидаты): +ru editor_header КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имён и терминов (в черновике термин исходника слева ДОЛЖЕН быть передан именно указанной формой справа — приводи к ней любые расхождения, склоняя по контексту; не вводи иных вариантов и не меняй ничего другого): +ru unverified_marker ⟨проверить⟩ +ru gender_male (муж. — мужские родовые формы) +ru gender_female (жен. — женские родовые формы) +ru gender_hidden (пол СКРЫТ до раскрытия — избегай родовых форм; при неизбежности — мужские) diff --git a/backend/internal/lang/data/refusal.txt b/backend/internal/lang/data/refusal.txt new file mode 100644 index 00000000..7cc874cb --- /dev/null +++ b/backend/internal/lang/data/refusal.txt @@ -0,0 +1,25 @@ +# Universal refusal blacklist (pair-14 §3, ported 1:1 from eval/refusal_bench.py). Engine safety data: +# a model can refuse in ANY language regardless of the book's pair, so EVERY pattern here fires for EVERY +# book (the ja→ru golden flags a Russian refusal with NO langpack). ONE regex per non-comment line, joined +# with "|" under (?im); order is immaterial to the boolean match but preserved. +# NOTE: the `# --- xx ---` markers are ORGANIZATIONAL comments, NOT functional scoping — RefusalPatterns() +# reads every line flat into the one universal set. Deleting a language's lines TRIMS that universal set +# (drops those phrases for ALL books); it does NOT "turn off" detection for one pair. Add a language's phrases +# = add lines under (or beside) its marker. +# --- en --- +I (?:can(?:'|no)t|cannot|won'?t|am (?:not able|unable)) (?:to )?(?:assist|help|translate|continue|provide|fulfill) +I'?m sorry,? but +against (?:my|our) (?:content )?polic +content polic(?:y|ies) +# --- ru --- +не могу (?:помочь|перевести|продолжить|выполнить) +я не буду переводить +противоречит (?:политике|правилам) +# --- ja --- +としてお答えできません +申し訳ありませんが +# --- zh --- +我(?:不能|无法)(?:协助|翻译|继续) +抱歉[,,] +# --- universal explicit marker --- +^\s*\[?(?:REFUSED|ОТКАЗ)\]?\s*$ diff --git a/backend/internal/lang/data/sentence-abbrev.txt b/backend/internal/lang/data/sentence-abbrev.txt new file mode 100644 index 00000000..c6aa2d4e --- /dev/null +++ b/backend/internal/lang/data/sentence-abbrev.txt @@ -0,0 +1,35 @@ +# Source-language sentence-splitter abbreviations (pair-14 §4): trailing tokens after which a lone "." is an +# abbreviation, not a sentence end (case-insensitive). Per SOURCE language — a CJK source uses 。 and needs +# none; `en` is the one that guards an ASCII period. `langtoken`. Add a source = add its section. +# NOTE: the chunker currently consumes the `en` section universally (its historical en-only design); making +# the splitter consume the BOOK's source section is a shallow follow-up (thread sourceLang into SplitChunks). +# --- en --- +en mr +en mrs +en ms +en dr +en prof +en st +en jr +en sr +en vs +en no +en vol +en ch +en fig +en col +en gen +en sgt +en capt +en lt +en rev +en gov +en sen +en rep +en etc +en inc +en ltd +en co +en mt +en ave +en rd diff --git a/backend/internal/lang/data/target-ru.txt b/backend/internal/lang/data/target-ru.txt new file mode 100644 index 00000000..13d9d514 --- /dev/null +++ b/backend/internal/lang/data/target-ru.txt @@ -0,0 +1,65 @@ +# Target-language (ru) readability checker DATA (pair-14 data-out): the wordlists and DETECTION patterns the +# TARGET-GENERAL checkers/sanitizer run on ANY →ru output (regardless of source pair). Engine/target data — +# a no-langpack →ru book (the ja→ru golden) still runs these — so it is embedded here, not a book pack; the +# checker/sanitizer ALGORITHMS stay generic Go. `categoryvalue`; value VERBATIM (a regex or a literal); +# lines are grouped by category IN ORDER (order matters for the sanitizer preamble preview). Add a target = +# add data/target-.txt + its accessor entry. +# +# broken-word: a Russian word ending in this suffix is a mangled infinitive (no valid word ends in «-йть»). +broken_suffix йть +# yofikator homograph whitelist: е-spellings that are DISTINCT words from their ё-counterpart (все≠всё, +# небо≠нёбо, берет≠берёт …), so the inconsistency rule must NOT flag «все»+«всё» as one word two ways. +yo_homograph все +yo_homograph всех +yo_homograph всем +yo_homograph всеми +yo_homograph небо +yo_homograph небом +yo_homograph узнаем +yo_homograph узнаете +yo_homograph узнает +yo_homograph падеж +yo_homograph падежа +yo_homograph совершенный +yo_homograph совершенное +yo_homograph совершенная +yo_homograph совершенно +yo_homograph совершенны +yo_homograph чем +yo_homograph тем +yo_homograph тема +yo_homograph теме +yo_homograph темы +yo_homograph берет +yo_homograph берете +yo_homograph берета +yo_homograph осел +yo_homograph осла +yo_homograph ослы +yo_homograph ослов +yo_homograph мел +yo_homograph мела +yo_homograph лен +yo_homograph лена +yo_homograph нес +yo_homograph несем +yo_homograph несет +yo_homograph вел +yo_homograph ведет +yo_homograph ведем +# magnitude_stem: Russian magnitude WORD stem → base-10 exponent, for the number-magnitude gate (a stem +# covers [base, base+2] since an unseen multiplier can lift it two orders). `magnitude_stemstemexp`. +magnitude_stem тысяч 3 +magnitude_stem миллион 6 +magnitude_stem миллиард 9 +magnitude_stem триллион 12 +# sanitizer DETECTION patterns (pair-14 data-out). The sanitizer is isRuTarget-gated, so these are +# ru-target data; the strip/detect ALGORITHM stays in sanitizer.go. `sanitizer_preamble` is ORDERED +# (detectLeadingPreamble returns the FIRST match's preview). Values VERBATIM (regexes). +sanitizer_preamble (?is)^\s*(?:вот|представляю|привожу|держите)\s+[^\n:]{0,40}?(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?: +sanitizer_preamble (?is)^\s*(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?: +sanitizer_preamble (?is)^\s*(?:вот\s+)?перевод\s+(?:фрагмента|текста|отрывка|главы|черновика)\s*: +sanitizer_preamble (?is)^\s*ниже\s+(?:привед[её]н|представлен|дан|след)[а-яё]*\s+[^\n:]{0,40}?(?:перевод|отредактированн|текст)[а-яё]*[^\n:]{0,20}?: +sanitizer_trailing_note (?i)^[*_>\s-]{0,4}(?:(?:примечани[ея]|заметк[аи]|комментари[йяю]|пояснени[ея]|сноск[аи])\s*:|(?:примечани[ея]|заметк[аи]|комментари[йяю])\s+(?:переводчик|редактор)[а-яё]*|прим\.\s*(?:перев|ред)\.?) +sanitizer_edit_meta (?im)^[*_>\s-]{0,4}(?:(?:внесённ|внесен)[а-яё]+\s+правк|основн[а-яё]+\s+правк[а-яё]*(?:\s+для\s+справк[а-яё]+|\s*:)|что\s+(?:было\s+)?(?:изменен|исправлен)|список\s+(?:правок|изменений)|список\s+внесённых) +sanitizer_invalid_sign (?i)(?:^|\P{L})[ъь][а-яё]|[ъь][ъь] diff --git a/backend/internal/lang/embedded.go b/backend/internal/lang/embedded.go new file mode 100644 index 00000000..8e73d044 --- /dev/null +++ b/backend/internal/lang/embedded.go @@ -0,0 +1,392 @@ +package lang + +import ( + "embed" + "fmt" + "sort" + "strconv" + "strings" + "sync" +) + +// embedded.go: language data the engine needs even for a book with NO langpack pack — engine-universal +// (a CJK numeral is a linguistic constant; a refusal phrase is engine safety behaviour) or target-generic +// (the Russian glossary-injection wire-text). It is versioned with the engine BINARY (like code), not with +// a book's pack, because a no-pack book still splits 第X章 chapters, still flags a refusal, and still injects +// a Russian glossary block (the ja→ru golden proves all three). Kept OUT of internal/pipeline so the +// data/algorithm boundary holds; embedded so it is available without a book's langpack_root. Each file is +// SECTIONED per language/target, so adding a language = add a section, removing one = delete its section. + +//go:embed data/cjk-section.txt data/refusal.txt data/injection.txt data/sentence-abbrev.txt data/target-ru.txt +var embeddedFS embed.FS + +// targetCheckFiles maps a TARGET language to its embedded readability-checker data file (pair-14 data-out). +// Add a target = add data/target-.txt to the go:embed line above and a row here. +var targetCheckFiles = map[string]string{"ru": "data/target-ru.txt"} + +// TargetChecks is the TARGET-language readability wordlists + DETECTION patterns (pair-14 data-out) the +// target-general checkers/sanitizer run on ANY →target output. A target with no file yields an empty value +// (HasData()==false → the consumers stay inert). Values are grouped per category in AUTHORED ORDER. +type TargetChecks struct{ byKey map[string][]string } + +// List returns the authored-order values for a category ("" categories → nil). +func (t TargetChecks) List(key string) []string { return t.byKey[key] } + +// HasData reports whether this target ships any checker data (used to gate the target-general checkers so a +// target without data flags nothing rather than mis-firing). +func (t TargetChecks) HasData() bool { return len(t.byKey) > 0 } + +var ( + targetChecksOnce sync.Once + targetChecksBy map[string]TargetChecks +) + +// TargetChecksFor returns the readability-checker data for a TARGET language, parsed once from the embedded +// per-target file. Unknown/absent target → empty TargetChecks. Panics on a corrupt embed. +func TargetChecksFor(tgt string) TargetChecks { + targetChecksOnce.Do(func() { + targetChecksBy = map[string]TargetChecks{} + for lang, file := range targetCheckFiles { + byKey, err := parseOrderedByCategory(mustEmbed(file)) + if err != nil { + panic(fmt.Sprintf("lang: embedded %s is corrupt: %v", file, err)) + } + targetChecksBy[lang] = TargetChecks{byKey: byKey} + } + }) + return targetChecksBy[tgt] +} + +// parseOrderedByCategory reads `categoryvalue` lines into category → ordered values. Value is VERBATIM +// (a regex may carry a trailing metachar), so only the whole line's \r is stripped; blank/`#`-comment lines +// (by their trimmed form) are dropped. +func parseOrderedByCategory(b []byte) (map[string][]string, error) { + out := map[string][]string{} + for i, raw := range strings.Split(string(b), "\n") { + line := strings.TrimRight(raw, "\r") + if strings.TrimSpace(line) == "" || strings.HasPrefix(strings.TrimSpace(line), "#") { + continue + } + f := strings.SplitN(line, "\t", 2) + if len(f) != 2 || strings.TrimSpace(f[0]) == "" || f[1] == "" { + return nil, fmt.Errorf("line %d: want `categoryvalue`, got %q", i+1, line) + } + out[f[0]] = append(out[f[0]], f[1]) + } + return out, nil +} + +// CJKSection is the universal chapter/section-numeral inventory shared by the ingest chapter-splitter and +// the chunker heading rule (pair-14 §4 + addendum-A: ONE source, no ingest↔chunker byte-duplication). The +// heading-numeral membership test derives from Digit ∪ Unit ∪ Zero — nothing is stored twice. +type CJKSection struct { + Digit map[rune]int // 一→1 … 九→9 (incl. 两/兩→2) + Unit map[rune]int // 十→10 百→100 千→1000 + Zero map[rune]bool // 〇 零 (positional zero: value*10 in a numeral run) + // ChapterUnit membership + ChapterUnitOrdered (authored order) — ingest's detectChapterUnit iterates in + // order and breaks a tie by first-seen, so the ORDER is behaviour, not just a set. + ChapterUnit map[rune]bool + ChapterUnitOrdered []rune + // Magnitude: the myriad-scale big units (万/億/兆 → value) for the number-magnitude gate. int64 (10^12 + // overflows int32). ONE source with Digit/Unit — the gate no longer keeps its own CJK numeral switch. + Magnitude map[rune]int64 + headingRunes string // sorted Zero∪Digit∪Unit, cached for the ingest chapter-numeral regex class +} + +// IsNumeralRune reports whether r can be part of a CJK numeral EXPRESSION (Zero∪Digit∪Unit∪Magnitude) — the +// alphabet of a magnitude candidate run (the gate adds ASCII digits itself). BigUnit returns a big unit's +// value (万/億/兆). Both read the single CJKSection source (pair-14 data-out dedup). +func (c *CJKSection) IsNumeralRune(r rune) bool { + if c.IsHeadingNumeral(r) { + return true + } + _, ok := c.Magnitude[r] + return ok +} +func (c *CJKSection) BigUnit(r rune) (int64, bool) { v, ok := c.Magnitude[r]; return v, ok } + +// IsHeadingNumeral reports whether r can be part of a CJK chapter-number run (Zero ∪ Digit ∪ Unit). Arabic +// digits are handled by the caller (they are not language data). +func (c *CJKSection) IsHeadingNumeral(r rune) bool { + if c.Zero[r] { + return true + } + if _, ok := c.Digit[r]; ok { + return true + } + _, ok := c.Unit[r] + return ok +} + +// HeadingNumeralClass returns the CJK heading-numeral runes (Zero∪Digit∪Unit) as a sorted string, for use +// INSIDE a regex character class (the ingest chapter-numeral pattern). Sorted for a stable pattern; order +// inside a class is immaterial to matching. None of the runes are class metacharacters, so no escaping. +func (c *CJKSection) HeadingNumeralClass() string { return c.headingRunes } + +var ( + cjkOnce sync.Once + cjkVal *CJKSection + cjkErr error +) + +// DefaultCJKSection returns the process-wide CJK section-numeral data, parsed once from the embedded file. +// It panics on a corrupt embed (a build-time asset the tests exercise, so a bad edit fails the suite, never +// production silently) — the parse is deterministic and input-free. +func DefaultCJKSection() *CJKSection { + cjkOnce.Do(func() { cjkVal, cjkErr = parseCJKSection(mustEmbed("data/cjk-section.txt")) }) + if cjkErr != nil { + panic(fmt.Sprintf("lang: embedded cjk-section.txt is corrupt: %v", cjkErr)) + } + return cjkVal +} + +func parseCJKSection(b []byte) (*CJKSection, error) { + c := &CJKSection{Digit: map[rune]int{}, Unit: map[rune]int{}, Zero: map[rune]bool{}, ChapterUnit: map[rune]bool{}, Magnitude: map[rune]int64{}} + valueRune := func(i int, f []string) (rune, int, error) { + if len(f) != 3 { + return 0, 0, fmt.Errorf("line %d: %s wants `%srunevalue`, got %q", i+1, f[0], f[0], strings.Join(f, "\t")) + } + r := []rune(f[1]) + if len(r) != 1 { + return 0, 0, fmt.Errorf("line %d: key %q must be a single rune", i+1, f[1]) + } + v, err := strconv.Atoi(strings.TrimSpace(f[2])) + if err != nil { + return 0, 0, fmt.Errorf("line %d: value %q: %w", i+1, f[2], err) + } + return r[0], v, nil + } + setRune := func(i int, f []string) (rune, error) { + if len(f) != 2 { + return 0, fmt.Errorf("line %d: %s wants `%srune`, got %q", i+1, f[0], f[0], strings.Join(f, "\t")) + } + r := []rune(f[1]) + if len(r) != 1 { + return 0, fmt.Errorf("line %d: key %q must be a single rune", i+1, f[1]) + } + return r[0], nil + } + for i, raw := range strings.Split(string(b), "\n") { + t := strings.TrimRight(raw, "\r") + if strings.TrimSpace(t) == "" || strings.HasPrefix(strings.TrimSpace(t), "#") { + continue + } + f := strings.Split(t, "\t") + switch f[0] { + case "digit": + r, v, err := valueRune(i, f) + if err != nil { + return nil, err + } + c.Digit[r] = v + case "unit": + r, v, err := valueRune(i, f) + if err != nil { + return nil, err + } + c.Unit[r] = v + case "zero": + r, err := setRune(i, f) + if err != nil { + return nil, err + } + c.Zero[r] = true + case "chapter_unit": + r, err := setRune(i, f) + if err != nil { + return nil, err + } + if !c.ChapterUnit[r] { + c.ChapterUnit[r] = true + c.ChapterUnitOrdered = append(c.ChapterUnitOrdered, r) + } + case "magnitude": + if len(f) != 3 { + return nil, fmt.Errorf("line %d: magnitude wants `magnituderunevalue`, got %q", i+1, t) + } + r := []rune(f[1]) + if len(r) != 1 { + return nil, fmt.Errorf("line %d: magnitude key %q must be a single rune", i+1, f[1]) + } + v, err := strconv.ParseInt(strings.TrimSpace(f[2]), 10, 64) + if err != nil { + return nil, fmt.Errorf("line %d: magnitude value %q: %w", i+1, f[2], err) + } + c.Magnitude[r[0]] = v + default: + return nil, fmt.Errorf("line %d: unknown category %q (want digit|unit|zero|chapter_unit|magnitude)", i+1, f[0]) + } + } + if len(c.Digit) == 0 || len(c.Unit) == 0 || len(c.Zero) == 0 || len(c.ChapterUnit) == 0 || len(c.Magnitude) == 0 { + return nil, fmt.Errorf("cjk-section needs non-empty digit, unit, zero, chapter_unit and magnitude sections") + } + // Cache the sorted heading-numeral class (Zero∪Digit∪Unit) for the ingest regex. + var runes []rune + for r := range c.Zero { + runes = append(runes, r) + } + for r := range c.Digit { + runes = append(runes, r) + } + for r := range c.Unit { + runes = append(runes, r) + } + sort.Slice(runes, func(i, j int) bool { return runes[i] < runes[j] }) + c.headingRunes = string(runes) + return c, nil +} + +// InjectionTexts is the TARGET-language glossary/injection wire-text (pair-14 §2): the Russian headers and +// annotations the memory renderers emit into a request. TARGET data (a Russian header for a →ru book), so +// the renderers gate on HasData — a target with no injection texts (e.g. →en) injects NOTHING rather than a +// stray Russian block. The ja→ru golden renders these with NO book pack, so they are engine/target data. +type InjectionTexts struct { + GlossaryHeader string // "ГЛОССАРИЙ (…):" — introduces the draft glossary block + EditorHeader string // "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ …:" — introduces the editor constraint block + UnverifiedMarker string // " ⟨проверить⟩" — tags an unverified draft candidate (leading spaces significant) + GenderMale string // " (муж. — …)" DC3 gender directive (leading space significant) + GenderFemale string // " (жен. — …)" + GenderHidden string // " (пол СКРЫТ …)" +} + +// HasData reports whether the target has injection texts (its header is present). The renderers use it to +// gate: a target without texts injects nothing, so a non-ru book never gets a Russian glossary block. +func (t InjectionTexts) HasData() bool { return t.GlossaryHeader != "" } + +var ( + injectionOnce sync.Once + injectionBy map[string]*InjectionTexts + injectionErr error +) + +// InjectionTextsFor returns the injection wire-text for a TARGET language (pair-14 §2), parsed once from the +// embedded file. A target with no rows returns a zero value (HasData()==false → no injection). Panics on a +// corrupt embed. VALUE bytes are verbatim (leading spaces in gender_*/unverified_marker are significant). +func InjectionTextsFor(targetLang string) InjectionTexts { + injectionOnce.Do(func() { injectionBy, injectionErr = parseInjection(mustEmbed("data/injection.txt")) }) + if injectionErr != nil { + panic(fmt.Sprintf("lang: embedded injection.txt is corrupt: %v", injectionErr)) + } + if t, ok := injectionBy[targetLang]; ok { + return *t + } + return InjectionTexts{} +} + +func parseInjection(b []byte) (map[string]*InjectionTexts, error) { + out := map[string]*InjectionTexts{} + for i, raw := range strings.Split(string(b), "\n") { + line := strings.TrimRight(raw, "\r") + if strings.TrimSpace(line) == "" || strings.HasPrefix(strings.TrimSpace(line), "#") { + continue + } + f := strings.SplitN(line, "\t", 3) // value (f[2]) VERBATIM — leading spaces are significant + if len(f) != 3 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" { + return nil, fmt.Errorf("line %d: want `targetkeyvalue`, got %q", i+1, line) + } + tgt, key, val := f[0], f[1], f[2] + if out[tgt] == nil { + out[tgt] = &InjectionTexts{} + } + t := out[tgt] + switch key { + case "glossary_header": + t.GlossaryHeader = val + case "editor_header": + t.EditorHeader = val + case "unverified_marker": + t.UnverifiedMarker = val + case "gender_male": + t.GenderMale = val + case "gender_female": + t.GenderFemale = val + case "gender_hidden": + t.GenderHidden = val + default: + return nil, fmt.Errorf("line %d: unknown injection key %q", i+1, key) + } + } + return out, nil +} + +var ( + refusalOnce sync.Once + refusalVal []string +) + +// RefusalPatterns returns the universal refusal-blacklist regex patterns (pair-14 §3), parsed once from the +// embedded sectioned file in authored order. Engine safety data: a model can refuse in any language for any +// pair, so this is not book-pack data — the caller joins the patterns into one case-insensitive regex. Each +// non-comment line is ONE pattern, taken verbatim (a regex may contain leading/trailing metacharacters, so +// it is NOT trimmed beyond the line's \r). Panics on a corrupt embed. +func RefusalPatterns() []string { + refusalOnce.Do(func() { refusalVal = embeddedLines(mustEmbed("data/refusal.txt")) }) + return refusalVal +} + +// embeddedLines returns the non-comment, non-blank lines of an embedded file VERBATIM (only the trailing \r +// is stripped; the content is NOT trimmed — significant for a regex pattern). A blank/`#`-comment line is +// dropped by its TRIMMED form, but a kept line keeps its own leading/trailing bytes. +func embeddedLines(b []byte) []string { + var out []string + for _, raw := range strings.Split(string(b), "\n") { + s := strings.TrimRight(raw, "\r") + if strings.TrimSpace(s) == "" || strings.HasPrefix(strings.TrimSpace(s), "#") { + continue + } + out = append(out, s) + } + return out +} + +var ( + abbrevOnce sync.Once + abbrevBy map[string]map[string]bool +) + +// SentenceAbbrev returns the lower-cased sentence-splitter abbreviation SET for a SOURCE language (pair-14 +// §4), parsed once from the embedded sectioned file. An unknown/absent language returns an empty set (a +// source with no ASCII-period abbreviations — e.g. a CJK source using 。). Panics on a corrupt embed. +func SentenceAbbrev(srcLang string) map[string]bool { + abbrevOnce.Do(func() { + var err error + abbrevBy, err = parseSectionedSet(mustEmbed("data/sentence-abbrev.txt")) + if err != nil { + panic(fmt.Sprintf("lang: embedded sentence-abbrev.txt is corrupt: %v", err)) + } + }) + if m, ok := abbrevBy[srcLang]; ok { + return m + } + return map[string]bool{} +} + +// parseSectionedSet reads `keymember` lines into per-key sets (keys lower-cased on read is NOT done +// here — the key is the language tag; members are stored verbatim). Used for the sentence-abbrev table. +func parseSectionedSet(b []byte) (map[string]map[string]bool, error) { + out := map[string]map[string]bool{} + for i, raw := range strings.Split(string(b), "\n") { + t := strings.TrimRight(raw, "\r") + if strings.TrimSpace(t) == "" || strings.HasPrefix(strings.TrimSpace(t), "#") { + continue + } + f := strings.Split(t, "\t") + if len(f) != 2 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" { + return nil, fmt.Errorf("line %d: want `keymember`, got %q", i+1, t) + } + key := strings.TrimSpace(f[0]) + if out[key] == nil { + out[key] = map[string]bool{} + } + out[key][strings.TrimSpace(f[1])] = true + } + return out, nil +} + +func mustEmbed(name string) []byte { + b, err := embeddedFS.ReadFile(name) + if err != nil { + panic(fmt.Sprintf("lang: missing embedded asset %q: %v", name, err)) + } + return b +} diff --git a/backend/internal/lang/langpack.go b/backend/internal/lang/langpack.go index 8b918809..6f635ca9 100644 --- a/backend/internal/lang/langpack.go +++ b/backend/internal/lang/langpack.go @@ -21,6 +21,7 @@ import ( "fmt" "os" "path/filepath" + "strconv" "strings" ) @@ -28,7 +29,9 @@ import ( // a schema change (a new file, a format change). The DATA content is versioned separately by hashing the // files into Version(), so editing a table also invalidates — you cannot forget to bump a version when // you change the data, because the data's bytes ARE the version (the memnorm.go drift-proofing). -const packAlgoVersion = "langpack-v1" +// v2 (pair-14): new pair/source files (title-formant, sentence-terminator, palladius-phonotactics, dc-checkers) +// + new formats (the pattern/category rows, the generic Palladius parser) — a schema change, so the tag bumps. +const packAlgoVersion = "langpack-v2" // Pack is a loaded, versioned language-data pack for one source→target pair. Fields are the DATA the // pipeline algorithms read; the zero value is unusable (load via Load). Maps are membership sets / lookup @@ -46,12 +49,25 @@ type Pack struct { GradePrefix map[rune]bool // grade/stem prefix chars 甲乙丙… Numeral map[rune]bool // CJK numeral chars AliasParticle map[rune]bool // trailing-particle set marking a boundary fragment + // TitleFormant: the rank/measure formant chars a miner classifies as a TITLE (patterns.formant_type + // 等/转/阶 → title). Pair-14: moved out of the pipeline formant switch, sitting beside TopoSuffix (its + // place-side neighbour). A rune SET. + TitleFormant map[rune]bool + // SentenceTerminator: the source sentence-ending marks the alias miner splits on (。!?, alias. + // cooccur_same_sentence). Pair-14: moved out of the pipeline literal. A rune SET; the miner adds "\n" + // (a structural newline, not a language mark) itself. + SentenceTerminator map[rune]bool - // - pair transliteration (configs/langpacks//) — the Palladius (Палладий) table. - PalladiusInitials map[string]string - PalladiusFinals map[string]string - PalladiusYW map[string]string - PalladiusSpecialI map[string]string + // Palladius is the - transliteration table (configs/langpacks//palladius.txt + + // palladius-phonotactics.txt) — the pinyin→Cyrillic maps and the syllable-generator's phonotactic + // constraints, as ONE typed value (owner addendum 24.07: a struct, not four+three parallel Pack fields). + Palladius Palladius + + // DCCheckers is the OPTIONAL pair data for the WS5 defect-class checkers (configs/langpacks// + // dc-checkers.txt, checkers_zh_ru.go). nil when the pair ships no file — the checker ALGORITHMS then read + // empty tables and fire 0 (a pair that does not opt in is never flagged). DATA only (pair-14 §6: the + // LOOKUP TABLES move out of the pipeline; the regex DETECTION patterns stay as the checker algorithm). + DCCheckers *DCCheckerData // Heading is the OPTIONAL chapter-heading rule (configs/langpacks//heading.txt). nil when the pair // carries no heading.txt — the chapter-title feature is then inert (the chunker keeps the source header @@ -63,6 +79,37 @@ type Pack struct { version string } +// Palladius is the pair transliteration table: the pinyin→Cyrillic maps and the syllable-generator's +// phonotactic constraint sets (owner addendum 24.07 — one typed struct instead of four+three parallel Pack +// fields). Populated from palladius.txt (Initials/Finals/YW/SpecialI) + palladius-phonotactics.txt +// (Retroflex/VFinal/VFinalInitial) by the generic category parser; which categories are REQUIRED is enforced +// by validate(), not the parser (an unknown category is collected, never a parse error). +type Palladius struct { + Initials, Finals, YW, SpecialI map[string]string // pinyin→Cyrillic + Retroflex, VFinal, VFinalInitial map[string]bool // phonotactic constraint sets (pinyin tokens) +} + +// newPalladius returns a Palladius with all tables allocated (so merge/union is nil-safe). +func newPalladius() Palladius { + return Palladius{ + Initials: map[string]string{}, Finals: map[string]string{}, YW: map[string]string{}, SpecialI: map[string]string{}, + Retroflex: map[string]bool{}, VFinal: map[string]bool{}, VFinalInitial: map[string]bool{}, + } +} + +// DCCheckerData is the per-pair lookup data the WS5 defect-class checkers read (checkers_zh_ru.go). The +// checker DETECTION regexes stay in the pipeline as the pair-scoped algorithm (§12.2); only the LOOKUP +// TABLES live here as data. Parsed from dc-checkers.txt. +type DCCheckerData struct { + Numeral map[rune]int // DC1: small CJK count → value (一→1 … 十→10, incl. 两→2) + RuHours map[string]int // DC1: Russian hours-count word → value (один→1 … шесть→6) + RegisterNeg []string // DC6: register negative-list lexemes (терем + case forms), authored order + // Patterns are the DETECTION patterns as DATA (pair-14 data-out): key → verbatim regex or literal probe. + // The `*_re` keys are compiled by the consumer; the rest are literal strings.Contains probes. A pair that + // ships no pattern for a key runs that sub-checker inert. + Patterns map[string]string +} + // HeadingRule is the per-pair data for the chapter-title policy: instead of letting the model render a // chapter heading (which drifted to «Раздел 2» / «Первая глава» / an orphaned « :» across models), the // chunker detects a source header (Marker + a numeral + a Unit rune), strips it from the model input, and @@ -78,13 +125,18 @@ type HeadingRule struct { func (p *Pack) Version() string { return p.version } // srcFiles are the source-morphology files (under configs/langpacks//), read in this fixed order. +// SCOPE (honest): this manifest is the zh-family NAME-MINER's morphology schema (百家姓 surnames, 甲乙丙 grade +// prefixes, 等/转/阶 title-formants, CJK numerals, …), and validate() forbids an empty table. So "add a pair = +// drop a directory, no recompile" holds for a source in THIS family (a second CJK source drops its files); a +// different source family that wants mining needs new Pack fields + parse + miner channels, not just a dir. var srcFiles = []string{ "surnames-single.txt", "surnames-compound.txt", "title-suffix.txt", "ordinal-title.txt", "rank-word.txt", "topo-suffix.txt", "grade-prefix.txt", "numeral.txt", "alias-particle.txt", + "title-formant.txt", "sentence-terminator.txt", } // pairFiles are the pair-transliteration files (under configs/langpacks//). -var pairFiles = []string{"palladius.txt"} +var pairFiles = []string{"palladius.txt", "palladius-phonotactics.txt"} // Load resolves and reads the pack for (sourceLang, targetLang) from root (e.g. "configs/langpacks"): the // source-morphology files under root// and the pair-transliteration files under root/-/. @@ -93,7 +145,7 @@ var pairFiles = []string{"palladius.txt"} // files in a fixed order, so it is deterministic and drift-proof. func Load(root, sourceLang, targetLang string) (*Pack, error) { pair := sourceLang + "-" + targetLang - p := &Pack{Pair: pair} + p := &Pack{Pair: pair, Palladius: newPalladius()} h := sha256.New() h.Write([]byte(packAlgoVersion)) @@ -144,6 +196,21 @@ func Load(root, sourceLang, targetLang string) (*Pack, error) { p.Heading = hr } + // Optional per-pair DC-checker tables (pair-14 §6). ABSENT → nil, the checkers read empty tables and fire + // 0 (a pair that does not opt in is never re-billed / flagged); PRESENT → bytes fold into the hash and it + // is parsed; CORRUPT → fail loud. Same optional contract as heading.txt. + if db, ok, derr := readOptional(root, pair, "dc-checkers.txt"); derr != nil { + return nil, fmt.Errorf("langpack %q dc-checkers.txt: %w", pair, derr) + } else if ok { + h.Write([]byte("\x00" + pair + "/dc-checkers.txt\x00")) + h.Write(db) + dc, perr := parseDCCheckers(db) + if perr != nil { + return nil, fmt.Errorf("langpack %q dc-checkers.txt: %w", pair, perr) + } + p.DCCheckers = dc + } + if err := p.validate(); err != nil { return nil, fmt.Errorf("langpack %q: %w", pair, err) } @@ -151,6 +218,138 @@ func Load(root, sourceLang, targetLang string) (*Pack, error) { return p, nil } +// LoadWithOverlay loads the shared pair pack (Load) and then UNIONS a book-scoped OVERLAY on top: a book's +// PRIVATE canon (a clan surname 古月, a sect term) that belongs to ONE book, not the shared pair langpack +// (pair-14 §1 — a book term in the shared pair layer is a leak). overlayRoot has the same layout as root +// (/… + /…); ONLY the files a book chooses to ship are present, each read OPTIONALLY and UNIONED +// (additive — sets gain members, ordered slices append; an overlay never removes a base entry). The overlay +// bytes fold into Version(), so a book-canon edit is a loud --resnapshot for THAT book while the shared pack +// stays byte-stable. overlayRoot == "" ⇒ identical to Load (same *Pack, same Version). +func LoadWithOverlay(root, src, tgt, overlayRoot string) (*Pack, error) { + p, err := Load(root, src, tgt) + if err != nil { + return nil, err + } + if overlayRoot == "" { + return p, nil + } + pair := src + "-" + tgt + // FAIL LOUD on a misnamed overlay file (review finding, pair-14 scale lens): the fold loop below reads + // ONLY the manifest names, so a typo (surname-compound.txt) or a non-mergeable file (heading.txt) shipped + // in an overlay would be SILENTLY ignored — the book's private canon never reaches the miner, recall + // degrades with no load-time signal. Enumerate the overlay dirs and refuse any unexpected file, keeping + // the "never silently empty / drift-proof" guarantee the shared loader makes. + mergeable := map[string]bool{} + for _, n := range srcFiles { + mergeable[src+"/"+n] = true + } + for _, n := range pairFiles { + mergeable[pair+"/"+n] = true + } + for _, dir := range []string{src, pair} { + names, derr := overlayDirFiles(filepath.Join(overlayRoot, dir)) + if derr != nil { + return nil, fmt.Errorf("langpack %q overlay %s: %w", pair, dir, derr) + } + for _, n := range names { + if !mergeable[dir+"/"+n] { + return nil, fmt.Errorf("langpack %q overlay: unexpected file %s/%s — an overlay merges only the source/pair manifest files (a typo, or a non-mergeable file like heading.txt/dc-checkers.txt, would be silently ignored)", pair, dir, n) + } + } + } + // Seed a fresh hash with the base version (which already uniquely encodes every base byte), then fold the + // overlay files in a fixed order — deterministic and drift-proof (edit the overlay → Version() moves). + h := sha256.New() + h.Write([]byte(p.version)) + merged := false + fold := func(dir, name string, pairFile bool) error { + b, ok, rerr := readOptional(overlayRoot, dir, name) + if rerr != nil { + return fmt.Errorf("langpack %q overlay %s/%s: %w", pair, dir, name, rerr) + } + if !ok { + return nil + } + h.Write([]byte("\x00" + dir + "/" + name + "\x00")) + h.Write(b) + merged = true + if pairFile { + return p.mergePair(name, b) + } + return p.mergeSrc(name, b) + } + for _, name := range srcFiles { + if err := fold(src, name, false); err != nil { + return nil, err + } + } + for _, name := range pairFiles { + if err := fold(pair, name, true); err != nil { + return nil, err + } + } + if merged { + p.version = packAlgoVersion + "-x" + hex.EncodeToString(h.Sum(nil))[:12] + } + return p, nil +} + +// mergeSrc unions an overlay source file into the loaded pack (additive; see LoadWithOverlay). Rune/string +// SETS gain members; ordered slices append (a book's extra title/rank tokens follow the shared ones). +func (p *Pack) mergeSrc(name string, b []byte) error { + switch name { + case "surnames-single.txt": + unionRuneSet(p.SurnamesSingle, runeSet(b)) + case "surnames-compound.txt": + unionStringSet(p.SurnamesCompound, stringSet(b)) + case "title-suffix.txt": + p.TitleSuffix = append(p.TitleSuffix, lines(b)...) + case "ordinal-title.txt": + p.OrdinalTitle = append(p.OrdinalTitle, lines(b)...) + case "rank-word.txt": + p.RankWord = append(p.RankWord, lines(b)...) + case "topo-suffix.txt": + unionRuneSet(p.TopoSuffix, runeSet(b)) + case "grade-prefix.txt": + unionRuneSet(p.GradePrefix, runeSet(b)) + case "numeral.txt": + unionRuneSet(p.Numeral, runeSet(b)) + case "alias-particle.txt": + unionRuneSet(p.AliasParticle, runeSet(b)) + case "title-formant.txt": + unionRuneSet(p.TitleFormant, runeSet(b)) + case "sentence-terminator.txt": + unionRuneSet(p.SentenceTerminator, runeSet(b)) + default: + return fmt.Errorf("unknown source file") + } + return nil +} + +// mergePair unions an overlay pair file into the loaded pack (additive). Both pair files carry Palladius +// categories, parsed generically and unioned into p.Palladius (same path as assignPair). +func (p *Pack) mergePair(name string, b []byte) error { + return p.assignPair(name, b) +} + +func unionRuneSet(dst, src map[rune]bool) { + for k := range src { + dst[k] = true + } +} + +func unionStringSet(dst, src map[string]bool) { + for k := range src { + dst[k] = true + } +} + +func unionStringMap(dst, src map[string]string) { + for k, v := range src { + dst[k] = v + } +} + // validate makes the "never silently empty" contract real: a present-but-empty or comment-only data file // (a fat-fingered edit once packs are hand-authored at R1) parses to an empty table with no error and would // silently disable a miner channel — recall degradation with no load-time signal. Every required table must @@ -172,10 +371,16 @@ func (p *Pack) validate() error { req("grade-prefix", len(p.GradePrefix)) req("numeral", len(p.Numeral)) req("alias-particle", len(p.AliasParticle)) - req("palladius/initials", len(p.PalladiusInitials)) - req("palladius/finals", len(p.PalladiusFinals)) - req("palladius/yw", len(p.PalladiusYW)) - req("palladius/special_i", len(p.PalladiusSpecialI)) + req("title-formant", len(p.TitleFormant)) + req("sentence-terminator", len(p.SentenceTerminator)) + // The Palladius REQUIRED-category list lives HERE (the consumer), not in the generic parser (addendum). + req("palladius/initials", len(p.Palladius.Initials)) + req("palladius/finals", len(p.Palladius.Finals)) + req("palladius/yw", len(p.Palladius.YW)) + req("palladius/special_i", len(p.Palladius.SpecialI)) + req("palladius-phonotactics/retroflex", len(p.Palladius.Retroflex)) + req("palladius-phonotactics/vfinal", len(p.Palladius.VFinal)) + req("palladius-phonotactics/vfinal_initial", len(p.Palladius.VFinalInitial)) if len(empty) > 0 { return fmt.Errorf("empty required table(s) %s — a present-but-empty/comment-only data file is a corrupt pack, not a valid one", strings.Join(empty, ", ")) } @@ -202,26 +407,49 @@ func (p *Pack) assignSrc(name string, b []byte) error { p.Numeral = runeSet(b) case "alias-particle.txt": p.AliasParticle = runeSet(b) + case "title-formant.txt": + p.TitleFormant = runeSet(b) + case "sentence-terminator.txt": + p.SentenceTerminator = runeSet(b) default: return fmt.Errorf("unknown source file") } return nil } +// assignPair reads a pair file into p.Palladius. Both pair files (palladius.txt, palladius-phonotactics.txt) +// carry `category…` rows, parsed by ONE generic category parser (owner addendum 24.07); the known +// Palladius categories are UNIONED into p.Palladius and an unknown category is simply ignored, NEVER a parse +// error — which categories are REQUIRED is enforced downstream by validate(), not here. func (p *Pack) assignPair(name string, b []byte) error { - switch name { - case "palladius.txt": - ini, fin, yw, si, err := parsePalladius(b) - if err != nil { - return err - } - p.PalladiusInitials, p.PalladiusFinals, p.PalladiusYW, p.PalladiusSpecialI = ini, fin, yw, si - default: - return fmt.Errorf("unknown pair file") + cats, err := parseCategoryRows(b) + if err != nil { + return err } + p.Palladius.merge(cats) return nil } +// overlayDirFiles lists the regular-file names directly in dir (an overlay's or subdir). A +// MISSING dir is fine (returns nil — a book may overlay only sources or only the pair). Any other read +// error is loud. Nested dirs are ignored (only top-level manifest files are mergeable). +func overlayDirFiles(dir string) ([]string, error) { + ents, err := os.ReadDir(dir) + if err != nil { + if os.IsNotExist(err) { + return nil, nil + } + return nil, err + } + var out []string + for _, e := range ents { + if !e.IsDir() { + out = append(out, e.Name()) + } + } + return out, nil +} + // readOptional reads an OPTIONAL pack file. A missing file returns (nil, false, nil) — the feature it // backs is simply inert — while any OTHER read error (permission, a directory) is a loud failure; a // present file returns (bytes, true, nil). Used for the pack-13 heading rule, which a pair opts into. @@ -314,34 +542,116 @@ func lines(b []byte) []string { return out } -// parsePalladius reads the `categorypinyincyrillic` table into the four maps. It scans the raw -// lines directly (not contentLines) so an error names the PHYSICAL file line — the point of the diagnostic -// is to send a human editing the table to the right line. -func parsePalladius(b []byte) (ini, fin, yw, si map[string]string, err error) { - ini, fin, yw, si = map[string]string{}, map[string]string{}, map[string]string{}, map[string]string{} +// parseCategoryRows is the GENERIC Palladius category parser (owner addendum 24.07): it reads a pair file's +// `categorykey[value]` rows into category → key → value, WITHOUT a per-category switch and WITHOUT +// treating an unknown category as an error (which categories are required is the CONSUMER's call — validate()). +// A 3-field row (initials/finals/yw/special_i) stores key→cyrillic; a 2-field row (retroflex/vfinal/…) stores +// key→"" (a set member). Scans raw lines so an error names the PHYSICAL file line; the only parse errors are a +// bad field count / an empty key. +func parseCategoryRows(b []byte) (map[string]map[string]string, error) { + cats := map[string]map[string]string{} for i, raw := range strings.Split(string(b), "\n") { t := strings.TrimSpace(strings.TrimRight(raw, "\r")) if t == "" || strings.HasPrefix(t, "#") { continue } f := strings.Split(t, "\t") - if len(f) != 3 { - return nil, nil, nil, nil, fmt.Errorf("line %d: want 3 tab-separated fields, got %d (%q)", i+1, len(f), t) + if len(f) < 2 || len(f) > 3 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" { + return nil, fmt.Errorf("line %d: want `categorykey[value]`, got %q", i+1, t) } + if cats[f[0]] == nil { + cats[f[0]] = map[string]string{} + } + val := "" // a 2-field row is a set member (value "") + if len(f) == 3 { + val = f[2] + } + cats[f[0]][f[1]] = val + } + return cats, nil +} + +// merge unions parsed categories into the Palladius table. KNOWN categories populate the typed fields (a map +// category takes key→cyrillic, a set category takes its keys); an UNKNOWN category is simply not consumed — +// validate() enforces that every REQUIRED category ended up non-empty. +func (pal *Palladius) merge(cats map[string]map[string]string) { + unionStringMap(pal.Initials, cats["initials"]) + unionStringMap(pal.Finals, cats["finals"]) + unionStringMap(pal.YW, cats["yw"]) + unionStringMap(pal.SpecialI, cats["special_i"]) + unionKeysAsSet(pal.Retroflex, cats["retroflex"]) + unionKeysAsSet(pal.VFinal, cats["vfinal"]) + unionKeysAsSet(pal.VFinalInitial, cats["vfinal_initial"]) +} + +// unionKeysAsSet adds the KEYS of a parsed category (a 2-field set) to dst. +func unionKeysAsSet(dst map[string]bool, src map[string]string) { + for k := range src { + dst[k] = true + } +} + +// parseDCCheckers reads the DC-checker pair tables from category-keyed lines (checkers_zh_ru.go data): +// +// cjk_numeralrunevalue | ru_hourwordvalue | register_neglexeme +// +// Scans raw lines so an error names the PHYSICAL file line. Fail-loud on a bad field count / non-integer +// value / unknown category (a malformed table is a corrupt pack). RegisterNeg keeps its authored order. +func parseDCCheckers(b []byte) (*DCCheckerData, error) { + d := &DCCheckerData{Numeral: map[rune]int{}, RuHours: map[string]int{}, Patterns: map[string]string{}} + for i, raw := range strings.Split(string(b), "\n") { + // A `pattern` line's VALUE is verbatim (a regex may carry trailing metachars), so only \r is stripped + // from the whole line, not the value; other categories tolerate the trimmed form. + line := strings.TrimRight(raw, "\r") + t := strings.TrimSpace(line) + if t == "" || strings.HasPrefix(t, "#") { + continue + } + if strings.HasPrefix(line, "pattern\t") { + f := strings.SplitN(line, "\t", 3) // value (f[2]) VERBATIM + if len(f) != 3 || strings.TrimSpace(f[1]) == "" || f[2] == "" { + return nil, fmt.Errorf("line %d: pattern wants `patternkeyvalue`, got %q", i+1, line) + } + d.Patterns[f[1]] = f[2] + continue + } + f := strings.Split(t, "\t") switch f[0] { - case "initials": - ini[f[1]] = f[2] - case "finals": - fin[f[1]] = f[2] - case "yw": - yw[f[1]] = f[2] - case "special_i": - si[f[1]] = f[2] + case "cjk_numeral": + if len(f) != 3 { + return nil, fmt.Errorf("line %d: cjk_numeral wants `cjk_numeralrunevalue`, got %q", i+1, t) + } + r := []rune(f[1]) + if len(r) != 1 { + return nil, fmt.Errorf("line %d: cjk_numeral key %q must be a single rune", i+1, f[1]) + } + v, err := strconv.Atoi(strings.TrimSpace(f[2])) + if err != nil { + return nil, fmt.Errorf("line %d: cjk_numeral value %q: %w", i+1, f[2], err) + } + d.Numeral[r[0]] = v + case "ru_hour": + if len(f) != 3 { + return nil, fmt.Errorf("line %d: ru_hour wants `ru_hourwordvalue`, got %q", i+1, t) + } + v, err := strconv.Atoi(strings.TrimSpace(f[2])) + if err != nil { + return nil, fmt.Errorf("line %d: ru_hour value %q: %w", i+1, f[2], err) + } + d.RuHours[f[1]] = v + case "register_neg": + if len(f) != 2 || strings.TrimSpace(f[1]) == "" { + return nil, fmt.Errorf("line %d: register_neg wants `register_neglexeme`, got %q", i+1, t) + } + d.RegisterNeg = append(d.RegisterNeg, f[1]) default: - return nil, nil, nil, nil, fmt.Errorf("line %d: unknown category %q", i+1, f[0]) + return nil, fmt.Errorf("line %d: unknown category %q (want cjk_numeral|ru_hour|register_neg)", i+1, f[0]) } } - return ini, fin, yw, si, nil + if len(d.Numeral) == 0 || len(d.RuHours) == 0 || len(d.RegisterNeg) == 0 { + return nil, fmt.Errorf("dc-checkers needs non-empty cjk_numeral, ru_hour and register_neg sections") + } + return d, nil } // contentLines splits into lines, dropping '#'-comment and blank lines. diff --git a/backend/internal/lang/langpack_test.go b/backend/internal/lang/langpack_test.go index 32906881..0373c505 100644 --- a/backend/internal/lang/langpack_test.go +++ b/backend/internal/lang/langpack_test.go @@ -29,8 +29,11 @@ func TestLoadResolvesRealZhRu(t *testing.T) { if p.SurnamesSingle['凝'] { t.Error("surnames-single must NOT contain 凝 (discarded — not a surname)") } - if !p.SurnamesCompound["古月"] || !p.SurnamesCompound["欧阳"] { - t.Error("surnames-compound missing 古月/欧阳") + if !p.SurnamesCompound["欧阳"] { + t.Error("surnames-compound missing 欧阳") + } + if p.SurnamesCompound["古月"] { + t.Error("surnames-compound must NOT contain 古月 (pair-14 §1: a 蛊真人 clan, moved to the book overlay)") } if len(p.TitleSuffix) == 0 || p.TitleSuffix[0] != "公子" { t.Errorf("title-suffix order not preserved: %v", p.TitleSuffix) @@ -38,12 +41,22 @@ func TestLoadResolvesRealZhRu(t *testing.T) { if !p.TopoSuffix['山'] || !p.Numeral['三'] || !p.GradePrefix['甲'] || !p.AliasParticle['的'] { t.Error("rune-set membership missing an expected char (山/三/甲/的)") } - // Pair transliteration spot-check. - if p.PalladiusInitials["b"] != "б" || p.PalladiusInitials["zh"] != "чж" { - t.Errorf("palladius initials wrong: b=%q zh=%q", p.PalladiusInitials["b"], p.PalladiusInitials["zh"]) + // pair-14 moves: title-formant + sentence-terminator rune sets, and the Palladius phonotactic sets. + if !p.TitleFormant['等'] || !p.TitleFormant['转'] || !p.TitleFormant['阶'] { + t.Error("title-formant missing 等/转/阶") } - if p.PalladiusSpecialI["zhi"] != "чжи" { - t.Errorf("palladius special_i zhi = %q, want чжи", p.PalladiusSpecialI["zhi"]) + if !p.SentenceTerminator['。'] || !p.SentenceTerminator['!'] || !p.SentenceTerminator['?'] { + t.Error("sentence-terminator missing 。/!/?") + } + if !p.Palladius.Retroflex["zh"] || !p.Palladius.VFinal["v"] || !p.Palladius.VFinalInitial["j"] { + t.Error("palladius phonotactics missing retroflex zh / vfinal v / vfinal_initial j") + } + // Pair transliteration spot-check. + if p.Palladius.Initials["b"] != "б" || p.Palladius.Initials["zh"] != "чж" { + t.Errorf("palladius initials wrong: b=%q zh=%q", p.Palladius.Initials["b"], p.Palladius.Initials["zh"]) + } + if p.Palladius.SpecialI["zhi"] != "чжи" { + t.Errorf("palladius special_i zhi = %q, want чжи", p.Palladius.SpecialI["zhi"]) } if !strings.HasPrefix(p.Version(), packAlgoVersion+"-") { t.Errorf("Version() = %q, want %s-", p.Version(), packAlgoVersion) @@ -82,6 +95,63 @@ func TestRoutesByPairToDifferentBytes(t *testing.T) { } } +// TestBookOverlayUnionsAndReVersions pins the pair-14 §1 book-overlay contract: an overlay UNIONS its +// private canon onto the shared pack (additive — the base entries survive, the overlay entry is added) and +// SHIFTS Version() (a book-canon edit is a loud --resnapshot for that book), while a no-overlay Load is +// byte-identical to before. Guards the exact mechanism the miner's {古月:22} parity rides. +func TestBookOverlayUnionsAndReVersions(t *testing.T) { + base, err := Load(realRoot, "zh", "ru") + if err != nil { + t.Fatalf("load base: %v", err) + } + if base.SurnamesCompound["古月"] { + t.Fatal("shared pack must not carry 古月 (it is a book clan)") + } + overlay := t.TempDir() + if err := os.MkdirAll(filepath.Join(overlay, "zh"), 0o755); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(filepath.Join(overlay, "zh", "surnames-compound.txt"), []byte("# book canon\n古月\n"), 0o644); err != nil { + t.Fatal(err) + } + ext, err := LoadWithOverlay(realRoot, "zh", "ru", overlay) + if err != nil { + t.Fatalf("load with overlay: %v", err) + } + if !ext.SurnamesCompound["古月"] { + t.Error("overlay must add 古月 to the effective pack") + } + if !ext.SurnamesCompound["欧阳"] { + t.Error("overlay must UNION (keep the shared 欧阳), not replace") + } + if ext.Version() == base.Version() { + t.Errorf("overlay must shift Version() (base=%s overlay=%s)", base.Version(), ext.Version()) + } + // An empty overlayRoot is exactly Load — same version, no re-hash. + same, err := LoadWithOverlay(realRoot, "zh", "ru", "") + if err != nil { + t.Fatal(err) + } + if same.Version() != base.Version() { + t.Errorf("empty overlay must equal Load (%s != %s)", same.Version(), base.Version()) + } + + // A MISNAMED overlay file must FAIL LOUD, not be silently ignored (review finding, scale lens): a typo'd + // canon file that never reaches the miner would degrade recall with no signal. + bad := t.TempDir() + if err := os.MkdirAll(filepath.Join(bad, "zh"), 0o755); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(filepath.Join(bad, "zh", "surname-compound.txt"), []byte("古月\n"), 0o644); err != nil { // typo: missing 's' + t.Fatal(err) + } + if _, err := LoadWithOverlay(realRoot, "zh", "ru", bad); err == nil { + t.Error("a misnamed overlay file must fail loud (silently-ignored canon would degrade recall)") + } else if !strings.Contains(err.Error(), "surname-compound.txt") { + t.Errorf("error must name the unexpected file, got: %v", err) + } +} + // TestFailsLoudOnMissingPair pins the fail-loud contract: a pair with no pack directory errors, naming the // missing file — never a silent empty pack (mirrors the prompt seam's PromptPathFor fail-loud). This is // the load-time invariant a live consumer relies on (fail before billing). @@ -120,16 +190,19 @@ func writeSyntheticPack(t *testing.T, root, src, tgt string) { t.Helper() pair := src + "-" + tgt files := map[string]string{ - filepath.Join(src, "surnames-single.txt"): "# synthetic\n甴甶甹\n", - filepath.Join(src, "surnames-compound.txt"): "# synthetic\n甲乙\n", - filepath.Join(src, "title-suffix.txt"): "# synthetic\n阁下\n", - filepath.Join(src, "ordinal-title.txt"): "# synthetic\n第甲\n", - filepath.Join(src, "rank-word.txt"): "# synthetic\n級\n", - filepath.Join(src, "topo-suffix.txt"): "# synthetic\n峰\n", - filepath.Join(src, "grade-prefix.txt"): "# synthetic\n子丑\n", - filepath.Join(src, "numeral.txt"): "# synthetic\n壹貳\n", - filepath.Join(src, "alias-particle.txt"): "# synthetic\n之乎\n", - filepath.Join(pair, "palladius.txt"): "# synthetic\ninitials\tb\tб\nfinals\ta\tа\nyw\tyi\tи\nspecial_i\tzhi\tчжи\n", + filepath.Join(src, "surnames-single.txt"): "# synthetic\n甴甶甹\n", + filepath.Join(src, "surnames-compound.txt"): "# synthetic\n甲乙\n", + filepath.Join(src, "title-suffix.txt"): "# synthetic\n阁下\n", + filepath.Join(src, "ordinal-title.txt"): "# synthetic\n第甲\n", + filepath.Join(src, "rank-word.txt"): "# synthetic\n級\n", + filepath.Join(src, "topo-suffix.txt"): "# synthetic\n峰\n", + filepath.Join(src, "grade-prefix.txt"): "# synthetic\n子丑\n", + filepath.Join(src, "numeral.txt"): "# synthetic\n壹貳\n", + filepath.Join(src, "alias-particle.txt"): "# synthetic\n之乎\n", + filepath.Join(src, "title-formant.txt"): "# synthetic\n甼\n", + filepath.Join(src, "sentence-terminator.txt"): "# synthetic\n。\n", + filepath.Join(pair, "palladius.txt"): "# synthetic\ninitials\tb\tб\nfinals\ta\tа\nyw\tyi\tи\nspecial_i\tzhi\tчжи\n", + filepath.Join(pair, "palladius-phonotactics.txt"): "# synthetic\nretroflex\tzh\nvfinal\tv\nvfinal_initial\tj\n", } for rel, body := range files { p := filepath.Join(root, rel) diff --git a/backend/internal/pipeline/cheapgates.go b/backend/internal/pipeline/cheapgates.go index f570bfa8..e9180621 100644 --- a/backend/internal/pipeline/cheapgates.go +++ b/backend/internal/pipeline/cheapgates.go @@ -5,6 +5,8 @@ import ( "sort" "strings" "unicode" + + "textmachine/backend/internal/lang" ) // cheapgates.go: four cheap, deterministic post-check flaggers on the FINAL chunk text (04-unhappy @@ -64,6 +66,11 @@ type cheapGateConfig struct { // OPT-IN observability flagger (draft→final length collapse + number drift) folded into this // result. Off → the two regression fields stay 0 and the output is byte-identical to before. regressionEnabled bool + // checkers is the compiled WS5/pack-13 checker spec (pair-14 data-out): the pair's DETECTION patterns + + // lookup tables (from the pair langpack) plus the target-general lists (from embedded target data), + // resolved once per run. nil for a bare config or a book with no data → the language-specific checkers + // run inert (fire 0, the no-pack golden path); the general Latin-residue check needs no spec. + checkers *dcCheckers } // cheapGateResult is the per-chunk outcome: a count per flagger plus human-readable detail lines @@ -103,33 +110,35 @@ func runCheapGates(source, draft, final string, cfg cheapGateConfig) cheapGateRe var r cheapGateResult n, det := lintDialogueDash(final) r.DialogueDash, r.Detail = n, append(r.Detail, det...) - n, det = lintYofikation(final, cfg.yoPolicy) + n, det = cfg.checkers.lintYofikation(final, cfg.yoPolicy) r.YoInconsistent = n r.Detail = append(r.Detail, det...) n, det = lintTranslitInterjections(final, cfg.allowlist) r.TranslitInterj = n r.Detail = append(r.Detail, det...) - n, det = lintNumberMagnitude(source, final) + n, det = cfg.checkers.lintNumberMagnitude(source, final) r.NumberMagnitude = n r.Detail = append(r.Detail, det...) - // WS5 defect-class checkers (DC1/DC2/DC6, checkers_zh_ru.go) — src↔target observability flaggers. - n, det = lintTimeUnits(source, final) + // WS5 defect-class checkers (DC1/DC2/DC6, checkers.go) — src↔target observability flaggers. All patterns + // + tables are pair langpack DATA (pair-14 data-out); the spec is nil/inert for a no-pack book → fire 0. + n, det = cfg.checkers.lintTimeUnits(source, final) r.DC1TimeUnits = n r.Detail = append(r.Detail, det...) - n, det = lintMagnitudeScale(source, final) + n, det = cfg.checkers.lintMagnitudeScale(source, final) r.DC2Magnitude = n r.Detail = append(r.Detail, det...) - n, det = lintRegisterLexicon(final) + n, det = cfg.checkers.lintRegisterLexicon(final) r.DC6Register = n r.Detail = append(r.Detail, det...) - // pack-13 general checkers (checkers_zh_ru.go) — src↔target / Russian-side observability flaggers. - n, det = lintPercentScale(source, final) + // pack-13 general checkers (checkers.go): percent scale (pair data), Latin residue (language-general), + // broken word (target data). + n, det = cfg.checkers.lintPercentScale(source, final) r.PercentScale = n r.Detail = append(r.Detail, det...) n, det = lintLatinResidue(final, cfg.allowlist) r.LatinResidue = n r.Detail = append(r.Detail, det...) - n, det = lintBrokenWord(final) + n, det = cfg.checkers.lintBrokenWord(final) r.BrokenWord = n r.Detail = append(r.Detail, det...) if cfg.regressionEnabled { @@ -248,33 +257,16 @@ func quoteShape(after []rune) bool { // --- 2. yofikator --------------------------------------------------------------- -// yoHomographEForms are е-spellings that are DISTINCT words from their ё-counterpart (все≠всё, -// небо≠нёбо, берет≠берёт, осел≠осёл, …). The inconsistency rule below would otherwise false-flag -// «все»+«всё» as one word spelled two ways. Expanded (self-review major) to cover the frequent -// distinct-word pairs and their common inflections; still not exhaustive — full disambiguation -// needs a ё-dictionary (deferred, B-tier). Names — the primary target (Пётр/Петр) — are never -// homographs, so they are always caught regardless. -var yoHomographEForms = map[string]bool{ - "все": true, "всех": true, "всем": true, "всеми": true, - "небо": true, "небом": true, - "узнаем": true, "узнаете": true, "узнает": true, - "падеж": true, "падежа": true, - "совершенный": true, "совершенное": true, "совершенная": true, "совершенно": true, "совершенны": true, - "чем": true, "тем": true, "тема": true, "теме": true, "темы": true, - "берет": true, "берете": true, "берета": true, - "осел": true, "осла": true, "ослы": true, "ослов": true, - "мел": true, "мела": true, - "лен": true, "лена": true, - "нес": true, "несем": true, "несет": true, - "вел": true, "ведет": true, "ведем": true, -} - // lintYofikation flags inconsistent ё. Default ("auto"): the same word appears BOTH with ё and, as -// a separate token, with its exact ё→е form (Пётр/Петр) — excluding the homograph traps above. -// Policy "all-e": any ё present is a violation (the project wants no ё). Policy "all-yo": an е-form -// of a word that ALSO appears somewhere with ё is flagged (the partial-yofikation case), same signal -// as auto — full "every word that SHOULD have ё" enforcement needs a ё-dictionary (deferred, B-tier). -func lintYofikation(text, policy string) (int, []string) { +// a separate token, with its exact ё→е form (Пётр/Петр) — excluding the target's homograph whitelist +// (е-spellings that are DISTINCT words from their ё-counterpart, все≠всё …; TARGET data, lang.TargetChecks +// "yo_homograph" — pair-14 data-out, so the ё↔е FOLD stays as generic orthography and only the wordlist is +// data). Policy "all-e": any ё present is a violation. Policy "all-yo": same signal as auto. Full "every +// word that SHOULD have ё" enforcement needs a ё-dictionary (deferred, B-tier). Inert if no target data. +func (c *dcCheckers) lintYofikation(text, policy string) (int, []string) { + if c == nil { + return 0, nil + } words := tokenizeCyrillic(text) if policy == "all-e" { seen := map[string]bool{} @@ -303,7 +295,7 @@ func lintYofikation(text, policy string) (int, []string) { continue } eForm := strings.ReplaceAll(w, "ё", "е") - if eForm == w || yoHomographEForms[eForm] { + if eForm == w || c.yoHomograph[eForm] { continue // no ё, or a distinct-word homograph (все/всё) — not an inconsistency } if present[eForm] && !seenPair[w] { @@ -404,7 +396,10 @@ func isWordRune(r rune) bool { // range for their possible multiplier, plus Arabic-number orders). It fires only when the source's // TOP order is outside every output range — a conservative, multiplier-tolerant signal that leaves // 三億→«триста миллионов» (8 within миллион's [6,8]) silent while catching 三万→«три миллиона». -func lintNumberMagnitude(source, final string) (int, []string) { +func (c *dcCheckers) lintNumberMagnitude(source, final string) (int, []string) { + if c == nil || len(c.magnitudeStem) == 0 { + return 0, nil // no target magnitude-word data → the gate can never confirm coverage; stay inert + } srcOrders := cjkMagnitudeOrders(source) if len(srcOrders) == 0 { return 0, nil @@ -427,7 +422,7 @@ func lintNumberMagnitude(source, final string) (int, []string) { return 0, nil // an Arabic figure of the source order confirms coverage (30000 for 三万) } } - wordRanges := magnitudeWordRanges(final) + wordRanges := c.magnitudeWordRanges(final) for _, rg := range wordRanges { if maxSrc >= rg[0] && maxSrc <= rg[1] { return 0, nil // a magnitude WORD covers the source order (三万 → «тридцать тысяч») @@ -444,14 +439,11 @@ func lintNumberMagnitude(source, final string) (int, []string) { return 1, []string{fmt.Sprintf("the source magnitude 10^%d (万/億) is not reflected in the translation's orders of magnitude — a possible magnitude error (e.g. 三万→«три миллиона»)", maxSrc)} } -// cjkNumeralRunes are the characters that can form a CJK numeral expression (digits, small units, -// big markers). A maximal run of these is one candidate number. -var cjkNumeralRunes = map[rune]bool{} - -func init() { - for _, r := range "0123456789〇零一二三四五六七八九十百千两兩万萬億亿兆" { - cjkNumeralRunes[r] = true - } +// isCJKNumeralRune reports whether r can form a CJK numeral expression: an ASCII digit or a shared +// lang.CJKSection numeral (Zero∪Digit∪Unit∪Magnitude, pair-14 data-out — the gate no longer keeps its own +// rune set). A maximal run of these is one candidate number. +func isCJKNumeralRune(r rune) bool { + return (r >= '0' && r <= '9') || lang.DefaultCJKSection().IsNumeralRune(r) } // cjkMagnitudeOrders returns the base-10 order of every CJK numeral run in text that contains a big @@ -460,12 +452,12 @@ func cjkMagnitudeOrders(text string) []int { var orders []int rs := []rune(text) for i := 0; i < len(rs); { - if !cjkNumeralRunes[rs[i]] { + if !isCJKNumeralRune(rs[i]) { i++ continue } j := i - for j < len(rs) && cjkNumeralRunes[rs[j]] { + for j < len(rs) && isCJKNumeralRune(rs[j]) { j++ } run := string(rs[i:j]) @@ -474,7 +466,7 @@ func cjkMagnitudeOrders(text string) []int { // 万一 (in case), 万分 (extremely), 万物 (all things), 千万 (by all means), 亿万 (myriads) — where 万/億 // is not preceded by a digit. It also drops bare-unit magnitudes (十万/百万) — an accepted recall // trade for not false-flagging the far more frequent idioms. - if strings.ContainsAny(run, "万萬億亿兆") && hasDigitBeforeBigMarker(run) { + if containsBigMarker(run) && hasDigitBeforeBigMarker(run) { if v, ok := parseCJKNumber(run); ok && v > 0 { orders = append(orders, orderOf(v)) } @@ -484,14 +476,27 @@ func cjkMagnitudeOrders(text string) []int { return orders } +// isBigMarker / containsBigMarker read the shared lang.CJKSection magnitude set (万/萬/億/亿/兆, pair-14 +// data-out) instead of a hard-coded rune string. +func isBigMarker(r rune) bool { _, ok := lang.DefaultCJKSection().BigUnit(r); return ok } +func containsBigMarker(run string) bool { + for _, r := range run { + if isBigMarker(r) { + return true + } + } + return false +} + // hasDigitBeforeBigMarker reports whether a digit (一-九 / 两 / 0-9) appears before the FIRST big // marker (万/億/兆) in the run — the signature of a real magnitude expression (三万) vs an idiom (万一). func hasDigitBeforeBigMarker(run string) bool { + sec := lang.DefaultCJKSection() for _, r := range run { - if strings.ContainsRune("万萬億亿兆", r) { + if isBigMarker(r) { return false // hit a big marker with no digit before it → idiom / bare unit } - if (r >= '0' && r <= '9') || strings.ContainsRune("一二三四五六七八九两兩", r) { + if _, isDigit := sec.Digit[r]; (r >= '0' && r <= '9') || isDigit { return true } } @@ -503,28 +508,29 @@ func hasDigitBeforeBigMarker(run string) bool { // <10^4 section; big units (万億兆) flush the section times the big unit into the total. Returns // ok=false on a shape it cannot parse (conservative — an unparseable run does not flag). func parseCJKNumber(s string) (int64, bool) { + sec := lang.DefaultCJKSection() var total, section, cur int64 sawBig := false for _, r := range s { switch { case r >= '0' && r <= '9': cur = cur*10 + int64(r-'0') - case r == '〇' || r == '零': + case sec.Zero[r]: cur = cur * 10 default: - if d, ok := cjkDigit(r); ok { - cur = cur*10 + d // positional accumulation (一二→12), matching the Arabic-digit branch (self-review) + if d, ok := sec.Digit[r]; ok { + cur = cur*10 + int64(d) // positional accumulation (一二→12), matching the Arabic-digit branch (self-review) continue } - if u, ok := cjkSmallUnit(r); ok { + if u, ok := sec.Unit[r]; ok { if cur == 0 { cur = 1 } - section += cur * u + section += cur * int64(u) cur = 0 continue } - if u, ok := cjkBigUnit(r); ok { + if u, ok := sec.BigUnit(r); ok { sawBig = true section += cur if section == 0 { @@ -544,54 +550,6 @@ func parseCJKNumber(s string) (int64, bool) { return total + section + cur, true } -func cjkDigit(r rune) (int64, bool) { - switch r { - case '一': - return 1, true - case '二', '两', '兩': - return 2, true - case '三': - return 3, true - case '四': - return 4, true - case '五': - return 5, true - case '六': - return 6, true - case '七': - return 7, true - case '八': - return 8, true - case '九': - return 9, true - } - return 0, false -} - -func cjkSmallUnit(r rune) (int64, bool) { - switch r { - case '十': - return 10, true - case '百': - return 100, true - case '千': - return 1000, true - } - return 0, false -} - -func cjkBigUnit(r rune) (int64, bool) { - switch r { - case '万', '萬': - return 10000, true - case '億', '亿': - return 100000000, true - case '兆': - return 1000000000000, true - } - return 0, false -} - func orderOf(v int64) int { o := 0 for v >= 10 { @@ -606,10 +564,10 @@ func orderOf(v int64) int { // two orders (триста миллионов = 3·10^8, base 6 → order 8). Arabic integers are handled separately // by the caller (they may only CONFIRM coverage, never trigger a mismatch — D20.4), so they are NOT // folded in here. -func magnitudeWordRanges(text string) [][2]int { +func (c *dcCheckers) magnitudeWordRanges(text string) [][2]int { low := strings.ToLower(text) var ranges [][2]int - for stem, base := range map[string]int{"тысяч": 3, "миллион": 6, "миллиард": 9, "триллион": 12} { + for stem, base := range c.magnitudeStem { // target-general data (pair-14 data-out) if strings.Contains(low, stem) { ranges = append(ranges, [2]int{base, base + 2}) } diff --git a/backend/internal/pipeline/cheapgates_test.go b/backend/internal/pipeline/cheapgates_test.go index 15858afc..c10eb909 100644 --- a/backend/internal/pipeline/cheapgates_test.go +++ b/backend/internal/pipeline/cheapgates_test.go @@ -56,9 +56,10 @@ func TestLintYofikation(t *testing.T) { {"all-e clean", "Петр ел мед.", "all-e", 0}, {"artem both spellings", "Артём и Артем — один человек.", "auto", 1}, } + dcc := testCheckers(t) for _, c := range cases { t.Run(c.name, func(t *testing.T) { - n, _ := lintYofikation(c.text, c.policy) + n, _ := dcc.lintYofikation(c.text, c.policy) if n != c.want { t.Errorf("lintYofikation(%q, %q) = %d, want %d", c.text, c.policy, n, c.want) } @@ -112,8 +113,8 @@ func TestParseCJKNumber(t *testing.T) { {"30万", 300000, true}, {"五億", 500000000, true}, {"一兆", 1000000000000, true}, - {"三", 0, false}, // no big marker - {"三千", 0, false}, // no big marker (千 is a small unit) + {"三", 0, false}, // no big marker + {"三千", 0, false}, // no big marker (千 is a small unit) {"", 0, false}, } for _, c := range cases { @@ -152,9 +153,10 @@ func TestLintNumberMagnitude(t *testing.T) { // output number must NOT flag. {"wan in a name plus stray year not flagged", "他叫赵三万,生于1985年。", "Его звали Чжао Саньвань, родился в 1985 году.", 0}, } + dcc := testCheckers(t) for _, c := range cases { t.Run(c.name, func(t *testing.T) { - n, _ := lintNumberMagnitude(c.source, c.final) + n, _ := dcc.lintNumberMagnitude(c.source, c.final) if n != c.want { t.Errorf("lintNumberMagnitude(%q → %q) = %d, want %d", c.source, c.final, n, c.want) } @@ -167,7 +169,7 @@ func TestLintNumberMagnitude(t *testing.T) { func TestRunCheapGatesCombined(t *testing.T) { src := "他有三万石粮食。" final := "- Ара-ара, — у Пётр было три миллиона мешков. Потом Петр ушёл." - cfg := cheapGateConfig{yoPolicy: "auto", allowlist: map[string]bool{}} + cfg := cheapGateConfig{yoPolicy: "auto", allowlist: map[string]bool{}, checkers: testCheckers(t)} r := runCheapGates(src, final, final, cfg) if r.DialogueDash == 0 { t.Error("expected a dialogue-dash flag (hyphen-led line)") diff --git a/backend/internal/pipeline/checkers.go b/backend/internal/pipeline/checkers.go new file mode 100644 index 00000000..c417eba4 --- /dev/null +++ b/backend/internal/pipeline/checkers.go @@ -0,0 +1,355 @@ +package pipeline + +import ( + "fmt" + "regexp" + "sort" + "strconv" + "strings" + "unicode" + + "textmachine/backend/internal/lang" +) + +// checkers.go: the WS5 defect-class checkers (DC1 时辰 double-hour units, DC2 千万/数十万 magnitude scale, DC6 +// register negative-list) + the pack-13 general checkers (percent scale, Latin residue, broken word) — +// deterministic, $0 OBSERVABILITY flaggers on the source↔FINAL text, ported from ws5_checkers_verify.py. +// Like the four cheap style gates they are NEVER a disposition (a hit is recorded, never drops a chunk), +// tuned PRECISION over recall. +// +// PAIR-AGNOSTIC (pair-14 data-out): this file no longer holds any language literal. Every DETECTION pattern, +// lookup table and wordlist is DATA — the SOURCE-gated pair checkers (DC1/DC2/percent: they need a zh source +// token and compare src↔tgt) read the pair pack configs/langpacks//dc-checkers.txt (lang.DCCheckerData); +// the TARGET-general ones (register, broken word: they run on ANY →target output) read the embedded target +// data (lang.TargetChecks). The ALGORITHM (compare counts, ×2 hours, suppress-if-ok, whole-word match) stays +// here. A pair/target that ships no data runs the relevant sub-checker inert (empty → 0, the no-pack golden +// path). Version rides the langpack Version() (data) + cheapGateVersion (algorithm) — a data or rule edit is +// a loud --resnapshot. + +// dcCheckers is the compiled, per-run checker spec: the pair's DETECTION patterns compiled ONCE + its lookup +// tables + the target-general lists, resolved from the langpack. A nil receiver, a nil pattern or an empty +// table leaves that sub-checker inert. Built once per run (compileCheckers), carried in cheapGateConfig. +type dcCheckers struct { + numeral map[rune]int + ruHours map[string]int + registerNeg []string + brokenSuffix []string // target-general (any →ru output) + yoHomograph map[string]bool // target-general: ё↔е homograph whitelist (cheapgates yofikator) + magnitudeStem map[string]int // target-general: ru magnitude word stem → base-10 exponent (cheapgates) + + shichenRE, ruHoursRE, chengRE, decimalFractionRE *regexp.Regexp + qianwanOKRE, shushiwanOKRE, shushiwanFireRE *regexp.Regexp + qianwanSrc, qianwanFireWord, qianwanVetoWord string + shushiwanSrc, percentWord string +} + +// compileCheckers resolves the checker spec from the pair pack (dc) and the target data (tc). A malformed +// regex in the pack is a corrupt pack → panic (deterministic, caught by the checker/golden tests). dc==nil +// (a no-pack book) → the pair sub-checkers are inert; tc still supplies the target-general lists. +func compileCheckers(dc *lang.DCCheckerData, tc lang.TargetChecks) *dcCheckers { + c := &dcCheckers{ + brokenSuffix: tc.List("broken_suffix"), + yoHomograph: listToSet(tc.List("yo_homograph")), + magnitudeStem: listToStemExp(tc.List("magnitude_stem")), + } + if dc != nil { + c.numeral, c.ruHours, c.registerNeg = dc.Numeral, dc.RuHours, dc.RegisterNeg + p := dc.Patterns + c.shichenRE = mustPairRE(p, "shichen_re") + c.ruHoursRE = mustPairRE(p, "ru_hours_re") + c.chengRE = mustPairRE(p, "cheng_re") + c.decimalFractionRE = mustPairRE(p, "decimal_fraction_re") + c.qianwanOKRE = mustPairRE(p, "qianwan_ok_re") + c.shushiwanOKRE = mustPairRE(p, "shushiwan_ok_re") + c.shushiwanFireRE = mustPairRE(p, "shushiwan_fire_re") + c.qianwanSrc, c.qianwanFireWord, c.qianwanVetoWord = p["qianwan_src"], p["qianwan_fire_word"], p["qianwan_veto_word"] + c.shushiwanSrc, c.percentWord = p["shushiwan_src"], p["percent_word"] + } + return c +} + +// listToSet turns an ordered value list into a membership set (target wordlists). +func listToSet(xs []string) map[string]bool { + m := make(map[string]bool, len(xs)) + for _, x := range xs { + m[x] = true + } + return m +} + +// listToStemExp parses `stemexp` values (the magnitude_stem list carries a second tab-field) into a +// stem→exponent map. A malformed value is a corrupt embed → panic (deterministic, caught by tests). +func listToStemExp(xs []string) map[string]int { + m := make(map[string]int, len(xs)) + for _, x := range xs { + f := strings.SplitN(x, "\t", 2) + if len(f) != 2 { + panic(fmt.Sprintf("pipeline: magnitude_stem wants `stemexp`, got %q", x)) + } + v, err := strconv.Atoi(strings.TrimSpace(f[1])) + if err != nil { + panic(fmt.Sprintf("pipeline: magnitude_stem exponent %q: %v", f[1], err)) + } + m[f[0]] = v + } + return m +} + +// dcCheckerData returns the pack's checker data, or nil when the book has no pack (nil-safe helper for the +// compile step, which runs even for a no-langpack book). +func dcCheckerData(p *lang.Pack) *lang.DCCheckerData { + if p != nil { + return p.DCCheckers + } + return nil +} + +// mustPairRE compiles a pair detection pattern by key; a missing key → nil (inert sub-checker), a malformed +// regex → panic (a corrupt pack, not a silent no-op — the same fail-loud discipline the loader keeps). +func mustPairRE(p map[string]string, key string) *regexp.Regexp { + s := p[key] + if s == "" { + return nil + } + re, err := regexp.Compile(s) + if err != nil { + panic(fmt.Sprintf("pipeline: langpack checker pattern %q is not a valid regex: %v", key, err)) + } + return re +} + +// --- DC-1: 时辰 (double-hour) unit checker (ws5.shichen_checker) ----------------------------------- + +// lintTimeUnits flags a 时辰 (=2h) unit error: N个时辰 rendered as N часов (the count copied as hours) +// instead of ~2N hours (三个时辰 → «три часа» should be ~6h). It fires ONLY on an explicit mismatch — a +// paraphrase with no hours count is a valid rendering, not a defect. Pure and deterministic. Inert when the +// pair ships no DC1 pattern (the shichen_re / ru_hours_re detection patterns are pair langpack DATA). +func (c *dcCheckers) lintTimeUnits(source, final string) (int, []string) { + if c == nil || c.shichenRE == nil || c.ruHoursRE == nil { + return 0, nil + } + m := c.shichenRE.FindStringSubmatch(source) + if m == nil { + return 0, nil + } + n, ok := dcParseCount(m[1], c.numeral) + if !ok { + return 0, nil + } + hm := c.ruHoursRE.FindStringSubmatch(final) + if hm == nil { + return 0, nil // no explicit hours rendering → a valid paraphrase, not a defect + } + ruNum, ok := dcParseRuHours(hm[1], c.ruHours) + if !ok { + return 0, nil + } + expectedHours := n * 2 + if ruNum == n && ruNum != expectedHours { + return 1, []string{fmt.Sprintf("DC1 时辰: %d个时辰 rendered as «%d час…» (counted as hours) instead of ~%d h (1 时辰 = 2 h)", n, ruNum, expectedHours)} + } + return 0, nil +} + +// dcParseCount parses the DC1 count group: an Arabic digit string or a single small CJK numeral (looked up +// in the pair's DCCheckerData.Numeral, empty for a no-pack book → CJK counts don't resolve, the check is inert). +func dcParseCount(s string, dcNum map[rune]int) (int, bool) { + if v, err := strconv.Atoi(s); err == nil { + return v, true + } + r := []rune(s) + if len(r) == 1 { + if v, ok := dcNum[r[0]]; ok { + return v, true + } + } + return 0, false +} + +// dcParseRuHours parses the DC1 hours group: an Arabic digit string or a Russian count word (pair data). +func dcParseRuHours(s string, dcRu map[string]int) (int, bool) { + if v, err := strconv.Atoi(s); err == nil { + return v, true + } + if v, ok := dcRu[s]; ok { + return v, true + } + return 0, false +} + +// --- DC-2: number-scale magnitude checker (ws5.magnitude_checker, 千万 / 数十万) -------------------- + +// lintMagnitudeScale flags a 千万 (10^7) / 数十万 (~several×10^5) magnitude rendered at a WRONG smaller +// scale. A CORRECT rendering anywhere in the chunk (the ok-suppressor) suppresses the flag (ws5 reference +// parity). The ok-suppressors are case-INsensitive (their data carries the (?i)); the FIRE predicates are +// case-SENSITIVE literal Contains (the reference does not pass re.I to the inner searches). ⚠ 千万 is also +// stock HYPERBOLE whose «тысячи» rendering is in-register (§5 A4). Observability only. All probes/patterns +// are pair langpack DATA — inert when the pair ships no DC2. +func (c *dcCheckers) lintMagnitudeScale(source, final string) (int, []string) { + if c == nil { + return 0, nil + } + var flags []string + if c.qianwanOKRE != nil && c.qianwanSrc != "" && strings.Contains(source, c.qianwanSrc) && !c.qianwanOKRE.MatchString(final) { + if strings.Contains(final, c.qianwanFireWord) && !strings.Contains(final, c.qianwanVetoWord) { // case-sensitive, per reference + flags = append(flags, "DC2 千万=10^7 rendered as «тысячи» (≈10000× under) — a possible magnitude error (hyperbole risk, §5-A4)") + } + } + if c.shushiwanOKRE != nil && c.shushiwanFireRE != nil && c.shushiwanSrc != "" && strings.Contains(source, c.shushiwanSrc) && !c.shushiwanOKRE.MatchString(final) { + if c.shushiwanFireRE.MatchString(final) { + flags = append(flags, "DC2 数十万≈several×10^5 rendered as «десятки тысяч» (≈10× under)") + } + } + return len(flags), flags +} + +// --- DC-6: register-lexicon negative-list (ws5.register_checker) ---------------------------------- + +// lintRegisterLexicon flags whole-word occurrences of a register negative-list lexeme in the FINAL text. +// The negative-list is pair langpack DATA (DCCheckerData.RegisterNeg): fairy-tale-Russian / chancery lexemes +// that break the xianxia register. Lower-cased; whole-word matched. Empty (a no-pack book) → nothing to flag. +func (c *dcCheckers) lintRegisterLexicon(final string) (int, []string) { + if c == nil || len(c.registerNeg) == 0 { + return 0, nil + } + low := []rune(strings.ToLower(final)) + hitSet := map[string]bool{} + for _, w := range c.registerNeg { + wr := []rune(w) + for i := 0; i+len(wr) <= len(low); i++ { + if !runesEqual(low[i:i+len(wr)], wr) { + continue + } + if (i == 0 || !isCyrLetter(low[i-1])) && (i+len(wr) == len(low) || !isCyrLetter(low[i+len(wr)])) { + hitSet[w] = true + } + } + } + if len(hitSet) == 0 { + return 0, nil + } + hits := make([]string, 0, len(hitSet)) + for w := range hitSet { + hits = append(hits, w) + } + sort.Strings(hits) + return len(hits), []string{"DC6 register: fairy-tale Russian lexis out of the xianxia genre: " + strings.Join(hits, ", ")} +} + +// isCyrLetter reports whether r is a Cyrillic letter (the word boundary for the register match). +func isCyrLetter(r rune) bool { return unicode.IsLetter(r) && unicode.Is(unicode.Cyrillic, r) } + +// --- percent-scale checker (成 = tenths) ----------------------------------------------------------- +// +// In Chinese, 成 is one tenth: 六成 = 60%, 六成六 = 66%. A common error renders this as a decimal FRACTION +// instead of a percentage (a ~100× scale error). Precision over recall: it fires only when the source has a +// «成[]» (cheng_re, pair data), the output has NO percent form (percent_word / «%»), AND the +// output carries a decimal-fraction cue (decimal_fraction_re). Inert when the pair ships no percent patterns. +func (c *dcCheckers) lintPercentScale(source, final string) (int, []string) { + if c == nil || c.chengRE == nil || c.decimalFractionRE == nil { + return 0, nil + } + m := c.chengRE.FindStringSubmatch(source) + if m == nil { + return 0, nil + } + low := strings.ToLower(final) + if (c.percentWord != "" && strings.Contains(low, c.percentWord)) || strings.Contains(final, "%") { + return 0, nil // the output uses a percent form — the scale is handled correctly + } + if !c.decimalFractionRE.MatchString(low) { + return 0, nil // no fraction cue — the magnitude was paraphrased, not mis-scaled + } + tens, _ := dcParseCount(m[1], c.numeral) + pct := tens * 10 + if m[2] != "" { + if ones, ok := dcParseCount(m[2], c.numeral); ok { + pct += ones + } + } + return 1, []string{fmt.Sprintf("成-percent: %s成%s = %d%% rendered as a decimal fraction instead of a percentage (~%d%%)", m[1], m[2], pct, pct)} +} + +// --- Latin residue in the Russian output ----------------------------------------------------------- +// +// A whole Latin WORD left untranslated in the output (e.g. «открыл их again, …»). Language-general (Latin is +// not pair data): it splits the output into maximal alphanumeric tokens and flags an all-lowercase Latin +// token (leaked prose is lowercase; a capital signals a proper noun/brand) with no digit, ≥ minLatinResidueLen +// letters, not a Roman numeral, not on the per-project allowlist. Precision over recall. +const minLatinResidueLen = 3 + +func lintLatinResidue(final string, allow map[string]bool) (int, []string) { + hits := map[string]bool{} + rs := []rune(final) + for i := 0; i < len(rs); { + if !isLatinLetterOrDigit(rs[i]) { + i++ + continue + } + j := i + reject := false // set on any digit or uppercase letter — not a lowercase leaked word + for j < len(rs) && isLatinLetterOrDigit(rs[j]) { + if (rs[j] >= '0' && rs[j] <= '9') || (rs[j] >= 'A' && rs[j] <= 'Z') { + reject = true + } + j++ + } + tok := string(rs[i:j]) + i = j + if !reject && len([]rune(tok)) >= minLatinResidueLen && !isRomanNumeral(tok) && !allow[tok] { + hits[tok] = true + } + } + if len(hits) == 0 { + return 0, nil + } + surfaces := make([]string, 0, len(hits)) + for s := range hits { + surfaces = append(surfaces, s) + } + sort.Strings(surfaces) + return len(surfaces), []string{"Latin word left untranslated in the Russian output: " + strings.Join(surfaces, ", ")} +} + +// isLatinLetterOrDigit reports whether r is an ASCII Latin letter or digit (the alphanumeric-token alphabet). +func isLatinLetterOrDigit(r rune) bool { + return (r >= 'A' && r <= 'Z') || (r >= 'a' && r <= 'z') || (r >= '0' && r <= '9') +} + +// isRomanNumeral reports whether a lowercase token is a Roman numeral (all chars in ivxlcdm) — «iii» reads +// as a numeral, not a leaked word; excluded to hold precision. (Uppercase «II» is already skipped as a cap.) +func isRomanNumeral(tok string) bool { + for _, r := range tok { + switch r { + case 'i', 'v', 'x', 'l', 'c', 'd', 'm': + default: + return false + } + } + return true +} + +// --- broken word forms (target-general) ------------------------------------------------------------ +// +// Flags a target word ending in a structurally-impossible suffix — for ru, «-йть», which no well-formed +// Russian word does (the shape of a mangled infinitive, «войть» for «войти»). The suffix set is TARGET data +// (lang.TargetChecks "broken_suffix"), so the ALGORITHM is language-general; a target with no suffix data +// flags nothing. Zero false-positive by construction (only a structural signature, no dictionary). +func (c *dcCheckers) lintBrokenWord(final string) (int, []string) { + if c == nil || len(c.brokenSuffix) == 0 { + return 0, nil + } + seen := map[string]bool{} + var det []string + for _, w := range tokenizeCyrillic(final) { + wl := len([]rune(w)) + for _, suf := range c.brokenSuffix { + if wl >= 4 && strings.HasSuffix(w, suf) && !seen[w] { + seen[w] = true + det = append(det, "malformed word ending in «-"+suf+"» (no valid Russian word does): "+w) + } + } + } + sort.Strings(det) + return len(det), det +} diff --git a/backend/internal/pipeline/checkers_pack13_test.go b/backend/internal/pipeline/checkers_pack13_test.go index 542f1d91..e3a861db 100644 --- a/backend/internal/pipeline/checkers_pack13_test.go +++ b/backend/internal/pipeline/checkers_pack13_test.go @@ -6,17 +6,18 @@ import "testing" // word lists): each defect shape fires and a clean counterpart / golden-shape Russian output stays 0 // (precision over recall). The 时辰/千万 cases confirm the pre-existing DC1/DC2 still catch the corpus shapes. func TestPack13Checkers(t *testing.T) { + dcc := testCheckers(t) // Percent scale (成 = tenths): a decimal-fraction rendering fires; a correct percent is suppressed. - if n, _ := lintPercentScale("море истинной ци 六成六 …", "море истинной ци — шесть десятых и шесть сотых"); n != 1 { + if n, _ := dcc.lintPercentScale("море истинной ци 六成六 …", "море истинной ци — шесть десятых и шесть сотых"); n != 1 { t.Errorf("percent: «шесть десятых и шесть сотых» should fire, got %d", n) } - if n, _ := lintPercentScale("六成六", "море истинной ци — шесть и шесть десятых"); n != 1 { + if n, _ := dcc.lintPercentScale("六成六", "море истинной ци — шесть и шесть десятых"); n != 1 { t.Errorf("percent: «шесть и шесть десятых» should fire, got %d", n) } - if n, _ := lintPercentScale("六成六", "заполнено на шестьдесят шесть процентов"); n != 0 { + if n, _ := dcc.lintPercentScale("六成六", "заполнено на шестьдесят шесть процентов"); n != 0 { t.Errorf("percent: a correct «процентов» rendering must be suppressed, got %d", n) } - if n, _ := lintPercentScale("нет числа", "шесть десятых чего-то"); n != 0 { + if n, _ := dcc.lintPercentScale("нет числа", "шесть десятых чего-то"); n != 0 { t.Errorf("percent: no 成 in source must not fire, got %d", n) } @@ -28,9 +29,9 @@ func TestPack13Checkers(t *testing.T) { for _, clean := range []string{ "ОТРЕДАКТИРОВАННЫЙ ПЕРЕВОД 5abc35ddfb65. Судзуки шёл по коридорам.", // a body-hash id (has digits) "Классы таланта А, Б, В и Г — от высшей к низшей.", // Cyrillic class letters - "Глава II начинается.", // a Roman numeral (uppercase) - "Он держал в руках iPhone и MacBook.", // brands — any capital → skipped - "Судзуки шёл в Академию.", // Cyrillic proper noun + "Глава II начинается.", // a Roman numeral (uppercase) + "Он держал в руках iPhone и MacBook.", // brands — any capital → skipped + "Судзуки шёл в Академию.", // Cyrillic proper noun } { if n, det := lintLatinResidue(clean, nil); n != 0 { t.Errorf("latin: clean text must not fire (%q): %d %v", clean, n, det) @@ -42,7 +43,7 @@ func TestPack13Checkers(t *testing.T) { } // Broken word — the general «-йть» rule only (no book-specific lists): «войть» fires, valid words do not. - if n, _ := lintBrokenWord("хотел тихонько войть и закрыть"); n != 1 { + if n, _ := dcc.lintBrokenWord("хотел тихонько войть и закрыть"); n != 1 { t.Errorf("broken: «войть» (-йть) should fire") } for _, clean := range []string{ @@ -50,16 +51,16 @@ func TestPack13Checkers(t *testing.T) { "глава клана Гуюэ поклонился", // valid prose (no -йть) "впереди идёт Фан Юань", // valid впереди } { - if n, det := lintBrokenWord(clean); n != 0 { + if n, det := dcc.lintBrokenWord(clean); n != 0 { t.Errorf("broken: clean text must not fire (%q): %d %v", clean, n, det) } } // The pre-existing DC1/DC2 still catch the corpus shapes (general Chinese units/idioms). - if n, _ := lintTimeUnits("僵持了三个时辰", "прошло три часа"); n != 1 { + if n, _ := dcc.lintTimeUnits("僵持了三个时辰", "прошло три часа"); n != 1 { t.Errorf("time-unit: 三个时辰→«три часа» should fire") } - if n, _ := lintMagnitudeScale("千万生灵", "погубил тысячи жизней"); n != 1 { + if n, _ := dcc.lintMagnitudeScale("千万生灵", "погубил тысячи жизней"); n != 1 { t.Errorf("magnitude: 千万→«тысячи» should fire") } } diff --git a/backend/internal/pipeline/checkers_zh_ru.go b/backend/internal/pipeline/checkers_zh_ru.go deleted file mode 100644 index 7ae7f77f..00000000 --- a/backend/internal/pipeline/checkers_zh_ru.go +++ /dev/null @@ -1,295 +0,0 @@ -package pipeline - -import ( - "fmt" - "regexp" - "sort" - "strconv" - "strings" - "unicode" -) - -// checkers_zh_ru.go: the WS5 defect-class checkers DC1 (时辰 double-hour units), DC2 (千万/数十万 magnitude -// scale) and DC6 (register negative-list) — deterministic, $0 flaggers on the source↔FINAL text, ported -// from the frozen ws5_checkers_verify.py. Like the four cheap style gates they are OBSERVABILITY, NEVER a -// disposition (research/20: judges too noisy → deterministic gates; a hit is recorded in the retrieval- -// state / report, never dropping a chunk). They are tuned PRECISION over recall — DC1/DC2 fire only on an -// EXPLICIT src↔target mismatch; the LANDING (whether the flag is trusted) is gated by a §5(д) false- -// positive measure on a fresh rerun, not by this build. Their rule VERSION rides cheapGateVersion (folded -// into the snapshot as StyleCheckVersion), so editing a rule / the pack is a loud --resnapshot. -// -// PER-PAIR PACK (layer-2 data, §5(в)): the unit table + negative-list below are the zh-ru convention pack. -// They are code consts versioned by cheapGateVersion (the same discipline the existing cheap-gate data -// uses), not a config-loaded per-pair file — a config-loaded versioned pack is a cleaner future form -// (noted residual). The checkers self-gate on source content (时辰/千万 are zh-specific → no fire on a -// non-zh source), and DC6 is target-side (gated behind ru like the other readability flaggers). - -// --- DC-1: 时辰 (double-hour) unit checker (ws5.shichen_checker) ----------------------------------- - -// dcCNNum maps the small CJK numerals the 时辰 pattern accepts to their value (1 时辰 = 2 modern hours). -var dcCNNum = map[rune]int{ - '一': 1, '二': 2, '两': 2, '三': 3, '四': 4, '五': 5, '六': 6, '七': 7, '八': 8, '九': 9, '十': 10, -} - -// dcShichenRE captures the count before 时辰 (an Arabic digit or a small CJK numeral), tolerating 个. -var dcShichenRE = regexp.MustCompile(`([0-9一二两三四五六七八九十])\s*个?\s*时辰`) - -// dcRuHoursRE captures a Russian " час…" rendering; the alternatives mirror the reference map. -var dcRuHoursRE = regexp.MustCompile(`(\d+|один|два|двух|три|трёх|трех|четыре|пять|шесть)\s+час`) - -// dcRuHourWord maps the Russian count words dcRuHoursRE captures to their value. -var dcRuHourWord = map[string]int{ - "один": 1, "два": 2, "двух": 2, "три": 3, "трёх": 3, "трех": 3, "четыре": 4, "пять": 5, "шесть": 6, -} - -// lintTimeUnits flags a 时辰 (=2h) unit error: N个时辰 rendered as N часов (the count copied as hours) -// instead of ~2N hours (三个时辰 → «три часа» should be ~6h). It fires ONLY on an explicit mismatch — a -// paraphrase with no hours count is a valid rendering, not a defect (the reference dropped the "no hours -// found" branch that false-flagged 1.8%). Pure and deterministic. -func lintTimeUnits(source, final string) (int, []string) { - m := dcShichenRE.FindStringSubmatch(source) - if m == nil { - return 0, nil - } - n, ok := dcParseCount(m[1]) - if !ok { - return 0, nil - } - hm := dcRuHoursRE.FindStringSubmatch(final) - if hm == nil { - return 0, nil // no explicit hours rendering → a valid paraphrase, not a defect - } - ruNum, ok := dcParseRuHours(hm[1]) - if !ok { - return 0, nil - } - expectedHours := n * 2 - if ruNum == n && ruNum != expectedHours { - return 1, []string{fmt.Sprintf("DC1 时辰: %d个时辰 rendered as «%d час…» (counted as hours) instead of ~%d h (1 时辰 = 2 h)", n, ruNum, expectedHours)} - } - return 0, nil -} - -// dcParseCount parses the DC1 count group: an Arabic digit string or a single small CJK numeral. -func dcParseCount(s string) (int, bool) { - if v, err := strconv.Atoi(s); err == nil { - return v, true - } - r := []rune(s) - if len(r) == 1 { - if v, ok := dcCNNum[r[0]]; ok { - return v, true - } - } - return 0, false -} - -// dcParseRuHours parses the DC1 hours group: an Arabic digit string or a Russian count word. -func dcParseRuHours(s string) (int, bool) { - if v, err := strconv.Atoi(s); err == nil { - return v, true - } - if v, ok := dcRuHourWord[s]; ok { - return v, true - } - return 0, false -} - -// --- DC-2: number-scale magnitude checker (ws5.magnitude_checker, 千万 / 数十万) -------------------- - -// DC2 regexes — a BYTE-FAITHFUL port of ws5.magnitude_checker (_MAG). The ok_re SUPPRESSION guards are -// case-INsensitive (the reference passes re.I); the FIRE predicates are case-SENSITIVE (the reference -// does NOT pass re.I to the inner тысяч/десятки-тысяч searches). \w is ported as \p{L}* (RE2's \w is -// ASCII; the Cyrillic word-continuation the reference intends is letters). -var ( - dcQianwanOkRE = regexp.MustCompile(`(?i)десят\p{L}* миллион|10\s*000\s*000|10000000`) // 千万 correct render → suppress - dcShushiwanOkRE = regexp.MustCompile(`(?i)сотн\p{L}* тысяч|нескольк\p{L}* сот\p{L}* тысяч|[1-9]00\s*000`) // 数十万 correct → suppress - dcDesyatkiTysRE = regexp.MustCompile(`десятк\p{L}* тысяч`) // 数十万 under-render (case-SENSITIVE, per reference) -) - -// lintMagnitudeScale flags a 千万 (10^7) / 数十万 (~several×10^5) magnitude rendered at a WRONG smaller -// scale (an order-of-magnitude error DC2 targets). 千万 → «тысячи» without «миллион» = 10000× under; 数十万 → «десятки -// тысяч» = 10× under. A CORRECT rendering anywhere in the chunk (the ok_re guard) SUPPRESSES the flag, so -// a chunk that renders 数十万 as «сотни тысяч» does not fire even if «десятки тысяч» appears elsewhere as -// an unrelated quantity (the ws5 reference parity — required so the §5(д) FP-measure taken against the -// frozen reference matches what ships). ⚠ 千万 is also a stock HYPERBOLE («несметно») whose «тысячи» -// rendering is in-register literary, NOT a hard error (§5 A4) — hyperbole-exposed, landing FP-gated §5(д). -// Observability only. (Word↔word fraction inversion — 四成四=44% via q4a rule_b1/b2 — is a residual -// sub-checker, not built here.) -func lintMagnitudeScale(source, final string) (int, []string) { - var flags []string - if strings.Contains(source, "千万") && !dcQianwanOkRE.MatchString(final) { - if strings.Contains(final, "тысяч") && !strings.Contains(final, "миллион") { // case-sensitive, per reference - flags = append(flags, "DC2 千万=10^7 rendered as «тысячи» (≈10000× under) — a possible magnitude error (hyperbole risk, §5-A4)") - } - } - if strings.Contains(source, "数十万") && !dcShushiwanOkRE.MatchString(final) { - if dcDesyatkiTysRE.MatchString(final) { - flags = append(flags, "DC2 数十万≈several×10^5 rendered as «десятки тысяч» (≈10× under)") - } - } - return len(flags), flags -} - -// --- DC-6: register-lexicon negative-list (ws5.register_checker) ---------------------------------- - -// dcRegisterNegList is the zh-ru register negative-list (layer-2 pack): fairy-tale-Russian / chancery -// lexemes that break the xianxia register («терем» in a cultivation novel is a hard register error). The -// owner extends this per corpus finding. Lower-cased; whole-word matched. -var dcRegisterNegList = []string{"терем", "терема", "тереме", "теремом", "терему", "теремах"} - -// lintRegisterLexicon flags whole-word occurrences of a register negative-list lexeme in the FINAL text. -func lintRegisterLexicon(final string) (int, []string) { - low := []rune(strings.ToLower(final)) - hitSet := map[string]bool{} - for _, w := range dcRegisterNegList { - wr := []rune(w) - for i := 0; i+len(wr) <= len(low); i++ { - if !runesEqual(low[i:i+len(wr)], wr) { - continue - } - if (i == 0 || !isCyrLetter(low[i-1])) && (i+len(wr) == len(low) || !isCyrLetter(low[i+len(wr)])) { - hitSet[w] = true - } - } - } - if len(hitSet) == 0 { - return 0, nil - } - hits := make([]string, 0, len(hitSet)) - for w := range hitSet { - hits = append(hits, w) - } - sort.Strings(hits) - return len(hits), []string{"DC6 register: fairy-tale Russian lexis out of the xianxia genre: " + strings.Join(hits, ", ")} -} - -// isCyrLetter reports whether r is a Cyrillic letter (the word boundary for the register match). -func isCyrLetter(r rune) bool { return unicode.IsLetter(r) && unicode.Is(unicode.Cyrillic, r) } - -// --- percent-scale checker (成 = tenths) ----------------------------------------------------------- -// -// In Chinese, 成 is one tenth: 六成 = 60%, 六成六 = 66%. A common translation error renders this as a -// decimal FRACTION instead of a percentage — «шесть десятых и шесть сотых» (0.66) or «шесть и шесть -// десятых» (6.6) for 六成六 — a ~100× scale error. This flags that. Precision over recall: it fires only -// when the source has «成[]», the output has NO percent form («процент»/«%»), AND the output -// carries a decimal-fraction cue (a «десятых»/«сотых» ordinal or a «N,N» number). So an output that renders -// the percent correctly is suppressed, and an output without a fraction cue (a paraphrase) stays silent. -// A general zh→ru unit convention, not tied to any book. - -// chengPercentRE matches a 成-percent expression: a count (CJK or Arabic), 成, and an optional second count. -var chengPercentRE = regexp.MustCompile(`([0-9一二三四五六七八九十])成([0-9一二三四五六七八九]?)`) - -// decimalFractionRE is the error cue: the tenths rendered as a Russian fraction ordinal or a decimal number. -var decimalFractionRE = regexp.MustCompile(`десят(?:ая|ых|ой|ые)|сот(?:ая|ых|ой|ые)|\d+[.,]\d`) - -// lintPercentScale flags a 成-percent rendered as a decimal fraction instead of a percentage. -func lintPercentScale(source, final string) (int, []string) { - m := chengPercentRE.FindStringSubmatch(source) - if m == nil { - return 0, nil - } - low := strings.ToLower(final) - if strings.Contains(low, "процент") || strings.Contains(final, "%") { - return 0, nil // the output uses a percent form — the scale is handled correctly - } - if !decimalFractionRE.MatchString(low) { - return 0, nil // no fraction cue — the magnitude was paraphrased, not mis-scaled - } - tens, _ := dcParseCount(m[1]) - pct := tens * 10 - if m[2] != "" { - if ones, ok := dcParseCount(m[2]); ok { - pct += ones - } - } - return 1, []string{fmt.Sprintf("成-percent: %s成%s = %d%% rendered as a decimal fraction instead of a percentage (~%d%%)", m[1], m[2], pct, pct)} -} - -// --- Latin residue in the Russian output ----------------------------------------------------------- -// -// A whole Latin WORD left untranslated in the Russian output (e.g. «открыл их again, …»). It splits the -// output into maximal alphanumeric tokens and flags a token that is worth reporting as leaked prose. The -// target is a lowercase mid-sentence English word, so the guards keep precision high: -// - the token must be ALL LOWERCASE Latin letters. Leaked prose is lowercase; a token with any capital -// is a proper noun / brand / acronym (iPhone, Google, Suzuki, «II») — legitimate in Russian text, and -// the sanitizer already treats capitals as a brand signal, so we skip them here too. -// - it must have NO digit: an id / hash like «5abc35ddfb65» is an alphanumeric token, not a word. -// - at least minLatinResidueLen letters, not a Roman numeral, and not on the per-project allowlist. -// A foreignizing brief that keeps intentional Latin (a motto, a scientific name) puts those on the -// allowlist. Tuned for precision over recall: a leaked word that is Capitalized, or a bare URL host -// («example.com» → «example»/«com»), is not caught — accepted for a low-noise observability signal. -const minLatinResidueLen = 3 - -func lintLatinResidue(final string, allow map[string]bool) (int, []string) { - hits := map[string]bool{} - rs := []rune(final) - for i := 0; i < len(rs); { - if !isLatinLetterOrDigit(rs[i]) { - i++ - continue - } - j := i - reject := false // set on any digit or uppercase letter — not a lowercase leaked word - for j < len(rs) && isLatinLetterOrDigit(rs[j]) { - if (rs[j] >= '0' && rs[j] <= '9') || (rs[j] >= 'A' && rs[j] <= 'Z') { - reject = true - } - j++ - } - tok := string(rs[i:j]) - i = j - if !reject && len([]rune(tok)) >= minLatinResidueLen && !isRomanNumeral(tok) && !allow[tok] { - hits[tok] = true - } - } - if len(hits) == 0 { - return 0, nil - } - surfaces := make([]string, 0, len(hits)) - for s := range hits { - surfaces = append(surfaces, s) - } - sort.Strings(surfaces) - return len(surfaces), []string{"Latin word left untranslated in the Russian output: " + strings.Join(surfaces, ", ")} -} - -// isLatinLetterOrDigit reports whether r is an ASCII Latin letter or digit (the alphanumeric-token alphabet). -func isLatinLetterOrDigit(r rune) bool { - return (r >= 'A' && r <= 'Z') || (r >= 'a' && r <= 'z') || (r >= '0' && r <= '9') -} - -// isRomanNumeral reports whether a lowercase token is a Roman numeral (all chars in ivxlcdm) — «iii» reads -// as a numeral, not a leaked word; excluded to hold precision. (Uppercase «II» is already skipped as a cap.) -func isRomanNumeral(tok string) bool { - for _, r := range tok { - switch r { - case 'i', 'v', 'x', 'l', 'c', 'd', 'm': - default: - return false - } - } - return true -} - -// --- broken Russian word forms --------------------------------------------------------------------- -// -// Deliberately GENERAL, with NO dictionary and NO book-specific word lists: it flags a Russian word ending -// in «-йть», which no well-formed Russian word does — the shape of a mangled infinitive (e.g. «войть» for -// «войти»). It is zero false-positive and applies to any book. Malformations WITHOUT such a structural -// signature — a plausible misspelling («Вперди» for «Впереди») or a case-agreement error («глава клан» -// for «главы клана») — are NOT detectable deterministically without a morphology/dictionary pass, which is -// out of scope here; the report states this recall limit honestly. The existing sanitizer broken-word class -// (invalid soft/hard-sign bigrams, script-mixed homoglyph tokens) is orthogonal and still runs. -func lintBrokenWord(final string) (int, []string) { - seen := map[string]bool{} - var det []string - for _, w := range tokenizeCyrillic(final) { - if len([]rune(w)) >= 4 && strings.HasSuffix(w, "йть") && !seen[w] { - seen[w] = true - det = append(det, "malformed word ending in «-йть» (no valid Russian word does): "+w) - } - } - sort.Strings(det) - return len(det), det -} diff --git a/backend/internal/pipeline/checkers_zh_ru_test.go b/backend/internal/pipeline/checkers_zh_ru_test.go index d63846a1..d40accda 100644 --- a/backend/internal/pipeline/checkers_zh_ru_test.go +++ b/backend/internal/pipeline/checkers_zh_ru_test.go @@ -3,8 +3,18 @@ package pipeline import ( "strings" "testing" + + "textmachine/backend/internal/lang" ) +// testCheckers builds the compiled checker spec from the REAL zh-ru pack + ru target data (pair-14 data-out): +// the fixtures exercise the checker ALGORITHM over the pack's own DETECTION patterns / tables / wordlists, so +// no pair literal is duplicated in the test — a pattern edit in the langpack flows straight into these cases. +func testCheckers(t *testing.T) *dcCheckers { + t.Helper() + return compileCheckers(testLangPack(t).DCCheckers, lang.TargetChecksFor("ru")) +} + // checkers_zh_ru_test.go: WS5 (г) — the DC1/DC2/DC6 checkers must CATCH the empirical trap positives // (ws5_checkers_verify.py §a) and stay SILENT on clean/in-register text (precision over recall). These // are the regression fixtures the plan requires (positives from exp15 §7.9 / exp14b). @@ -20,9 +30,10 @@ func TestDC1TimeUnits(t *testing.T) { {"no-shichen", "他走了三里。", "Он прошёл три ли.", false}, {"arabic-count", "过了2个时辰。", "Прошло 2 часа.", true}, // 2时辰=4h rendered «2 часа» } + dcc := testCheckers(t) for _, c := range cases { t.Run(c.name, func(t *testing.T) { - n, _ := lintTimeUnits(c.src, c.tgt) + n, _ := dcc.lintTimeUnits(c.src, c.tgt) if (n > 0) != c.wantFlag { t.Fatalf("lintTimeUnits(%q,%q) fired=%v, want %v", c.src, c.tgt, n > 0, c.wantFlag) } @@ -47,9 +58,10 @@ func TestDC2MagnitudeScale(t *testing.T) { {"case-sensitive-capitalized", "聚集了数十万人。", "Десятки тысяч человек собрались.", false}, {"no-magnitude", "他有三个朋友。", "У него три друга.", false}, } + dcc := testCheckers(t) for _, c := range cases { t.Run(c.name, func(t *testing.T) { - n, _ := lintMagnitudeScale(c.src, c.tgt) + n, _ := dcc.lintMagnitudeScale(c.src, c.tgt) if (n > 0) != c.wantFlag { t.Fatalf("lintMagnitudeScale(%q,%q) fired=%v, want %v", c.src, c.tgt, n > 0, c.wantFlag) } @@ -67,9 +79,10 @@ func TestDC6RegisterLexicon(t *testing.T) { {"clean", "Он вошёл в высокий зал павильона.", 0}, {"substring-not-word", "Термин был странным.", 0}, // «терм» inside «термин» must NOT fire (whole-word) } + dcc := testCheckers(t) for _, c := range cases { t.Run(c.name, func(t *testing.T) { - n, det := lintRegisterLexicon(c.tgt) + n, det := dcc.lintRegisterLexicon(c.tgt) if n != c.wantN { t.Fatalf("lintRegisterLexicon(%q) = %d (%v), want %d", c.tgt, n, det, c.wantN) } @@ -92,15 +105,15 @@ func TestDC3GenderInjection(t *testing.T) { mk := func(gender string) []pickedEntry { return []pickedEntry{{entry: &memoryEntry{src: "方源", dst: "Фан Юань", status: "approved", gender: gender}, via: "方源", disp: memConfirmed}} } - male := renderEditorConstraintBlock(mk("male")) + male := renderEditorConstraintBlock(mk("male"), ruTX()) if !strings.Contains(male, "方源 → «Фан Юань» (муж.") { t.Fatalf("male term must carry a masculine directive, got:\n%s", male) } - hidden := renderEditorConstraintBlock(mk("hidden")) + hidden := renderEditorConstraintBlock(mk("hidden"), ruTX()) if !strings.Contains(hidden, "пол СКРЫТ") { t.Fatalf("hidden term must carry the gender-avoidance mandate, got:\n%s", hidden) } - none := renderEditorConstraintBlock(mk("")) + none := renderEditorConstraintBlock(mk(""), ruTX()) if !strings.HasSuffix(none, "«Фан Юань»") { t.Fatalf("a genderless term must have NO gender note (line ends at the dst), got:\n%s", none) } diff --git a/backend/internal/pipeline/chunker.go b/backend/internal/pipeline/chunker.go index 5241e7d8..8fcd3d03 100644 --- a/backend/internal/pipeline/chunker.go +++ b/backend/internal/pipeline/chunker.go @@ -198,13 +198,14 @@ func matchHeaderLine(line string, hr *lang.HeadingRule) (n int, subtitle string, } // isHeadingNumeral reports whether a rune can be part of a chapter-number run: an Arabic digit -// (half/fullwidth) or a CJK numeral character. +// (half/fullwidth) or a CJK numeral character (the shared lang.CJKSection — pair-14 §4, the SAME source +// ingest's chapter-numeral regex reads, so the two can never byte-drift apart). func isHeadingNumeral(r rune) bool { switch { case r >= '0' && r <= '9', r >= '0' && r <= '9': return true } - return strings.ContainsRune("〇零一二三四五六七八九十百千两兩", r) + return lang.DefaultCJKSection().IsHeadingNumeral(r) } // isHeaderContentRune reports whether a rune is CONTENT (a letter/digit/ideograph/kana) rather than a @@ -239,7 +240,7 @@ func parseSectionNumeral(s string) (int, bool) { case r >= '0' && r <= '9': num = num*10 + int(r-'0') any = true - case r == '〇' || r == '零': + case lang.DefaultCJKSection().Zero[r]: num = num * 10 any = true default: @@ -267,40 +268,16 @@ func parseSectionNumeral(s string) (int, bool) { return v, true } +// cjkSectionDigit / cjkSectionUnit read the shared lang.CJKSection (pair-14 §4): the SAME digit/unit value +// tables the ingest chapter-numeral inventory derives from, so there is ONE source of truth. func cjkSectionDigit(r rune) (int, bool) { - switch r { - case '一': - return 1, true - case '二', '两', '兩': - return 2, true - case '三': - return 3, true - case '四': - return 4, true - case '五': - return 5, true - case '六': - return 6, true - case '七': - return 7, true - case '八': - return 8, true - case '九': - return 9, true - } - return 0, false + v, ok := lang.DefaultCJKSection().Digit[r] + return v, ok } func cjkSectionUnit(r rune) (int, bool) { - switch r { - case '十': - return 10, true - case '百': - return 100, true - case '千': - return 1000, true - } - return 0, false + v, ok := lang.DefaultCJKSection().Unit[r] + return v, ok } // chapterDraftChunks packs one chapter's paragraphs into DRAFT chunks (the fine tiling). Rule: @@ -621,17 +598,13 @@ func isASCIISpace(r rune) bool { return false } -// sourceAbbrevs are trailing tokens after which a lone "." is treated as an -// abbreviation, not a sentence end (case-insensitive). A pragmatic set for prose; -// the acceptance path is ja→ru where CJK terminators dominate, so this only guards -// the en source. Single-letter initials are handled separately (isAbbrevBefore). -var sourceAbbrevs = map[string]bool{ - "mr": true, "mrs": true, "ms": true, "dr": true, "prof": true, "st": true, - "jr": true, "sr": true, "vs": true, "no": true, "vol": true, "ch": true, - "fig": true, "col": true, "gen": true, "sgt": true, "capt": true, "lt": true, - "rev": true, "gov": true, "sen": true, "rep": true, "etc": true, "inc": true, - "ltd": true, "co": true, "mt": true, "ave": true, "rd": true, -} +// sourceAbbrevs are trailing tokens after which a lone "." is treated as an abbreviation, not a sentence +// end (case-insensitive). DATA lives in internal/lang (embedded, sectioned per SOURCE language, pair-14 §4); +// the splitter consumes the "en" section — the one source whose ASCII period needs the guard (the acceptance +// path is zh/ja→ru where 。 terminators dominate, so those sources ship no abbreviations). Making the splitter +// consume the BOOK's source section is a shallow follow-up (thread sourceLang into SplitChunks). Single- +// letter initials are handled separately (isAbbrevBefore). +var sourceAbbrevs = lang.SentenceAbbrev("en") // isAbbrevBefore reports whether the text immediately before a lone "." ends in a // known abbreviation or a single-letter initial (so the "." is not a boundary). diff --git a/backend/internal/pipeline/chunkrun.go b/backend/internal/pipeline/chunkrun.go index 110c8bb9..a11fa38c 100644 --- a/backend/internal/pipeline/chunkrun.go +++ b/backend/internal/pipeline/chunkrun.go @@ -5,6 +5,8 @@ import ( "maps" "slices" "strings" + + "textmachine/backend/internal/lang" "textmachine/backend/internal/store" ) @@ -77,8 +79,9 @@ func (r *Runner) classifyOutput(role, source, output, finish string, isFinal boo return cls, "" } -// injectionRenderer serializes a chunk's selected memory records into a role's injection message. -type injectionRenderer func(injected []pickedEntry) string +// injectionRenderer serializes a chunk's selected memory records into a role's injection message, using the +// target-language wire-text (lang.InjectionTexts, pair-14 §2) — a target with no texts renders nothing. +type injectionRenderer func(injected []pickedEntry, tx lang.InjectionTexts) string // roleInjectionRenderers maps a stage ROLE to the memory-injection renderer it consumes — a registry // instead of a hand-edited switch (D39 layer 7, L8-role-injection-hardcoded-switch), so a new Phase-2 @@ -171,5 +174,6 @@ func (r *Runner) cheapGateConfig() cheapGateConfig { yoPolicy: r.Book.YoPolicy, allowlist: allow, regressionEnabled: r.Pipeline.Gates.RegressionGuard.Enabled, + checkers: r.checkers, // compiled once in openRunner (pair-14 data-out); nil-inert for a no-pack book } } diff --git a/backend/internal/pipeline/disposition.go b/backend/internal/pipeline/disposition.go index 11189622..cadd0cd1 100644 --- a/backend/internal/pipeline/disposition.go +++ b/backend/internal/pipeline/disposition.go @@ -7,6 +7,7 @@ import ( "unicode" "unicode/utf8" + "textmachine/backend/internal/lang" "textmachine/backend/internal/llm" ) @@ -231,23 +232,12 @@ func classify(in classifyInput) classification { } // --- refusal blacklist (en/ru/zh/ja) — ported 1:1 from eval/refusal_bench.py --- - -var refusalPatterns = []string{ - `I (?:can(?:'|no)t|cannot|won'?t|am (?:not able|unable)) (?:to )?(?:assist|help|translate|continue|provide|fulfill)`, - `I'?m sorry,? but`, - `against (?:my|our) (?:content )?polic`, - `content polic(?:y|ies)`, - `не могу (?:помочь|перевести|продолжить|выполнить)`, - `я не буду переводить`, - `противоречит (?:политике|правилам)`, - `としてお答えできません`, - `申し訳ありませんが`, - `我(?:不能|无法)(?:协助|翻译|继续)`, - `抱歉[,,]`, - `^\s*\[?(?:REFUSED|ОТКАЗ)\]?\s*$`, -} - -var refusalRE = regexp.MustCompile("(?im)" + strings.Join(refusalPatterns, "|")) +// +// The patterns are UNIVERSAL engine safety data (a model can refuse in any language regardless of the +// book's pair), so they live in internal/lang as an embedded, per-language-sectioned file (pair-14 §3), NOT +// a book pack — a no-langpack book (the ja→ru golden) still flags a refusal. Joined here into one +// case-insensitive regex, byte-identically to the ported reference. +var refusalRE = regexp.MustCompile("(?im)" + strings.Join(lang.RefusalPatterns(), "|")) var thinkRE = regexp.MustCompile(`(?s).*?\s*`) diff --git a/backend/internal/pipeline/fewshot_test.go b/backend/internal/pipeline/fewshot_test.go index 1e9ee729..c494a23f 100644 --- a/backend/internal/pipeline/fewshot_test.go +++ b/backend/internal/pipeline/fewshot_test.go @@ -64,7 +64,7 @@ func TestFewShotEnabledDefault(t *testing.T) { // the 物是人非 chengyu-atom), and few_shot OFF keeps the discourse core (incl. the chengyu-atom) // but drops the examples — never losing the meaning-preservation instructions. func TestEditorFewShotToggleRealFile(t *testing.T) { - tpl, err := LoadPromptTemplate(filepath.Join("..", "..", "prompts", "editor.md")) + tpl, err := LoadPromptTemplate(filepath.Join("..", "..", "prompts", "zh-ru", "editor.md")) if err != nil { t.Fatalf("load editor.md: %v", err) } diff --git a/backend/internal/pipeline/ingest.go b/backend/internal/pipeline/ingest.go index 8e25e442..1bfc5951 100644 --- a/backend/internal/pipeline/ingest.go +++ b/backend/internal/pipeline/ingest.go @@ -11,9 +11,12 @@ import ( "path" "regexp" "strings" + "sync" "unicode" "unicode/utf8" + "textmachine/backend/internal/lang" + "golang.org/x/text/encoding/simplifiedchinese" xunicode "golang.org/x/text/encoding/unicode" "golang.org/x/text/transform" @@ -112,12 +115,20 @@ func ingestTXT(p, encoding, sourceLang string) (*Document, error) { // --- CJK chapter-header splitting (D18: real zh/ja txt mark chapters as «第N章/节/回») --------- -// chapterUnitRunes are the section-level chapter markers auto-detected in a txt. 卷 (volume) is -// deliberately EXCLUDED — it is coarser than a chapter and would carve a tiny title-only "chapter". -var chapterUnitRunes = []rune{'章', '节', '節', '回'} +// chapterNumeralRE matches the numeral run of a chapter header: Arabic (half/fullwidth) or CJK. Built ONCE +// from the shared lang.CJKSection (pair-14 §4 + addendum-A): the section markers (章节節回) and the CJK +// numeral class are the SAME data the chunker heading rule reads — ONE source, no ingest↔chunker byte-drift. +// A CONSTANT (the CJK numeral system is language-invariant), so it is available even to a no-langpack book +// (the ja→ru golden splits 第X章 here). Arabic ranges stay in code (not language data). +var chapterNumeralReOnce sync.Once +var chapterNumeralReVal *regexp.Regexp -// chapterNumeralRE matches the numeral run of a chapter header: Arabic (half/fullwidth) or CJK. -var chapterNumeralRE = regexp.MustCompile(`^\s*第[0-90-9〇零一二三四五六七八九十百千两兩]+`) +func chapterNumeralRE() *regexp.Regexp { + chapterNumeralReOnce.Do(func() { + chapterNumeralReVal = regexp.MustCompile(`^\s*第[0-90-9` + lang.DefaultCJKSection().HeadingNumeralClass() + `]+`) + }) + return chapterNumeralReVal +} // chapterHeaderMaxRunes bounds a header line so a prose sentence that merely opens with «第三节…» // (a longer line) is not mistaken for a header. The 蛊真人 headers are ≤23 runes; 60 leaves room for @@ -138,7 +149,7 @@ func isCJKChapterHeader(line string, unit rune) bool { if utf8.RuneCountInString(t) == 0 || utf8.RuneCountInString(t) > chapterHeaderMaxRunes { return false } - loc := chapterNumeralRE.FindStringIndex(t) + loc := chapterNumeralRE().FindStringIndex(t) if loc == nil { return false } @@ -170,7 +181,9 @@ func isHeaderSeparator(r rune) bool { // single stray match). Returns 0 when no unit qualifies (→ the part stays a single chapter). func detectChapterUnit(lines []string) rune { best, bestN := rune(0), 0 - for _, unit := range chapterUnitRunes { + // The section markers (章节節回) are shared lang.CJKSection data (pair-14 addendum-A); iterated in AUTHORED + // order so a tie between two units breaks deterministically (first-seen), as the fixed slice once did. + for _, unit := range lang.DefaultCJKSection().ChapterUnitOrdered { n := 0 for _, ln := range lines { if isCJKChapterHeader(ln, unit) { diff --git a/backend/internal/pipeline/memory.go b/backend/internal/pipeline/memory.go index e96c7331..96c744b8 100644 --- a/backend/internal/pipeline/memory.go +++ b/backend/internal/pipeline/memory.go @@ -11,6 +11,7 @@ import ( "strings" "unicode" + "textmachine/backend/internal/lang" "textmachine/backend/internal/store" ) @@ -498,17 +499,17 @@ func priorityRank(p pickedEntry) int { return rank } -// glossaryBlockHeader introduces the injected glossary block. The layout mirrors the -// SakuraLLM/GalTransl convention "src → dst", with unverified (ambiguous) records -// tagged so the model — and a human reviewer — see they are not authoritative (A2). -const glossaryBlockHeader = "ГЛОССАРИЙ (используй эти утверждённые переводы имён и терминов последовательно; строки с пометкой ⟨проверить⟩ — неподтверждённые кандидаты):" - // renderGlossaryBlock serializes the selected records into the injection message // (§C). Records with no dst yet (ruby candidates) are skipped — a "src → " line // carries nothing. Returns "" when nothing renders, so an empty selection injects NO // message at all ("better nothing than garbage"). Deterministic: the injected order is the -// budget priority order fixed by Select. -func renderGlossaryBlock(injected []pickedEntry) string { +// budget priority order fixed by Select. tx is the TARGET-language wire-text (pair-14 §2): a +// target with no injection texts (HasData()==false) injects nothing (a non-ru book gets no +// Russian block), so the whole render is gated on it. +func renderGlossaryBlock(injected []pickedEntry, tx lang.InjectionTexts) string { + if !tx.HasData() { + return "" + } var lines []string for _, p := range injected { if strings.TrimSpace(p.entry.dst) == "" { @@ -516,21 +517,21 @@ func renderGlossaryBlock(injected []pickedEntry) string { } line := p.entry.src + " → " + p.entry.dst if p.disp != memConfirmed { - line += " ⟨проверить⟩" + line += tx.UnverifiedMarker } else { // pack-13 injection-completeness fix (D39.21 owner directive: «род должен доезжать»): the // gender of a CONFIRMED named term (蛊 Надежды = ж.р., a cicada gu, …) now reaches the DRAFT // wire too, not only the editor — so the TRANSLATOR renders the right родовые формы FIRST, // instead of leaving the editor to repair a wrong gender. "" for a genderless term (the common // case → byte-identical). Mirrors the editor block's confirmed-only gender (renderEditorConstraintBlock). - line += genderConstraintNote(p.entry.gender) + line += genderConstraintNote(p.entry.gender, tx) } lines = append(lines, line) } if len(lines) == 0 { return "" } - return glossaryBlockHeader + "\n" + strings.Join(lines, "\n") + return tx.GlossaryHeader + "\n" + strings.Join(lines, "\n") } // renderFormatVersion versions the FORMAT of the role-injection renderers whose output is NOT @@ -547,11 +548,10 @@ func renderGlossaryBlock(injected []pickedEntry) string { // draft injected bytes → the wire, so a loud --resnapshot; a genderless bank is byte-identical to v2. const renderFormatVersion = "renderfmt-v3-draft-gender+editor-src2dst+dc3-gender" -// editorConstraintHeader introduces the editor's canonical-constraint block. It gives the BILINGUAL -// editor (D30.1) the approved src→dst bindings as consistency constraints — it must render each -// listed source term with EXACTLY the paired canonical form (inflecting for context) and touch -// nothing else. -const editorConstraintHeader = "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имён и терминов (в черновике термин исходника слева ДОЛЖЕН быть передан именно указанной формой справа — приводи к ней любые расхождения, склоняя по контексту; не вводи иных вариантов и не меняй ничего другого):" +// The editor's canonical-constraint block header and the glossary header are TARGET-language wire-text +// (lang.InjectionTexts, pair-14 §2) — they gave the BILINGUAL editor (D30.1) the approved src→dst bindings +// as consistency constraints. Relocated out of the pipeline (a Russian header rendered for a →en book was a +// leak); now gated on the target and sourced from embedded per-target data. // renderEditorConstraintBlock serializes the selected records' CONFIRMED renderings into the // editor's injection as a src→dst MAPPING (WS2 §2в resolves the D30.1-open question): "源термин → @@ -568,7 +568,10 @@ const editorConstraintHeader = "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имё // separate token budget. The src→dst FORMAT is snapshot-folded via renderFormatVersion (a format // edit is a loud --resnapshot); the injection is a message, so it also enters request_hash // directly (no silent false-hit). -func renderEditorConstraintBlock(injected []pickedEntry) string { +func renderEditorConstraintBlock(injected []pickedEntry, tx lang.InjectionTexts) string { + if !tx.HasData() { + return "" + } var lines []string seen := map[[2]string]bool{} for _, p := range injected { @@ -585,12 +588,12 @@ func renderEditorConstraintBlock(injected []pickedEntry) string { continue } seen[key] = true - lines = append(lines, "- "+src+" → «"+dst+"»"+genderConstraintNote(p.entry.gender)) + lines = append(lines, "- "+src+" → «"+dst+"»"+genderConstraintNote(p.entry.gender, tx)) } if len(lines) == 0 { return "" } - return editorConstraintHeader + "\n" + strings.Join(lines, "\n") + return tx.EditorHeader + "\n" + strings.Join(lines, "\n") } // genderConstraintNote is the DC3 gender directive appended to a CONFIRMED editor-constraint line (WS5 @@ -598,14 +601,14 @@ func renderEditorConstraintBlock(injected []pickedEntry) string { // the Bai Ninbing class — no coreference needed. male/female ⇒ hard gender forms; hidden ⇒ a mandate to // AVOID gender-marking constructions until the reveal (a masculine default when unavoidable, D19.3). "" // for a term with no gender datum (the common case → the line is unchanged, byte-identical to before). -func genderConstraintNote(gender string) string { +func genderConstraintNote(gender string, tx lang.InjectionTexts) string { switch gender { case "male", "m": - return " (муж. — мужские родовые формы)" + return tx.GenderMale case "female", "f": - return " (жен. — женские родовые формы)" + return tx.GenderFemale case "hidden": - return " (пол СКРЫТ до раскрытия — избегай родовых форм; при неизбежности — мужские)" + return tx.GenderHidden } return "" } diff --git a/backend/internal/pipeline/memory_test.go b/backend/internal/pipeline/memory_test.go index 7fced217..ac1cd8c6 100644 --- a/backend/internal/pipeline/memory_test.go +++ b/backend/internal/pipeline/memory_test.go @@ -23,7 +23,9 @@ func declJSON(invariant bool, forms ...string) string { return string(b) } -func alias(a string) store.GlossaryAlias { return store.GlossaryAlias{Alias: a, AliasType: "прозвище"} } +func alias(a string) store.GlossaryAlias { + return store.GlossaryAlias{Alias: a, AliasType: "прозвище"} +} // specGlossary mirrors eval/memory_hotpath.py's G (with decl forms replacing the // toy accept-regexps). @@ -237,7 +239,7 @@ func TestSingleKeyBanAndAllowShort(t *testing.T) { func TestBudgetEviction(t *testing.T) { entries := []store.GlossaryEntry{ gl("甲甲", "Альфа", "", "approved"), - gl("乙乙", "Бета", "", "auto"), // ambiguous → lower priority + gl("乙乙", "Бета", "", "auto"), // ambiguous → lower priority gl("丙丙", "Гамма", "", "approved"), } b := bankFrom(entries) @@ -295,7 +297,7 @@ func TestEmptyDstNotMatchable(t *testing.T) { if len(sel.injected) != 1 || sel.injected[0].entry.src != "乙乙" { t.Fatalf("empty-dst phantom polluted selection: injected=%v", injMap(sel)) } - if renderGlossaryBlock(sel.injected) == "" { + if renderGlossaryBlock(sel.injected, ruTX()) == "" { t.Error("a renderable line was evicted by an empty-dst phantom") } // Unbounded: the phantom never appears as an exact hit. @@ -368,8 +370,8 @@ func TestPerLanguageMinKeyAndCollisionDisposition(t *testing.T) { gl("鈴木", "Судзуки", "", "approved"), // 2 Han → fires, CONFIRMED (ideographic anchor) gl("すずき", "Судзуки-х", "s2", "approved"), // 3 kana → fires but AMBIGUOUS (collision-prone) gl("ながいなまえ", "Длинное имя", "", "approved"), // 6 kana → fires, CONFIRMED (long enough) - gl("リン", "Рин", "", "approved"), // 2 kana → BANNED (phonetic min 3) - gl("ai", "ИИ", "", "approved"), // 2 latin → BANNED (phonetic min 3) + gl("リン", "Рин", "", "approved"), // 2 kana → BANNED (phonetic min 3) + gl("ai", "ИИ", "", "approved"), // 2 latin → BANNED (phonetic min 3) } b := bankFrom(entries) sel := b.Select("鈴木とすずきとながいなまえが会った。", 1, nil, 0) @@ -548,7 +550,7 @@ func TestRenderEditorConstraintBlock(t *testing.T) { {entry: tanakaDup, via: "田中", disp: memConfirmed}, {entry: homonym, via: "済", disp: memConfirmed}, } - block := renderEditorConstraintBlock(sel) + block := renderEditorConstraintBlock(sel, ruTX()) // src→dst mapping present under the editor's own header. if !strings.Contains(block, "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ") { @@ -570,7 +572,7 @@ func TestRenderEditorConstraintBlock(t *testing.T) { t.Errorf("homonymic dst must stay distinguishable by source term: %q", block) } // An all-AMBIGUOUS (or empty) selection yields no block at all. - if got := renderEditorConstraintBlock([]pickedEntry{{entry: cand, disp: memAmbiguous}}); got != "" { + if got := renderEditorConstraintBlock([]pickedEntry{{entry: cand, disp: memAmbiguous}}, ruTX()); got != "" { t.Errorf("an all-AMBIGUOUS selection must yield no editor block, got %q", got) } } diff --git a/backend/internal/pipeline/memory_trustgate_test.go b/backend/internal/pipeline/memory_trustgate_test.go index ad45b95e..5d20f2dc 100644 --- a/backend/internal/pipeline/memory_trustgate_test.go +++ b/backend/internal/pipeline/memory_trustgate_test.go @@ -43,7 +43,7 @@ func TestSuppressorTrustGateFires(t *testing.T) { t.Errorf("draft 四代族长 must still inject AMBIGUOUS alongside, got %q", m["四代族长"]) } // Editor constraint block: the approved term is now present (was empty before the fix). - if blk := renderEditorConstraintBlock(sel.injected); !strings.Contains(blk, "глава клана") { + if blk := renderEditorConstraintBlock(sel.injected, ruTX()); !strings.Contains(blk, "глава клана") { t.Errorf("editor constraint block must carry the approved «глава клана», got %q", blk) } // Loud record of the refused suppression. diff --git a/backend/internal/pipeline/miner_alias.go b/backend/internal/pipeline/miner_alias.go index 8cc496b3..1a1ddd3a 100644 --- a/backend/internal/pipeline/miner_alias.go +++ b/backend/internal/pipeline/miner_alias.go @@ -47,10 +47,10 @@ func isFragment(src string, seedSurfaces map[string]bool, p *lang.Pack) bool { } // cooccurSameSentence reports whether a and b appear in one source sentence anywhere (alias. -// cooccur_same_sentence: split on 。!?\n). -func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string) bool { +// cooccur_same_sentence: split on the pack's sentence terminators + \n). +func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string, p *lang.Pack) bool { for _, c := range chunks { - for _, sent := range splitMinerSentences(c.NSource) { + for _, sent := range splitMinerSentences(c.NSource, p) { if strings.Contains(sent, aNorm) && strings.Contains(sent, bNorm) { return true } @@ -59,9 +59,11 @@ func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string) bool { return false } -func splitMinerSentences(s string) []string { +// splitMinerSentences splits on the pair's source sentence terminators (langpack DATA, pair-14) plus the +// structural newline (a layout mark, not a language one, so it stays in code). +func splitMinerSentences(s string, p *lang.Pack) []string { return strings.FieldsFunc(s, func(r rune) bool { - return r == '。' || r == '!' || r == '?' || r == '\n' + return r == '\n' || p.SentenceTerminator[r] }) } @@ -112,7 +114,7 @@ func proposeAliasEdges(surfaces map[string]aliasSurface, chunks []MinerChunk, se givenA := trimRunePrefix(a.src, surA) givenB := trimRunePrefix(b.src, surB) if givenA != givenB && givenA != "" && givenB != "" { // R4-i different given names → family only - if cooccurSameSentence(chunks, a.src, b.src) { + if cooccurSameSentence(chunks, a.src, b.src, p) { family = append(family, aliasEdge{a.src, b.src, "R2+R4iv", "family_copresent"}) } else { family = append(family, aliasEdge{a.src, b.src, "R2", "family"}) diff --git a/backend/internal/pipeline/miner_palladius.go b/backend/internal/pipeline/miner_palladius.go index 84072a4e..5285f9c8 100644 --- a/backend/internal/pipeline/miner_palladius.go +++ b/backend/internal/pipeline/miner_palladius.go @@ -21,32 +21,28 @@ import ( // name-shape signal a future faithful-A variant (with a Go morphology backend) would consume, and the // pattern P4 transliteration-by-pair seam (§B5). -var palladiusJQX = map[string]bool{"j": true, "q": true, "x": true} - // buildPalladiusCyrSyllables derives the distinct Cyrillic syllable set, longest-first, for greedy // segmentation (palladius._CYR_SYL) from the pack's pinyin→Cyrillic table. Ties in length never both match // a position (distinct strings), so the order among equal-length forms is immaterial — sorted (len desc, -// then value) for a stable artifact. The table (initials/finals/Y_W/SPECIAL_I) is langpack DATA. +// then value) for a stable artifact. The table (initials/finals/Y_W/SPECIAL_I) AND the phonotactic +// constraints (retroflex / ü-finals / their legal initials) are langpack DATA (pair-14). func buildPalladiusCyrSyllables(p *lang.Pack) []string { set := map[string]bool{} - for _, v := range p.PalladiusYW { + pal := p.Palladius + for _, v := range pal.YW { set[v] = true } - for _, v := range p.PalladiusSpecialI { + for _, v := range pal.SpecialI { set[v] = true } - retroflex := map[string]bool{"zh": true, "ch": true, "sh": true, "r": true, "z": true, "c": true, "s": true} - for pi, ci := range p.PalladiusInitials { - for pf, cf := range p.PalladiusFinals { - // ü finals (written v) are valid only after j/q/x (pinyin writes ü as plain u) or l/n. - switch pf { - case "v", "ve", "van", "vn": - if !palladiusJQX[pi] && pi != "l" && pi != "n" { - continue - } + for pi, ci := range pal.Initials { + for pf, cf := range pal.Finals { + // ü finals (written v) are valid only after a vfinal_initial (pinyin writes ü as plain u). + if pal.VFinal[pf] && !pal.VFinalInitial[pi] { + continue } // skip the retroflex/sibilant + bare i (handled by SPECIAL_I). - if pf == "i" && retroflex[pi] { + if pf == "i" && pal.Retroflex[pi] { continue } set[ci+cf] = true diff --git a/backend/internal/pipeline/miner_patterns.go b/backend/internal/pipeline/miner_patterns.go index 05050e24..75429dcd 100644 --- a/backend/internal/pipeline/miner_patterns.go +++ b/backend/internal/pipeline/miner_patterns.go @@ -56,13 +56,14 @@ func runeHasPrefixAt(run []rune, i int, p []rune) bool { return true } -// formantType maps a formant char to a candidate type (patterns.formant_type): topo→place, 等/转/阶→ -// title, else term. +// formantType maps a formant char to a candidate type (patterns.formant_type): topo→place, title-formant +// (等/转/阶)→title, else term. Both the place-side (TopoSuffix) and title-side (TitleFormant) inventories are +// langpack DATA (pair-14: the title-formant set moved out of this switch, beside its TopoSuffix neighbour). func formantType(c rune, p *lang.Pack) string { if p.TopoSuffix[c] { return "place" } - if c == '等' || c == '转' || c == '阶' { + if p.TitleFormant[c] { return "title" } return "term" diff --git a/backend/internal/pipeline/miner_test.go b/backend/internal/pipeline/miner_test.go index fc4101e2..9a93c4fc 100644 --- a/backend/internal/pipeline/miner_test.go +++ b/backend/internal/pipeline/miner_test.go @@ -14,7 +14,10 @@ import ( // exercised over the exact same data. func testLangPack(t *testing.T) *lang.Pack { t.Helper() - p, err := lang.Load("../../configs/langpacks", "zh", "ru") + // Load the shared zh-ru pack PLUS the 蛊真人 book-scoped overlay (pair-14 §1): 古月 is the book's clan, + // carried in the overlay, not the shared pair langpack. Every miner fixture (and the stand parity test) + // runs on this book's effective pack, so the surname-anchor cases (古月方源) and {古月:22} stay EXACT. + p, err := lang.LoadWithOverlay("../../configs/langpacks", "zh", "ru", "testdata/langpack-overlay-guzhenren") if err != nil { t.Fatalf("load langpack: %v", err) } diff --git a/backend/internal/pipeline/render_memory_test.go b/backend/internal/pipeline/render_memory_test.go index 0dcdabbd..57242194 100644 --- a/backend/internal/pipeline/render_memory_test.go +++ b/backend/internal/pipeline/render_memory_test.go @@ -3,8 +3,14 @@ package pipeline import ( "strings" "testing" + + "textmachine/backend/internal/lang" ) +// ruTX is the ru target injection wire-text (pair-14 §2), for the renderer fixtures that assert the +// Russian glossary/editor block bytes (the values that moved to lang/data/injection.txt). +func ruTX() lang.InjectionTexts { return lang.InjectionTextsFor("ru") } + func TestMessagesInjectionLayout(t *testing.T) { tpl := writeTemplate(t, "Стабильный системный префикс.\n---USER---\nПереведи: {{text}}") v := RenderVars{Book: testBook(), Text: "исходный чанк"} @@ -73,7 +79,7 @@ func TestRenderGlossaryBlock(t *testing.T) { ambiguous := pickedEntry{entry: &memoryEntry{src: "小D", dst: "Малыш Дэ", status: "auto"}, disp: memAmbiguous} emptyDst := pickedEntry{entry: &memoryEntry{src: "鈴木", dst: "", status: "auto"}, disp: memAmbiguous} - block := renderGlossaryBlock([]pickedEntry{confirmed, ambiguous, emptyDst}) + block := renderGlossaryBlock([]pickedEntry{confirmed, ambiguous, emptyDst}, ruTX()) if !strings.Contains(block, "阿Q → А-кью") { t.Errorf("confirmed line missing: %q", block) } @@ -84,10 +90,10 @@ func TestRenderGlossaryBlock(t *testing.T) { t.Errorf("empty-dst candidate should be skipped (nothing to inject): %q", block) } // Nothing to inject → empty block (better nothing). - if renderGlossaryBlock(nil) != "" { + if renderGlossaryBlock(nil, ruTX()) != "" { t.Error("empty selection must render an empty block") } - if renderGlossaryBlock([]pickedEntry{emptyDst}) != "" { + if renderGlossaryBlock([]pickedEntry{emptyDst}, ruTX()) != "" { t.Error("only-empty-dst selection must render an empty block") } } diff --git a/backend/internal/pipeline/runner.go b/backend/internal/pipeline/runner.go index 5469d38f..2ac3164c 100644 --- a/backend/internal/pipeline/runner.go +++ b/backend/internal/pipeline/runner.go @@ -69,6 +69,11 @@ type Runner struct { // enriched `memory`. nil ⇔ memory is nil (materialized together in seedGlossary). baseMemory *MemoryBank + // checkers is the compiled WS5/pack-13 observability checker spec (pair-14 data-out): the pair's + // DETECTION patterns + tables (from pack) + the target-general lists (from embedded target data), + // compiled once in loadLangPack. nil-safe: a no-pack / no-target-data book runs the checkers inert. + checkers *dcCheckers + // pack is the book's language-data pack (internal/lang), loaded once in openRunner from // book.LangpackRoot (D39.15/16). The the bank-mining stop bank-miner reads its tables; pack.Version() is folded into // the snapshot (a pack edit is a loud --resnapshot). nil when the book declares no langpack_root, or @@ -191,6 +196,10 @@ func openRunner(bookPath string, logger *slog.Logger, forWrite bool) (*Runner, e // catalog — or a book with no langpack_root — runs with a nil pack (the miner is inert, the bank-mining stop auto-continues, // and the snapshot fold is omitted). Presence of the pair directory is the "catalog exists" signal. func (r *Runner) loadLangPack() error { + // The observability checkers compile from the pair pack (DETECTION patterns + tables, may be nil) AND the + // embedded target data (target-general lists, keyed by target lang). Built here so it exists even for a + // no-langpack book (a nil pack → inert pair checkers; the target lists still load). r.pack is set below. + defer func() { r.checkers = compileCheckers(dcCheckerData(r.pack), lang.TargetChecksFor(r.Book.TargetLang)) }() if r.Book.LangpackRoot == "" { return nil // no langpack declared → nil pack, miner inert } @@ -199,7 +208,9 @@ func (r *Runner) loadLangPack() error { // No catalog for this pair → nil-and-run (a ja book against a zh-only root just runs without mining). return nil } - pack, err := lang.Load(r.Book.LangpackRoot, r.Book.SourceLang, r.Book.TargetLang) + // A book-scoped overlay (LangpackExtend) unions the book's PRIVATE canon onto the shared pair pack + // (pair-14 §1: 古月 is a 蛊真人 clan, not a shared 百家姓 surname). "" ⇒ plain Load. Folds into Version(). + pack, err := lang.LoadWithOverlay(r.Book.LangpackRoot, r.Book.SourceLang, r.Book.TargetLang, r.Book.LangpackExtend) if err != nil { return fmt.Errorf("pipeline: load langpack for %s: %w", r.Book.LangPair(), err) } diff --git a/backend/internal/pipeline/runner_memory_test.go b/backend/internal/pipeline/runner_memory_test.go index 77013a58..0d1d4db8 100644 --- a/backend/internal/pipeline/runner_memory_test.go +++ b/backend/internal/pipeline/runner_memory_test.go @@ -316,7 +316,7 @@ func TestRunnerMemoryResnapshotOnApprovedChange(t *testing.T) { // ({{draft}}). The monolingual variant is PRESERVED as prompts/editor-mono.md (the D13.1 // confirming pilot arm) — it must NOT reference {{text}}. func TestEditorPromptIsBilingual(t *testing.T) { - raw, err := os.ReadFile(filepath.Join("..", "..", "prompts", "editor.md")) + raw, err := os.ReadFile(filepath.Join("..", "..", "prompts", "zh-ru", "editor.md")) if err != nil { t.Fatalf("read editor.md: %v", err) } @@ -326,7 +326,7 @@ func TestEditorPromptIsBilingual(t *testing.T) { if !strings.Contains(string(raw), "{{draft}}") { t.Error("editor.md must reference {{draft}} (it edits the draft)") } - mono, err := os.ReadFile(filepath.Join("..", "..", "prompts", "editor-mono.md")) + mono, err := os.ReadFile(filepath.Join("..", "..", "prompts", "zh-ru", "editor-mono.md")) if err != nil { t.Fatalf("read editor-mono.md: %v (D30.1: the monolingual variant must be preserved for the D13.1 pilot arm)", err) } diff --git a/backend/internal/pipeline/sanitizer.go b/backend/internal/pipeline/sanitizer.go index d4896a81..a8fe4c03 100644 --- a/backend/internal/pipeline/sanitizer.go +++ b/backend/internal/pipeline/sanitizer.go @@ -7,10 +7,39 @@ import ( "strings" "unicode" + "textmachine/backend/internal/lang" + "golang.org/x/text/unicode/norm" "golang.org/x/text/width" ) +// ruSanitizer holds the ru-target sanitizer DETECTION patterns as DATA (pair-14 data-out): the sanitizer is +// isRuTarget-gated (chunkrun/waverun), so its preamble / trailing-note / edit-meta / invalid-sign patterns +// are ru-target data (lang.TargetChecks), loaded once at package init. The strip/detect ALGORITHM stays in +// this file. The "ru" key mirrors the isRuTarget call-site gate; threading the target lang so a future non-ru +// sanitizer loads its own patterns is a follow-up (like sourceAbbrevs). +var ruSanitizer = lang.TargetChecksFor("ru") + +// compileSanitizerREs compiles an ORDERED list of sanitizer patterns by key (e.g. the 4 preamble shapes). A +// malformed pattern panics (a corrupt embed, caught by the sanitizer/golden tests). +func compileSanitizerREs(key string) []*regexp.Regexp { + vals := ruSanitizer.List(key) + out := make([]*regexp.Regexp, len(vals)) + for i, v := range vals { + out[i] = regexp.MustCompile(v) + } + return out +} + +// compileSanitizerRE compiles a single sanitizer pattern by key; the key must carry exactly one value. +func compileSanitizerRE(key string) *regexp.Regexp { + vals := ruSanitizer.List(key) + if len(vals) != 1 { + panic(fmt.Sprintf("pipeline: sanitizer pattern %q wants exactly 1 value, got %d", key, len(vals))) + } + return regexp.MustCompile(vals[0]) +} + // sanitizer.go: the output-sanitizer (D30.3) — a deterministic verdict-axis gate on the // FINAL chunk text that catches the "instant unreadability" defect classes NO existing // gate detects (exp12 root-cause + flagman §5). It is OPT-IN (Gates.Sanitizer.Enabled), @@ -206,18 +235,11 @@ func sanitizeOutput(text string) sanitizerResult { // вариант noun, or «перевод фрагмента/текста/…», or «ниже приведён … перевод/текст». // // Go RE2 \w/\b are ASCII-only, so Cyrillic uses [а-яё]. -var leadingPreamblePatterns = []*regexp.Regexp{ - // «Вот/Представляю/Привожу … отредактированный/исправленный перевод/текст/вариант …:» - // (the edit-adjective front-gate holds precision, so the tail-to-colon may be long — the - // gemini leak carries «… с соблюдением всех терминов из глоссария:»). - regexp.MustCompile(`(?is)^\s*(?:вот|представляю|привожу|держите)\s+[^\n:]{0,40}?(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:`), - // «Отредактированный/Исправленный перевод/текст/вариант …:» (label at the very start) - regexp.MustCompile(`(?is)^\s*(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:`), - // «(Вот) перевод фрагмента/текста/отрывка/главы/черновика:» - regexp.MustCompile(`(?is)^\s*(?:вот\s+)?перевод\s+(?:фрагмента|текста|отрывка|главы|черновика)\s*:`), - // «Ниже приведён/представлен … перевод/отредактированный текст:» - regexp.MustCompile(`(?is)^\s*ниже\s+(?:привед[её]н|представлен|дан|след)[а-яё]*\s+[^\n:]{0,40}?(?:перевод|отредактированн|текст)[а-яё]*[^\n:]{0,20}?:`), -} +// Pair-14 data-out: the actual ru patterns are DATA (lang.TargetChecks "sanitizer_preamble", ORDERED — the +// preview returns the FIRST match). The 4 shapes: (1) «Вот… отредактированный перевод…:» edit-adjective front- +// gate; (2) «Отредактированный перевод…:» label at the start; (3) «(Вот) перевод фрагмента/…:»; (4) «Ниже +// приведён … перевод/текст:». +var leadingPreamblePatterns = compileSanitizerREs("sanitizer_preamble") func detectLeadingPreamble(text string) string { t := strings.TrimSpace(text) @@ -237,7 +259,7 @@ func detectLeadingPreamble(text string) string { // «Сноска объясняла обычай.» are ordinary narrative nouns/gerunds — NOT headers). So the header // word MUST be followed by a COLON, or be «примечание/заметки/комментарий + переводчика/редактора», // or be the «(прим. перев./ред.)» marker. Only counted inside the trailing region (below). -var trailingNoteRE = regexp.MustCompile(`(?i)^[*_>\s-]{0,4}(?:(?:примечани[ея]|заметк[аи]|комментари[йяю]|пояснени[ея]|сноск[аи])\s*:|(?:примечани[ея]|заметк[аи]|комментари[йяю])\s+(?:переводчик|редактор)[а-яё]*|прим\.\s*(?:перев|ред)\.?)`) +var trailingNoteRE = compileSanitizerRE("sanitizer_trailing_note") // editMetaAnywhereRE matches UNAMBIGUOUS editor meta-commentary that is a defect wherever it // appears (not only trailing): «Внесённые правки:», «Что изменено:», «Список правок:», and the @@ -248,7 +270,7 @@ var trailingNoteRE = regexp.MustCompile(`(?i)^[*_>\s-]{0,4}(?:(?:примеча // are common narrative openers), keeping only editor-specific compounds that never occur in prose. // «основн…правк» is gated to «… для справки» OR a colon so a bare narrative «Основные правки внёс // редактор» (no summary label) does not fire. -var editMetaAnywhereRE = regexp.MustCompile(`(?im)^[*_>\s-]{0,4}(?:(?:внесённ|внесен)[а-яё]+\s+правк|основн[а-яё]+\s+правк[а-яё]*(?:\s+для\s+справк[а-яё]+|\s*:)|что\s+(?:было\s+)?(?:изменен|исправлен)|список\s+(?:правок|изменений)|список\s+внесённых)`) +var editMetaAnywhereRE = compileSanitizerRE("sanitizer_edit_meta") func detectTrailingNote(text string) string { if loc := editMetaAnywhereRE.FindStringIndex(text); loc != nil { @@ -359,7 +381,7 @@ func hasUpperLatin(s string) bool { // metalinguistic mention of the letter itself, «Буква ъ называлась ером.» (adversarial // review nit). High precision otherwise: word-initial-sign-then-letter and doubled signs are // impossible in well-formed text. \P{L} (non-letter), not RE2's ASCII-only \b. -var invalidSignRE = regexp.MustCompile(`(?i)(?:^|\P{L})[ъь][а-яё]|[ъь][ъь]`) +var invalidSignRE = compileSanitizerRE("sanitizer_invalid_sign") func detectBrokenWords(text string) (int, []string) { var det []string diff --git a/backend/internal/pipeline/testdata/langpack-overlay-guzhenren/zh/surnames-compound.txt b/backend/internal/pipeline/testdata/langpack-overlay-guzhenren/zh/surnames-compound.txt new file mode 100644 index 00000000..f034936d --- /dev/null +++ b/backend/internal/pipeline/testdata/langpack-overlay-guzhenren/zh/surnames-compound.txt @@ -0,0 +1,4 @@ +# 蛊真人 (Reverend Insanity) BOOK-SCOPED langpack overlay (pair-14 §1). 古月 is the protagonist's CLAN, +# not a 百家姓 compound surname — it lives here, unioned onto the shared pack FOR THIS BOOK, so the +# shared zh langpack stays a clean pair layer while the miner still anchors 古月方源 for 蛊真人. +古月 diff --git a/backend/internal/pipeline/waverun.go b/backend/internal/pipeline/waverun.go index 5bce1a66..737a0d41 100644 --- a/backend/internal/pipeline/waverun.go +++ b/backend/internal/pipeline/waverun.go @@ -8,6 +8,7 @@ import ( "sync" "textmachine/backend/internal/config" + "textmachine/backend/internal/lang" "textmachine/backend/internal/store" ) @@ -331,8 +332,9 @@ func (r *Runner) runDraftChunk(ctx context.Context, draftSnapshot string, ch Chu // Serialize the per-role injection blocks via the role→renderer registry (D39 layer 7) from the // precomputed base-bank selection: the translator gets its src→dst glossary block; any other // role's block is rendered too but consumed only if a stage of that role runs in this wave. + tx := lang.InjectionTextsFor(r.Book.TargetLang) for role, render := range roleInjectionRenderers { - injectionByRole[role] = render(memSel.injected) // pure → map-order-independent + injectionByRole[role] = render(memSel.injected, tx) // pure → map-order-independent } if len(memSel.trustGated) > 0 { r.Log.WarnContext(ctx, "memory: lower-trust longer key refused from suppressing a higher-trust nested key (approved term preserved; reconcile the seed)", @@ -442,8 +444,9 @@ func (r *Runner) runEditUnit(ctx context.Context, editSnapshot string, unit edit // Render the per-role injection from the ENRICHED unit selection via the registry (D39 layer 7): the // editor gets its CONFIRMED-dst constraint block; other roles' blocks are rendered but consumed // only if a stage of that role runs in the edit wave. + tx := lang.InjectionTextsFor(r.Book.TargetLang) for role, render := range roleInjectionRenderers { - injectionByRole[role] = render(editSel.injected) + injectionByRole[role] = render(editSel.injected, tx) } } seq, err := r.runStageSequence(ctx, editStages, editSnapshot, leader, unitDraft, injectionByRole) diff --git a/docs/PROGRESS.md b/docs/PROGRESS.md index 28b05ad8..5c0d49ab 100644 --- a/docs/PROGRESS.md +++ b/docs/PROGRESS.md @@ -9,6 +9,10 @@ > - **Ждём от владельца:** тачпойнты exp16 (карта подписи/пол · мини-голд алиасов · precision@30, `books/gu-zhenren/exp16/`, ~30–40 мин — для сид-дельты к пере-прогону) · развилки плана §10 (W1.5-UX · DC5-стих-политика · Edit-ceiling · и др.) · реплика ja→ru (тест общности §B5) · **ре-чек прайса DeepSeek у катовера слагов 24.07** (до него платных deepseek-прогонов нет) · чтение пере-прогона (после R1+resnapshot) · старые висящие: планка запуска (лучше-фана/гибрид/издательский) · билингв-якорь пилота (D25.9-Q1) · контаминация пилот-корпуса (D27.4) · publishable/waiver (D25.1) · FN-bound L3 (D25.4) · юр-пакет · провенанс 12-*-доков. > - Архивы хроники: `archive/PROGRESS-2026-07-04-10.md` (D31) · `archive/PROGRESS-2026-07-10-13.md` (D39.6-гигиена). Записи ниже — живая эра D39. +## Оркестратор №7 — ПАК-14 «генеральность-пасс» ПРИНЯТ (воркфлоу 11 линз + доприёмка 4) и залендён; пак-15 «структура+форма» готовится, 24.07 + +Сессия исполнила 9 пунктов + оба аддендума + **расширение по директиве владельца «full data-out»** (детект-регексы чекеров/гейтов/санитайзера → данные, не только lookup-таблицы; `checkers_zh_ru.go` → пар-агностичный `checkers.go`). Два дома данных: `configs/langpacks/`+книго-overlay (`book.yaml: langpack_extend`, аддитивный union, fail-loud на чужой файл) для пар/книго-данных · `internal/lang/data/` go:embed для того, что нужно движку БЕЗ пака (golden ja→ru без langpack это доказал: refusal/CJK-нарезка/→ru-инъекция). **Приёмка исполнением:** 0 блокеров; байт-сверка каждого переноса против HEAD; парити EXACT (`{方源:0 蛊:1 蛊师:2 古月:22}`, 古月 теперь из base∪overlay — общий zh-пак чист); golden байт-нетронут; тесты грузят фикстуры из реального пака. Блокер доприёмки (аддендум parsePalladius не был исполнен) закрыт дельтой: generic-парсер категорий + один типизированный `Pack.Palladius`, required валидирует потребитель; + c2-segmentation явно, packAlgoVersion v2. **Поправка отчёту (ревью-шапка):** Приложение A «distinct/matches 8» — греп-артефакт (0 кодовых ссылок; 35→32/14), хойст-вывод не задет. **Гигиена-инцидент мой:** `77dd9b8` утащил staged `git mv` промптов сессии (git commit коммитит весь индекс) — одна неконсистентная bisect-точка, вылечена этим лендингом; правило: перед лендинг-коммитом `git diff --cached`. Прод-overlay 古月 подключён в оба стендовых book.yaml. **Пак-15 (готовится):** хойст `internal/text`+value-типов → сплит miner/checks/membank/chunk (карта связности+экспорт-поверхности = приложения A/B отчёта пака-14, с моей поправкой) + форм-фиксы Export/RequestHash/resolveChunkState + конфиг-слой пары (пар-калибровки из ран-конфигов, промпт-резолв конвенцией) + перенесённый остаток (тесты Палладия · fail-loud оверлея · unionStringMap-семантика · empty-probe гард · терминаторы chunker/coverage · `第`). + ## Оркестратор №7 — research/22 (доменные харнессы) ПРИНЯТ и залендён; вердикт: калибрующий, не переворачивающий, 24.07 Сессия сдала ревизию-2 (после СОБСТВЕННОГО 9-агентного адверсариала синтеза, поймавшего её факт-ошибки — прецедент мандата 12.07 в лучшей форме). Мой спот-чек 5/5 (detectCJKLeak/cheapgate-v4/ledger наши · ProofreadTask/GenDic их). **Оценка:** ядровые ставки ПОДТВЕРЖДЕНЫ рынком+академией (fused-майнинг>bolt-on доказан смертью KeywordGacha; DelTA-роадмап = наш пройденный путь; TransAgents назвали нашу D39.20-проблему и не решили; north-star «измеренная редактура» НЕ роадмапится никем — окно открыто). **Мы безоговорочно впереди:** COGS-телеметрия · 18+ · общность · измеренность (recall 0.932, вклад ролей). **Позади:** Q4 выпуск/epub (разрыв РАСТЁТ — bbm/Immersive активно куют) · gate+repair-петля QA (они чинят typed-error→адресный ре-перевод, мы флагаем). **В планы:** слой-3 diff-редактор получает прецедент AiNiee/LinguaGacha (приоритет↑, пак-15/16) · epub-tag-rewrite не держать до Ф3 бесконечно (кандидат после пилота) · Q7-леджер дополнить (глагольный вид/addition/voice-flattening) · Problem.py-классы reflow-aware (char-freq, control-code) · LinguaGacha = бенчмарк-кандидат №1 · SakuraLLM мониторить (канал B/ja→ru). «D39.7-коллизия» разрешена: не коллизия, ссылка на стандарт. Промт → архив. diff --git a/docs/README.md b/docs/README.md index 73090dc9..afb0c40e 100644 --- a/docs/README.md +++ b/docs/README.md @@ -11,17 +11,17 @@ ## Структура -- `architecture/` — синтез. **Источник истины по решениям — [`05-decisions-log.md`](architecture/05-decisions-log.md) (D1–D39.16); при конфликте с любым доком он выше.** +- `architecture/` — синтез. **Источник истины по решениям — [`05-decisions-log.md`](architecture/05-decisions-log.md) (D1–D39.22+); при конфликте с любым доком он выше.** - `01-decisions.md` — принципы Р1–Р10; `02-mvp-plan.md` — фазы и приёмка (v3, 09.07); `03-implementation-notes.md` — контракты Фазы 0; `04-unhappy-paths.md` — ~70 режимов отказа → механизм; `06-memory-risk-registry.md` — реестр рисков банка памяти; **[`07-strategic-review.md`](architecture/07-strategic-review.md) — стратегический аудит (09.07): вердикт end-to-end, топ-риски, курс-коррекции**; **[`08-sync-audit-ledger.md`](architecture/08-sync-audit-ledger.md) — верифицированный ледджер синк-аудита «сломанного телефона» (65 находок, D39)** · **[`09-target-architecture.md`](architecture/09-target-architecture.md) — целевая 7-слойная архитектура + фазовый план арх-ресета (D39; инвариант общности §0.1)** · [`10-prompt-architecture.md`](architecture/10-prompt-architecture.md) — консолидированная промпт-заметка (концерн 4); **[`11-implementation-plan.md`](architecture/11-implementation-plan.md) — РАТИФИЦИРОВАННЫЙ план стройки пере-прогонного стека (D39.12; дизайн-оф-рекорд бэкенд-пака, три рубежа верификации)**; `components.puml`/`pipeline.puml` — диаграммы v3 (владелец смотрит PlantUML-расширением VS Code; вручную НЕ рендерить). - `experiments/` — эмпирика «Полигона»: `00-provider-quirks` (читать перед любым вызовом провайдера), `01-token-calibration`, `02-refusal-benchmark`, `03-local-stand`, `04-editor-quality`, `06-local-extraction`, `07-coverage-precision`, `08-cost-model-v2` (актуальная денежная модель), `09-pilot-protocol` (пилот Ф2.5 + поправки D13), `10-explicit-benchmark` (18+ violence-рука канала B), `11-erotica-benchmark` (erotica по трём парам/регистрам — закрытие D14.4, D22). - `research/` — фактура исследований 04–05.07: `01–10` базовые, `11-gap-*` добор критиком, `12-*` режимы отказа/отзывы/таксономии (+ два внешних материала с провенанс-шапками), `13` валидация памяти, `14` адаптивная память, `15` голос и состояние (принят, D21), **`16` ридер-IDE (принят с ревью-шапкой, D29)**, **`17` внешняя критика GPT-5.6 (принят с ревью-шапкой, D25)** — у 16/17 читать шапку прежде тела. **`18` рычаги качества (два отчёта, D36)** · **`19` нарезка+когезия+контракт t/e (D39.1)** · **`20` банк-майнинг W1.5 (D39.6)** · **`21` обзор LLM-транспорта чужих харнессов (23.07: наш транспорт опережает/вровень со всеми 11)** · **`22` доменные харнессы перевода (24.07: калибрующий — ядро подтверждено, впереди COGS/18+/общность/измеренность, позади Q4-выпуск и gate+repair-петля; сиквел `02`/`05`)** — у всех ревью-шапки. ⚠ Часть под superseded-баннерами (01/02/03/04/05/09 и gap-1/2/5) — **читай баннер прежде содержимого**. - `PROGRESS.md` — **журнал** (CURRENT-STATE сверху, ниже хронология; НЕ источник решений). -- **Активные хендофф-промты (пост-паки-12/13, 24.07):** [`BACKEND_GENERALITY_PASS_SESSION_PROMPT.md`](BACKEND_GENERALITY_PASS_SESSION_PROMPT.md) (**пак-14 генеральность-пасс — ТЕКУЩИЙ**: подтверждённые утечки пара/книга-данных из Go в langpack/сид; байт-точность+парити EXACT; норматив — [`architecture/12-go-style-notes.md`](architecture/12-go-style-notes.md)) · [`ORCHESTRATOR_SESSION_PROMPT.md`](ORCHESTRATOR_SESSION_PROMPT.md) (хендофф №5) · [`POLYGON_PACKAGE4_SESSION_PROMPT.md`](POLYGON_PACKAGE4_SESSION_PROMPT.md) (residual пилот/18+/echo). **Паки 12/13 ЗАЛЕНДЕНЫ 24.07** (`7fe0a5b`/`2f91b04`, приёмка 25-агентным воркфлоу; их промты + research/22-промт → архив после ревью research/22). Дальше: пак-15 = канал B вживую + хвосты слоя 7 → пак-16 = слой 3 diff-based редактор → голос/состояние D21 + C2/C3. Закрытые — в `archive/prompts/`. +- **Активные хендофф-промты (пост-пак-14, 24.07):** [`ORCHESTRATOR_SESSION_PROMPT.md`](ORCHESTRATOR_SESSION_PROMPT.md) (хендофф №5) · [`POLYGON_PACKAGE4_SESSION_PROMPT.md`](POLYGON_PACKAGE4_SESSION_PROMPT.md) (residual пилот/18+/echo). **Паки 12/13/14 ЗАЛЕНДЕНЫ 24.07** (`7fe0a5b`/`2f91b04`/пак-14 — приёмка воркфлоу-линзами; отчёт пака-14 с ревью-шапкой — [`archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md`](archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md); норматив общности — [`architecture/12-go-style-notes.md`](architecture/12-go-style-notes.md)). Дальше: **пак-15 «структура+форма»** (хойст text/model-базы → сплит miner/checks/membank/chunk + Export/RequestHash/resolveChunkState + конфиг-слой пары; промт готовится) → пак-16 = канал B вживую / слой 3 diff-based редактор (прецедент research/22) → голос/состояние D21 + C2/C3. Закрытые — в `archive/prompts/`. - `archive/` — закрытые сессионные промты (только история, инструкции оттуда не исполнять). ## Статус (2026-07-23, пост-D39.20 — механизация выпускного качества) -Фаза 0 ✅; Ф1-инфра ✅ (D20–D28). **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ:** 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1–D39.10) → план (D39.12, `11-*`) → стройка пака-11 (D39.13–D39.19: волновой исполнитель · чанкер output-бюджет · src→dst-редактор · Go-майнер паритет-EXACT · банкнота · `internal/lang`+langpacks · reject-set · echo-сплит). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07):** операционка вся подтверждена живьём (ОДИН resnapshot · **$0-резюм драфта ×3 арма** · echo 0 · $6.63<$15 · майнинг-стоп→подпись→reject-set); **планка ≤2 НЕ пройдена** — но дефект-классы полностью механизируемы. Три сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232, floor-шум 0 · два слепых чтения): **анти-корреляция гладкость↔верность доказана трижды**; **mistral исключён как несущий редактор** (голос = north-star промпта); dspro-vs-glm → итерация №2 после механизации (драфт $0). research/21: транспорт наш опережает/вровень со всеми 11, рефакторинг отвергнут. **Эксп ЗАКРЫТ (D39.22), итерация №2 отложена; редактор = deepseek-v4-pro ИНТЕРИМ (glm-резерв); стайл-каноны D39.21 ратифицированы.** **Очередь (курс — разработка бэкенда):** пак-12 транспорт-гигиена (лендится первым) → пак-13 «выпускной QA» (wire-двигающий; вкл. дефолт-редактор dspro) → сид-дельта (元/赤城/学堂家老/春秋蝉) → масштаб целой книги + канал B вживую + Ф2-механизмы (D21 голос/состояние, native-Gemini судья) + ja→ru §B5 → пилот Ф2.5 (блокеры прежние: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Судья 18+ = grok-фолбэк (квирк D22.6 подтверждён живьём). Ключ xAI: единый, data-sharing, off перед продом (D27). +Фаза 0 ✅; Ф1-инфра ✅ (D20–D28). **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ:** 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1–D39.10) → план (D39.12, `11-*`) → стройка пака-11 (D39.13–D39.19: волновой исполнитель · чанкер output-бюджет · src→dst-редактор · Go-майнер паритет-EXACT · банкнота · `internal/lang`+langpacks · reject-set · echo-сплит). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07):** операционка вся подтверждена живьём (ОДИН resnapshot · **$0-резюм драфта ×3 арма** · echo 0 · $6.63<$15 · майнинг-стоп→подпись→reject-set); **планка ≤2 НЕ пройдена** — но дефект-классы полностью механизируемы. Три сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232, floor-шум 0 · два слепых чтения): **анти-корреляция гладкость↔верность доказана трижды**; **mistral исключён как несущий редактор** (голос = north-star промпта); dspro-vs-glm → итерация №2 после механизации (драфт $0). research/21: транспорт наш опережает/вровень со всеми 11, рефакторинг отвергнут. **Эксп ЗАКРЫТ (D39.22), итерация №2 отложена; редактор = deepseek-v4-pro ИНТЕРИМ (glm-резерв); стайл-каноны D39.21 ратифицированы.** **Очередь (курс — разработка бэкенда):** паки 12/13/14 залендены (транспорт-гигиена · выпускной QA · генеральность: данные→langpack/embedded, книго-overlay 古月, пар-агностичные чекеры) → пак-15 «структура+форма» → сид-дельта (元/赤城/学堂家老/春秋蝉) → канал B вживую + слой-3 diff-редактор + Ф2-механизмы (D21 голос/состояние, native-Gemini судья) + ja→ru §B5 → пилот Ф2.5 (блокеры прежние: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Судья 18+ = grok-фолбэк (квирк D22.6 подтверждён живьём). Ключ xAI: единый, data-sharing, off перед продом (D27). ## Доступные ключи от моделей DEEPSEEK_API_KEY, ZAI_API_KEY, KIMI_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, XAI_API_KEY, MISTRAL_API_KEY diff --git a/docs/BACKEND_GENERALITY_PASS_SESSION_PROMPT.md b/docs/archive/prompts/BACKEND_GENERALITY_PASS_SESSION_PROMPT_2026-07-24.md similarity index 100% rename from docs/BACKEND_GENERALITY_PASS_SESSION_PROMPT.md rename to docs/archive/prompts/BACKEND_GENERALITY_PASS_SESSION_PROMPT_2026-07-24.md diff --git a/docs/BACKEND_PACK13_QA_SESSION_PROMPT.md b/docs/archive/prompts/BACKEND_PACK13_QA_SESSION_PROMPT_2026-07-24.md similarity index 100% rename from docs/BACKEND_PACK13_QA_SESSION_PROMPT.md rename to docs/archive/prompts/BACKEND_PACK13_QA_SESSION_PROMPT_2026-07-24.md diff --git a/docs/BACKEND_TRANSPORT_PACK_SESSION_PROMPT.md b/docs/archive/prompts/BACKEND_TRANSPORT_PACK_SESSION_PROMPT_2026-07-24.md similarity index 100% rename from docs/BACKEND_TRANSPORT_PACK_SESSION_PROMPT.md rename to docs/archive/prompts/BACKEND_TRANSPORT_PACK_SESSION_PROMPT_2026-07-24.md diff --git a/docs/archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md b/docs/archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md new file mode 100644 index 00000000..d20075e1 --- /dev/null +++ b/docs/archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md @@ -0,0 +1,253 @@ +# PACK-14 «генеральность-пасс» — приёмочный отчёт (2026-07-24) + +> **Ревью-шапка оркестратора №7 (24.07, приёмка исполнением, два воркфлоу: 11 линз ~970k ток. + доприёмка дельты 4 линзы; 0 блокеров).** Байт-точность КАЖДОГО переноса независимо сверена против git HEAD (refusal 12/12 · injection 6/6 с ведущими пробелами · dc-регексы · санитайзер · CJK-числительные · фонотактика · форманты — ноль расхождений); парити EXACT свежими `-count=1` прогонами; golden байт-нетронут; тесты не ослаблены (2 удалённых ассерта — законные замены). **Поправки к телу отчёта:** (а) Приложение A — ведро «distinct/matches (8 refs)» = греп-артефакт: все 8 в комментариях, кодовых ссылок 0; заголовок «miner→membank 35» реально 32 полных вхождений / 14 кодовых; остальные вёдра точны, вывод «хойст первым» НЕ задет; (б) ёфикатор = 38 слов (строка «40 ru-слов» — обсчёт в прозе); (в) свип-47 воспроизводится как 41 (точный класс) + 6 строк ё/кана (класс сессии был шире заявленного регекса), скрытых пар-данных нет; класс «пунктуация-сепараторы» в 47 не представлен (перенос из классификации-100). **Дельта доприёмки принята** (generic-Палладий по аддендуму, c2-segmentation, README-пути, want-list magnitude, packAlgoVersion v2). **Остаток в пак-15 (зафиксирован):** тесты unknown-category-толерантности и required-fail Палладия · fail-loud оверлея уровнем выше (пропавший overlayRoot/чужой подкаталог/категория-опечатка в overlay-файле — тихие) · семантика override unionStringMap · isRuTarget/'ru'-швы · empty-probe гард lintMagnitudeScale · дуп терминаторов chunker/coverage · `第`. Прод-оверлей 古月 подключён оркестратором в оба стендовых book.yaml (rerun2). + +Бэкенд-сессия. Источник: `docs/BACKEND_GENERALITY_PASS_SESSION_PROMPT.md` (9 пунктов) + два аддендума +оркестратора (A: ingest `章节節回`+CJK-числительные → та же ед. точки истины, что у чанкера; B: форманты +`等/转/阶`). Норматив: `docs/architecture/12-go-style-notes.md` §0. **Перенос ДАННЫХ, не редизайн алгоритмов.** +**Сессия НЕ коммитит.** + +## Вердикт: ✅ жёсткий инвариант выполнен + +- **golden БАЙТ-ИДЕНТИЧЕН** — `testdata/golden/capture.golden` не тронут (`git diff` пуст). Сильнее допущенного + «version-only»: у golden-книги нет langpack (ja→ru, `langpack_root` не задан → `pack==nil`), поэтому + `packVersion()==""` и её снапшот НЕ двигается ни от одного НОВОГО langpack-файла; `TestGolden` сверяет + capture побайтно и зелёный. +- **майнер-парити EXACT** — `TM_MINER_PARITY=1`: `n=13618 catastrophe{方源:0 蛊:1 蛊师:2 古月:22} + recall@proposed=0.9655` — тождественно дореформенному (вкл. 古月-механику через книго-расширение). +- **`go build/vet/test -race` зелёные** (весь модуль, вкл. новый `internal/lang`). + +## Ключевой факт, определивший архитектуру + +golden — ja→ru книга БЕЗ langpack (`pack==nil`), но она живьём: (1) флагает `hard_refusal` (RU «Не могу +помочь»), (2) инъектит русские заголовки глоссария в wire (16×), (3) режет главы `第X章`, (4) style_flags=0. +Значит рефузал-паттерны, CJK-числительные глав и →ru-инъекцию НЕЛЬЗЯ держать в книжном паке (nil для golden). +`request_hash` фолдит `snapshotID` (render.go) — любое движение снапшота golden-книги сдвинуло бы wire. → + +## Два дома данных (правило, выдержавшее «100 языков» линзу владельца) + +1. **`configs/langpacks/|/` + книжный overlay** — пер-исходник / пер-пара / пер-книга данные. + Добавить язык = **положить каталог, без перекомпиляции**; убрать = удалить каталог; загрузчик **падает + ГРОМКО, называя файл** (никогда тихо-пусто). Сюда: item 1 (surnames/古月), 5 (терминаторы, Палладий- + фонотактика), 6 (DC-таблицы), addB (форманты). Overlay: добавить/убрать КНИГУ = ±один каталог; + DC-таблицы OPT-IN (убрать файл → чекеры инертны, не болтаются). +2. **`internal/lang/data/` (go:embed, секционировано по языку/таргету)** — данные, нужные движку БЕЗ книжного + пака (golden доказал: рефузал, CJK-нарезка, →ru-инъекция — все без langpack). Добавить язык = добавить + секцию в один data-файл; убрать = удалить секцию. Перекомпиляция — приемлемо: это инженерные КОНСТАНТЫ + (версионируются с бинарём, как код), а байт-точность no-pack golden запрещает фолдить их в книжный пак. + Сюда: item 3 (refusal.txt), 4+addA (cjk-section.txt), 2 (injection.txt), sourceAbbrevs (sentence-abbrev.txt). + +## По пунктам (что сделано, чем доказана нейтральность) + +**1. 古月 ВОН из общего zh-langpack.** Удалён из `configs/langpacks/zh/surnames-compound.txt`. Механизм — +книго-скоупный overlay: новое поле `book.yaml: langpack_extend` (каталог той же раскладки), `lang.LoadWithOverlay` +делает АДДИТИВНЫЙ union (множества дополняются, слайсы append; overlay ничего не удаляет), байты overlay +фолдятся в `Version()` (правка книжного канона = громкий --resnapshot ТОЛЬКО этой книги). Парити: `testLangPack` +грузит `testdata/langpack-overlay-guzhenren` → эффективный пак = старый (base ∪ {古月}) → {古月:22} EXACT. Стенд- +overlay создан `/home/ubuntu/books/gu-zhenren/langpack-extend/zh/surnames-compound.txt`. Прод-book.yaml — +см. «Хвосты владельцу». + +**2. Инъекция-тексты → таргет-данные, с гейтом.** `glossaryBlockHeader`/`editorConstraintHeader`/маркер +`⟨проверить⟩`/3 gender-ноты вынесены в `internal/lang/data/injection.txt` (target-keyed, `lang.InjectionTextsFor`), +рендереры гейтятся на `tx.HasData()` — →en книга получит ПУСТО, не русский блок (это и есть фикс утечки). +Значимые ведущие пробелы (` (муж.…`, ` ⟨проверить⟩`) сохранены дословно (парсер НЕ триммит value). `renderFormatVersion` +НЕ тронут → снапшот golden не двигается → wire байт-идентичен (проверено golden). Langpack-ФАЙЛ (а не embedded) +отложен: golden без пака потребовал бы target-loader вне book-langpack_root, что сдвинуло бы снапшот golden — +запрещено инвариантом. + +**3. Рефузал-блэклист → универсальные embedded-данные.** `internal/lang/data/refusal.txt` (секции en/ru/ja/zh ++ маркер), `disposition.go` строит `refusalRE` из `lang.RefusalPatterns()`. Байт-в-байт 12 паттернов в том же +порядке (сверено скриптом). golden флагает `hard_refusal` тождественно. + +**4 + addendum-A. CJK-числительные — ОДНА точка истины.** `internal/lang/data/cjk-section.txt` (digit/unit +значения, zero, chapter_unit) — читают И чанкер (`isHeadingNumeral`/`cjkSectionDigit`/`cjkSectionUnit`, путь +заголовков, nil для golden), И ingest (`chapterNumeralRE` строится из `HeadingNumeralClass()`, `detectChapterUnit` +итерит `ChapterUnitOrdered` — авторский порядок сохранён для детерминированного тай-брейка). Устранён +ingest↔chunker байт-дубль. golden режет `第X章` тождественно (ingest живёт для golden — верифицировано TestGolden). +`sourceAbbrevs` → `internal/lang/data/sentence-abbrev.txt` (секции по ИСХОДНИКУ, en=29 токенов); чанкер потребляет +секцию `en` (как и раньше — универсально; исходник-параметризация сплиттера = мелкий отложенный шаг, чтобы не +трогать 18 колл-сайтов SplitChunks на не-приёмочном пути). + +**5. Терминаторы `。!?` + Палладий-фонотактика.** `zh/sentence-terminator.txt` (`p.SentenceTerminator`, +`splitMinerSentences`); фонотактика jqx/ретрофлексы/ü-финалы → `zh-ru/palladius-phonotactics.txt` +(`PalladiusRetroflex/VFinal/VFinalInitial`), l/n-исключение стало данными `vfinal_initial`. Майнер-only, +парити EXACT. + +**6. DC-таблицы → пар-данные, регексы = алгоритм.** `dcCNNum`+`dcRuHourWord`+`dcRegisterNegList` → OPTIONAL +`zh-ru/dc-checkers.txt` (`lang.DCCheckerData`), прокинуто через `cheapGateConfig` (nil пак → пусто → чекеры +0 → golden style_flags=0). ДЕТЕКТ-регексы (`时辰`/`千万`/`成`/`терем`-паттерны) остались как пар-скоупный +алгоритм чекера (§12.2 — «данные наружу, алгоритм остаётся»). + +**7. Калибровка (config/pipeline.go).** Только пере-документирование (числа НЕ тронуты): дефолты 1797/3200/ +1.1978/0.3852 помечены как «zh-ru пар-калибровка = последний generic-fallback», шиппинг-конфиги (c1/армы) ставят +блок ЯВНО (пар-конфиг = источник истины). Релокация в langpack ОТВЕРГНУТА (была бы мёртвой для всех живых путей: +конфиги ставят явно, у golden нет langpack; least-mechanism §12.1). Числа держатся EXACT: no-langpack книга +(golden) чанкается на этом fallback — пере-выведенное число сдвинуло бы её границы→wire. + +**addendum-B. Форманты `等/转/阶`.** `zh/title-formant.txt` (`p.TitleFormant`, `formantType`). Рядом с TopoSuffix +(сосед уже data-driven). Майнер-only, парити EXACT. + +**9. Промпты → `prompts/zh-ru/`.** `backend/prompts/{translator,editor,editor-mono,judge-selector}.md` → +`prompts/zh-ru/` (git mv, БАЙТЫ не тронуты — снапшот фолдит `PromptSHA256` от СОДЕРЖИМОГО, не пути). Обновлены 5 +шиппинг-конфигов (`../prompts/zh-ru/…`) + 4 тест-ссылки на реальные файлы. Тесты, грузящие реальные c1/арм-конфиги +(prompt_pack_test, echo_mine_test), зелёные. + +## Item 8 — свип-добор (полный grep non-test `.go`, построчно) + +Метод: изолирован **non-comment** grep `[\x{4e00}-\x{9fff}]|[а-яА-Я]` (строчные комментарии срезаны). Было 133 +дата-литерала → **перенесено 33** (items 1–6+addA/B) → **осталось 100**, классификация: + +- **SQL-DDL комментарии (migrate.go, 26):** русская документация ВНУТРИ raw-string схемы. Не пар-данные, 0 + поведения. OK / вне скоупа (это перевод комментариев, отдельная забота). +- **Диагностика/сообщения-флагов (~11):** `book.go:188`(Р7), `models.go:184`(Р4)/`231`, `pipeline.go:411`(Р2), + `checkers_zh_ru.go:63/119/124/200`, `cheapgates.go:451/294/319`, `checkers:285` — operator-диагностика / + observability-строки (англ.-первичные, цитируют матч для человека). НЕ wire, НЕ пар-данные. OK. +- **Пар-скоупный ДЕТЕКТ-алгоритм (регексы), §12.2 (данные уже вынесены, регекс остаётся):** + `checkers_zh_ru.go:34/38/101/102/103/117/118/122/175/178/187/283` (DC1/DC2/percent/broken-word детект); + `sanitizer.go:213/215/217/219/240/251/362` (ru-таргет editorial-preamble/note/invalid-sign детект). OK. +- **Алгоритм-инвариантные литералы (OK-generic):** `chunker.go:222` (класс СЕПАРАТОРОВ пунктуации :、,。…); + `memnorm.go:180-181` (ё→е фолд ключа памяти, 1-символьное правило); `memseed.go:570` (ー/・ в кана-ридингах, + ja-норм); `miner_palladius.go:71` (ъ-drop перед сегментацией). `ingest.go:128` — мой новый регекс, `第` + остался литералом (кандидат — см. ниже). +- **Ратифицированный DEFER (НЕ тронут):** `cheapgates.go:361` translit-interj блоклист (ара-ара/маа/…) — + D39.16 явно (ja-филлеры, мис-кей-ловушка). ✅ оставлен. + +### Кандидаты СЛЕДУЮЩЕГО пака (тот же класс, НЕ в подтверждённом фикс-листе — решение оркестратора) +Одного класса с перенесёнными, но не названы в 9 пунктах/аддендумах; байт-нейтральны (golden фаерит 0), но +переносить их сейчас = расширять скоуп сверх санкционированного + повтор D39.16-концерна «живой пара-агностик». +Флагаю для след. пака: +- **`cheapgates.go:265-276` yofikator homograph-whitelist (40 ru-слов)** — ru-таргет readability-данные (класс + DC6). Нужен ru-TARGET langpack-слот (сейчас нет). +- **`cheapgates.go:459-619` 万/億 magnitude-gate** — zh-исходник числ-значения (万萬億亿兆 — ДРУГОЙ инвентарь, не + section) + ru-стемы (тысяч/миллион/…). Кандидат: расширить `lang.CJKSection` категорией `magnitude` + ru-слот. +- **`sanitizer.go` ru-фрагменты** — data-экстракция лексем-детекторов в ru-таргет-слот (регексы = алгоритм). +- **`ingest.go:128` `第`** — маркер главы; можно добавить в `cjk-section.txt` рядом с chapter_unit (1 руна). + +## «100 языков» линза (ответ на вопрос владельца) + +- **Добавить пару:** положить `configs/langpacks//`(11 required src-файлов)+`-/`(2 pair + опц. + heading/dc-checkers) + book.yaml. НИ строки под `internal/pipeline`. Загрузчик fail-loud называет отсутствующий + файл. Смелл: `title-formant`/`surnames-*` осмысленны только для CJK-исходника (весь минер CJK-центричен — + ПРЕ-существующее свойство, не ухудшено; для не-CJK исходника миннер сам подлежит переосмыслению). +- **Добавить таргет:** секция в `injection.txt`/`refusal.txt` — правка ДАННЫХ + перекомпиляция (go:embed). + Оправдано: инженерные универсалии/таргет-генерики. +- **Убрать язык:** удалить каталог пака / секцию embedded-файла — ничего не болтается (DC opt-in инертен; + overlay опционален; загрузчик громкий). Тесты не ломаются на удалении пары (fail-loud по требованию). + +## Хвосты владельцу (операционка, вне зоны backend/) + +- **Прод-`book.yaml` gu-zhenren** (стенд `rerun2/book.yaml`, `book-mistral.yaml`): добавить + `langpack_extend: /home/ubuntu/books/gu-zhenren/langpack-extend` — иначе прод-майнинг потеряет 古月 (overlay- + каталог уже создан на стенде). Не тронул сам: зона книжных данных владельца + закрытый эксперимент rerun2. + Любой ре-ран и так --resnapshot (langpack Version() сдвинут новыми файлами). +- **Стенд pipeline-конфиги** (`rerun2/pipeline-*.yaml`) и **docs**, ссылающиеся на старый путь `backend/prompts/…` + (истор. записи в archive/experiments) — обновить путь на `prompts/zh-ru/…` при следующем касании (зона + оркестратора/владельца). + +## Само-проверка (мандат 12.07) + адверсариальный 4-линзовый ревью + +Ревью ИСПОЛНЕНИЕМ: каждый item верифицирован golden+parity сразу после кода (не в конце). Линзы: байт-точность / +парити / общность-ja-ru / **масштаб-100-языков (доп. линза владельца)**. + +### Итог адверсариального 4-линзового ревью (воркфлоу, 4 независимых агента, 289k токенов) + +- **БАЙТ-ТОЧНОСТЬ: CLEAN** — не опровергнуто. Агент независимо байт-сверил КАЖДЫЙ перенесённый литерал против + git HEAD: injection (6 значений вкл. ведущие пробелы: unverified_marker len=13, gender len=31), refusal + (12 паттернов, joined_equal=True), cjk-section (digit/unit/zero MATCH, chapter_unit ПОРЯДОК [章 节 節 回], + heading-класс symmetric-diff пуст), sentence-abbrev (en=29 exact), dc-checkers (11+9+6 rows match), Палладий- + фонотактика/форманты/терминаторы MATCH, промпты = чистые rename (0 байт). golden capture пуст в git, TestGolden + байт-сверка PASS. +- **МАЙНЕР-ПАРИТИ: CLEAN** — не опровергнуто (свежий -count=1). Overlay-union доказан аддитивным (base−古月 ∪ + {古月} = старый 16-сет); l/n-исключение фонотактики логически тождественно старому switch. +- **ОБЩНОСТЬ: CONCERNS** — 1 MINOR + 3 NOTE (ниже, все либо пофикшены, либо документированы). +- **МАСШТАБ-100-ЯЗЫКОВ: CONCERNS** — 1 MAJOR (пофикшен) + 3 MINOR/NOTE (документированы). + +**Пофикшено по ревью (пост-ревью, верифицировано golden+parity):** +- **MAJOR (масштаб):** `LoadWithOverlay` ТИХО игнорировал overlay-файл вне манифеста (опечатка `surname-compound.txt` + или книжный `heading.txt` в overlay → приватный канон не доезжает до майнера, recall падает БЕЗ сигнала). Фикс: + `overlayDirFiles` сканирует overlay-каталоги и **падает громко**, называя неожиданный файл. + тест + (misnamed-overlay fails loud). Восстанавливает fail-loud-гарантию на масштабе десятков книг-overlay. +- **NOTE (масштаб):** `refusal.txt` — комментарий-заголовки `# --- xx ---` вводили в заблуждение («удали секцию = + отключи детект для пары»). На деле `RefusalPatterns()` читает плоско, ВСЕ паттерны фаерят для ВСЕХ книг. Комментарий + переписан: секции ОРГАНИЗАЦИОННЫЕ; удаление триммит УНИВЕРСАЛЬНЫЙ сет (только `#`-строки — байты паттернов, golden не сдвинут). +- **MINOR (масштаб):** header `srcFiles` — уточнён скоуп: «drop a directory» верно ВНУТРИ zh-family name-miner + морфо-схемы; другое семейство исходников требует новых Pack-полей (не просто каталог). + +**Документированные остатки (не фикшу — конфликт с инвариантом/мандатом/пре-существующее):** +- **MINOR (общность): калибровка = ТИХИЙ zh-ru дефолт** при пропуске блока (ломает fail-loud). НЕ фиксится: + item 7 предписал «оставить механизм», а golden (ja→ru) ПРОПУСКАЕТ блок и опирается на этот fallback — падение + громко на пропуске сломало бы golden. Тайтен возможен лишь когда golden сам явно проставит сегментацию (будущий + пак, который и так пере-снимет golden). +- **NOTE (общность+масштаб): `sourceAbbrevs` захардкожен на `en`** — исходник-параметризация сплиттера отложена (18 + колл-сайтов SplitChunks, не-приёмочный путь). Секция не-en = мёртвые данные до треда `sourceLang`. +- **NOTE (общность+масштаб): DC2 `lintMagnitudeScale` держит НЕЗАВИСИМЫЙ CJK-числ-парсер** (万/億/兆), дублирующий + `lang.CJKSection` — пре-существующий, self-gating (инертен на не-zh), не тронут этим диффом. Кандидат след. пака + (в списке item-8). +- **NOTE (общность): DC ДЕТЕКТ-регексы пар-скоупны в Go** (§12.2 by-design: данные вынесены, детект-алгоритм остаётся; + self-gating → добавление ПАРЫ не требует Go, только СВОИ чекеры пары — да). +- **NOTE (масштаб): двух-домовая когезия — шов.** →ru target-wire-text в ДВУХ домах: injection (embedded, recompile) + vs heading.txt «Глава {n}» (pack, data-drop). Причина — байт-точность no-pack golden: то, что golden НУЖНО без пака + → embedded; пар-гейтнутое, что golden НЕ трогает → pack. Честный шов, задокументирован; унификация потребовала бы + target-loader вне book-langpack_root (сдвиг снапшота golden — запрещён). + сиблинг target/source-данные ещё в Go + (cheapgates ё-политика/оному/magnitude) — граница data/algorithm проведена неравномерно (кандидаты item-8). + +**Финальная верификация (пост-фиксы):** `go build/vet` clean; `TM_MINER_PARITY=1 go test -race ./...` весь зелёный; +парити `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}` EXACT; `git status` golden capture пуст (байт-идентичен). + +--- + +# ДОБАВКА: полный data-out чекеров/гейтов/санитайзера (директива владельца + оркестратора) + +**Триггер:** владелец (после core-пака) — «магические названия остались в `checkers_zh_ru.go`/`checkers_pack13_test.go`; что при +100 языках? это мелкие примеры, ты всё отловил?». Честный ответ: НЕТ — core-пак вынес LOOKUP-таблицы, но оставил DETECTION-регексы + сиблинг-гейты как «алгоритм/DEFER» (слишком мягко для планки 100 языков). Решение оркестратора: **доводи data-out, структуру НЕ трогай** (сплит — след. пак с форм-фиксами). Выбор владельца: «Full: все паттерны → данные, Go generic» + «тесты грузят фикстуры из пака». + +## Внешняя сверка (харнесы `/home/ubuntu/projects/tmp`, 5-агентный воркфлоу) +- **crush** (Charm, Go AI-агент, ~300 .go) — **~71 фокус-пакет**, один каталог = один концерн; провайдеры = ДАННЫЕ (внешний SDK + JSON-каталог, 0 per-provider Go); промпты = embedded `.md`; пост-обработка РАЗНЕСЕНА по фокус-пакетам, НЕ «checkers-куча». Сильное свидетельство против нашей 38-файловой кучи И за «volatile-ось = данные». +- **anthropic/openai-go** — плоский root = кодоген-артефакт (не образец); но ручные части = МАЛЫЕ single-purpose пакеты (`apijson`/`apiquery`/…). +- **GalTransl/AiNiee** (домен, что мы зеркалим) — валидируют data-мандат (глоссарий TSV, промпты per-язык, WHICH-чеки = config-список, control-codes `regex.json` = ДАННЫЕ). НО обе ТЕКУТ языком в ТЕЛА чекеров (残留日文/独白男他 gated on target_lang; per-lang Unicode-регексы) — ровно наша утечка; AiNiee `regex.json` = чистый контр-пример = наш фикс. + +## Что вынесено (Full data-out) +**`checkers_zh_ru.go` → `checkers.go`, ПОЛНОСТЬЮ пар-агностичен.** DC1/DC2/percent DETECTION-паттерны (`shichen_re`/`ru_hours_re`/`qianwan_*`/`shushiwan_*`/`cheng_re`/`decimal_fraction_re`/`percent_word`) → пар-пак `dc-checkers.txt` (`patternkeyvalue`, verbatim); компилируются ОДИН раз в `compileCheckers` → `*dcCheckers` спек, живёт в `r.checkers` (openRunner) → `cheapGateConfig`. Target-general (broken-word `-йть`) → embedded `internal/lang/data/target-ru.txt`. dc*-символы стали ключами данных; чекеры — методы generic-спека, nil-инертны для no-pack. +**`cheapgates.go`:** yofikator homograph-whitelist (38 слов) → target-ru `yo_homograph`; 万億 magnitude — CJK-парсер (digit/unit/**magnitude 万萬億亿兆**) **ДЕДУП в `lang.CJKSection`** (закрыл reviewer-NOTE о дубле), ru-стемы (тысяч/миллион/…) → target-ru `magnitude_stem`. `lintYofikation`/`lintNumberMagnitude` — методы спека. **Оставлено (документировано):** ё↔е фолд (1-символьная орфография = алгоритм), dialogue-dash (пунктуация), **translit-interj `ара-ара`/… = РАТИФИЦИРОВАННЫЙ DEFER** (D39.16, prompt NOT-do — единственный ru-output-блоклист, придержан). +**`sanitizer.go`:** 4 preamble + trailing-note + edit-meta + invalid-sign регексы → target-ru (извлечены из исходника СКРИПТОМ, round-trip byte-match — не перепечатка). Санитайзер уже `isRuTarget`-гейтнут → паттерны = ru-target данные, загружаются `ruSanitizer = lang.TargetChecksFor("ru")`; strip/detect-АЛГОРИТМ остаётся. «ru»-ключ зеркалит гейт (тред таргета = follow-up, как sourceAbbrevs). +**Тесты грузят из пака:** `testCheckers(t) = compileCheckers(pack.DCCheckers, TargetChecksFor("ru"))`; checkers/cheapgates-фикстуры прогоняют АЛГОРИТМ над паттернами пака — 0 дублирования пар-значений (входные примеры-сценарии остаются, §0.3 легитимно). + +**Дома по ГЕЙТИНГУ:** source-gated пара-чекеры (DC1/DC2/percent, нужен zh-исходник) → пар-пак (nil для golden→инертны); target-general (register/broken/yofikator/magnitude-стемы/sanitizer, бегут на любом →ru) → embedded target-ru (нужны no-pack golden'у); source-CJK (magnitude-руны) → embedded cjk-section. + +**Байт-точность (пост-каждый-инкремент):** golden БАЙТ-ИДЕНТИЧЕН (`capture.golden` не тронут — вкл. ch6, где санитайзер СРЕЗАЕТ ru-преамбулу → доказано, что паттерны из данных = те же байты); парити EXACT `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}`; `-race` весь зелёный. **Свип: 100→47** остаток (26 SQL-коммент migrate.go · flag-МESSAGE-диагностики checkers/cheapgates · 1-символьные фолды ё→е/ъ/kana · пунктуация-сепараторы · translit-DEFER · `第`-маркер ingest — кандидат). Все классифицированы/легитимны. + +## Приложение A — карта связности подсистем (измерено grep'ом, вкл. тесты) — ДЛЯ ПРОМТА СЛЕД. ПАКА + +Кандидаты-пакеты и их файлы: **miner** (miner*.go, mining.go — 8) · **membank** (memory/memnorm/mempostcheck/memseed/seeding — 5) · **chunk** (chunker/ingest — 2) · **checks** (cheapgates/checkers/sanitizer/coverage/quality/regressionguard — 6). Драйвер (runner/bookrun/waverun/wave/stagerun/chunkrun/resume) + инфра (snapshot/render/export/status/banknote/disposition/escalation) = композит-корень. + +**Кто ссылается ВНУТРЬ каждого (non-test · TEST):** +- **miner** ← non-test: runner:2, waverun:2, seeding:2 (узко!) · test: miner_test:63, miner_parity:7, seedlint:4, waverun:2. +- **membank** ← non-test: waverun:18, wave:14, runner:13, snapshot:13, +многие (горячий путь) · test: memory:78, memseed:54, memnorm:38. +- **chunk** ← non-test: waverun:12, wave:8, runner:5, status:5, stagerun:5, memseed:4 · test: chunker:35, ingest_encoding:28, ingest:23. +- **checks** ← non-test: chunkrun:17, waverun:16, status:16, export:9, disposition:9 · test: sanitizer:90, coverage:29, regressionguard:19, cheapgates:15. + +**Латеральная связность — НЕ «ноль» (уточнение к survey «star topology»):** есть ОБЩИЙ СУБСТРАТ, который надо ХОЙСТНУТЬ ПЕРВЫМ, иначе сплит потечёт: +- **miner → membank: 35 refs**, но это ШАРЕД-утилиты, не membank-API: `normalizeSourceKey`@memnorm (12 — текст-нормализация, общая для майнера И банка), `loadGlossarySeed`/`seedTerm`/`seedAlias`/`seedFile`@memseed (9 — майнер читает СИД как GT/entity-guard), `runesEqual`@mempostcheck (3), `distinct`/`matches`@memory (8 — мелкие дженерик-хелперы). +- **membank → chunk: 7** (`Chunk`-тип), **checks → membank: 19** (`memoryEntry`/`pickedEntry`/`tokenizeCyrillic`), **chunk/miner → checks: 6+6**. + +**Вывод для сплита (порядок):** (1) ХОЙСТ шаред-базы ПЕРЕД подсистемами — `Chunk`/`MinerChunk` value-типы + текст-нормализация (`normalizeSourceKey`/`runesEqual`/`tokenizeCyrillic`) + сид-типы (`SeedTerm`/`loadGlossarySeed`) в базовый пакет (напр. `internal/pipeline/model` + `internal/text`); (2) затем miner (самый узкий: 6 внешних не-тест ссылок, импортит только lang+store+base) → checks (чистые ф-ии над строками, импортят config+lang+base) → membank → chunk. Драйвер импортит всё (композит-корень, как crush `internal/agent`). Форм-фиксы оркестратора (Export/RequestHash/resolveChunkState) едут в тот же пак — golden ревьюится ОДНИМ слоем. + +## Приложение B — предлагаемая экспорт-поверхность (что капитализируется / что прячется) + +- **`internal/text` (база):** `NormalizeSourceKey`, `RunesEqual`, `TokenizeCyrillic`. (сейчас memnorm/mempostcheck; общие для miner+membank+checks). +- **`internal/pipeline/model` (база):** `Chunk`, `MinerChunk`, `SegBudget`, `Document` (value-типы, вокабуляр). +- **`miner`:** ЭКСПОРТ `MineBank`, `MinerConfig`, `MinedTerm`, `Contrast`, `LoadContrast`; ПРЯЧЕТ mineDetect/proposeAliasEdges/buildPalladiusCyrSyllables/patternCandidates/emissionEligible. Импорт: lang, store, text, model. +- **`membank`:** ЭКСПОРТ `MemoryBank`, `Select`, `Version`, `BaseVersion`, `MemoryEntry`, инъекц-рендереры; ПРЯЧЕТ pickedEntry/Aho-Corasick/suppressContained. Импорт: lang, store, text. +- **`chunk`:** ЭКСПОРТ `SplitChunks`, `Ingest`, `IngestEncoded`; ПРЯЧЕТ chapterDraftChunks/packSentences/splitSourceSentences. Импорт: lang, model. +- **`checks`:** ЭКСПОРТ `RunCheapGates`(+`CheapGateConfig`/`CheapGateResult`), `SanitizeOutput`(+`SanitizerResult`), `CompileCheckers`(+`Checkers`); ПРЯЧЕТ lint*-методы/dcCheckers-внутренности. Импорт: config, lang, text. + +--- + +# ДОПРИЁМОЧНАЯ ДЕЛЬТА (лендинг-блокер оркестратора, 24.07) + +**(1) ОБЯЗАТЕЛЬНО — аддендум `parsePalladius` ВЫПОЛНЕН.** (а) generic-парсер `parseCategoryRows` → `map[категория]map[key]value` (3-field key→cyr, 2-field key→"" set-member); НЕТ per-category switch, незнакомая категория ≠ ошибка парсера. (б) ОДИН типизированный `Palladius`-struct (Initials/Finals/YW/SpecialI + Retroflex/VFinal/VFinalInitial) вместо 4-картежа + 7 параллельных Pack-полей → `Pack.Palladius Palladius`. (в) required-категории валидирует ПОТРЕБИТЕЛЬ (`validate()`), не парсер. Удалены `parsePalladius`+`parsePalladiusPhonotactics`; `assignPair`/`mergePair` схлопнуты в parse+merge (оба пар-файла через generic-путь, union нил-сейф через `newPalladius()`); потребители (miner_palladius.go, langpack_test.go) обновлены. **Hash-нейтрально** (байты .txt не тронуты); `packAlgoVersion` v1→v2 (схемо-изменение по политике константы — выбрал БАМП; сдвиг Version() = version-only, golden без пака не тронут). + +**(2) Хвосты:** `pipeline-c2.yaml` — явный segmentation-блок (1797/3200/1.1978/0.3852, байт-нейтрально, закрыл обещание коммента pipeline.go «все шиппинг ставят явно») · `backend/README.md:59` пути промптов → `prompts/zh-ru/` · `embedded.go` parseCJKSection unknown-category want-list → `+magnitude` · `packAlgoVersion` бамп (см. выше). + +**(3) НЕ тронуто (кандидаты пака-15, по указанию оркестратора):** fail-loud оверлея уровнем выше · override-семантика unionStringMap · isRuTarget/'ru'-швы · empty-probe гард lintMagnitudeScale · дуп терминаторов chunker/coverage · 第. + +**Верификация дельты:** `go build/vet/test -race` зелёные · golden БАЙТ-ИДЕНТИЧЕН (`capture.golden` не тронут) · парити EXACT `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}`. Сессия НЕ коммитит.