Land pack-14 generality pass: detection patterns and pair data to langpacks and embedded target data, book-scoped overlay, pair-agnostic checkers, generic palladius parser, prompts under zh-ru
This commit is contained in:
parent
77dd9b8156
commit
952469b278
53 changed files with 2019 additions and 651 deletions
|
|
@ -39,6 +39,6 @@ Go-бэкенд издательского художественного пер
|
|||
2. `docs/architecture/05-decisions-log.md` целиком (контракт).
|
||||
3. По роли: Бэкенд — `backend/README.md` + `03-implementation-notes.md`; Полигон — `eval/README.md` + `experiments/00` + `09-pilot-protocol.md`; всем — `07-strategic-review.md` §6–9 (вердикт/риски/курс).
|
||||
|
||||
## Текущее состояние (2026-07-23, пост-D39.20 — эра «механизация выпускного качества»)
|
||||
## Текущее состояние (2026-07-24, пост-D39.22/пак-14 — эра «механизация выпускного качества»)
|
||||
|
||||
Фаза 0 ✅; **Ф1-инфра ✅** (D20–D28; golden = инвариант детерминизма, №8 README = экспорт-контракт). Арка «качество-первым» D26→D38.5 ЗАКРЫТА; **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ**: 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1–D39.10) → план (D39.12) → **стройка пака-11 завершена** (D39.13–D39.19: волновой исполнитель + чанкер output-бюджет + src→dst-редактор + Go-майнер паритет-EXACT + банкнота + `internal/lang`+langpacks + reject-set + echo-сплит; хроника — D-лог). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07, `books/gu-zhenren/rerun2/`):** вся операционка подтверждена живьём — ОДИН resnapshot, **$0-резюм драфта на всех 3 армах**, echo 0, $6.63<$15; майнинг-стоп→подпись→reject-set отработали. **Вердикт качества:** планка ≤2 претензии НЕ пройдена (пилот не открыт); три независимых сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232 при floor-шуме 0 · два слепых чтения) → **анти-корреляция гладкость↔верность доказана трижды** (чисто-художественные выборы обоих читателей идентичны = mistral 4/5 — арм с критическими смысловыми); **mistral исключён как несущий редактор** (голос = north-star промпта), **dspro-vs-glm решает итерация №2 после механизации** (драфт $0 → ~$0.08/арм + судья-lite). **Эксп ЗАКРЫТ решением владельца (D39.22), итерация №2 отложена** (её вопросы — довесок следующего платного прогона); **редактор = deepseek-v4-pro ИНТЕРИМ** (судья №1 · проза/глоссарий · дефекты чинибельны · ×2 дешевле; glm-5 резерв, отмена = один арм-конфиг); стайл-решения D39.21: меры → единицы читателя (时辰=2ч) · generic-«гу» средний род, именные по-термово · стихи смыслом · 资质=«талант». research/21: транспорт наш опережает/вровень, рефакторинг ОТВЕРГНУТ. **Курс — разработка бэкенда:** пак-12 транспорт-гигиена (выдан, гейт открыт, лендится первым) → пак-13 «выпускной QA» (выдан: заголовок-шаблон · юнит/число-чекеры · broken-word/latin · mined_delta-крэш-фикс · аудит банк-инъекции · стайл-канон · дефолт-редактор dspro; wire-двигающий, один resnapshot) → сид-дельта (元-корень · 赤城 · 学堂家老 · 春秋蝉) → **масштаб целой книги** (волны/майнинг-каденс/потолки на сотни глав) · канал B (18+) вживую · Ф2-механизмы (голос/состояние D21 · native-Gemini судья) · ja→ru §B5 → пилот Ф2.5. Стек: draft flash thinking-ON → **editor deepseek-v4-pro** → судья gemini-preview + grok-фолбэк на 18+ (~$0.85/ранобэ D30.4). **Висит на владельце:** мини-каноны сид-дельты (元-корень · 春秋蝉) · ja→ru-реплика (§B5) · старые (пилот Ф2.5: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Ключ xAI: единый, data-sharing, off перед продом (D27). Книга на стенде `/home/ubuntu/books/gu-zhenren/` (GB18030; текст и производные ВНЕ git). Ключи: DEEPSEEK, ZAI, KIMI, OPENAI, GEMINI, XAI, MISTRAL. Стенд: WSL2, GTX 1070 8GB (localhost из-под прокси = 403 — no-proxy транспорт для local).
|
||||
Фаза 0 ✅; **Ф1-инфра ✅** (D20–D28; golden = инвариант детерминизма, №8 README = экспорт-контракт). Арка «качество-первым» D26→D38.5 ЗАКРЫТА; **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ**: 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1–D39.10) → план (D39.12) → **стройка пака-11 завершена** (D39.13–D39.19: волновой исполнитель + чанкер output-бюджет + src→dst-редактор + Go-майнер паритет-EXACT + банкнота + `internal/lang`+langpacks + reject-set + echo-сплит; хроника — D-лог). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07, `books/gu-zhenren/rerun2/`):** вся операционка подтверждена живьём — ОДИН resnapshot, **$0-резюм драфта на всех 3 армах**, echo 0, $6.63<$15; майнинг-стоп→подпись→reject-set отработали. **Вердикт качества:** планка ≤2 претензии НЕ пройдена (пилот не открыт); три независимых сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232 при floor-шуме 0 · два слепых чтения) → **анти-корреляция гладкость↔верность доказана трижды** (чисто-художественные выборы обоих читателей идентичны = mistral 4/5 — арм с критическими смысловыми); **mistral исключён как несущий редактор** (голос = north-star промпта), **dspro-vs-glm решает итерация №2 после механизации** (драфт $0 → ~$0.08/арм + судья-lite). **Эксп ЗАКРЫТ решением владельца (D39.22), итерация №2 отложена** (её вопросы — довесок следующего платного прогона); **редактор = deepseek-v4-pro ИНТЕРИМ** (судья №1 · проза/глоссарий · дефекты чинибельны · ×2 дешевле; glm-5 резерв, отмена = один арм-конфиг); стайл-решения D39.21: меры → единицы читателя (时辰=2ч) · generic-«гу» средний род, именные по-термово · стихи смыслом · 资质=«талант». research/21: транспорт наш опережает/вровень, рефакторинг ОТВЕРГНУТ. **Курс — разработка бэкенда:** паки 12/13/14 ЗАЛЕНДЕНЫ 24.07 (транспорт-гигиена `7fe0a5b` · выпускной QA `2f91b04` · **генеральность-пасс**: пара/книга-данные → langpack+`internal/lang/data` embedded, книго-overlay `langpack_extend` (古月 вон из общего zh-пака), пар-агностичный `checkers.go`, generic-Палладий; приёмка воркфлоу 11+4 линз, байт-сверка против HEAD, отчёт с ревью-шапкой в `docs/archive/reports/`) → **пак-15 «структура+форма»** (хойст `internal/text`+value-типов → сплит miner/checks/membank/chunk по карте связности отчёта пака-14 · Export/RequestHash/resolveChunkState · конфиг-слой пары · fail-loud-хвосты оверлея) → сид-дельта (元-корень · 赤城 · 学堂家老 · 春秋蝉) → канал B (18+) вживую · слой-3 diff-редактор (прецедент AiNiee/LinguaGacha, research/22) · Ф2-механизмы (голос/состояние D21 · native-Gemini судья) · ja→ru §B5 → пилот Ф2.5. Стек: draft flash thinking-ON → **editor deepseek-v4-pro** → судья gemini-preview + grok-фолбэк на 18+ (~$0.85/ранобэ D30.4). **Висит на владельце:** мини-каноны сид-дельты (元-корень · 春秋蝉) · ja→ru-реплика (§B5) · старые (пилот Ф2.5: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Ключ xAI: единый, data-sharing, off перед продом (D27). Книга на стенде `/home/ubuntu/books/gu-zhenren/` (GB18030; текст и производные ВНЕ git). Ключи: DEEPSEEK, ZAI, KIMI, OPENAI, GEMINI, XAI, MISTRAL. Стенд: WSL2, GTX 1070 8GB (localhost из-под прокси = 403 — no-proxy транспорт для local).
|
||||
|
|
|
|||
|
|
@ -56,5 +56,5 @@ set -a; . ./.env; set +a; TM_LIVE=1 go test -tags live -run TestLive -v ./intern
|
|||
- **D15.2:** READ-половина `tmctl export` ПОСТРОЕНА (D39.5, минимальная форма); annotation/override-половина и content-addressed resume — отложены (v3.1-спека = дизайн-оф-рекорд, `docs/D15.2-…spec.md`).
|
||||
- **Retries.RegenerateBeforeEscalate НЕ в снапшоте** (pre-existing, D39.4-LOW): в крэш-окне пониженный ретрай-бюджет молча бросает оплаченные OK-чекпоинты. Дизайн-фикс при следующем заходе в stagerun.
|
||||
- F3 at-most-once (после D15.2); Escal.Chains — валидируемый мёртвый конфиг (раннер читает только `escalate_to`); флоры min_max_tokens = схема (D24.3).
|
||||
- `prompts/editor-mono.md` — референс-вариант D13.1 (не боевой, тестом закреплён); боевой = `editor.md` v3-discourse (БИЛИНГВ). Глосс-капы санитайзера (6 CJK / 24 лат-токен) — tunable-эвристики (D39.5), ревизия по чтению пере-прогона.
|
||||
- `prompts/zh-ru/editor-mono.md` — референс-вариант D13.1 (не боевой, тестом закреплён); боевой = `prompts/zh-ru/editor.md` v3-discourse (БИЛИНГВ) (промпты pair-keyed под `prompts/zh-ru/`, pair-14 §9). Глосс-капы санитайзера (6 CJK / 24 лат-токен) — tunable-эвристики (D39.5), ревизия по чтению пере-прогона.
|
||||
- Хвост LOW/NOTE-находок адверсариала D39.4 — в ledger-очереди (golden-ячейка глосса, escalated+stripped тест, trustGateEvent первый-отказ и пр.).
|
||||
|
|
|
|||
51
backend/configs/langpacks/zh-ru/dc-checkers.txt
Normal file
51
backend/configs/langpacks/zh-ru/dc-checkers.txt
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
# WS5 defect-class checker DATA for zh→ru (checkers.go). Pair-14 data-out: the DETECTION PATTERNS (regexes +
|
||||
# literal probes) are DATA too, not engine code — a new pair ships its own. `category<TAB>key[<TAB>value]`.
|
||||
# pattern values are VERBATIM (a regex or a literal probe); the checker ALGORITHM (compare counts, ×2 hours,
|
||||
# suppress-if-ok) stays generic Go. The `*_re` keys are compiled; the others are literal strings.Contains probes.
|
||||
#
|
||||
# DC1 时辰 (double-hour): small CJK count → value (1 时辰 = 2 modern hours).
|
||||
cjk_numeral 一 1
|
||||
cjk_numeral 二 2
|
||||
cjk_numeral 两 2
|
||||
cjk_numeral 三 3
|
||||
cjk_numeral 四 4
|
||||
cjk_numeral 五 5
|
||||
cjk_numeral 六 6
|
||||
cjk_numeral 七 7
|
||||
cjk_numeral 八 8
|
||||
cjk_numeral 九 9
|
||||
cjk_numeral 十 10
|
||||
# DC1: Russian hours-count word → value (the alternatives ru_hours_re captures).
|
||||
ru_hour один 1
|
||||
ru_hour два 2
|
||||
ru_hour двух 2
|
||||
ru_hour три 3
|
||||
ru_hour трёх 3
|
||||
ru_hour трех 3
|
||||
ru_hour четыре 4
|
||||
ru_hour пять 5
|
||||
ru_hour шесть 6
|
||||
# DC6 register negative-list: fairy-tale/chancery lexemes out of the xianxia register (терем + case forms).
|
||||
register_neg терем
|
||||
register_neg терема
|
||||
register_neg тереме
|
||||
register_neg теремом
|
||||
register_neg терему
|
||||
register_neg теремах
|
||||
# --- DETECTION patterns (pair-14 data-out). `pattern<TAB>key<TAB>value`; value VERBATIM. ---
|
||||
# DC1: count before 时辰 (src); Russian «<count> час…» rendering (final).
|
||||
pattern shichen_re ([0-9一二两三四五六七八九十])\s*个?\s*时辰
|
||||
pattern ru_hours_re (\d+|один|два|двух|три|трёх|трех|четыре|пять|шесть)\s+час
|
||||
# DC2 千万=10^7: src probe, ok-suppressor (case-insensitive), fire/veto words (case-sensitive Contains).
|
||||
pattern qianwan_src 千万
|
||||
pattern qianwan_ok_re (?i)десят\p{L}* миллион|10\s*000\s*000|10000000
|
||||
pattern qianwan_fire_word тысяч
|
||||
pattern qianwan_veto_word миллион
|
||||
# DC2 数十万: src probe, ok-suppressor, under-render fire (case-sensitive, per reference).
|
||||
pattern shushiwan_src 数十万
|
||||
pattern shushiwan_ok_re (?i)сотн\p{L}* тысяч|нескольк\p{L}* сот\p{L}* тысяч|[1-9]00\s*000
|
||||
pattern shushiwan_fire_re десятк\p{L}* тысяч
|
||||
# percent 成=tenths: src expression; error-cue fraction; percent-form suppressor word.
|
||||
pattern cheng_re ([0-9一二三四五六七八九十])成([0-9一二三四五六七八九]?)
|
||||
pattern decimal_fraction_re десят(?:ая|ых|ой|ые)|сот(?:ая|ых|ой|ые)|\d+[.,]\d
|
||||
pattern percent_word процент
|
||||
20
backend/configs/langpacks/zh-ru/palladius-phonotactics.txt
Normal file
20
backend/configs/langpacks/zh-ru/palladius-phonotactics.txt
Normal file
|
|
@ -0,0 +1,20 @@
|
|||
# Palladius syllable-generator phonotactic constraints (miner_palladius.buildPalladiusCyrSyllables). `category<TAB>member`, pinyin tokens.
|
||||
# retroflex: retroflex/sibilant initials that take SPECIAL_I, not a bare "i" final.
|
||||
retroflex zh
|
||||
retroflex ch
|
||||
retroflex sh
|
||||
retroflex r
|
||||
retroflex z
|
||||
retroflex c
|
||||
retroflex s
|
||||
# vfinal: the ü-finals (pinyin writes ü as plain u); legal only after a vfinal_initial.
|
||||
vfinal v
|
||||
vfinal ve
|
||||
vfinal van
|
||||
vfinal vn
|
||||
# vfinal_initial: the initials after which a ü-final is legal (j/q/x + l/n).
|
||||
vfinal_initial j
|
||||
vfinal_initial q
|
||||
vfinal_initial x
|
||||
vfinal_initial l
|
||||
vfinal_initial n
|
||||
2
backend/configs/langpacks/zh/sentence-terminator.txt
Normal file
2
backend/configs/langpacks/zh/sentence-terminator.txt
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
# Source sentence-ending marks the alias miner splits on (alias.cooccur_same_sentence 。!?). A rune SET; stored sorted. The miner adds "\n" itself.
|
||||
。!?
|
||||
|
|
@ -3,7 +3,6 @@
|
|||
东方
|
||||
公孙
|
||||
南宫
|
||||
古月
|
||||
司徒
|
||||
司马
|
||||
夏侯
|
||||
|
|
|
|||
2
backend/configs/langpacks/zh/title-formant.txt
Normal file
2
backend/configs/langpacks/zh/title-formant.txt
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
# Rank/measure formant chars a candidate treats as a TITLE (patterns.formant_type 等/转/阶 → title). A rune SET; stored sorted.
|
||||
等转阶
|
||||
|
|
@ -37,7 +37,7 @@ stages:
|
|||
role: translator
|
||||
model: deepseek-v4-flash
|
||||
prompts:
|
||||
zh-ru: ../prompts/translator.md
|
||||
zh-ru: ../prompts/zh-ru/translator.md
|
||||
prompt_version: v1-reflow
|
||||
temperature: 0.3
|
||||
reasoning: "off"
|
||||
|
|
@ -48,7 +48,7 @@ stages:
|
|||
# через ReasoningNone no-op). Пиннут (editor не эскалирует — D12).
|
||||
model: deepseek-v4-pro
|
||||
prompts:
|
||||
zh-ru: ../prompts/editor.md
|
||||
zh-ru: ../prompts/zh-ru/editor.md
|
||||
prompt_version: v3-discourse-reflow
|
||||
temperature: 0.4
|
||||
reasoning: "off" # NO-OP на deepseek-v4-pro (ReasoningNone) → thinking ON; НЕ вооружает эхо-мину
|
||||
|
|
|
|||
|
|
@ -36,7 +36,7 @@ stages:
|
|||
role: translator
|
||||
model: deepseek-v4-flash
|
||||
prompts:
|
||||
zh-ru: ../prompts/translator.md
|
||||
zh-ru: ../prompts/zh-ru/translator.md
|
||||
prompt_version: v1-reflow
|
||||
temperature: 0.3
|
||||
reasoning: "off"
|
||||
|
|
@ -47,7 +47,7 @@ stages:
|
|||
# P1a-дискурс few-shot держится (glm не reasoning-модель, CoT-конфликта нет, в отличие от dspro-арма).
|
||||
model: glm-5
|
||||
prompts:
|
||||
zh-ru: ../prompts/editor.md
|
||||
zh-ru: ../prompts/zh-ru/editor.md
|
||||
prompt_version: v3-discourse-reflow
|
||||
temperature: 0.4
|
||||
reasoning: "off" # glm thinking:disabled — таймауты ×3 (полигон)
|
||||
|
|
|
|||
|
|
@ -35,7 +35,7 @@ stages:
|
|||
role: translator
|
||||
model: deepseek-v4-flash # черновик — тот же, что базлайн (свап только редактора)
|
||||
prompts:
|
||||
zh-ru: ../prompts/translator.md
|
||||
zh-ru: ../prompts/zh-ru/translator.md
|
||||
prompt_version: v1-reflow
|
||||
temperature: 0.3
|
||||
reasoning: "off"
|
||||
|
|
@ -46,7 +46,7 @@ stages:
|
|||
# WS6(г); не reasoning-модель → эхо-мины нет). Идёт ТОЛЬКО за rate-guard (models.yaml rate_limit).
|
||||
model: mistral-large-2512
|
||||
prompts:
|
||||
zh-ru: ../prompts/editor.md
|
||||
zh-ru: ../prompts/zh-ru/editor.md
|
||||
prompt_version: v3-discourse-reflow
|
||||
temperature: 0.4
|
||||
reasoning: "off"
|
||||
|
|
|
|||
|
|
@ -51,7 +51,7 @@ stages:
|
|||
# loadTemplates (закрывает латентный баг «ja едет через китайский паратаксис»). Язык промпта —
|
||||
# свойство пакета. Прод zh→ru резолвит тот же translator.md → тот же SHA (поведение-нейтрально).
|
||||
prompts:
|
||||
zh-ru: ../prompts/translator.md
|
||||
zh-ru: ../prompts/zh-ru/translator.md
|
||||
# D30.2 reflow: из translator.md убрано «Сохраняй разбивку на абзацы» (построчная
|
||||
# вёрстка исходника деструктивна для ru — корень нечитаемости exp12 §2.5-S). Bump
|
||||
# v0-draft→v1-reflow: сменил SHA промпта, осознанный --resnapshot.
|
||||
|
|
@ -76,11 +76,11 @@ stages:
|
|||
model: deepseek-v4-pro
|
||||
# Слой 2 (D39): pair-keyed пакет конвенций (см. draft-стадию) — наполнена zh→ru.
|
||||
prompts:
|
||||
zh-ru: ../prompts/editor.md
|
||||
zh-ru: ../prompts/zh-ru/editor.md
|
||||
# Bump v2-bilingual-reflow→v3-discourse-reflow: в editor.md вшито валидированное P1a-ядро
|
||||
# ДИСКУРС-ПЕРЕВЁРСТКИ + few-shot (exp14/D37 §2а — претензия-1 «рубленые абзацы» = рычаг ПРОМПТА,
|
||||
# не модель; P1a реверстает дефект у всех семей). Меняет SHA промпта → осознанный --resnapshot.
|
||||
# Билингв-каркас (D30.1) и omission-осторожность (D34.3) сохранены. Моно-вариант — prompts/editor-mono.md.
|
||||
# Билингв-каркас (D30.1) и omission-осторожность (D34.3) сохранены. Моно-вариант — prompts/zh-ru/editor-mono.md.
|
||||
prompt_version: v3-discourse-reflow
|
||||
temperature: 0.4
|
||||
reasoning: "off" # NO-OP на deepseek-v4-pro (ReasoningNone) → thinking ON; НЕ вооружает эхо-мину
|
||||
|
|
|
|||
|
|
@ -15,6 +15,15 @@ context:
|
|||
cache_ttl: "5m"
|
||||
# STMDepth/OverlapTokens сняты (WS2 §2а — carryover не строим).
|
||||
|
||||
# Явный zh-ru пар-калибровочный блок (pair-14 §7 — все шиппинг-конфиги ставят его ЯВНО, пар-конфиг =
|
||||
# источник истины; значения = дореформенный Go-fallback → байт-нейтрально, но не полагаемся на fallback).
|
||||
segmentation:
|
||||
draft_budget_out: 1797
|
||||
edit_ceiling_out: 3200
|
||||
fertility:
|
||||
cjk: 1.1978
|
||||
other: 0.3852
|
||||
|
||||
retries:
|
||||
regenerate_before_escalate: 1
|
||||
|
||||
|
|
@ -23,7 +32,7 @@ stages:
|
|||
role: translator
|
||||
model: deepseek-v4-flash
|
||||
prompts:
|
||||
zh-ru: ../prompts/translator.md
|
||||
zh-ru: ../prompts/zh-ru/translator.md
|
||||
prompt_version: v1-reflow # D30.2 reflow (тот же translator.md, что в C1) — label-SHA дисциплина
|
||||
temperature: 0.8 # C2: кандидаты сэмплируются горячее (Р2: T=0.6–1.0)
|
||||
reasoning: "off"
|
||||
|
|
@ -35,7 +44,7 @@ stages:
|
|||
role: judge
|
||||
model: glm-5
|
||||
prompts:
|
||||
zh-ru: ../prompts/judge-selector.md
|
||||
zh-ru: ../prompts/zh-ru/judge-selector.md
|
||||
prompt_version: v0-skeleton
|
||||
temperature: 0
|
||||
reasoning: "off"
|
||||
|
|
@ -47,7 +56,7 @@ stages:
|
|||
# select выше — judge/glm-5, Gemini-слот Ф2 — НЕ трогаем: это судья, не редактор.)
|
||||
model: deepseek-v4-pro
|
||||
prompts:
|
||||
zh-ru: ../prompts/editor.md
|
||||
zh-ru: ../prompts/zh-ru/editor.md
|
||||
# Бамп v2-bilingual-reflow→v3-discourse-reflow (label-SHA дисциплина: тот же editor.md, что в C1 —
|
||||
# вшито P1a-ядро ДИСКУРС-ПЕРЕВЁРСТКИ, D37 §2а; лейбл обязан следовать за новым SHA файла).
|
||||
prompt_version: v3-discourse-reflow
|
||||
|
|
|
|||
|
|
@ -65,6 +65,13 @@ type Book struct {
|
|||
// set ⇒ the runner loads the pack in W0, failing LOUD when the pair's catalog dir exists but a file is
|
||||
// missing/corrupt, and folds pack.Version() into the snapshot so a pack edit is a loud --resnapshot (R1).
|
||||
LangpackRoot string `yaml:"langpack_root"`
|
||||
// LangpackExtend is the optional root of a BOOK-SCOPED langpack overlay (pair-14 §1): a book's PRIVATE
|
||||
// canon — a clan surname (古月 in 蛊真人 is a book clan, not a 百家姓 family), a sect term — that must reach
|
||||
// the miner FOR THIS BOOK without polluting the shared pair langpack. Same layout as langpack_root; only
|
||||
// the files a book ships are present, each UNIONED onto the shared pack (additive) with its bytes folded
|
||||
// into pack.Version() (a book-canon edit is a loud --resnapshot for this book only). Resolved relative to
|
||||
// book.yaml. Empty ⇒ no overlay (the shared pack is used as-is).
|
||||
LangpackExtend string `yaml:"langpack_extend"`
|
||||
// MinedDelta is the optional path to the owner-curated mined-term delta YAML (the W1.5 sign artifact,
|
||||
// R1). Its terms are loaded as Source:"mined" (NOT Source:"seed"), so they fold into the ENRICHED bank
|
||||
// version but NOT the base — adding them moves only snapshot_W2 (the pay-once invariant). Same seedTerm
|
||||
|
|
@ -108,6 +115,7 @@ func LoadBook(path string) (*Book, error) {
|
|||
b.SourceFile = resolve(b.SourceFile)
|
||||
b.GlossarySeed = resolve(b.GlossarySeed)
|
||||
b.LangpackRoot = resolve(b.LangpackRoot)
|
||||
b.LangpackExtend = resolve(b.LangpackExtend)
|
||||
b.MinedDelta = resolve(b.MinedDelta)
|
||||
b.MinedRejects = resolve(b.MinedRejects)
|
||||
if b.ProjectDB == "" {
|
||||
|
|
|
|||
|
|
@ -361,8 +361,16 @@ func LoadPipeline(path string, models *Models) (*Pipeline, error) {
|
|||
if p.Waves.Workers <= 0 {
|
||||
p.Waves.Workers = 1
|
||||
}
|
||||
// WS2 segmentation defaults (ratified zh-ru, §2в): a config that omits the block gets the
|
||||
// output-token budget + independently-re-derived fertility. Zero/negative → default.
|
||||
// Segmentation calibration (pair-14 §7). These numbers are the zh-ru PAIR CALIBRATION — the budgets
|
||||
// tuned for zh→ru and the fertility (est_out per source char-class) independently re-derived on the zh-ru
|
||||
// rerun (R²=0.96). They are NOT a language-neutral engine constant: a new pair must set its OWN
|
||||
// calibration in its pair-config, so every SHIPPING pipeline (pipeline-c1 / the arm yamls) sets the whole
|
||||
// block explicitly — the pair-config is the source of truth ("brать из пар-конфига"). The literals below
|
||||
// are only the last-resort GENERIC FALLBACK for a config that omits the block. The values are held EXACT
|
||||
// on purpose: a book with no langpack chunks entirely on these (the ja→ru golden fixture relies on this
|
||||
// fallback — a re-derived number would shift its chunk boundaries → the wire). Relocating the canonical
|
||||
// zh-ru calibration into the langpack was considered and NOT done: it would be dead for every live path
|
||||
// (shipping configs set the block; the golden has no langpack to read it from) — least-mechanism §12.1.
|
||||
if p.Segmentation.DraftBudgetOut <= 0 {
|
||||
p.Segmentation.DraftBudgetOut = 1797
|
||||
}
|
||||
|
|
|
|||
|
|
@ -132,8 +132,8 @@ func TestBoevoyPipelineC1ResolvesZhRu(t *testing.T) {
|
|||
for _, st := range p.Stages {
|
||||
if st.Role == "translator" {
|
||||
zh, _ := st.PromptPathFor("zh-ru")
|
||||
if !strings.HasSuffix(zh, filepath.Join("prompts", "translator.md")) {
|
||||
t.Errorf("draft zh-ru must be prompts/translator.md, got %q", zh)
|
||||
if !strings.HasSuffix(zh, filepath.Join("prompts", "zh-ru", "translator.md")) {
|
||||
t.Errorf("draft zh-ru must be prompts/zh-ru/translator.md, got %q", zh)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
34
backend/internal/lang/data/cjk-section.txt
Normal file
34
backend/internal/lang/data/cjk-section.txt
Normal file
|
|
@ -0,0 +1,34 @@
|
|||
# Universal CJK section-numeral data (pair-14 §4 + addendum-A): the ONE source of truth for the chapter/
|
||||
# section numeral inventory shared by the ingest chapter-splitter (章节節回 headers) and the chunker heading
|
||||
# rule. A linguistic CONSTANT (the CJK numeral system does not vary by book/pair), so it is embedded in the
|
||||
# language-data package and available even to a book with NO langpack (the ja→ru golden splits 第X章 here).
|
||||
# `category<TAB>rune[<TAB>value]`. digit/unit carry a value; zero/chapter_unit are membership only.
|
||||
digit 一 1
|
||||
digit 二 2
|
||||
digit 两 2
|
||||
digit 兩 2
|
||||
digit 三 3
|
||||
digit 四 4
|
||||
digit 五 5
|
||||
digit 六 6
|
||||
digit 七 7
|
||||
digit 八 8
|
||||
digit 九 9
|
||||
unit 十 10
|
||||
unit 百 100
|
||||
unit 千 1000
|
||||
zero 〇
|
||||
zero 零
|
||||
# chapter_unit: the section-level markers auto-detected in a source txt (第N章/节/節/回). 卷 (volume) is
|
||||
# deliberately excluded (coarser than a chapter).
|
||||
chapter_unit 章
|
||||
chapter_unit 节
|
||||
chapter_unit 節
|
||||
chapter_unit 回
|
||||
# magnitude: the myriad-scale BIG units (万/億/兆) → value, for the number-magnitude gate (cheapgates). ONE
|
||||
# source with digit/unit above — the gate's CJK numeral parser reads this section instead of its own switch.
|
||||
magnitude 万 10000
|
||||
magnitude 萬 10000
|
||||
magnitude 億 100000000
|
||||
magnitude 亿 100000000
|
||||
magnitude 兆 1000000000000
|
||||
11
backend/internal/lang/data/injection.txt
Normal file
11
backend/internal/lang/data/injection.txt
Normal file
|
|
@ -0,0 +1,11 @@
|
|||
# Target-language glossary/injection wire-text (pair-14 §2). Rendered into the model request, so it is
|
||||
# TARGET-language data (a Russian header for a →ru book); gated on the target so a non-ru book never
|
||||
# gets a Russian block. The ja→ru golden injects these with NO langpack, so they are engine/target data
|
||||
# here, not a book pack. `target<TAB>key<TAB>value`; VALUE IS VERBATIM (leading spaces are significant).
|
||||
# Add a target = add its rows; remove one = delete its rows.
|
||||
ru glossary_header ГЛОССАРИЙ (используй эти утверждённые переводы имён и терминов последовательно; строки с пометкой ⟨проверить⟩ — неподтверждённые кандидаты):
|
||||
ru editor_header КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имён и терминов (в черновике термин исходника слева ДОЛЖЕН быть передан именно указанной формой справа — приводи к ней любые расхождения, склоняя по контексту; не вводи иных вариантов и не меняй ничего другого):
|
||||
ru unverified_marker ⟨проверить⟩
|
||||
ru gender_male (муж. — мужские родовые формы)
|
||||
ru gender_female (жен. — женские родовые формы)
|
||||
ru gender_hidden (пол СКРЫТ до раскрытия — избегай родовых форм; при неизбежности — мужские)
|
||||
25
backend/internal/lang/data/refusal.txt
Normal file
25
backend/internal/lang/data/refusal.txt
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
# Universal refusal blacklist (pair-14 §3, ported 1:1 from eval/refusal_bench.py). Engine safety data:
|
||||
# a model can refuse in ANY language regardless of the book's pair, so EVERY pattern here fires for EVERY
|
||||
# book (the ja→ru golden flags a Russian refusal with NO langpack). ONE regex per non-comment line, joined
|
||||
# with "|" under (?im); order is immaterial to the boolean match but preserved.
|
||||
# NOTE: the `# --- xx ---` markers are ORGANIZATIONAL comments, NOT functional scoping — RefusalPatterns()
|
||||
# reads every line flat into the one universal set. Deleting a language's lines TRIMS that universal set
|
||||
# (drops those phrases for ALL books); it does NOT "turn off" detection for one pair. Add a language's phrases
|
||||
# = add lines under (or beside) its marker.
|
||||
# --- en ---
|
||||
I (?:can(?:'|no)t|cannot|won'?t|am (?:not able|unable)) (?:to )?(?:assist|help|translate|continue|provide|fulfill)
|
||||
I'?m sorry,? but
|
||||
against (?:my|our) (?:content )?polic
|
||||
content polic(?:y|ies)
|
||||
# --- ru ---
|
||||
не могу (?:помочь|перевести|продолжить|выполнить)
|
||||
я не буду переводить
|
||||
противоречит (?:политике|правилам)
|
||||
# --- ja ---
|
||||
としてお答えできません
|
||||
申し訳ありませんが
|
||||
# --- zh ---
|
||||
我(?:不能|无法)(?:协助|翻译|继续)
|
||||
抱歉[,,]
|
||||
# --- universal explicit marker ---
|
||||
^\s*\[?(?:REFUSED|ОТКАЗ)\]?\s*$
|
||||
35
backend/internal/lang/data/sentence-abbrev.txt
Normal file
35
backend/internal/lang/data/sentence-abbrev.txt
Normal file
|
|
@ -0,0 +1,35 @@
|
|||
# Source-language sentence-splitter abbreviations (pair-14 §4): trailing tokens after which a lone "." is an
|
||||
# abbreviation, not a sentence end (case-insensitive). Per SOURCE language — a CJK source uses 。 and needs
|
||||
# none; `en` is the one that guards an ASCII period. `lang<TAB>token`. Add a source = add its section.
|
||||
# NOTE: the chunker currently consumes the `en` section universally (its historical en-only design); making
|
||||
# the splitter consume the BOOK's source section is a shallow follow-up (thread sourceLang into SplitChunks).
|
||||
# --- en ---
|
||||
en mr
|
||||
en mrs
|
||||
en ms
|
||||
en dr
|
||||
en prof
|
||||
en st
|
||||
en jr
|
||||
en sr
|
||||
en vs
|
||||
en no
|
||||
en vol
|
||||
en ch
|
||||
en fig
|
||||
en col
|
||||
en gen
|
||||
en sgt
|
||||
en capt
|
||||
en lt
|
||||
en rev
|
||||
en gov
|
||||
en sen
|
||||
en rep
|
||||
en etc
|
||||
en inc
|
||||
en ltd
|
||||
en co
|
||||
en mt
|
||||
en ave
|
||||
en rd
|
||||
65
backend/internal/lang/data/target-ru.txt
Normal file
65
backend/internal/lang/data/target-ru.txt
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
# Target-language (ru) readability checker DATA (pair-14 data-out): the wordlists and DETECTION patterns the
|
||||
# TARGET-GENERAL checkers/sanitizer run on ANY →ru output (regardless of source pair). Engine/target data —
|
||||
# a no-langpack →ru book (the ja→ru golden) still runs these — so it is embedded here, not a book pack; the
|
||||
# checker/sanitizer ALGORITHMS stay generic Go. `category<TAB>value`; value VERBATIM (a regex or a literal);
|
||||
# lines are grouped by category IN ORDER (order matters for the sanitizer preamble preview). Add a target =
|
||||
# add data/target-<tgt>.txt + its accessor entry.
|
||||
#
|
||||
# broken-word: a Russian word ending in this suffix is a mangled infinitive (no valid word ends in «-йть»).
|
||||
broken_suffix йть
|
||||
# yofikator homograph whitelist: е-spellings that are DISTINCT words from their ё-counterpart (все≠всё,
|
||||
# небо≠нёбо, берет≠берёт …), so the inconsistency rule must NOT flag «все»+«всё» as one word two ways.
|
||||
yo_homograph все
|
||||
yo_homograph всех
|
||||
yo_homograph всем
|
||||
yo_homograph всеми
|
||||
yo_homograph небо
|
||||
yo_homograph небом
|
||||
yo_homograph узнаем
|
||||
yo_homograph узнаете
|
||||
yo_homograph узнает
|
||||
yo_homograph падеж
|
||||
yo_homograph падежа
|
||||
yo_homograph совершенный
|
||||
yo_homograph совершенное
|
||||
yo_homograph совершенная
|
||||
yo_homograph совершенно
|
||||
yo_homograph совершенны
|
||||
yo_homograph чем
|
||||
yo_homograph тем
|
||||
yo_homograph тема
|
||||
yo_homograph теме
|
||||
yo_homograph темы
|
||||
yo_homograph берет
|
||||
yo_homograph берете
|
||||
yo_homograph берета
|
||||
yo_homograph осел
|
||||
yo_homograph осла
|
||||
yo_homograph ослы
|
||||
yo_homograph ослов
|
||||
yo_homograph мел
|
||||
yo_homograph мела
|
||||
yo_homograph лен
|
||||
yo_homograph лена
|
||||
yo_homograph нес
|
||||
yo_homograph несем
|
||||
yo_homograph несет
|
||||
yo_homograph вел
|
||||
yo_homograph ведет
|
||||
yo_homograph ведем
|
||||
# magnitude_stem: Russian magnitude WORD stem → base-10 exponent, for the number-magnitude gate (a stem
|
||||
# covers [base, base+2] since an unseen multiplier can lift it two orders). `magnitude_stem<TAB>stem<TAB>exp`.
|
||||
magnitude_stem тысяч 3
|
||||
magnitude_stem миллион 6
|
||||
magnitude_stem миллиард 9
|
||||
magnitude_stem триллион 12
|
||||
# sanitizer DETECTION patterns (pair-14 data-out). The sanitizer is isRuTarget-gated, so these are
|
||||
# ru-target data; the strip/detect ALGORITHM stays in sanitizer.go. `sanitizer_preamble` is ORDERED
|
||||
# (detectLeadingPreamble returns the FIRST match's preview). Values VERBATIM (regexes).
|
||||
sanitizer_preamble (?is)^\s*(?:вот|представляю|привожу|держите)\s+[^\n:]{0,40}?(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:
|
||||
sanitizer_preamble (?is)^\s*(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:
|
||||
sanitizer_preamble (?is)^\s*(?:вот\s+)?перевод\s+(?:фрагмента|текста|отрывка|главы|черновика)\s*:
|
||||
sanitizer_preamble (?is)^\s*ниже\s+(?:привед[её]н|представлен|дан|след)[а-яё]*\s+[^\n:]{0,40}?(?:перевод|отредактированн|текст)[а-яё]*[^\n:]{0,20}?:
|
||||
sanitizer_trailing_note (?i)^[*_>\s-]{0,4}(?:(?:примечани[ея]|заметк[аи]|комментари[йяю]|пояснени[ея]|сноск[аи])\s*:|(?:примечани[ея]|заметк[аи]|комментари[йяю])\s+(?:переводчик|редактор)[а-яё]*|прим\.\s*(?:перев|ред)\.?)
|
||||
sanitizer_edit_meta (?im)^[*_>\s-]{0,4}(?:(?:внесённ|внесен)[а-яё]+\s+правк|основн[а-яё]+\s+правк[а-яё]*(?:\s+для\s+справк[а-яё]+|\s*:)|что\s+(?:было\s+)?(?:изменен|исправлен)|список\s+(?:правок|изменений)|список\s+внесённых)
|
||||
sanitizer_invalid_sign (?i)(?:^|\P{L})[ъь][а-яё]|[ъь][ъь]
|
||||
392
backend/internal/lang/embedded.go
Normal file
392
backend/internal/lang/embedded.go
Normal file
|
|
@ -0,0 +1,392 @@
|
|||
package lang
|
||||
|
||||
import (
|
||||
"embed"
|
||||
"fmt"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
)
|
||||
|
||||
// embedded.go: language data the engine needs even for a book with NO langpack pack — engine-universal
|
||||
// (a CJK numeral is a linguistic constant; a refusal phrase is engine safety behaviour) or target-generic
|
||||
// (the Russian glossary-injection wire-text). It is versioned with the engine BINARY (like code), not with
|
||||
// a book's pack, because a no-pack book still splits 第X章 chapters, still flags a refusal, and still injects
|
||||
// a Russian glossary block (the ja→ru golden proves all three). Kept OUT of internal/pipeline so the
|
||||
// data/algorithm boundary holds; embedded so it is available without a book's langpack_root. Each file is
|
||||
// SECTIONED per language/target, so adding a language = add a section, removing one = delete its section.
|
||||
|
||||
//go:embed data/cjk-section.txt data/refusal.txt data/injection.txt data/sentence-abbrev.txt data/target-ru.txt
|
||||
var embeddedFS embed.FS
|
||||
|
||||
// targetCheckFiles maps a TARGET language to its embedded readability-checker data file (pair-14 data-out).
|
||||
// Add a target = add data/target-<tgt>.txt to the go:embed line above and a row here.
|
||||
var targetCheckFiles = map[string]string{"ru": "data/target-ru.txt"}
|
||||
|
||||
// TargetChecks is the TARGET-language readability wordlists + DETECTION patterns (pair-14 data-out) the
|
||||
// target-general checkers/sanitizer run on ANY →target output. A target with no file yields an empty value
|
||||
// (HasData()==false → the consumers stay inert). Values are grouped per category in AUTHORED ORDER.
|
||||
type TargetChecks struct{ byKey map[string][]string }
|
||||
|
||||
// List returns the authored-order values for a category ("" categories → nil).
|
||||
func (t TargetChecks) List(key string) []string { return t.byKey[key] }
|
||||
|
||||
// HasData reports whether this target ships any checker data (used to gate the target-general checkers so a
|
||||
// target without data flags nothing rather than mis-firing).
|
||||
func (t TargetChecks) HasData() bool { return len(t.byKey) > 0 }
|
||||
|
||||
var (
|
||||
targetChecksOnce sync.Once
|
||||
targetChecksBy map[string]TargetChecks
|
||||
)
|
||||
|
||||
// TargetChecksFor returns the readability-checker data for a TARGET language, parsed once from the embedded
|
||||
// per-target file. Unknown/absent target → empty TargetChecks. Panics on a corrupt embed.
|
||||
func TargetChecksFor(tgt string) TargetChecks {
|
||||
targetChecksOnce.Do(func() {
|
||||
targetChecksBy = map[string]TargetChecks{}
|
||||
for lang, file := range targetCheckFiles {
|
||||
byKey, err := parseOrderedByCategory(mustEmbed(file))
|
||||
if err != nil {
|
||||
panic(fmt.Sprintf("lang: embedded %s is corrupt: %v", file, err))
|
||||
}
|
||||
targetChecksBy[lang] = TargetChecks{byKey: byKey}
|
||||
}
|
||||
})
|
||||
return targetChecksBy[tgt]
|
||||
}
|
||||
|
||||
// parseOrderedByCategory reads `category<TAB>value` lines into category → ordered values. Value is VERBATIM
|
||||
// (a regex may carry a trailing metachar), so only the whole line's \r is stripped; blank/`#`-comment lines
|
||||
// (by their trimmed form) are dropped.
|
||||
func parseOrderedByCategory(b []byte) (map[string][]string, error) {
|
||||
out := map[string][]string{}
|
||||
for i, raw := range strings.Split(string(b), "\n") {
|
||||
line := strings.TrimRight(raw, "\r")
|
||||
if strings.TrimSpace(line) == "" || strings.HasPrefix(strings.TrimSpace(line), "#") {
|
||||
continue
|
||||
}
|
||||
f := strings.SplitN(line, "\t", 2)
|
||||
if len(f) != 2 || strings.TrimSpace(f[0]) == "" || f[1] == "" {
|
||||
return nil, fmt.Errorf("line %d: want `category<TAB>value`, got %q", i+1, line)
|
||||
}
|
||||
out[f[0]] = append(out[f[0]], f[1])
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// CJKSection is the universal chapter/section-numeral inventory shared by the ingest chapter-splitter and
|
||||
// the chunker heading rule (pair-14 §4 + addendum-A: ONE source, no ingest↔chunker byte-duplication). The
|
||||
// heading-numeral membership test derives from Digit ∪ Unit ∪ Zero — nothing is stored twice.
|
||||
type CJKSection struct {
|
||||
Digit map[rune]int // 一→1 … 九→9 (incl. 两/兩→2)
|
||||
Unit map[rune]int // 十→10 百→100 千→1000
|
||||
Zero map[rune]bool // 〇 零 (positional zero: value*10 in a numeral run)
|
||||
// ChapterUnit membership + ChapterUnitOrdered (authored order) — ingest's detectChapterUnit iterates in
|
||||
// order and breaks a tie by first-seen, so the ORDER is behaviour, not just a set.
|
||||
ChapterUnit map[rune]bool
|
||||
ChapterUnitOrdered []rune
|
||||
// Magnitude: the myriad-scale big units (万/億/兆 → value) for the number-magnitude gate. int64 (10^12
|
||||
// overflows int32). ONE source with Digit/Unit — the gate no longer keeps its own CJK numeral switch.
|
||||
Magnitude map[rune]int64
|
||||
headingRunes string // sorted Zero∪Digit∪Unit, cached for the ingest chapter-numeral regex class
|
||||
}
|
||||
|
||||
// IsNumeralRune reports whether r can be part of a CJK numeral EXPRESSION (Zero∪Digit∪Unit∪Magnitude) — the
|
||||
// alphabet of a magnitude candidate run (the gate adds ASCII digits itself). BigUnit returns a big unit's
|
||||
// value (万/億/兆). Both read the single CJKSection source (pair-14 data-out dedup).
|
||||
func (c *CJKSection) IsNumeralRune(r rune) bool {
|
||||
if c.IsHeadingNumeral(r) {
|
||||
return true
|
||||
}
|
||||
_, ok := c.Magnitude[r]
|
||||
return ok
|
||||
}
|
||||
func (c *CJKSection) BigUnit(r rune) (int64, bool) { v, ok := c.Magnitude[r]; return v, ok }
|
||||
|
||||
// IsHeadingNumeral reports whether r can be part of a CJK chapter-number run (Zero ∪ Digit ∪ Unit). Arabic
|
||||
// digits are handled by the caller (they are not language data).
|
||||
func (c *CJKSection) IsHeadingNumeral(r rune) bool {
|
||||
if c.Zero[r] {
|
||||
return true
|
||||
}
|
||||
if _, ok := c.Digit[r]; ok {
|
||||
return true
|
||||
}
|
||||
_, ok := c.Unit[r]
|
||||
return ok
|
||||
}
|
||||
|
||||
// HeadingNumeralClass returns the CJK heading-numeral runes (Zero∪Digit∪Unit) as a sorted string, for use
|
||||
// INSIDE a regex character class (the ingest chapter-numeral pattern). Sorted for a stable pattern; order
|
||||
// inside a class is immaterial to matching. None of the runes are class metacharacters, so no escaping.
|
||||
func (c *CJKSection) HeadingNumeralClass() string { return c.headingRunes }
|
||||
|
||||
var (
|
||||
cjkOnce sync.Once
|
||||
cjkVal *CJKSection
|
||||
cjkErr error
|
||||
)
|
||||
|
||||
// DefaultCJKSection returns the process-wide CJK section-numeral data, parsed once from the embedded file.
|
||||
// It panics on a corrupt embed (a build-time asset the tests exercise, so a bad edit fails the suite, never
|
||||
// production silently) — the parse is deterministic and input-free.
|
||||
func DefaultCJKSection() *CJKSection {
|
||||
cjkOnce.Do(func() { cjkVal, cjkErr = parseCJKSection(mustEmbed("data/cjk-section.txt")) })
|
||||
if cjkErr != nil {
|
||||
panic(fmt.Sprintf("lang: embedded cjk-section.txt is corrupt: %v", cjkErr))
|
||||
}
|
||||
return cjkVal
|
||||
}
|
||||
|
||||
func parseCJKSection(b []byte) (*CJKSection, error) {
|
||||
c := &CJKSection{Digit: map[rune]int{}, Unit: map[rune]int{}, Zero: map[rune]bool{}, ChapterUnit: map[rune]bool{}, Magnitude: map[rune]int64{}}
|
||||
valueRune := func(i int, f []string) (rune, int, error) {
|
||||
if len(f) != 3 {
|
||||
return 0, 0, fmt.Errorf("line %d: %s wants `%s<TAB>rune<TAB>value`, got %q", i+1, f[0], f[0], strings.Join(f, "\t"))
|
||||
}
|
||||
r := []rune(f[1])
|
||||
if len(r) != 1 {
|
||||
return 0, 0, fmt.Errorf("line %d: key %q must be a single rune", i+1, f[1])
|
||||
}
|
||||
v, err := strconv.Atoi(strings.TrimSpace(f[2]))
|
||||
if err != nil {
|
||||
return 0, 0, fmt.Errorf("line %d: value %q: %w", i+1, f[2], err)
|
||||
}
|
||||
return r[0], v, nil
|
||||
}
|
||||
setRune := func(i int, f []string) (rune, error) {
|
||||
if len(f) != 2 {
|
||||
return 0, fmt.Errorf("line %d: %s wants `%s<TAB>rune`, got %q", i+1, f[0], f[0], strings.Join(f, "\t"))
|
||||
}
|
||||
r := []rune(f[1])
|
||||
if len(r) != 1 {
|
||||
return 0, fmt.Errorf("line %d: key %q must be a single rune", i+1, f[1])
|
||||
}
|
||||
return r[0], nil
|
||||
}
|
||||
for i, raw := range strings.Split(string(b), "\n") {
|
||||
t := strings.TrimRight(raw, "\r")
|
||||
if strings.TrimSpace(t) == "" || strings.HasPrefix(strings.TrimSpace(t), "#") {
|
||||
continue
|
||||
}
|
||||
f := strings.Split(t, "\t")
|
||||
switch f[0] {
|
||||
case "digit":
|
||||
r, v, err := valueRune(i, f)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
c.Digit[r] = v
|
||||
case "unit":
|
||||
r, v, err := valueRune(i, f)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
c.Unit[r] = v
|
||||
case "zero":
|
||||
r, err := setRune(i, f)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
c.Zero[r] = true
|
||||
case "chapter_unit":
|
||||
r, err := setRune(i, f)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if !c.ChapterUnit[r] {
|
||||
c.ChapterUnit[r] = true
|
||||
c.ChapterUnitOrdered = append(c.ChapterUnitOrdered, r)
|
||||
}
|
||||
case "magnitude":
|
||||
if len(f) != 3 {
|
||||
return nil, fmt.Errorf("line %d: magnitude wants `magnitude<TAB>rune<TAB>value`, got %q", i+1, t)
|
||||
}
|
||||
r := []rune(f[1])
|
||||
if len(r) != 1 {
|
||||
return nil, fmt.Errorf("line %d: magnitude key %q must be a single rune", i+1, f[1])
|
||||
}
|
||||
v, err := strconv.ParseInt(strings.TrimSpace(f[2]), 10, 64)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("line %d: magnitude value %q: %w", i+1, f[2], err)
|
||||
}
|
||||
c.Magnitude[r[0]] = v
|
||||
default:
|
||||
return nil, fmt.Errorf("line %d: unknown category %q (want digit|unit|zero|chapter_unit|magnitude)", i+1, f[0])
|
||||
}
|
||||
}
|
||||
if len(c.Digit) == 0 || len(c.Unit) == 0 || len(c.Zero) == 0 || len(c.ChapterUnit) == 0 || len(c.Magnitude) == 0 {
|
||||
return nil, fmt.Errorf("cjk-section needs non-empty digit, unit, zero, chapter_unit and magnitude sections")
|
||||
}
|
||||
// Cache the sorted heading-numeral class (Zero∪Digit∪Unit) for the ingest regex.
|
||||
var runes []rune
|
||||
for r := range c.Zero {
|
||||
runes = append(runes, r)
|
||||
}
|
||||
for r := range c.Digit {
|
||||
runes = append(runes, r)
|
||||
}
|
||||
for r := range c.Unit {
|
||||
runes = append(runes, r)
|
||||
}
|
||||
sort.Slice(runes, func(i, j int) bool { return runes[i] < runes[j] })
|
||||
c.headingRunes = string(runes)
|
||||
return c, nil
|
||||
}
|
||||
|
||||
// InjectionTexts is the TARGET-language glossary/injection wire-text (pair-14 §2): the Russian headers and
|
||||
// annotations the memory renderers emit into a request. TARGET data (a Russian header for a →ru book), so
|
||||
// the renderers gate on HasData — a target with no injection texts (e.g. →en) injects NOTHING rather than a
|
||||
// stray Russian block. The ja→ru golden renders these with NO book pack, so they are engine/target data.
|
||||
type InjectionTexts struct {
|
||||
GlossaryHeader string // "ГЛОССАРИЙ (…):" — introduces the draft glossary block
|
||||
EditorHeader string // "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ …:" — introduces the editor constraint block
|
||||
UnverifiedMarker string // " ⟨проверить⟩" — tags an unverified draft candidate (leading spaces significant)
|
||||
GenderMale string // " (муж. — …)" DC3 gender directive (leading space significant)
|
||||
GenderFemale string // " (жен. — …)"
|
||||
GenderHidden string // " (пол СКРЫТ …)"
|
||||
}
|
||||
|
||||
// HasData reports whether the target has injection texts (its header is present). The renderers use it to
|
||||
// gate: a target without texts injects nothing, so a non-ru book never gets a Russian glossary block.
|
||||
func (t InjectionTexts) HasData() bool { return t.GlossaryHeader != "" }
|
||||
|
||||
var (
|
||||
injectionOnce sync.Once
|
||||
injectionBy map[string]*InjectionTexts
|
||||
injectionErr error
|
||||
)
|
||||
|
||||
// InjectionTextsFor returns the injection wire-text for a TARGET language (pair-14 §2), parsed once from the
|
||||
// embedded file. A target with no rows returns a zero value (HasData()==false → no injection). Panics on a
|
||||
// corrupt embed. VALUE bytes are verbatim (leading spaces in gender_*/unverified_marker are significant).
|
||||
func InjectionTextsFor(targetLang string) InjectionTexts {
|
||||
injectionOnce.Do(func() { injectionBy, injectionErr = parseInjection(mustEmbed("data/injection.txt")) })
|
||||
if injectionErr != nil {
|
||||
panic(fmt.Sprintf("lang: embedded injection.txt is corrupt: %v", injectionErr))
|
||||
}
|
||||
if t, ok := injectionBy[targetLang]; ok {
|
||||
return *t
|
||||
}
|
||||
return InjectionTexts{}
|
||||
}
|
||||
|
||||
func parseInjection(b []byte) (map[string]*InjectionTexts, error) {
|
||||
out := map[string]*InjectionTexts{}
|
||||
for i, raw := range strings.Split(string(b), "\n") {
|
||||
line := strings.TrimRight(raw, "\r")
|
||||
if strings.TrimSpace(line) == "" || strings.HasPrefix(strings.TrimSpace(line), "#") {
|
||||
continue
|
||||
}
|
||||
f := strings.SplitN(line, "\t", 3) // value (f[2]) VERBATIM — leading spaces are significant
|
||||
if len(f) != 3 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" {
|
||||
return nil, fmt.Errorf("line %d: want `target<TAB>key<TAB>value`, got %q", i+1, line)
|
||||
}
|
||||
tgt, key, val := f[0], f[1], f[2]
|
||||
if out[tgt] == nil {
|
||||
out[tgt] = &InjectionTexts{}
|
||||
}
|
||||
t := out[tgt]
|
||||
switch key {
|
||||
case "glossary_header":
|
||||
t.GlossaryHeader = val
|
||||
case "editor_header":
|
||||
t.EditorHeader = val
|
||||
case "unverified_marker":
|
||||
t.UnverifiedMarker = val
|
||||
case "gender_male":
|
||||
t.GenderMale = val
|
||||
case "gender_female":
|
||||
t.GenderFemale = val
|
||||
case "gender_hidden":
|
||||
t.GenderHidden = val
|
||||
default:
|
||||
return nil, fmt.Errorf("line %d: unknown injection key %q", i+1, key)
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
var (
|
||||
refusalOnce sync.Once
|
||||
refusalVal []string
|
||||
)
|
||||
|
||||
// RefusalPatterns returns the universal refusal-blacklist regex patterns (pair-14 §3), parsed once from the
|
||||
// embedded sectioned file in authored order. Engine safety data: a model can refuse in any language for any
|
||||
// pair, so this is not book-pack data — the caller joins the patterns into one case-insensitive regex. Each
|
||||
// non-comment line is ONE pattern, taken verbatim (a regex may contain leading/trailing metacharacters, so
|
||||
// it is NOT trimmed beyond the line's \r). Panics on a corrupt embed.
|
||||
func RefusalPatterns() []string {
|
||||
refusalOnce.Do(func() { refusalVal = embeddedLines(mustEmbed("data/refusal.txt")) })
|
||||
return refusalVal
|
||||
}
|
||||
|
||||
// embeddedLines returns the non-comment, non-blank lines of an embedded file VERBATIM (only the trailing \r
|
||||
// is stripped; the content is NOT trimmed — significant for a regex pattern). A blank/`#`-comment line is
|
||||
// dropped by its TRIMMED form, but a kept line keeps its own leading/trailing bytes.
|
||||
func embeddedLines(b []byte) []string {
|
||||
var out []string
|
||||
for _, raw := range strings.Split(string(b), "\n") {
|
||||
s := strings.TrimRight(raw, "\r")
|
||||
if strings.TrimSpace(s) == "" || strings.HasPrefix(strings.TrimSpace(s), "#") {
|
||||
continue
|
||||
}
|
||||
out = append(out, s)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
var (
|
||||
abbrevOnce sync.Once
|
||||
abbrevBy map[string]map[string]bool
|
||||
)
|
||||
|
||||
// SentenceAbbrev returns the lower-cased sentence-splitter abbreviation SET for a SOURCE language (pair-14
|
||||
// §4), parsed once from the embedded sectioned file. An unknown/absent language returns an empty set (a
|
||||
// source with no ASCII-period abbreviations — e.g. a CJK source using 。). Panics on a corrupt embed.
|
||||
func SentenceAbbrev(srcLang string) map[string]bool {
|
||||
abbrevOnce.Do(func() {
|
||||
var err error
|
||||
abbrevBy, err = parseSectionedSet(mustEmbed("data/sentence-abbrev.txt"))
|
||||
if err != nil {
|
||||
panic(fmt.Sprintf("lang: embedded sentence-abbrev.txt is corrupt: %v", err))
|
||||
}
|
||||
})
|
||||
if m, ok := abbrevBy[srcLang]; ok {
|
||||
return m
|
||||
}
|
||||
return map[string]bool{}
|
||||
}
|
||||
|
||||
// parseSectionedSet reads `key<TAB>member` lines into per-key sets (keys lower-cased on read is NOT done
|
||||
// here — the key is the language tag; members are stored verbatim). Used for the sentence-abbrev table.
|
||||
func parseSectionedSet(b []byte) (map[string]map[string]bool, error) {
|
||||
out := map[string]map[string]bool{}
|
||||
for i, raw := range strings.Split(string(b), "\n") {
|
||||
t := strings.TrimRight(raw, "\r")
|
||||
if strings.TrimSpace(t) == "" || strings.HasPrefix(strings.TrimSpace(t), "#") {
|
||||
continue
|
||||
}
|
||||
f := strings.Split(t, "\t")
|
||||
if len(f) != 2 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" {
|
||||
return nil, fmt.Errorf("line %d: want `key<TAB>member`, got %q", i+1, t)
|
||||
}
|
||||
key := strings.TrimSpace(f[0])
|
||||
if out[key] == nil {
|
||||
out[key] = map[string]bool{}
|
||||
}
|
||||
out[key][strings.TrimSpace(f[1])] = true
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
func mustEmbed(name string) []byte {
|
||||
b, err := embeddedFS.ReadFile(name)
|
||||
if err != nil {
|
||||
panic(fmt.Sprintf("lang: missing embedded asset %q: %v", name, err))
|
||||
}
|
||||
return b
|
||||
}
|
||||
|
|
@ -21,6 +21,7 @@ import (
|
|||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
|
|
@ -28,7 +29,9 @@ import (
|
|||
// a schema change (a new file, a format change). The DATA content is versioned separately by hashing the
|
||||
// files into Version(), so editing a table also invalidates — you cannot forget to bump a version when
|
||||
// you change the data, because the data's bytes ARE the version (the memnorm.go drift-proofing).
|
||||
const packAlgoVersion = "langpack-v1"
|
||||
// v2 (pair-14): new pair/source files (title-formant, sentence-terminator, palladius-phonotactics, dc-checkers)
|
||||
// + new formats (the pattern/category rows, the generic Palladius parser) — a schema change, so the tag bumps.
|
||||
const packAlgoVersion = "langpack-v2"
|
||||
|
||||
// Pack is a loaded, versioned language-data pack for one source→target pair. Fields are the DATA the
|
||||
// pipeline algorithms read; the zero value is unusable (load via Load). Maps are membership sets / lookup
|
||||
|
|
@ -46,12 +49,25 @@ type Pack struct {
|
|||
GradePrefix map[rune]bool // grade/stem prefix chars 甲乙丙…
|
||||
Numeral map[rune]bool // CJK numeral chars
|
||||
AliasParticle map[rune]bool // trailing-particle set marking a boundary fragment
|
||||
// TitleFormant: the rank/measure formant chars a miner classifies as a TITLE (patterns.formant_type
|
||||
// 等/转/阶 → title). Pair-14: moved out of the pipeline formant switch, sitting beside TopoSuffix (its
|
||||
// place-side neighbour). A rune SET.
|
||||
TitleFormant map[rune]bool
|
||||
// SentenceTerminator: the source sentence-ending marks the alias miner splits on (。!?, alias.
|
||||
// cooccur_same_sentence). Pair-14: moved out of the pipeline literal. A rune SET; the miner adds "\n"
|
||||
// (a structural newline, not a language mark) itself.
|
||||
SentenceTerminator map[rune]bool
|
||||
|
||||
// <src>-<tgt> pair transliteration (configs/langpacks/<pair>/) — the Palladius (Палладий) table.
|
||||
PalladiusInitials map[string]string
|
||||
PalladiusFinals map[string]string
|
||||
PalladiusYW map[string]string
|
||||
PalladiusSpecialI map[string]string
|
||||
// Palladius is the <src>-<tgt> transliteration table (configs/langpacks/<pair>/palladius.txt +
|
||||
// palladius-phonotactics.txt) — the pinyin→Cyrillic maps and the syllable-generator's phonotactic
|
||||
// constraints, as ONE typed value (owner addendum 24.07: a struct, not four+three parallel Pack fields).
|
||||
Palladius Palladius
|
||||
|
||||
// DCCheckers is the OPTIONAL pair data for the WS5 defect-class checkers (configs/langpacks/<pair>/
|
||||
// dc-checkers.txt, checkers_zh_ru.go). nil when the pair ships no file — the checker ALGORITHMS then read
|
||||
// empty tables and fire 0 (a pair that does not opt in is never flagged). DATA only (pair-14 §6: the
|
||||
// LOOKUP TABLES move out of the pipeline; the regex DETECTION patterns stay as the checker algorithm).
|
||||
DCCheckers *DCCheckerData
|
||||
|
||||
// Heading is the OPTIONAL chapter-heading rule (configs/langpacks/<pair>/heading.txt). nil when the pair
|
||||
// carries no heading.txt — the chapter-title feature is then inert (the chunker keeps the source header
|
||||
|
|
@ -63,6 +79,37 @@ type Pack struct {
|
|||
version string
|
||||
}
|
||||
|
||||
// Palladius is the pair transliteration table: the pinyin→Cyrillic maps and the syllable-generator's
|
||||
// phonotactic constraint sets (owner addendum 24.07 — one typed struct instead of four+three parallel Pack
|
||||
// fields). Populated from palladius.txt (Initials/Finals/YW/SpecialI) + palladius-phonotactics.txt
|
||||
// (Retroflex/VFinal/VFinalInitial) by the generic category parser; which categories are REQUIRED is enforced
|
||||
// by validate(), not the parser (an unknown category is collected, never a parse error).
|
||||
type Palladius struct {
|
||||
Initials, Finals, YW, SpecialI map[string]string // pinyin→Cyrillic
|
||||
Retroflex, VFinal, VFinalInitial map[string]bool // phonotactic constraint sets (pinyin tokens)
|
||||
}
|
||||
|
||||
// newPalladius returns a Palladius with all tables allocated (so merge/union is nil-safe).
|
||||
func newPalladius() Palladius {
|
||||
return Palladius{
|
||||
Initials: map[string]string{}, Finals: map[string]string{}, YW: map[string]string{}, SpecialI: map[string]string{},
|
||||
Retroflex: map[string]bool{}, VFinal: map[string]bool{}, VFinalInitial: map[string]bool{},
|
||||
}
|
||||
}
|
||||
|
||||
// DCCheckerData is the per-pair lookup data the WS5 defect-class checkers read (checkers_zh_ru.go). The
|
||||
// checker DETECTION regexes stay in the pipeline as the pair-scoped algorithm (§12.2); only the LOOKUP
|
||||
// TABLES live here as data. Parsed from dc-checkers.txt.
|
||||
type DCCheckerData struct {
|
||||
Numeral map[rune]int // DC1: small CJK count → value (一→1 … 十→10, incl. 两→2)
|
||||
RuHours map[string]int // DC1: Russian hours-count word → value (один→1 … шесть→6)
|
||||
RegisterNeg []string // DC6: register negative-list lexemes (терем + case forms), authored order
|
||||
// Patterns are the DETECTION patterns as DATA (pair-14 data-out): key → verbatim regex or literal probe.
|
||||
// The `*_re` keys are compiled by the consumer; the rest are literal strings.Contains probes. A pair that
|
||||
// ships no pattern for a key runs that sub-checker inert.
|
||||
Patterns map[string]string
|
||||
}
|
||||
|
||||
// HeadingRule is the per-pair data for the chapter-title policy: instead of letting the model render a
|
||||
// chapter heading (which drifted to «Раздел 2» / «Первая глава» / an orphaned « :» across models), the
|
||||
// chunker detects a source header (Marker + a numeral + a Unit rune), strips it from the model input, and
|
||||
|
|
@ -78,13 +125,18 @@ type HeadingRule struct {
|
|||
func (p *Pack) Version() string { return p.version }
|
||||
|
||||
// srcFiles are the source-morphology files (under configs/langpacks/<src>/), read in this fixed order.
|
||||
// SCOPE (honest): this manifest is the zh-family NAME-MINER's morphology schema (百家姓 surnames, 甲乙丙 grade
|
||||
// prefixes, 等/转/阶 title-formants, CJK numerals, …), and validate() forbids an empty table. So "add a pair =
|
||||
// drop a directory, no recompile" holds for a source in THIS family (a second CJK source drops its files); a
|
||||
// different source family that wants mining needs new Pack fields + parse + miner channels, not just a dir.
|
||||
var srcFiles = []string{
|
||||
"surnames-single.txt", "surnames-compound.txt", "title-suffix.txt", "ordinal-title.txt",
|
||||
"rank-word.txt", "topo-suffix.txt", "grade-prefix.txt", "numeral.txt", "alias-particle.txt",
|
||||
"title-formant.txt", "sentence-terminator.txt",
|
||||
}
|
||||
|
||||
// pairFiles are the pair-transliteration files (under configs/langpacks/<pair>/).
|
||||
var pairFiles = []string{"palladius.txt"}
|
||||
var pairFiles = []string{"palladius.txt", "palladius-phonotactics.txt"}
|
||||
|
||||
// Load resolves and reads the pack for (sourceLang, targetLang) from root (e.g. "configs/langpacks"): the
|
||||
// source-morphology files under root/<src>/ and the pair-transliteration files under root/<src>-<tgt>/.
|
||||
|
|
@ -93,7 +145,7 @@ var pairFiles = []string{"palladius.txt"}
|
|||
// files in a fixed order, so it is deterministic and drift-proof.
|
||||
func Load(root, sourceLang, targetLang string) (*Pack, error) {
|
||||
pair := sourceLang + "-" + targetLang
|
||||
p := &Pack{Pair: pair}
|
||||
p := &Pack{Pair: pair, Palladius: newPalladius()}
|
||||
h := sha256.New()
|
||||
h.Write([]byte(packAlgoVersion))
|
||||
|
||||
|
|
@ -144,6 +196,21 @@ func Load(root, sourceLang, targetLang string) (*Pack, error) {
|
|||
p.Heading = hr
|
||||
}
|
||||
|
||||
// Optional per-pair DC-checker tables (pair-14 §6). ABSENT → nil, the checkers read empty tables and fire
|
||||
// 0 (a pair that does not opt in is never re-billed / flagged); PRESENT → bytes fold into the hash and it
|
||||
// is parsed; CORRUPT → fail loud. Same optional contract as heading.txt.
|
||||
if db, ok, derr := readOptional(root, pair, "dc-checkers.txt"); derr != nil {
|
||||
return nil, fmt.Errorf("langpack %q dc-checkers.txt: %w", pair, derr)
|
||||
} else if ok {
|
||||
h.Write([]byte("\x00" + pair + "/dc-checkers.txt\x00"))
|
||||
h.Write(db)
|
||||
dc, perr := parseDCCheckers(db)
|
||||
if perr != nil {
|
||||
return nil, fmt.Errorf("langpack %q dc-checkers.txt: %w", pair, perr)
|
||||
}
|
||||
p.DCCheckers = dc
|
||||
}
|
||||
|
||||
if err := p.validate(); err != nil {
|
||||
return nil, fmt.Errorf("langpack %q: %w", pair, err)
|
||||
}
|
||||
|
|
@ -151,6 +218,138 @@ func Load(root, sourceLang, targetLang string) (*Pack, error) {
|
|||
return p, nil
|
||||
}
|
||||
|
||||
// LoadWithOverlay loads the shared pair pack (Load) and then UNIONS a book-scoped OVERLAY on top: a book's
|
||||
// PRIVATE canon (a clan surname 古月, a sect term) that belongs to ONE book, not the shared pair langpack
|
||||
// (pair-14 §1 — a book term in the shared pair layer is a leak). overlayRoot has the same layout as root
|
||||
// (<src>/… + <pair>/…); ONLY the files a book chooses to ship are present, each read OPTIONALLY and UNIONED
|
||||
// (additive — sets gain members, ordered slices append; an overlay never removes a base entry). The overlay
|
||||
// bytes fold into Version(), so a book-canon edit is a loud --resnapshot for THAT book while the shared pack
|
||||
// stays byte-stable. overlayRoot == "" ⇒ identical to Load (same *Pack, same Version).
|
||||
func LoadWithOverlay(root, src, tgt, overlayRoot string) (*Pack, error) {
|
||||
p, err := Load(root, src, tgt)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if overlayRoot == "" {
|
||||
return p, nil
|
||||
}
|
||||
pair := src + "-" + tgt
|
||||
// FAIL LOUD on a misnamed overlay file (review finding, pair-14 scale lens): the fold loop below reads
|
||||
// ONLY the manifest names, so a typo (surname-compound.txt) or a non-mergeable file (heading.txt) shipped
|
||||
// in an overlay would be SILENTLY ignored — the book's private canon never reaches the miner, recall
|
||||
// degrades with no load-time signal. Enumerate the overlay dirs and refuse any unexpected file, keeping
|
||||
// the "never silently empty / drift-proof" guarantee the shared loader makes.
|
||||
mergeable := map[string]bool{}
|
||||
for _, n := range srcFiles {
|
||||
mergeable[src+"/"+n] = true
|
||||
}
|
||||
for _, n := range pairFiles {
|
||||
mergeable[pair+"/"+n] = true
|
||||
}
|
||||
for _, dir := range []string{src, pair} {
|
||||
names, derr := overlayDirFiles(filepath.Join(overlayRoot, dir))
|
||||
if derr != nil {
|
||||
return nil, fmt.Errorf("langpack %q overlay %s: %w", pair, dir, derr)
|
||||
}
|
||||
for _, n := range names {
|
||||
if !mergeable[dir+"/"+n] {
|
||||
return nil, fmt.Errorf("langpack %q overlay: unexpected file %s/%s — an overlay merges only the source/pair manifest files (a typo, or a non-mergeable file like heading.txt/dc-checkers.txt, would be silently ignored)", pair, dir, n)
|
||||
}
|
||||
}
|
||||
}
|
||||
// Seed a fresh hash with the base version (which already uniquely encodes every base byte), then fold the
|
||||
// overlay files in a fixed order — deterministic and drift-proof (edit the overlay → Version() moves).
|
||||
h := sha256.New()
|
||||
h.Write([]byte(p.version))
|
||||
merged := false
|
||||
fold := func(dir, name string, pairFile bool) error {
|
||||
b, ok, rerr := readOptional(overlayRoot, dir, name)
|
||||
if rerr != nil {
|
||||
return fmt.Errorf("langpack %q overlay %s/%s: %w", pair, dir, name, rerr)
|
||||
}
|
||||
if !ok {
|
||||
return nil
|
||||
}
|
||||
h.Write([]byte("\x00" + dir + "/" + name + "\x00"))
|
||||
h.Write(b)
|
||||
merged = true
|
||||
if pairFile {
|
||||
return p.mergePair(name, b)
|
||||
}
|
||||
return p.mergeSrc(name, b)
|
||||
}
|
||||
for _, name := range srcFiles {
|
||||
if err := fold(src, name, false); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
for _, name := range pairFiles {
|
||||
if err := fold(pair, name, true); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
if merged {
|
||||
p.version = packAlgoVersion + "-x" + hex.EncodeToString(h.Sum(nil))[:12]
|
||||
}
|
||||
return p, nil
|
||||
}
|
||||
|
||||
// mergeSrc unions an overlay source file into the loaded pack (additive; see LoadWithOverlay). Rune/string
|
||||
// SETS gain members; ordered slices append (a book's extra title/rank tokens follow the shared ones).
|
||||
func (p *Pack) mergeSrc(name string, b []byte) error {
|
||||
switch name {
|
||||
case "surnames-single.txt":
|
||||
unionRuneSet(p.SurnamesSingle, runeSet(b))
|
||||
case "surnames-compound.txt":
|
||||
unionStringSet(p.SurnamesCompound, stringSet(b))
|
||||
case "title-suffix.txt":
|
||||
p.TitleSuffix = append(p.TitleSuffix, lines(b)...)
|
||||
case "ordinal-title.txt":
|
||||
p.OrdinalTitle = append(p.OrdinalTitle, lines(b)...)
|
||||
case "rank-word.txt":
|
||||
p.RankWord = append(p.RankWord, lines(b)...)
|
||||
case "topo-suffix.txt":
|
||||
unionRuneSet(p.TopoSuffix, runeSet(b))
|
||||
case "grade-prefix.txt":
|
||||
unionRuneSet(p.GradePrefix, runeSet(b))
|
||||
case "numeral.txt":
|
||||
unionRuneSet(p.Numeral, runeSet(b))
|
||||
case "alias-particle.txt":
|
||||
unionRuneSet(p.AliasParticle, runeSet(b))
|
||||
case "title-formant.txt":
|
||||
unionRuneSet(p.TitleFormant, runeSet(b))
|
||||
case "sentence-terminator.txt":
|
||||
unionRuneSet(p.SentenceTerminator, runeSet(b))
|
||||
default:
|
||||
return fmt.Errorf("unknown source file")
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// mergePair unions an overlay pair file into the loaded pack (additive). Both pair files carry Palladius
|
||||
// categories, parsed generically and unioned into p.Palladius (same path as assignPair).
|
||||
func (p *Pack) mergePair(name string, b []byte) error {
|
||||
return p.assignPair(name, b)
|
||||
}
|
||||
|
||||
func unionRuneSet(dst, src map[rune]bool) {
|
||||
for k := range src {
|
||||
dst[k] = true
|
||||
}
|
||||
}
|
||||
|
||||
func unionStringSet(dst, src map[string]bool) {
|
||||
for k := range src {
|
||||
dst[k] = true
|
||||
}
|
||||
}
|
||||
|
||||
func unionStringMap(dst, src map[string]string) {
|
||||
for k, v := range src {
|
||||
dst[k] = v
|
||||
}
|
||||
}
|
||||
|
||||
// validate makes the "never silently empty" contract real: a present-but-empty or comment-only data file
|
||||
// (a fat-fingered edit once packs are hand-authored at R1) parses to an empty table with no error and would
|
||||
// silently disable a miner channel — recall degradation with no load-time signal. Every required table must
|
||||
|
|
@ -172,10 +371,16 @@ func (p *Pack) validate() error {
|
|||
req("grade-prefix", len(p.GradePrefix))
|
||||
req("numeral", len(p.Numeral))
|
||||
req("alias-particle", len(p.AliasParticle))
|
||||
req("palladius/initials", len(p.PalladiusInitials))
|
||||
req("palladius/finals", len(p.PalladiusFinals))
|
||||
req("palladius/yw", len(p.PalladiusYW))
|
||||
req("palladius/special_i", len(p.PalladiusSpecialI))
|
||||
req("title-formant", len(p.TitleFormant))
|
||||
req("sentence-terminator", len(p.SentenceTerminator))
|
||||
// The Palladius REQUIRED-category list lives HERE (the consumer), not in the generic parser (addendum).
|
||||
req("palladius/initials", len(p.Palladius.Initials))
|
||||
req("palladius/finals", len(p.Palladius.Finals))
|
||||
req("palladius/yw", len(p.Palladius.YW))
|
||||
req("palladius/special_i", len(p.Palladius.SpecialI))
|
||||
req("palladius-phonotactics/retroflex", len(p.Palladius.Retroflex))
|
||||
req("palladius-phonotactics/vfinal", len(p.Palladius.VFinal))
|
||||
req("palladius-phonotactics/vfinal_initial", len(p.Palladius.VFinalInitial))
|
||||
if len(empty) > 0 {
|
||||
return fmt.Errorf("empty required table(s) %s — a present-but-empty/comment-only data file is a corrupt pack, not a valid one", strings.Join(empty, ", "))
|
||||
}
|
||||
|
|
@ -202,26 +407,49 @@ func (p *Pack) assignSrc(name string, b []byte) error {
|
|||
p.Numeral = runeSet(b)
|
||||
case "alias-particle.txt":
|
||||
p.AliasParticle = runeSet(b)
|
||||
case "title-formant.txt":
|
||||
p.TitleFormant = runeSet(b)
|
||||
case "sentence-terminator.txt":
|
||||
p.SentenceTerminator = runeSet(b)
|
||||
default:
|
||||
return fmt.Errorf("unknown source file")
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// assignPair reads a pair file into p.Palladius. Both pair files (palladius.txt, palladius-phonotactics.txt)
|
||||
// carry `category<TAB>…` rows, parsed by ONE generic category parser (owner addendum 24.07); the known
|
||||
// Palladius categories are UNIONED into p.Palladius and an unknown category is simply ignored, NEVER a parse
|
||||
// error — which categories are REQUIRED is enforced downstream by validate(), not here.
|
||||
func (p *Pack) assignPair(name string, b []byte) error {
|
||||
switch name {
|
||||
case "palladius.txt":
|
||||
ini, fin, yw, si, err := parsePalladius(b)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
p.PalladiusInitials, p.PalladiusFinals, p.PalladiusYW, p.PalladiusSpecialI = ini, fin, yw, si
|
||||
default:
|
||||
return fmt.Errorf("unknown pair file")
|
||||
cats, err := parseCategoryRows(b)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
p.Palladius.merge(cats)
|
||||
return nil
|
||||
}
|
||||
|
||||
// overlayDirFiles lists the regular-file names directly in dir (an overlay's <src> or <pair> subdir). A
|
||||
// MISSING dir is fine (returns nil — a book may overlay only sources or only the pair). Any other read
|
||||
// error is loud. Nested dirs are ignored (only top-level manifest files are mergeable).
|
||||
func overlayDirFiles(dir string) ([]string, error) {
|
||||
ents, err := os.ReadDir(dir)
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
return nil, nil
|
||||
}
|
||||
return nil, err
|
||||
}
|
||||
var out []string
|
||||
for _, e := range ents {
|
||||
if !e.IsDir() {
|
||||
out = append(out, e.Name())
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// readOptional reads an OPTIONAL pack file. A missing file returns (nil, false, nil) — the feature it
|
||||
// backs is simply inert — while any OTHER read error (permission, a directory) is a loud failure; a
|
||||
// present file returns (bytes, true, nil). Used for the pack-13 heading rule, which a pair opts into.
|
||||
|
|
@ -314,34 +542,116 @@ func lines(b []byte) []string {
|
|||
return out
|
||||
}
|
||||
|
||||
// parsePalladius reads the `category<TAB>pinyin<TAB>cyrillic` table into the four maps. It scans the raw
|
||||
// lines directly (not contentLines) so an error names the PHYSICAL file line — the point of the diagnostic
|
||||
// is to send a human editing the table to the right line.
|
||||
func parsePalladius(b []byte) (ini, fin, yw, si map[string]string, err error) {
|
||||
ini, fin, yw, si = map[string]string{}, map[string]string{}, map[string]string{}, map[string]string{}
|
||||
// parseCategoryRows is the GENERIC Palladius category parser (owner addendum 24.07): it reads a pair file's
|
||||
// `category<TAB>key[<TAB>value]` rows into category → key → value, WITHOUT a per-category switch and WITHOUT
|
||||
// treating an unknown category as an error (which categories are required is the CONSUMER's call — validate()).
|
||||
// A 3-field row (initials/finals/yw/special_i) stores key→cyrillic; a 2-field row (retroflex/vfinal/…) stores
|
||||
// key→"" (a set member). Scans raw lines so an error names the PHYSICAL file line; the only parse errors are a
|
||||
// bad field count / an empty key.
|
||||
func parseCategoryRows(b []byte) (map[string]map[string]string, error) {
|
||||
cats := map[string]map[string]string{}
|
||||
for i, raw := range strings.Split(string(b), "\n") {
|
||||
t := strings.TrimSpace(strings.TrimRight(raw, "\r"))
|
||||
if t == "" || strings.HasPrefix(t, "#") {
|
||||
continue
|
||||
}
|
||||
f := strings.Split(t, "\t")
|
||||
if len(f) != 3 {
|
||||
return nil, nil, nil, nil, fmt.Errorf("line %d: want 3 tab-separated fields, got %d (%q)", i+1, len(f), t)
|
||||
if len(f) < 2 || len(f) > 3 || strings.TrimSpace(f[0]) == "" || strings.TrimSpace(f[1]) == "" {
|
||||
return nil, fmt.Errorf("line %d: want `category<TAB>key[<TAB>value]`, got %q", i+1, t)
|
||||
}
|
||||
if cats[f[0]] == nil {
|
||||
cats[f[0]] = map[string]string{}
|
||||
}
|
||||
val := "" // a 2-field row is a set member (value "")
|
||||
if len(f) == 3 {
|
||||
val = f[2]
|
||||
}
|
||||
cats[f[0]][f[1]] = val
|
||||
}
|
||||
return cats, nil
|
||||
}
|
||||
|
||||
// merge unions parsed categories into the Palladius table. KNOWN categories populate the typed fields (a map
|
||||
// category takes key→cyrillic, a set category takes its keys); an UNKNOWN category is simply not consumed —
|
||||
// validate() enforces that every REQUIRED category ended up non-empty.
|
||||
func (pal *Palladius) merge(cats map[string]map[string]string) {
|
||||
unionStringMap(pal.Initials, cats["initials"])
|
||||
unionStringMap(pal.Finals, cats["finals"])
|
||||
unionStringMap(pal.YW, cats["yw"])
|
||||
unionStringMap(pal.SpecialI, cats["special_i"])
|
||||
unionKeysAsSet(pal.Retroflex, cats["retroflex"])
|
||||
unionKeysAsSet(pal.VFinal, cats["vfinal"])
|
||||
unionKeysAsSet(pal.VFinalInitial, cats["vfinal_initial"])
|
||||
}
|
||||
|
||||
// unionKeysAsSet adds the KEYS of a parsed category (a 2-field set) to dst.
|
||||
func unionKeysAsSet(dst map[string]bool, src map[string]string) {
|
||||
for k := range src {
|
||||
dst[k] = true
|
||||
}
|
||||
}
|
||||
|
||||
// parseDCCheckers reads the DC-checker pair tables from category-keyed lines (checkers_zh_ru.go data):
|
||||
//
|
||||
// cjk_numeral<TAB>rune<TAB>value | ru_hour<TAB>word<TAB>value | register_neg<TAB>lexeme
|
||||
//
|
||||
// Scans raw lines so an error names the PHYSICAL file line. Fail-loud on a bad field count / non-integer
|
||||
// value / unknown category (a malformed table is a corrupt pack). RegisterNeg keeps its authored order.
|
||||
func parseDCCheckers(b []byte) (*DCCheckerData, error) {
|
||||
d := &DCCheckerData{Numeral: map[rune]int{}, RuHours: map[string]int{}, Patterns: map[string]string{}}
|
||||
for i, raw := range strings.Split(string(b), "\n") {
|
||||
// A `pattern` line's VALUE is verbatim (a regex may carry trailing metachars), so only \r is stripped
|
||||
// from the whole line, not the value; other categories tolerate the trimmed form.
|
||||
line := strings.TrimRight(raw, "\r")
|
||||
t := strings.TrimSpace(line)
|
||||
if t == "" || strings.HasPrefix(t, "#") {
|
||||
continue
|
||||
}
|
||||
if strings.HasPrefix(line, "pattern\t") {
|
||||
f := strings.SplitN(line, "\t", 3) // value (f[2]) VERBATIM
|
||||
if len(f) != 3 || strings.TrimSpace(f[1]) == "" || f[2] == "" {
|
||||
return nil, fmt.Errorf("line %d: pattern wants `pattern<TAB>key<TAB>value`, got %q", i+1, line)
|
||||
}
|
||||
d.Patterns[f[1]] = f[2]
|
||||
continue
|
||||
}
|
||||
f := strings.Split(t, "\t")
|
||||
switch f[0] {
|
||||
case "initials":
|
||||
ini[f[1]] = f[2]
|
||||
case "finals":
|
||||
fin[f[1]] = f[2]
|
||||
case "yw":
|
||||
yw[f[1]] = f[2]
|
||||
case "special_i":
|
||||
si[f[1]] = f[2]
|
||||
case "cjk_numeral":
|
||||
if len(f) != 3 {
|
||||
return nil, fmt.Errorf("line %d: cjk_numeral wants `cjk_numeral<TAB>rune<TAB>value`, got %q", i+1, t)
|
||||
}
|
||||
r := []rune(f[1])
|
||||
if len(r) != 1 {
|
||||
return nil, fmt.Errorf("line %d: cjk_numeral key %q must be a single rune", i+1, f[1])
|
||||
}
|
||||
v, err := strconv.Atoi(strings.TrimSpace(f[2]))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("line %d: cjk_numeral value %q: %w", i+1, f[2], err)
|
||||
}
|
||||
d.Numeral[r[0]] = v
|
||||
case "ru_hour":
|
||||
if len(f) != 3 {
|
||||
return nil, fmt.Errorf("line %d: ru_hour wants `ru_hour<TAB>word<TAB>value`, got %q", i+1, t)
|
||||
}
|
||||
v, err := strconv.Atoi(strings.TrimSpace(f[2]))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("line %d: ru_hour value %q: %w", i+1, f[2], err)
|
||||
}
|
||||
d.RuHours[f[1]] = v
|
||||
case "register_neg":
|
||||
if len(f) != 2 || strings.TrimSpace(f[1]) == "" {
|
||||
return nil, fmt.Errorf("line %d: register_neg wants `register_neg<TAB>lexeme`, got %q", i+1, t)
|
||||
}
|
||||
d.RegisterNeg = append(d.RegisterNeg, f[1])
|
||||
default:
|
||||
return nil, nil, nil, nil, fmt.Errorf("line %d: unknown category %q", i+1, f[0])
|
||||
return nil, fmt.Errorf("line %d: unknown category %q (want cjk_numeral|ru_hour|register_neg)", i+1, f[0])
|
||||
}
|
||||
}
|
||||
return ini, fin, yw, si, nil
|
||||
if len(d.Numeral) == 0 || len(d.RuHours) == 0 || len(d.RegisterNeg) == 0 {
|
||||
return nil, fmt.Errorf("dc-checkers needs non-empty cjk_numeral, ru_hour and register_neg sections")
|
||||
}
|
||||
return d, nil
|
||||
}
|
||||
|
||||
// contentLines splits into lines, dropping '#'-comment and blank lines.
|
||||
|
|
|
|||
|
|
@ -29,8 +29,11 @@ func TestLoadResolvesRealZhRu(t *testing.T) {
|
|||
if p.SurnamesSingle['凝'] {
|
||||
t.Error("surnames-single must NOT contain 凝 (discarded — not a surname)")
|
||||
}
|
||||
if !p.SurnamesCompound["古月"] || !p.SurnamesCompound["欧阳"] {
|
||||
t.Error("surnames-compound missing 古月/欧阳")
|
||||
if !p.SurnamesCompound["欧阳"] {
|
||||
t.Error("surnames-compound missing 欧阳")
|
||||
}
|
||||
if p.SurnamesCompound["古月"] {
|
||||
t.Error("surnames-compound must NOT contain 古月 (pair-14 §1: a 蛊真人 clan, moved to the book overlay)")
|
||||
}
|
||||
if len(p.TitleSuffix) == 0 || p.TitleSuffix[0] != "公子" {
|
||||
t.Errorf("title-suffix order not preserved: %v", p.TitleSuffix)
|
||||
|
|
@ -38,12 +41,22 @@ func TestLoadResolvesRealZhRu(t *testing.T) {
|
|||
if !p.TopoSuffix['山'] || !p.Numeral['三'] || !p.GradePrefix['甲'] || !p.AliasParticle['的'] {
|
||||
t.Error("rune-set membership missing an expected char (山/三/甲/的)")
|
||||
}
|
||||
// Pair transliteration spot-check.
|
||||
if p.PalladiusInitials["b"] != "б" || p.PalladiusInitials["zh"] != "чж" {
|
||||
t.Errorf("palladius initials wrong: b=%q zh=%q", p.PalladiusInitials["b"], p.PalladiusInitials["zh"])
|
||||
// pair-14 moves: title-formant + sentence-terminator rune sets, and the Palladius phonotactic sets.
|
||||
if !p.TitleFormant['等'] || !p.TitleFormant['转'] || !p.TitleFormant['阶'] {
|
||||
t.Error("title-formant missing 等/转/阶")
|
||||
}
|
||||
if p.PalladiusSpecialI["zhi"] != "чжи" {
|
||||
t.Errorf("palladius special_i zhi = %q, want чжи", p.PalladiusSpecialI["zhi"])
|
||||
if !p.SentenceTerminator['。'] || !p.SentenceTerminator['!'] || !p.SentenceTerminator['?'] {
|
||||
t.Error("sentence-terminator missing 。/!/?")
|
||||
}
|
||||
if !p.Palladius.Retroflex["zh"] || !p.Palladius.VFinal["v"] || !p.Palladius.VFinalInitial["j"] {
|
||||
t.Error("palladius phonotactics missing retroflex zh / vfinal v / vfinal_initial j")
|
||||
}
|
||||
// Pair transliteration spot-check.
|
||||
if p.Palladius.Initials["b"] != "б" || p.Palladius.Initials["zh"] != "чж" {
|
||||
t.Errorf("palladius initials wrong: b=%q zh=%q", p.Palladius.Initials["b"], p.Palladius.Initials["zh"])
|
||||
}
|
||||
if p.Palladius.SpecialI["zhi"] != "чжи" {
|
||||
t.Errorf("palladius special_i zhi = %q, want чжи", p.Palladius.SpecialI["zhi"])
|
||||
}
|
||||
if !strings.HasPrefix(p.Version(), packAlgoVersion+"-") {
|
||||
t.Errorf("Version() = %q, want %s-<hash>", p.Version(), packAlgoVersion)
|
||||
|
|
@ -82,6 +95,63 @@ func TestRoutesByPairToDifferentBytes(t *testing.T) {
|
|||
}
|
||||
}
|
||||
|
||||
// TestBookOverlayUnionsAndReVersions pins the pair-14 §1 book-overlay contract: an overlay UNIONS its
|
||||
// private canon onto the shared pack (additive — the base entries survive, the overlay entry is added) and
|
||||
// SHIFTS Version() (a book-canon edit is a loud --resnapshot for that book), while a no-overlay Load is
|
||||
// byte-identical to before. Guards the exact mechanism the miner's {古月:22} parity rides.
|
||||
func TestBookOverlayUnionsAndReVersions(t *testing.T) {
|
||||
base, err := Load(realRoot, "zh", "ru")
|
||||
if err != nil {
|
||||
t.Fatalf("load base: %v", err)
|
||||
}
|
||||
if base.SurnamesCompound["古月"] {
|
||||
t.Fatal("shared pack must not carry 古月 (it is a book clan)")
|
||||
}
|
||||
overlay := t.TempDir()
|
||||
if err := os.MkdirAll(filepath.Join(overlay, "zh"), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(overlay, "zh", "surnames-compound.txt"), []byte("# book canon\n古月\n"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
ext, err := LoadWithOverlay(realRoot, "zh", "ru", overlay)
|
||||
if err != nil {
|
||||
t.Fatalf("load with overlay: %v", err)
|
||||
}
|
||||
if !ext.SurnamesCompound["古月"] {
|
||||
t.Error("overlay must add 古月 to the effective pack")
|
||||
}
|
||||
if !ext.SurnamesCompound["欧阳"] {
|
||||
t.Error("overlay must UNION (keep the shared 欧阳), not replace")
|
||||
}
|
||||
if ext.Version() == base.Version() {
|
||||
t.Errorf("overlay must shift Version() (base=%s overlay=%s)", base.Version(), ext.Version())
|
||||
}
|
||||
// An empty overlayRoot is exactly Load — same version, no re-hash.
|
||||
same, err := LoadWithOverlay(realRoot, "zh", "ru", "")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if same.Version() != base.Version() {
|
||||
t.Errorf("empty overlay must equal Load (%s != %s)", same.Version(), base.Version())
|
||||
}
|
||||
|
||||
// A MISNAMED overlay file must FAIL LOUD, not be silently ignored (review finding, scale lens): a typo'd
|
||||
// canon file that never reaches the miner would degrade recall with no signal.
|
||||
bad := t.TempDir()
|
||||
if err := os.MkdirAll(filepath.Join(bad, "zh"), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(bad, "zh", "surname-compound.txt"), []byte("古月\n"), 0o644); err != nil { // typo: missing 's'
|
||||
t.Fatal(err)
|
||||
}
|
||||
if _, err := LoadWithOverlay(realRoot, "zh", "ru", bad); err == nil {
|
||||
t.Error("a misnamed overlay file must fail loud (silently-ignored canon would degrade recall)")
|
||||
} else if !strings.Contains(err.Error(), "surname-compound.txt") {
|
||||
t.Errorf("error must name the unexpected file, got: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestFailsLoudOnMissingPair pins the fail-loud contract: a pair with no pack directory errors, naming the
|
||||
// missing file — never a silent empty pack (mirrors the prompt seam's PromptPathFor fail-loud). This is
|
||||
// the load-time invariant a live consumer relies on (fail before billing).
|
||||
|
|
@ -120,16 +190,19 @@ func writeSyntheticPack(t *testing.T, root, src, tgt string) {
|
|||
t.Helper()
|
||||
pair := src + "-" + tgt
|
||||
files := map[string]string{
|
||||
filepath.Join(src, "surnames-single.txt"): "# synthetic\n甴甶甹\n",
|
||||
filepath.Join(src, "surnames-compound.txt"): "# synthetic\n甲乙\n",
|
||||
filepath.Join(src, "title-suffix.txt"): "# synthetic\n阁下\n",
|
||||
filepath.Join(src, "ordinal-title.txt"): "# synthetic\n第甲\n",
|
||||
filepath.Join(src, "rank-word.txt"): "# synthetic\n級\n",
|
||||
filepath.Join(src, "topo-suffix.txt"): "# synthetic\n峰\n",
|
||||
filepath.Join(src, "grade-prefix.txt"): "# synthetic\n子丑\n",
|
||||
filepath.Join(src, "numeral.txt"): "# synthetic\n壹貳\n",
|
||||
filepath.Join(src, "alias-particle.txt"): "# synthetic\n之乎\n",
|
||||
filepath.Join(pair, "palladius.txt"): "# synthetic\ninitials\tb\tб\nfinals\ta\tа\nyw\tyi\tи\nspecial_i\tzhi\tчжи\n",
|
||||
filepath.Join(src, "surnames-single.txt"): "# synthetic\n甴甶甹\n",
|
||||
filepath.Join(src, "surnames-compound.txt"): "# synthetic\n甲乙\n",
|
||||
filepath.Join(src, "title-suffix.txt"): "# synthetic\n阁下\n",
|
||||
filepath.Join(src, "ordinal-title.txt"): "# synthetic\n第甲\n",
|
||||
filepath.Join(src, "rank-word.txt"): "# synthetic\n級\n",
|
||||
filepath.Join(src, "topo-suffix.txt"): "# synthetic\n峰\n",
|
||||
filepath.Join(src, "grade-prefix.txt"): "# synthetic\n子丑\n",
|
||||
filepath.Join(src, "numeral.txt"): "# synthetic\n壹貳\n",
|
||||
filepath.Join(src, "alias-particle.txt"): "# synthetic\n之乎\n",
|
||||
filepath.Join(src, "title-formant.txt"): "# synthetic\n甼\n",
|
||||
filepath.Join(src, "sentence-terminator.txt"): "# synthetic\n。\n",
|
||||
filepath.Join(pair, "palladius.txt"): "# synthetic\ninitials\tb\tб\nfinals\ta\tа\nyw\tyi\tи\nspecial_i\tzhi\tчжи\n",
|
||||
filepath.Join(pair, "palladius-phonotactics.txt"): "# synthetic\nretroflex\tzh\nvfinal\tv\nvfinal_initial\tj\n",
|
||||
}
|
||||
for rel, body := range files {
|
||||
p := filepath.Join(root, rel)
|
||||
|
|
|
|||
|
|
@ -5,6 +5,8 @@ import (
|
|||
"sort"
|
||||
"strings"
|
||||
"unicode"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
)
|
||||
|
||||
// cheapgates.go: four cheap, deterministic post-check flaggers on the FINAL chunk text (04-unhappy
|
||||
|
|
@ -64,6 +66,11 @@ type cheapGateConfig struct {
|
|||
// OPT-IN observability flagger (draft→final length collapse + number drift) folded into this
|
||||
// result. Off → the two regression fields stay 0 and the output is byte-identical to before.
|
||||
regressionEnabled bool
|
||||
// checkers is the compiled WS5/pack-13 checker spec (pair-14 data-out): the pair's DETECTION patterns +
|
||||
// lookup tables (from the pair langpack) plus the target-general lists (from embedded target data),
|
||||
// resolved once per run. nil for a bare config or a book with no data → the language-specific checkers
|
||||
// run inert (fire 0, the no-pack golden path); the general Latin-residue check needs no spec.
|
||||
checkers *dcCheckers
|
||||
}
|
||||
|
||||
// cheapGateResult is the per-chunk outcome: a count per flagger plus human-readable detail lines
|
||||
|
|
@ -103,33 +110,35 @@ func runCheapGates(source, draft, final string, cfg cheapGateConfig) cheapGateRe
|
|||
var r cheapGateResult
|
||||
n, det := lintDialogueDash(final)
|
||||
r.DialogueDash, r.Detail = n, append(r.Detail, det...)
|
||||
n, det = lintYofikation(final, cfg.yoPolicy)
|
||||
n, det = cfg.checkers.lintYofikation(final, cfg.yoPolicy)
|
||||
r.YoInconsistent = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
n, det = lintTranslitInterjections(final, cfg.allowlist)
|
||||
r.TranslitInterj = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
n, det = lintNumberMagnitude(source, final)
|
||||
n, det = cfg.checkers.lintNumberMagnitude(source, final)
|
||||
r.NumberMagnitude = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
// WS5 defect-class checkers (DC1/DC2/DC6, checkers_zh_ru.go) — src↔target observability flaggers.
|
||||
n, det = lintTimeUnits(source, final)
|
||||
// WS5 defect-class checkers (DC1/DC2/DC6, checkers.go) — src↔target observability flaggers. All patterns
|
||||
// + tables are pair langpack DATA (pair-14 data-out); the spec is nil/inert for a no-pack book → fire 0.
|
||||
n, det = cfg.checkers.lintTimeUnits(source, final)
|
||||
r.DC1TimeUnits = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
n, det = lintMagnitudeScale(source, final)
|
||||
n, det = cfg.checkers.lintMagnitudeScale(source, final)
|
||||
r.DC2Magnitude = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
n, det = lintRegisterLexicon(final)
|
||||
n, det = cfg.checkers.lintRegisterLexicon(final)
|
||||
r.DC6Register = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
// pack-13 general checkers (checkers_zh_ru.go) — src↔target / Russian-side observability flaggers.
|
||||
n, det = lintPercentScale(source, final)
|
||||
// pack-13 general checkers (checkers.go): percent scale (pair data), Latin residue (language-general),
|
||||
// broken word (target data).
|
||||
n, det = cfg.checkers.lintPercentScale(source, final)
|
||||
r.PercentScale = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
n, det = lintLatinResidue(final, cfg.allowlist)
|
||||
r.LatinResidue = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
n, det = lintBrokenWord(final)
|
||||
n, det = cfg.checkers.lintBrokenWord(final)
|
||||
r.BrokenWord = n
|
||||
r.Detail = append(r.Detail, det...)
|
||||
if cfg.regressionEnabled {
|
||||
|
|
@ -248,33 +257,16 @@ func quoteShape(after []rune) bool {
|
|||
|
||||
// --- 2. yofikator ---------------------------------------------------------------
|
||||
|
||||
// yoHomographEForms are е-spellings that are DISTINCT words from their ё-counterpart (все≠всё,
|
||||
// небо≠нёбо, берет≠берёт, осел≠осёл, …). The inconsistency rule below would otherwise false-flag
|
||||
// «все»+«всё» as one word spelled two ways. Expanded (self-review major) to cover the frequent
|
||||
// distinct-word pairs and their common inflections; still not exhaustive — full disambiguation
|
||||
// needs a ё-dictionary (deferred, B-tier). Names — the primary target (Пётр/Петр) — are never
|
||||
// homographs, so they are always caught regardless.
|
||||
var yoHomographEForms = map[string]bool{
|
||||
"все": true, "всех": true, "всем": true, "всеми": true,
|
||||
"небо": true, "небом": true,
|
||||
"узнаем": true, "узнаете": true, "узнает": true,
|
||||
"падеж": true, "падежа": true,
|
||||
"совершенный": true, "совершенное": true, "совершенная": true, "совершенно": true, "совершенны": true,
|
||||
"чем": true, "тем": true, "тема": true, "теме": true, "темы": true,
|
||||
"берет": true, "берете": true, "берета": true,
|
||||
"осел": true, "осла": true, "ослы": true, "ослов": true,
|
||||
"мел": true, "мела": true,
|
||||
"лен": true, "лена": true,
|
||||
"нес": true, "несем": true, "несет": true,
|
||||
"вел": true, "ведет": true, "ведем": true,
|
||||
}
|
||||
|
||||
// lintYofikation flags inconsistent ё. Default ("auto"): the same word appears BOTH with ё and, as
|
||||
// a separate token, with its exact ё→е form (Пётр/Петр) — excluding the homograph traps above.
|
||||
// Policy "all-e": any ё present is a violation (the project wants no ё). Policy "all-yo": an е-form
|
||||
// of a word that ALSO appears somewhere with ё is flagged (the partial-yofikation case), same signal
|
||||
// as auto — full "every word that SHOULD have ё" enforcement needs a ё-dictionary (deferred, B-tier).
|
||||
func lintYofikation(text, policy string) (int, []string) {
|
||||
// a separate token, with its exact ё→е form (Пётр/Петр) — excluding the target's homograph whitelist
|
||||
// (е-spellings that are DISTINCT words from their ё-counterpart, все≠всё …; TARGET data, lang.TargetChecks
|
||||
// "yo_homograph" — pair-14 data-out, so the ё↔е FOLD stays as generic orthography and only the wordlist is
|
||||
// data). Policy "all-e": any ё present is a violation. Policy "all-yo": same signal as auto. Full "every
|
||||
// word that SHOULD have ё" enforcement needs a ё-dictionary (deferred, B-tier). Inert if no target data.
|
||||
func (c *dcCheckers) lintYofikation(text, policy string) (int, []string) {
|
||||
if c == nil {
|
||||
return 0, nil
|
||||
}
|
||||
words := tokenizeCyrillic(text)
|
||||
if policy == "all-e" {
|
||||
seen := map[string]bool{}
|
||||
|
|
@ -303,7 +295,7 @@ func lintYofikation(text, policy string) (int, []string) {
|
|||
continue
|
||||
}
|
||||
eForm := strings.ReplaceAll(w, "ё", "е")
|
||||
if eForm == w || yoHomographEForms[eForm] {
|
||||
if eForm == w || c.yoHomograph[eForm] {
|
||||
continue // no ё, or a distinct-word homograph (все/всё) — not an inconsistency
|
||||
}
|
||||
if present[eForm] && !seenPair[w] {
|
||||
|
|
@ -404,7 +396,10 @@ func isWordRune(r rune) bool {
|
|||
// range for their possible multiplier, plus Arabic-number orders). It fires only when the source's
|
||||
// TOP order is outside every output range — a conservative, multiplier-tolerant signal that leaves
|
||||
// 三億→«триста миллионов» (8 within миллион's [6,8]) silent while catching 三万→«три миллиона».
|
||||
func lintNumberMagnitude(source, final string) (int, []string) {
|
||||
func (c *dcCheckers) lintNumberMagnitude(source, final string) (int, []string) {
|
||||
if c == nil || len(c.magnitudeStem) == 0 {
|
||||
return 0, nil // no target magnitude-word data → the gate can never confirm coverage; stay inert
|
||||
}
|
||||
srcOrders := cjkMagnitudeOrders(source)
|
||||
if len(srcOrders) == 0 {
|
||||
return 0, nil
|
||||
|
|
@ -427,7 +422,7 @@ func lintNumberMagnitude(source, final string) (int, []string) {
|
|||
return 0, nil // an Arabic figure of the source order confirms coverage (30000 for 三万)
|
||||
}
|
||||
}
|
||||
wordRanges := magnitudeWordRanges(final)
|
||||
wordRanges := c.magnitudeWordRanges(final)
|
||||
for _, rg := range wordRanges {
|
||||
if maxSrc >= rg[0] && maxSrc <= rg[1] {
|
||||
return 0, nil // a magnitude WORD covers the source order (三万 → «тридцать тысяч»)
|
||||
|
|
@ -444,14 +439,11 @@ func lintNumberMagnitude(source, final string) (int, []string) {
|
|||
return 1, []string{fmt.Sprintf("the source magnitude 10^%d (万/億) is not reflected in the translation's orders of magnitude — a possible magnitude error (e.g. 三万→«три миллиона»)", maxSrc)}
|
||||
}
|
||||
|
||||
// cjkNumeralRunes are the characters that can form a CJK numeral expression (digits, small units,
|
||||
// big markers). A maximal run of these is one candidate number.
|
||||
var cjkNumeralRunes = map[rune]bool{}
|
||||
|
||||
func init() {
|
||||
for _, r := range "0123456789〇零一二三四五六七八九十百千两兩万萬億亿兆" {
|
||||
cjkNumeralRunes[r] = true
|
||||
}
|
||||
// isCJKNumeralRune reports whether r can form a CJK numeral expression: an ASCII digit or a shared
|
||||
// lang.CJKSection numeral (Zero∪Digit∪Unit∪Magnitude, pair-14 data-out — the gate no longer keeps its own
|
||||
// rune set). A maximal run of these is one candidate number.
|
||||
func isCJKNumeralRune(r rune) bool {
|
||||
return (r >= '0' && r <= '9') || lang.DefaultCJKSection().IsNumeralRune(r)
|
||||
}
|
||||
|
||||
// cjkMagnitudeOrders returns the base-10 order of every CJK numeral run in text that contains a big
|
||||
|
|
@ -460,12 +452,12 @@ func cjkMagnitudeOrders(text string) []int {
|
|||
var orders []int
|
||||
rs := []rune(text)
|
||||
for i := 0; i < len(rs); {
|
||||
if !cjkNumeralRunes[rs[i]] {
|
||||
if !isCJKNumeralRune(rs[i]) {
|
||||
i++
|
||||
continue
|
||||
}
|
||||
j := i
|
||||
for j < len(rs) && cjkNumeralRunes[rs[j]] {
|
||||
for j < len(rs) && isCJKNumeralRune(rs[j]) {
|
||||
j++
|
||||
}
|
||||
run := string(rs[i:j])
|
||||
|
|
@ -474,7 +466,7 @@ func cjkMagnitudeOrders(text string) []int {
|
|||
// 万一 (in case), 万分 (extremely), 万物 (all things), 千万 (by all means), 亿万 (myriads) — where 万/億
|
||||
// is not preceded by a digit. It also drops bare-unit magnitudes (十万/百万) — an accepted recall
|
||||
// trade for not false-flagging the far more frequent idioms.
|
||||
if strings.ContainsAny(run, "万萬億亿兆") && hasDigitBeforeBigMarker(run) {
|
||||
if containsBigMarker(run) && hasDigitBeforeBigMarker(run) {
|
||||
if v, ok := parseCJKNumber(run); ok && v > 0 {
|
||||
orders = append(orders, orderOf(v))
|
||||
}
|
||||
|
|
@ -484,14 +476,27 @@ func cjkMagnitudeOrders(text string) []int {
|
|||
return orders
|
||||
}
|
||||
|
||||
// isBigMarker / containsBigMarker read the shared lang.CJKSection magnitude set (万/萬/億/亿/兆, pair-14
|
||||
// data-out) instead of a hard-coded rune string.
|
||||
func isBigMarker(r rune) bool { _, ok := lang.DefaultCJKSection().BigUnit(r); return ok }
|
||||
func containsBigMarker(run string) bool {
|
||||
for _, r := range run {
|
||||
if isBigMarker(r) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// hasDigitBeforeBigMarker reports whether a digit (一-九 / 两 / 0-9) appears before the FIRST big
|
||||
// marker (万/億/兆) in the run — the signature of a real magnitude expression (三万) vs an idiom (万一).
|
||||
func hasDigitBeforeBigMarker(run string) bool {
|
||||
sec := lang.DefaultCJKSection()
|
||||
for _, r := range run {
|
||||
if strings.ContainsRune("万萬億亿兆", r) {
|
||||
if isBigMarker(r) {
|
||||
return false // hit a big marker with no digit before it → idiom / bare unit
|
||||
}
|
||||
if (r >= '0' && r <= '9') || strings.ContainsRune("一二三四五六七八九两兩", r) {
|
||||
if _, isDigit := sec.Digit[r]; (r >= '0' && r <= '9') || isDigit {
|
||||
return true
|
||||
}
|
||||
}
|
||||
|
|
@ -503,28 +508,29 @@ func hasDigitBeforeBigMarker(run string) bool {
|
|||
// <10^4 section; big units (万億兆) flush the section times the big unit into the total. Returns
|
||||
// ok=false on a shape it cannot parse (conservative — an unparseable run does not flag).
|
||||
func parseCJKNumber(s string) (int64, bool) {
|
||||
sec := lang.DefaultCJKSection()
|
||||
var total, section, cur int64
|
||||
sawBig := false
|
||||
for _, r := range s {
|
||||
switch {
|
||||
case r >= '0' && r <= '9':
|
||||
cur = cur*10 + int64(r-'0')
|
||||
case r == '〇' || r == '零':
|
||||
case sec.Zero[r]:
|
||||
cur = cur * 10
|
||||
default:
|
||||
if d, ok := cjkDigit(r); ok {
|
||||
cur = cur*10 + d // positional accumulation (一二→12), matching the Arabic-digit branch (self-review)
|
||||
if d, ok := sec.Digit[r]; ok {
|
||||
cur = cur*10 + int64(d) // positional accumulation (一二→12), matching the Arabic-digit branch (self-review)
|
||||
continue
|
||||
}
|
||||
if u, ok := cjkSmallUnit(r); ok {
|
||||
if u, ok := sec.Unit[r]; ok {
|
||||
if cur == 0 {
|
||||
cur = 1
|
||||
}
|
||||
section += cur * u
|
||||
section += cur * int64(u)
|
||||
cur = 0
|
||||
continue
|
||||
}
|
||||
if u, ok := cjkBigUnit(r); ok {
|
||||
if u, ok := sec.BigUnit(r); ok {
|
||||
sawBig = true
|
||||
section += cur
|
||||
if section == 0 {
|
||||
|
|
@ -544,54 +550,6 @@ func parseCJKNumber(s string) (int64, bool) {
|
|||
return total + section + cur, true
|
||||
}
|
||||
|
||||
func cjkDigit(r rune) (int64, bool) {
|
||||
switch r {
|
||||
case '一':
|
||||
return 1, true
|
||||
case '二', '两', '兩':
|
||||
return 2, true
|
||||
case '三':
|
||||
return 3, true
|
||||
case '四':
|
||||
return 4, true
|
||||
case '五':
|
||||
return 5, true
|
||||
case '六':
|
||||
return 6, true
|
||||
case '七':
|
||||
return 7, true
|
||||
case '八':
|
||||
return 8, true
|
||||
case '九':
|
||||
return 9, true
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
func cjkSmallUnit(r rune) (int64, bool) {
|
||||
switch r {
|
||||
case '十':
|
||||
return 10, true
|
||||
case '百':
|
||||
return 100, true
|
||||
case '千':
|
||||
return 1000, true
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
func cjkBigUnit(r rune) (int64, bool) {
|
||||
switch r {
|
||||
case '万', '萬':
|
||||
return 10000, true
|
||||
case '億', '亿':
|
||||
return 100000000, true
|
||||
case '兆':
|
||||
return 1000000000000, true
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
func orderOf(v int64) int {
|
||||
o := 0
|
||||
for v >= 10 {
|
||||
|
|
@ -606,10 +564,10 @@ func orderOf(v int64) int {
|
|||
// two orders (триста миллионов = 3·10^8, base 6 → order 8). Arabic integers are handled separately
|
||||
// by the caller (they may only CONFIRM coverage, never trigger a mismatch — D20.4), so they are NOT
|
||||
// folded in here.
|
||||
func magnitudeWordRanges(text string) [][2]int {
|
||||
func (c *dcCheckers) magnitudeWordRanges(text string) [][2]int {
|
||||
low := strings.ToLower(text)
|
||||
var ranges [][2]int
|
||||
for stem, base := range map[string]int{"тысяч": 3, "миллион": 6, "миллиард": 9, "триллион": 12} {
|
||||
for stem, base := range c.magnitudeStem { // target-general data (pair-14 data-out)
|
||||
if strings.Contains(low, stem) {
|
||||
ranges = append(ranges, [2]int{base, base + 2})
|
||||
}
|
||||
|
|
|
|||
|
|
@ -56,9 +56,10 @@ func TestLintYofikation(t *testing.T) {
|
|||
{"all-e clean", "Петр ел мед.", "all-e", 0},
|
||||
{"artem both spellings", "Артём и Артем — один человек.", "auto", 1},
|
||||
}
|
||||
dcc := testCheckers(t)
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
n, _ := lintYofikation(c.text, c.policy)
|
||||
n, _ := dcc.lintYofikation(c.text, c.policy)
|
||||
if n != c.want {
|
||||
t.Errorf("lintYofikation(%q, %q) = %d, want %d", c.text, c.policy, n, c.want)
|
||||
}
|
||||
|
|
@ -112,8 +113,8 @@ func TestParseCJKNumber(t *testing.T) {
|
|||
{"30万", 300000, true},
|
||||
{"五億", 500000000, true},
|
||||
{"一兆", 1000000000000, true},
|
||||
{"三", 0, false}, // no big marker
|
||||
{"三千", 0, false}, // no big marker (千 is a small unit)
|
||||
{"三", 0, false}, // no big marker
|
||||
{"三千", 0, false}, // no big marker (千 is a small unit)
|
||||
{"", 0, false},
|
||||
}
|
||||
for _, c := range cases {
|
||||
|
|
@ -152,9 +153,10 @@ func TestLintNumberMagnitude(t *testing.T) {
|
|||
// output number must NOT flag.
|
||||
{"wan in a name plus stray year not flagged", "他叫赵三万,生于1985年。", "Его звали Чжао Саньвань, родился в 1985 году.", 0},
|
||||
}
|
||||
dcc := testCheckers(t)
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
n, _ := lintNumberMagnitude(c.source, c.final)
|
||||
n, _ := dcc.lintNumberMagnitude(c.source, c.final)
|
||||
if n != c.want {
|
||||
t.Errorf("lintNumberMagnitude(%q → %q) = %d, want %d", c.source, c.final, n, c.want)
|
||||
}
|
||||
|
|
@ -167,7 +169,7 @@ func TestLintNumberMagnitude(t *testing.T) {
|
|||
func TestRunCheapGatesCombined(t *testing.T) {
|
||||
src := "他有三万石粮食。"
|
||||
final := "- Ара-ара, — у Пётр было три миллиона мешков. Потом Петр ушёл."
|
||||
cfg := cheapGateConfig{yoPolicy: "auto", allowlist: map[string]bool{}}
|
||||
cfg := cheapGateConfig{yoPolicy: "auto", allowlist: map[string]bool{}, checkers: testCheckers(t)}
|
||||
r := runCheapGates(src, final, final, cfg)
|
||||
if r.DialogueDash == 0 {
|
||||
t.Error("expected a dialogue-dash flag (hyphen-led line)")
|
||||
|
|
|
|||
355
backend/internal/pipeline/checkers.go
Normal file
355
backend/internal/pipeline/checkers.go
Normal file
|
|
@ -0,0 +1,355 @@
|
|||
package pipeline
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"unicode"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
)
|
||||
|
||||
// checkers.go: the WS5 defect-class checkers (DC1 时辰 double-hour units, DC2 千万/数十万 magnitude scale, DC6
|
||||
// register negative-list) + the pack-13 general checkers (percent scale, Latin residue, broken word) —
|
||||
// deterministic, $0 OBSERVABILITY flaggers on the source↔FINAL text, ported from ws5_checkers_verify.py.
|
||||
// Like the four cheap style gates they are NEVER a disposition (a hit is recorded, never drops a chunk),
|
||||
// tuned PRECISION over recall.
|
||||
//
|
||||
// PAIR-AGNOSTIC (pair-14 data-out): this file no longer holds any language literal. Every DETECTION pattern,
|
||||
// lookup table and wordlist is DATA — the SOURCE-gated pair checkers (DC1/DC2/percent: they need a zh source
|
||||
// token and compare src↔tgt) read the pair pack configs/langpacks/<pair>/dc-checkers.txt (lang.DCCheckerData);
|
||||
// the TARGET-general ones (register, broken word: they run on ANY →target output) read the embedded target
|
||||
// data (lang.TargetChecks). The ALGORITHM (compare counts, ×2 hours, suppress-if-ok, whole-word match) stays
|
||||
// here. A pair/target that ships no data runs the relevant sub-checker inert (empty → 0, the no-pack golden
|
||||
// path). Version rides the langpack Version() (data) + cheapGateVersion (algorithm) — a data or rule edit is
|
||||
// a loud --resnapshot.
|
||||
|
||||
// dcCheckers is the compiled, per-run checker spec: the pair's DETECTION patterns compiled ONCE + its lookup
|
||||
// tables + the target-general lists, resolved from the langpack. A nil receiver, a nil pattern or an empty
|
||||
// table leaves that sub-checker inert. Built once per run (compileCheckers), carried in cheapGateConfig.
|
||||
type dcCheckers struct {
|
||||
numeral map[rune]int
|
||||
ruHours map[string]int
|
||||
registerNeg []string
|
||||
brokenSuffix []string // target-general (any →ru output)
|
||||
yoHomograph map[string]bool // target-general: ё↔е homograph whitelist (cheapgates yofikator)
|
||||
magnitudeStem map[string]int // target-general: ru magnitude word stem → base-10 exponent (cheapgates)
|
||||
|
||||
shichenRE, ruHoursRE, chengRE, decimalFractionRE *regexp.Regexp
|
||||
qianwanOKRE, shushiwanOKRE, shushiwanFireRE *regexp.Regexp
|
||||
qianwanSrc, qianwanFireWord, qianwanVetoWord string
|
||||
shushiwanSrc, percentWord string
|
||||
}
|
||||
|
||||
// compileCheckers resolves the checker spec from the pair pack (dc) and the target data (tc). A malformed
|
||||
// regex in the pack is a corrupt pack → panic (deterministic, caught by the checker/golden tests). dc==nil
|
||||
// (a no-pack book) → the pair sub-checkers are inert; tc still supplies the target-general lists.
|
||||
func compileCheckers(dc *lang.DCCheckerData, tc lang.TargetChecks) *dcCheckers {
|
||||
c := &dcCheckers{
|
||||
brokenSuffix: tc.List("broken_suffix"),
|
||||
yoHomograph: listToSet(tc.List("yo_homograph")),
|
||||
magnitudeStem: listToStemExp(tc.List("magnitude_stem")),
|
||||
}
|
||||
if dc != nil {
|
||||
c.numeral, c.ruHours, c.registerNeg = dc.Numeral, dc.RuHours, dc.RegisterNeg
|
||||
p := dc.Patterns
|
||||
c.shichenRE = mustPairRE(p, "shichen_re")
|
||||
c.ruHoursRE = mustPairRE(p, "ru_hours_re")
|
||||
c.chengRE = mustPairRE(p, "cheng_re")
|
||||
c.decimalFractionRE = mustPairRE(p, "decimal_fraction_re")
|
||||
c.qianwanOKRE = mustPairRE(p, "qianwan_ok_re")
|
||||
c.shushiwanOKRE = mustPairRE(p, "shushiwan_ok_re")
|
||||
c.shushiwanFireRE = mustPairRE(p, "shushiwan_fire_re")
|
||||
c.qianwanSrc, c.qianwanFireWord, c.qianwanVetoWord = p["qianwan_src"], p["qianwan_fire_word"], p["qianwan_veto_word"]
|
||||
c.shushiwanSrc, c.percentWord = p["shushiwan_src"], p["percent_word"]
|
||||
}
|
||||
return c
|
||||
}
|
||||
|
||||
// listToSet turns an ordered value list into a membership set (target wordlists).
|
||||
func listToSet(xs []string) map[string]bool {
|
||||
m := make(map[string]bool, len(xs))
|
||||
for _, x := range xs {
|
||||
m[x] = true
|
||||
}
|
||||
return m
|
||||
}
|
||||
|
||||
// listToStemExp parses `stem<TAB>exp` values (the magnitude_stem list carries a second tab-field) into a
|
||||
// stem→exponent map. A malformed value is a corrupt embed → panic (deterministic, caught by tests).
|
||||
func listToStemExp(xs []string) map[string]int {
|
||||
m := make(map[string]int, len(xs))
|
||||
for _, x := range xs {
|
||||
f := strings.SplitN(x, "\t", 2)
|
||||
if len(f) != 2 {
|
||||
panic(fmt.Sprintf("pipeline: magnitude_stem wants `stem<TAB>exp`, got %q", x))
|
||||
}
|
||||
v, err := strconv.Atoi(strings.TrimSpace(f[1]))
|
||||
if err != nil {
|
||||
panic(fmt.Sprintf("pipeline: magnitude_stem exponent %q: %v", f[1], err))
|
||||
}
|
||||
m[f[0]] = v
|
||||
}
|
||||
return m
|
||||
}
|
||||
|
||||
// dcCheckerData returns the pack's checker data, or nil when the book has no pack (nil-safe helper for the
|
||||
// compile step, which runs even for a no-langpack book).
|
||||
func dcCheckerData(p *lang.Pack) *lang.DCCheckerData {
|
||||
if p != nil {
|
||||
return p.DCCheckers
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// mustPairRE compiles a pair detection pattern by key; a missing key → nil (inert sub-checker), a malformed
|
||||
// regex → panic (a corrupt pack, not a silent no-op — the same fail-loud discipline the loader keeps).
|
||||
func mustPairRE(p map[string]string, key string) *regexp.Regexp {
|
||||
s := p[key]
|
||||
if s == "" {
|
||||
return nil
|
||||
}
|
||||
re, err := regexp.Compile(s)
|
||||
if err != nil {
|
||||
panic(fmt.Sprintf("pipeline: langpack checker pattern %q is not a valid regex: %v", key, err))
|
||||
}
|
||||
return re
|
||||
}
|
||||
|
||||
// --- DC-1: 时辰 (double-hour) unit checker (ws5.shichen_checker) -----------------------------------
|
||||
|
||||
// lintTimeUnits flags a 时辰 (=2h) unit error: N个时辰 rendered as N часов (the count copied as hours)
|
||||
// instead of ~2N hours (三个时辰 → «три часа» should be ~6h). It fires ONLY on an explicit mismatch — a
|
||||
// paraphrase with no hours count is a valid rendering, not a defect. Pure and deterministic. Inert when the
|
||||
// pair ships no DC1 pattern (the shichen_re / ru_hours_re detection patterns are pair langpack DATA).
|
||||
func (c *dcCheckers) lintTimeUnits(source, final string) (int, []string) {
|
||||
if c == nil || c.shichenRE == nil || c.ruHoursRE == nil {
|
||||
return 0, nil
|
||||
}
|
||||
m := c.shichenRE.FindStringSubmatch(source)
|
||||
if m == nil {
|
||||
return 0, nil
|
||||
}
|
||||
n, ok := dcParseCount(m[1], c.numeral)
|
||||
if !ok {
|
||||
return 0, nil
|
||||
}
|
||||
hm := c.ruHoursRE.FindStringSubmatch(final)
|
||||
if hm == nil {
|
||||
return 0, nil // no explicit hours rendering → a valid paraphrase, not a defect
|
||||
}
|
||||
ruNum, ok := dcParseRuHours(hm[1], c.ruHours)
|
||||
if !ok {
|
||||
return 0, nil
|
||||
}
|
||||
expectedHours := n * 2
|
||||
if ruNum == n && ruNum != expectedHours {
|
||||
return 1, []string{fmt.Sprintf("DC1 时辰: %d个时辰 rendered as «%d час…» (counted as hours) instead of ~%d h (1 时辰 = 2 h)", n, ruNum, expectedHours)}
|
||||
}
|
||||
return 0, nil
|
||||
}
|
||||
|
||||
// dcParseCount parses the DC1 count group: an Arabic digit string or a single small CJK numeral (looked up
|
||||
// in the pair's DCCheckerData.Numeral, empty for a no-pack book → CJK counts don't resolve, the check is inert).
|
||||
func dcParseCount(s string, dcNum map[rune]int) (int, bool) {
|
||||
if v, err := strconv.Atoi(s); err == nil {
|
||||
return v, true
|
||||
}
|
||||
r := []rune(s)
|
||||
if len(r) == 1 {
|
||||
if v, ok := dcNum[r[0]]; ok {
|
||||
return v, true
|
||||
}
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// dcParseRuHours parses the DC1 hours group: an Arabic digit string or a Russian count word (pair data).
|
||||
func dcParseRuHours(s string, dcRu map[string]int) (int, bool) {
|
||||
if v, err := strconv.Atoi(s); err == nil {
|
||||
return v, true
|
||||
}
|
||||
if v, ok := dcRu[s]; ok {
|
||||
return v, true
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// --- DC-2: number-scale magnitude checker (ws5.magnitude_checker, 千万 / 数十万) --------------------
|
||||
|
||||
// lintMagnitudeScale flags a 千万 (10^7) / 数十万 (~several×10^5) magnitude rendered at a WRONG smaller
|
||||
// scale. A CORRECT rendering anywhere in the chunk (the ok-suppressor) suppresses the flag (ws5 reference
|
||||
// parity). The ok-suppressors are case-INsensitive (their data carries the (?i)); the FIRE predicates are
|
||||
// case-SENSITIVE literal Contains (the reference does not pass re.I to the inner searches). ⚠ 千万 is also
|
||||
// stock HYPERBOLE whose «тысячи» rendering is in-register (§5 A4). Observability only. All probes/patterns
|
||||
// are pair langpack DATA — inert when the pair ships no DC2.
|
||||
func (c *dcCheckers) lintMagnitudeScale(source, final string) (int, []string) {
|
||||
if c == nil {
|
||||
return 0, nil
|
||||
}
|
||||
var flags []string
|
||||
if c.qianwanOKRE != nil && c.qianwanSrc != "" && strings.Contains(source, c.qianwanSrc) && !c.qianwanOKRE.MatchString(final) {
|
||||
if strings.Contains(final, c.qianwanFireWord) && !strings.Contains(final, c.qianwanVetoWord) { // case-sensitive, per reference
|
||||
flags = append(flags, "DC2 千万=10^7 rendered as «тысячи» (≈10000× under) — a possible magnitude error (hyperbole risk, §5-A4)")
|
||||
}
|
||||
}
|
||||
if c.shushiwanOKRE != nil && c.shushiwanFireRE != nil && c.shushiwanSrc != "" && strings.Contains(source, c.shushiwanSrc) && !c.shushiwanOKRE.MatchString(final) {
|
||||
if c.shushiwanFireRE.MatchString(final) {
|
||||
flags = append(flags, "DC2 数十万≈several×10^5 rendered as «десятки тысяч» (≈10× under)")
|
||||
}
|
||||
}
|
||||
return len(flags), flags
|
||||
}
|
||||
|
||||
// --- DC-6: register-lexicon negative-list (ws5.register_checker) ----------------------------------
|
||||
|
||||
// lintRegisterLexicon flags whole-word occurrences of a register negative-list lexeme in the FINAL text.
|
||||
// The negative-list is pair langpack DATA (DCCheckerData.RegisterNeg): fairy-tale-Russian / chancery lexemes
|
||||
// that break the xianxia register. Lower-cased; whole-word matched. Empty (a no-pack book) → nothing to flag.
|
||||
func (c *dcCheckers) lintRegisterLexicon(final string) (int, []string) {
|
||||
if c == nil || len(c.registerNeg) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
low := []rune(strings.ToLower(final))
|
||||
hitSet := map[string]bool{}
|
||||
for _, w := range c.registerNeg {
|
||||
wr := []rune(w)
|
||||
for i := 0; i+len(wr) <= len(low); i++ {
|
||||
if !runesEqual(low[i:i+len(wr)], wr) {
|
||||
continue
|
||||
}
|
||||
if (i == 0 || !isCyrLetter(low[i-1])) && (i+len(wr) == len(low) || !isCyrLetter(low[i+len(wr)])) {
|
||||
hitSet[w] = true
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(hitSet) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
hits := make([]string, 0, len(hitSet))
|
||||
for w := range hitSet {
|
||||
hits = append(hits, w)
|
||||
}
|
||||
sort.Strings(hits)
|
||||
return len(hits), []string{"DC6 register: fairy-tale Russian lexis out of the xianxia genre: " + strings.Join(hits, ", ")}
|
||||
}
|
||||
|
||||
// isCyrLetter reports whether r is a Cyrillic letter (the word boundary for the register match).
|
||||
func isCyrLetter(r rune) bool { return unicode.IsLetter(r) && unicode.Is(unicode.Cyrillic, r) }
|
||||
|
||||
// --- percent-scale checker (成 = tenths) -----------------------------------------------------------
|
||||
//
|
||||
// In Chinese, 成 is one tenth: 六成 = 60%, 六成六 = 66%. A common error renders this as a decimal FRACTION
|
||||
// instead of a percentage (a ~100× scale error). Precision over recall: it fires only when the source has a
|
||||
// «<count>成[<count>]» (cheng_re, pair data), the output has NO percent form (percent_word / «%»), AND the
|
||||
// output carries a decimal-fraction cue (decimal_fraction_re). Inert when the pair ships no percent patterns.
|
||||
func (c *dcCheckers) lintPercentScale(source, final string) (int, []string) {
|
||||
if c == nil || c.chengRE == nil || c.decimalFractionRE == nil {
|
||||
return 0, nil
|
||||
}
|
||||
m := c.chengRE.FindStringSubmatch(source)
|
||||
if m == nil {
|
||||
return 0, nil
|
||||
}
|
||||
low := strings.ToLower(final)
|
||||
if (c.percentWord != "" && strings.Contains(low, c.percentWord)) || strings.Contains(final, "%") {
|
||||
return 0, nil // the output uses a percent form — the scale is handled correctly
|
||||
}
|
||||
if !c.decimalFractionRE.MatchString(low) {
|
||||
return 0, nil // no fraction cue — the magnitude was paraphrased, not mis-scaled
|
||||
}
|
||||
tens, _ := dcParseCount(m[1], c.numeral)
|
||||
pct := tens * 10
|
||||
if m[2] != "" {
|
||||
if ones, ok := dcParseCount(m[2], c.numeral); ok {
|
||||
pct += ones
|
||||
}
|
||||
}
|
||||
return 1, []string{fmt.Sprintf("成-percent: %s成%s = %d%% rendered as a decimal fraction instead of a percentage (~%d%%)", m[1], m[2], pct, pct)}
|
||||
}
|
||||
|
||||
// --- Latin residue in the Russian output -----------------------------------------------------------
|
||||
//
|
||||
// A whole Latin WORD left untranslated in the output (e.g. «открыл их again, …»). Language-general (Latin is
|
||||
// not pair data): it splits the output into maximal alphanumeric tokens and flags an all-lowercase Latin
|
||||
// token (leaked prose is lowercase; a capital signals a proper noun/brand) with no digit, ≥ minLatinResidueLen
|
||||
// letters, not a Roman numeral, not on the per-project allowlist. Precision over recall.
|
||||
const minLatinResidueLen = 3
|
||||
|
||||
func lintLatinResidue(final string, allow map[string]bool) (int, []string) {
|
||||
hits := map[string]bool{}
|
||||
rs := []rune(final)
|
||||
for i := 0; i < len(rs); {
|
||||
if !isLatinLetterOrDigit(rs[i]) {
|
||||
i++
|
||||
continue
|
||||
}
|
||||
j := i
|
||||
reject := false // set on any digit or uppercase letter — not a lowercase leaked word
|
||||
for j < len(rs) && isLatinLetterOrDigit(rs[j]) {
|
||||
if (rs[j] >= '0' && rs[j] <= '9') || (rs[j] >= 'A' && rs[j] <= 'Z') {
|
||||
reject = true
|
||||
}
|
||||
j++
|
||||
}
|
||||
tok := string(rs[i:j])
|
||||
i = j
|
||||
if !reject && len([]rune(tok)) >= minLatinResidueLen && !isRomanNumeral(tok) && !allow[tok] {
|
||||
hits[tok] = true
|
||||
}
|
||||
}
|
||||
if len(hits) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
surfaces := make([]string, 0, len(hits))
|
||||
for s := range hits {
|
||||
surfaces = append(surfaces, s)
|
||||
}
|
||||
sort.Strings(surfaces)
|
||||
return len(surfaces), []string{"Latin word left untranslated in the Russian output: " + strings.Join(surfaces, ", ")}
|
||||
}
|
||||
|
||||
// isLatinLetterOrDigit reports whether r is an ASCII Latin letter or digit (the alphanumeric-token alphabet).
|
||||
func isLatinLetterOrDigit(r rune) bool {
|
||||
return (r >= 'A' && r <= 'Z') || (r >= 'a' && r <= 'z') || (r >= '0' && r <= '9')
|
||||
}
|
||||
|
||||
// isRomanNumeral reports whether a lowercase token is a Roman numeral (all chars in ivxlcdm) — «iii» reads
|
||||
// as a numeral, not a leaked word; excluded to hold precision. (Uppercase «II» is already skipped as a cap.)
|
||||
func isRomanNumeral(tok string) bool {
|
||||
for _, r := range tok {
|
||||
switch r {
|
||||
case 'i', 'v', 'x', 'l', 'c', 'd', 'm':
|
||||
default:
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// --- broken word forms (target-general) ------------------------------------------------------------
|
||||
//
|
||||
// Flags a target word ending in a structurally-impossible suffix — for ru, «-йть», which no well-formed
|
||||
// Russian word does (the shape of a mangled infinitive, «войть» for «войти»). The suffix set is TARGET data
|
||||
// (lang.TargetChecks "broken_suffix"), so the ALGORITHM is language-general; a target with no suffix data
|
||||
// flags nothing. Zero false-positive by construction (only a structural signature, no dictionary).
|
||||
func (c *dcCheckers) lintBrokenWord(final string) (int, []string) {
|
||||
if c == nil || len(c.brokenSuffix) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
seen := map[string]bool{}
|
||||
var det []string
|
||||
for _, w := range tokenizeCyrillic(final) {
|
||||
wl := len([]rune(w))
|
||||
for _, suf := range c.brokenSuffix {
|
||||
if wl >= 4 && strings.HasSuffix(w, suf) && !seen[w] {
|
||||
seen[w] = true
|
||||
det = append(det, "malformed word ending in «-"+suf+"» (no valid Russian word does): "+w)
|
||||
}
|
||||
}
|
||||
}
|
||||
sort.Strings(det)
|
||||
return len(det), det
|
||||
}
|
||||
|
|
@ -6,17 +6,18 @@ import "testing"
|
|||
// word lists): each defect shape fires and a clean counterpart / golden-shape Russian output stays 0
|
||||
// (precision over recall). The 时辰/千万 cases confirm the pre-existing DC1/DC2 still catch the corpus shapes.
|
||||
func TestPack13Checkers(t *testing.T) {
|
||||
dcc := testCheckers(t)
|
||||
// Percent scale (成 = tenths): a decimal-fraction rendering fires; a correct percent is suppressed.
|
||||
if n, _ := lintPercentScale("море истинной ци 六成六 …", "море истинной ци — шесть десятых и шесть сотых"); n != 1 {
|
||||
if n, _ := dcc.lintPercentScale("море истинной ци 六成六 …", "море истинной ци — шесть десятых и шесть сотых"); n != 1 {
|
||||
t.Errorf("percent: «шесть десятых и шесть сотых» should fire, got %d", n)
|
||||
}
|
||||
if n, _ := lintPercentScale("六成六", "море истинной ци — шесть и шесть десятых"); n != 1 {
|
||||
if n, _ := dcc.lintPercentScale("六成六", "море истинной ци — шесть и шесть десятых"); n != 1 {
|
||||
t.Errorf("percent: «шесть и шесть десятых» should fire, got %d", n)
|
||||
}
|
||||
if n, _ := lintPercentScale("六成六", "заполнено на шестьдесят шесть процентов"); n != 0 {
|
||||
if n, _ := dcc.lintPercentScale("六成六", "заполнено на шестьдесят шесть процентов"); n != 0 {
|
||||
t.Errorf("percent: a correct «процентов» rendering must be suppressed, got %d", n)
|
||||
}
|
||||
if n, _ := lintPercentScale("нет числа", "шесть десятых чего-то"); n != 0 {
|
||||
if n, _ := dcc.lintPercentScale("нет числа", "шесть десятых чего-то"); n != 0 {
|
||||
t.Errorf("percent: no 成 in source must not fire, got %d", n)
|
||||
}
|
||||
|
||||
|
|
@ -28,9 +29,9 @@ func TestPack13Checkers(t *testing.T) {
|
|||
for _, clean := range []string{
|
||||
"ОТРЕДАКТИРОВАННЫЙ ПЕРЕВОД 5abc35ddfb65. Судзуки шёл по коридорам.", // a body-hash id (has digits)
|
||||
"Классы таланта А, Б, В и Г — от высшей к низшей.", // Cyrillic class letters
|
||||
"Глава II начинается.", // a Roman numeral (uppercase)
|
||||
"Он держал в руках iPhone и MacBook.", // brands — any capital → skipped
|
||||
"Судзуки шёл в Академию.", // Cyrillic proper noun
|
||||
"Глава II начинается.", // a Roman numeral (uppercase)
|
||||
"Он держал в руках iPhone и MacBook.", // brands — any capital → skipped
|
||||
"Судзуки шёл в Академию.", // Cyrillic proper noun
|
||||
} {
|
||||
if n, det := lintLatinResidue(clean, nil); n != 0 {
|
||||
t.Errorf("latin: clean text must not fire (%q): %d %v", clean, n, det)
|
||||
|
|
@ -42,7 +43,7 @@ func TestPack13Checkers(t *testing.T) {
|
|||
}
|
||||
|
||||
// Broken word — the general «-йть» rule only (no book-specific lists): «войть» fires, valid words do not.
|
||||
if n, _ := lintBrokenWord("хотел тихонько войть и закрыть"); n != 1 {
|
||||
if n, _ := dcc.lintBrokenWord("хотел тихонько войть и закрыть"); n != 1 {
|
||||
t.Errorf("broken: «войть» (-йть) should fire")
|
||||
}
|
||||
for _, clean := range []string{
|
||||
|
|
@ -50,16 +51,16 @@ func TestPack13Checkers(t *testing.T) {
|
|||
"глава клана Гуюэ поклонился", // valid prose (no -йть)
|
||||
"впереди идёт Фан Юань", // valid впереди
|
||||
} {
|
||||
if n, det := lintBrokenWord(clean); n != 0 {
|
||||
if n, det := dcc.lintBrokenWord(clean); n != 0 {
|
||||
t.Errorf("broken: clean text must not fire (%q): %d %v", clean, n, det)
|
||||
}
|
||||
}
|
||||
|
||||
// The pre-existing DC1/DC2 still catch the corpus shapes (general Chinese units/idioms).
|
||||
if n, _ := lintTimeUnits("僵持了三个时辰", "прошло три часа"); n != 1 {
|
||||
if n, _ := dcc.lintTimeUnits("僵持了三个时辰", "прошло три часа"); n != 1 {
|
||||
t.Errorf("time-unit: 三个时辰→«три часа» should fire")
|
||||
}
|
||||
if n, _ := lintMagnitudeScale("千万生灵", "погубил тысячи жизней"); n != 1 {
|
||||
if n, _ := dcc.lintMagnitudeScale("千万生灵", "погубил тысячи жизней"); n != 1 {
|
||||
t.Errorf("magnitude: 千万→«тысячи» should fire")
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -1,295 +0,0 @@
|
|||
package pipeline
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"unicode"
|
||||
)
|
||||
|
||||
// checkers_zh_ru.go: the WS5 defect-class checkers DC1 (时辰 double-hour units), DC2 (千万/数十万 magnitude
|
||||
// scale) and DC6 (register negative-list) — deterministic, $0 flaggers on the source↔FINAL text, ported
|
||||
// from the frozen ws5_checkers_verify.py. Like the four cheap style gates they are OBSERVABILITY, NEVER a
|
||||
// disposition (research/20: judges too noisy → deterministic gates; a hit is recorded in the retrieval-
|
||||
// state / report, never dropping a chunk). They are tuned PRECISION over recall — DC1/DC2 fire only on an
|
||||
// EXPLICIT src↔target mismatch; the LANDING (whether the flag is trusted) is gated by a §5(д) false-
|
||||
// positive measure on a fresh rerun, not by this build. Their rule VERSION rides cheapGateVersion (folded
|
||||
// into the snapshot as StyleCheckVersion), so editing a rule / the pack is a loud --resnapshot.
|
||||
//
|
||||
// PER-PAIR PACK (layer-2 data, §5(в)): the unit table + negative-list below are the zh-ru convention pack.
|
||||
// They are code consts versioned by cheapGateVersion (the same discipline the existing cheap-gate data
|
||||
// uses), not a config-loaded per-pair file — a config-loaded versioned pack is a cleaner future form
|
||||
// (noted residual). The checkers self-gate on source content (时辰/千万 are zh-specific → no fire on a
|
||||
// non-zh source), and DC6 is target-side (gated behind ru like the other readability flaggers).
|
||||
|
||||
// --- DC-1: 时辰 (double-hour) unit checker (ws5.shichen_checker) -----------------------------------
|
||||
|
||||
// dcCNNum maps the small CJK numerals the 时辰 pattern accepts to their value (1 时辰 = 2 modern hours).
|
||||
var dcCNNum = map[rune]int{
|
||||
'一': 1, '二': 2, '两': 2, '三': 3, '四': 4, '五': 5, '六': 6, '七': 7, '八': 8, '九': 9, '十': 10,
|
||||
}
|
||||
|
||||
// dcShichenRE captures the count before 时辰 (an Arabic digit or a small CJK numeral), tolerating 个.
|
||||
var dcShichenRE = regexp.MustCompile(`([0-9一二两三四五六七八九十])\s*个?\s*时辰`)
|
||||
|
||||
// dcRuHoursRE captures a Russian "<count> час…" rendering; the alternatives mirror the reference map.
|
||||
var dcRuHoursRE = regexp.MustCompile(`(\d+|один|два|двух|три|трёх|трех|четыре|пять|шесть)\s+час`)
|
||||
|
||||
// dcRuHourWord maps the Russian count words dcRuHoursRE captures to their value.
|
||||
var dcRuHourWord = map[string]int{
|
||||
"один": 1, "два": 2, "двух": 2, "три": 3, "трёх": 3, "трех": 3, "четыре": 4, "пять": 5, "шесть": 6,
|
||||
}
|
||||
|
||||
// lintTimeUnits flags a 时辰 (=2h) unit error: N个时辰 rendered as N часов (the count copied as hours)
|
||||
// instead of ~2N hours (三个时辰 → «три часа» should be ~6h). It fires ONLY on an explicit mismatch — a
|
||||
// paraphrase with no hours count is a valid rendering, not a defect (the reference dropped the "no hours
|
||||
// found" branch that false-flagged 1.8%). Pure and deterministic.
|
||||
func lintTimeUnits(source, final string) (int, []string) {
|
||||
m := dcShichenRE.FindStringSubmatch(source)
|
||||
if m == nil {
|
||||
return 0, nil
|
||||
}
|
||||
n, ok := dcParseCount(m[1])
|
||||
if !ok {
|
||||
return 0, nil
|
||||
}
|
||||
hm := dcRuHoursRE.FindStringSubmatch(final)
|
||||
if hm == nil {
|
||||
return 0, nil // no explicit hours rendering → a valid paraphrase, not a defect
|
||||
}
|
||||
ruNum, ok := dcParseRuHours(hm[1])
|
||||
if !ok {
|
||||
return 0, nil
|
||||
}
|
||||
expectedHours := n * 2
|
||||
if ruNum == n && ruNum != expectedHours {
|
||||
return 1, []string{fmt.Sprintf("DC1 时辰: %d个时辰 rendered as «%d час…» (counted as hours) instead of ~%d h (1 时辰 = 2 h)", n, ruNum, expectedHours)}
|
||||
}
|
||||
return 0, nil
|
||||
}
|
||||
|
||||
// dcParseCount parses the DC1 count group: an Arabic digit string or a single small CJK numeral.
|
||||
func dcParseCount(s string) (int, bool) {
|
||||
if v, err := strconv.Atoi(s); err == nil {
|
||||
return v, true
|
||||
}
|
||||
r := []rune(s)
|
||||
if len(r) == 1 {
|
||||
if v, ok := dcCNNum[r[0]]; ok {
|
||||
return v, true
|
||||
}
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// dcParseRuHours parses the DC1 hours group: an Arabic digit string or a Russian count word.
|
||||
func dcParseRuHours(s string) (int, bool) {
|
||||
if v, err := strconv.Atoi(s); err == nil {
|
||||
return v, true
|
||||
}
|
||||
if v, ok := dcRuHourWord[s]; ok {
|
||||
return v, true
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// --- DC-2: number-scale magnitude checker (ws5.magnitude_checker, 千万 / 数十万) --------------------
|
||||
|
||||
// DC2 regexes — a BYTE-FAITHFUL port of ws5.magnitude_checker (_MAG). The ok_re SUPPRESSION guards are
|
||||
// case-INsensitive (the reference passes re.I); the FIRE predicates are case-SENSITIVE (the reference
|
||||
// does NOT pass re.I to the inner тысяч/десятки-тысяч searches). \w is ported as \p{L}* (RE2's \w is
|
||||
// ASCII; the Cyrillic word-continuation the reference intends is letters).
|
||||
var (
|
||||
dcQianwanOkRE = regexp.MustCompile(`(?i)десят\p{L}* миллион|10\s*000\s*000|10000000`) // 千万 correct render → suppress
|
||||
dcShushiwanOkRE = regexp.MustCompile(`(?i)сотн\p{L}* тысяч|нескольк\p{L}* сот\p{L}* тысяч|[1-9]00\s*000`) // 数十万 correct → suppress
|
||||
dcDesyatkiTysRE = regexp.MustCompile(`десятк\p{L}* тысяч`) // 数十万 under-render (case-SENSITIVE, per reference)
|
||||
)
|
||||
|
||||
// lintMagnitudeScale flags a 千万 (10^7) / 数十万 (~several×10^5) magnitude rendered at a WRONG smaller
|
||||
// scale (an order-of-magnitude error DC2 targets). 千万 → «тысячи» without «миллион» = 10000× under; 数十万 → «десятки
|
||||
// тысяч» = 10× under. A CORRECT rendering anywhere in the chunk (the ok_re guard) SUPPRESSES the flag, so
|
||||
// a chunk that renders 数十万 as «сотни тысяч» does not fire even if «десятки тысяч» appears elsewhere as
|
||||
// an unrelated quantity (the ws5 reference parity — required so the §5(д) FP-measure taken against the
|
||||
// frozen reference matches what ships). ⚠ 千万 is also a stock HYPERBOLE («несметно») whose «тысячи»
|
||||
// rendering is in-register literary, NOT a hard error (§5 A4) — hyperbole-exposed, landing FP-gated §5(д).
|
||||
// Observability only. (Word↔word fraction inversion — 四成四=44% via q4a rule_b1/b2 — is a residual
|
||||
// sub-checker, not built here.)
|
||||
func lintMagnitudeScale(source, final string) (int, []string) {
|
||||
var flags []string
|
||||
if strings.Contains(source, "千万") && !dcQianwanOkRE.MatchString(final) {
|
||||
if strings.Contains(final, "тысяч") && !strings.Contains(final, "миллион") { // case-sensitive, per reference
|
||||
flags = append(flags, "DC2 千万=10^7 rendered as «тысячи» (≈10000× under) — a possible magnitude error (hyperbole risk, §5-A4)")
|
||||
}
|
||||
}
|
||||
if strings.Contains(source, "数十万") && !dcShushiwanOkRE.MatchString(final) {
|
||||
if dcDesyatkiTysRE.MatchString(final) {
|
||||
flags = append(flags, "DC2 数十万≈several×10^5 rendered as «десятки тысяч» (≈10× under)")
|
||||
}
|
||||
}
|
||||
return len(flags), flags
|
||||
}
|
||||
|
||||
// --- DC-6: register-lexicon negative-list (ws5.register_checker) ----------------------------------
|
||||
|
||||
// dcRegisterNegList is the zh-ru register negative-list (layer-2 pack): fairy-tale-Russian / chancery
|
||||
// lexemes that break the xianxia register («терем» in a cultivation novel is a hard register error). The
|
||||
// owner extends this per corpus finding. Lower-cased; whole-word matched.
|
||||
var dcRegisterNegList = []string{"терем", "терема", "тереме", "теремом", "терему", "теремах"}
|
||||
|
||||
// lintRegisterLexicon flags whole-word occurrences of a register negative-list lexeme in the FINAL text.
|
||||
func lintRegisterLexicon(final string) (int, []string) {
|
||||
low := []rune(strings.ToLower(final))
|
||||
hitSet := map[string]bool{}
|
||||
for _, w := range dcRegisterNegList {
|
||||
wr := []rune(w)
|
||||
for i := 0; i+len(wr) <= len(low); i++ {
|
||||
if !runesEqual(low[i:i+len(wr)], wr) {
|
||||
continue
|
||||
}
|
||||
if (i == 0 || !isCyrLetter(low[i-1])) && (i+len(wr) == len(low) || !isCyrLetter(low[i+len(wr)])) {
|
||||
hitSet[w] = true
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(hitSet) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
hits := make([]string, 0, len(hitSet))
|
||||
for w := range hitSet {
|
||||
hits = append(hits, w)
|
||||
}
|
||||
sort.Strings(hits)
|
||||
return len(hits), []string{"DC6 register: fairy-tale Russian lexis out of the xianxia genre: " + strings.Join(hits, ", ")}
|
||||
}
|
||||
|
||||
// isCyrLetter reports whether r is a Cyrillic letter (the word boundary for the register match).
|
||||
func isCyrLetter(r rune) bool { return unicode.IsLetter(r) && unicode.Is(unicode.Cyrillic, r) }
|
||||
|
||||
// --- percent-scale checker (成 = tenths) -----------------------------------------------------------
|
||||
//
|
||||
// In Chinese, 成 is one tenth: 六成 = 60%, 六成六 = 66%. A common translation error renders this as a
|
||||
// decimal FRACTION instead of a percentage — «шесть десятых и шесть сотых» (0.66) or «шесть и шесть
|
||||
// десятых» (6.6) for 六成六 — a ~100× scale error. This flags that. Precision over recall: it fires only
|
||||
// when the source has «<count>成[<count>]», the output has NO percent form («процент»/«%»), AND the output
|
||||
// carries a decimal-fraction cue (a «десятых»/«сотых» ordinal or a «N,N» number). So an output that renders
|
||||
// the percent correctly is suppressed, and an output without a fraction cue (a paraphrase) stays silent.
|
||||
// A general zh→ru unit convention, not tied to any book.
|
||||
|
||||
// chengPercentRE matches a 成-percent expression: a count (CJK or Arabic), 成, and an optional second count.
|
||||
var chengPercentRE = regexp.MustCompile(`([0-9一二三四五六七八九十])成([0-9一二三四五六七八九]?)`)
|
||||
|
||||
// decimalFractionRE is the error cue: the tenths rendered as a Russian fraction ordinal or a decimal number.
|
||||
var decimalFractionRE = regexp.MustCompile(`десят(?:ая|ых|ой|ые)|сот(?:ая|ых|ой|ые)|\d+[.,]\d`)
|
||||
|
||||
// lintPercentScale flags a 成-percent rendered as a decimal fraction instead of a percentage.
|
||||
func lintPercentScale(source, final string) (int, []string) {
|
||||
m := chengPercentRE.FindStringSubmatch(source)
|
||||
if m == nil {
|
||||
return 0, nil
|
||||
}
|
||||
low := strings.ToLower(final)
|
||||
if strings.Contains(low, "процент") || strings.Contains(final, "%") {
|
||||
return 0, nil // the output uses a percent form — the scale is handled correctly
|
||||
}
|
||||
if !decimalFractionRE.MatchString(low) {
|
||||
return 0, nil // no fraction cue — the magnitude was paraphrased, not mis-scaled
|
||||
}
|
||||
tens, _ := dcParseCount(m[1])
|
||||
pct := tens * 10
|
||||
if m[2] != "" {
|
||||
if ones, ok := dcParseCount(m[2]); ok {
|
||||
pct += ones
|
||||
}
|
||||
}
|
||||
return 1, []string{fmt.Sprintf("成-percent: %s成%s = %d%% rendered as a decimal fraction instead of a percentage (~%d%%)", m[1], m[2], pct, pct)}
|
||||
}
|
||||
|
||||
// --- Latin residue in the Russian output -----------------------------------------------------------
|
||||
//
|
||||
// A whole Latin WORD left untranslated in the Russian output (e.g. «открыл их again, …»). It splits the
|
||||
// output into maximal alphanumeric tokens and flags a token that is worth reporting as leaked prose. The
|
||||
// target is a lowercase mid-sentence English word, so the guards keep precision high:
|
||||
// - the token must be ALL LOWERCASE Latin letters. Leaked prose is lowercase; a token with any capital
|
||||
// is a proper noun / brand / acronym (iPhone, Google, Suzuki, «II») — legitimate in Russian text, and
|
||||
// the sanitizer already treats capitals as a brand signal, so we skip them here too.
|
||||
// - it must have NO digit: an id / hash like «5abc35ddfb65» is an alphanumeric token, not a word.
|
||||
// - at least minLatinResidueLen letters, not a Roman numeral, and not on the per-project allowlist.
|
||||
// A foreignizing brief that keeps intentional Latin (a motto, a scientific name) puts those on the
|
||||
// allowlist. Tuned for precision over recall: a leaked word that is Capitalized, or a bare URL host
|
||||
// («example.com» → «example»/«com»), is not caught — accepted for a low-noise observability signal.
|
||||
const minLatinResidueLen = 3
|
||||
|
||||
func lintLatinResidue(final string, allow map[string]bool) (int, []string) {
|
||||
hits := map[string]bool{}
|
||||
rs := []rune(final)
|
||||
for i := 0; i < len(rs); {
|
||||
if !isLatinLetterOrDigit(rs[i]) {
|
||||
i++
|
||||
continue
|
||||
}
|
||||
j := i
|
||||
reject := false // set on any digit or uppercase letter — not a lowercase leaked word
|
||||
for j < len(rs) && isLatinLetterOrDigit(rs[j]) {
|
||||
if (rs[j] >= '0' && rs[j] <= '9') || (rs[j] >= 'A' && rs[j] <= 'Z') {
|
||||
reject = true
|
||||
}
|
||||
j++
|
||||
}
|
||||
tok := string(rs[i:j])
|
||||
i = j
|
||||
if !reject && len([]rune(tok)) >= minLatinResidueLen && !isRomanNumeral(tok) && !allow[tok] {
|
||||
hits[tok] = true
|
||||
}
|
||||
}
|
||||
if len(hits) == 0 {
|
||||
return 0, nil
|
||||
}
|
||||
surfaces := make([]string, 0, len(hits))
|
||||
for s := range hits {
|
||||
surfaces = append(surfaces, s)
|
||||
}
|
||||
sort.Strings(surfaces)
|
||||
return len(surfaces), []string{"Latin word left untranslated in the Russian output: " + strings.Join(surfaces, ", ")}
|
||||
}
|
||||
|
||||
// isLatinLetterOrDigit reports whether r is an ASCII Latin letter or digit (the alphanumeric-token alphabet).
|
||||
func isLatinLetterOrDigit(r rune) bool {
|
||||
return (r >= 'A' && r <= 'Z') || (r >= 'a' && r <= 'z') || (r >= '0' && r <= '9')
|
||||
}
|
||||
|
||||
// isRomanNumeral reports whether a lowercase token is a Roman numeral (all chars in ivxlcdm) — «iii» reads
|
||||
// as a numeral, not a leaked word; excluded to hold precision. (Uppercase «II» is already skipped as a cap.)
|
||||
func isRomanNumeral(tok string) bool {
|
||||
for _, r := range tok {
|
||||
switch r {
|
||||
case 'i', 'v', 'x', 'l', 'c', 'd', 'm':
|
||||
default:
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// --- broken Russian word forms ---------------------------------------------------------------------
|
||||
//
|
||||
// Deliberately GENERAL, with NO dictionary and NO book-specific word lists: it flags a Russian word ending
|
||||
// in «-йть», which no well-formed Russian word does — the shape of a mangled infinitive (e.g. «войть» for
|
||||
// «войти»). It is zero false-positive and applies to any book. Malformations WITHOUT such a structural
|
||||
// signature — a plausible misspelling («Вперди» for «Впереди») or a case-agreement error («глава клан»
|
||||
// for «главы клана») — are NOT detectable deterministically without a morphology/dictionary pass, which is
|
||||
// out of scope here; the report states this recall limit honestly. The existing sanitizer broken-word class
|
||||
// (invalid soft/hard-sign bigrams, script-mixed homoglyph tokens) is orthogonal and still runs.
|
||||
func lintBrokenWord(final string) (int, []string) {
|
||||
seen := map[string]bool{}
|
||||
var det []string
|
||||
for _, w := range tokenizeCyrillic(final) {
|
||||
if len([]rune(w)) >= 4 && strings.HasSuffix(w, "йть") && !seen[w] {
|
||||
seen[w] = true
|
||||
det = append(det, "malformed word ending in «-йть» (no valid Russian word does): "+w)
|
||||
}
|
||||
}
|
||||
sort.Strings(det)
|
||||
return len(det), det
|
||||
}
|
||||
|
|
@ -3,8 +3,18 @@ package pipeline
|
|||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
)
|
||||
|
||||
// testCheckers builds the compiled checker spec from the REAL zh-ru pack + ru target data (pair-14 data-out):
|
||||
// the fixtures exercise the checker ALGORITHM over the pack's own DETECTION patterns / tables / wordlists, so
|
||||
// no pair literal is duplicated in the test — a pattern edit in the langpack flows straight into these cases.
|
||||
func testCheckers(t *testing.T) *dcCheckers {
|
||||
t.Helper()
|
||||
return compileCheckers(testLangPack(t).DCCheckers, lang.TargetChecksFor("ru"))
|
||||
}
|
||||
|
||||
// checkers_zh_ru_test.go: WS5 (г) — the DC1/DC2/DC6 checkers must CATCH the empirical trap positives
|
||||
// (ws5_checkers_verify.py §a) and stay SILENT on clean/in-register text (precision over recall). These
|
||||
// are the regression fixtures the plan requires (positives from exp15 §7.9 / exp14b).
|
||||
|
|
@ -20,9 +30,10 @@ func TestDC1TimeUnits(t *testing.T) {
|
|||
{"no-shichen", "他走了三里。", "Он прошёл три ли.", false},
|
||||
{"arabic-count", "过了2个时辰。", "Прошло 2 часа.", true}, // 2时辰=4h rendered «2 часа»
|
||||
}
|
||||
dcc := testCheckers(t)
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
n, _ := lintTimeUnits(c.src, c.tgt)
|
||||
n, _ := dcc.lintTimeUnits(c.src, c.tgt)
|
||||
if (n > 0) != c.wantFlag {
|
||||
t.Fatalf("lintTimeUnits(%q,%q) fired=%v, want %v", c.src, c.tgt, n > 0, c.wantFlag)
|
||||
}
|
||||
|
|
@ -47,9 +58,10 @@ func TestDC2MagnitudeScale(t *testing.T) {
|
|||
{"case-sensitive-capitalized", "聚集了数十万人。", "Десятки тысяч человек собрались.", false},
|
||||
{"no-magnitude", "他有三个朋友。", "У него три друга.", false},
|
||||
}
|
||||
dcc := testCheckers(t)
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
n, _ := lintMagnitudeScale(c.src, c.tgt)
|
||||
n, _ := dcc.lintMagnitudeScale(c.src, c.tgt)
|
||||
if (n > 0) != c.wantFlag {
|
||||
t.Fatalf("lintMagnitudeScale(%q,%q) fired=%v, want %v", c.src, c.tgt, n > 0, c.wantFlag)
|
||||
}
|
||||
|
|
@ -67,9 +79,10 @@ func TestDC6RegisterLexicon(t *testing.T) {
|
|||
{"clean", "Он вошёл в высокий зал павильона.", 0},
|
||||
{"substring-not-word", "Термин был странным.", 0}, // «терм» inside «термин» must NOT fire (whole-word)
|
||||
}
|
||||
dcc := testCheckers(t)
|
||||
for _, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
n, det := lintRegisterLexicon(c.tgt)
|
||||
n, det := dcc.lintRegisterLexicon(c.tgt)
|
||||
if n != c.wantN {
|
||||
t.Fatalf("lintRegisterLexicon(%q) = %d (%v), want %d", c.tgt, n, det, c.wantN)
|
||||
}
|
||||
|
|
@ -92,15 +105,15 @@ func TestDC3GenderInjection(t *testing.T) {
|
|||
mk := func(gender string) []pickedEntry {
|
||||
return []pickedEntry{{entry: &memoryEntry{src: "方源", dst: "Фан Юань", status: "approved", gender: gender}, via: "方源", disp: memConfirmed}}
|
||||
}
|
||||
male := renderEditorConstraintBlock(mk("male"))
|
||||
male := renderEditorConstraintBlock(mk("male"), ruTX())
|
||||
if !strings.Contains(male, "方源 → «Фан Юань» (муж.") {
|
||||
t.Fatalf("male term must carry a masculine directive, got:\n%s", male)
|
||||
}
|
||||
hidden := renderEditorConstraintBlock(mk("hidden"))
|
||||
hidden := renderEditorConstraintBlock(mk("hidden"), ruTX())
|
||||
if !strings.Contains(hidden, "пол СКРЫТ") {
|
||||
t.Fatalf("hidden term must carry the gender-avoidance mandate, got:\n%s", hidden)
|
||||
}
|
||||
none := renderEditorConstraintBlock(mk(""))
|
||||
none := renderEditorConstraintBlock(mk(""), ruTX())
|
||||
if !strings.HasSuffix(none, "«Фан Юань»") {
|
||||
t.Fatalf("a genderless term must have NO gender note (line ends at the dst), got:\n%s", none)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -198,13 +198,14 @@ func matchHeaderLine(line string, hr *lang.HeadingRule) (n int, subtitle string,
|
|||
}
|
||||
|
||||
// isHeadingNumeral reports whether a rune can be part of a chapter-number run: an Arabic digit
|
||||
// (half/fullwidth) or a CJK numeral character.
|
||||
// (half/fullwidth) or a CJK numeral character (the shared lang.CJKSection — pair-14 §4, the SAME source
|
||||
// ingest's chapter-numeral regex reads, so the two can never byte-drift apart).
|
||||
func isHeadingNumeral(r rune) bool {
|
||||
switch {
|
||||
case r >= '0' && r <= '9', r >= '0' && r <= '9':
|
||||
return true
|
||||
}
|
||||
return strings.ContainsRune("〇零一二三四五六七八九十百千两兩", r)
|
||||
return lang.DefaultCJKSection().IsHeadingNumeral(r)
|
||||
}
|
||||
|
||||
// isHeaderContentRune reports whether a rune is CONTENT (a letter/digit/ideograph/kana) rather than a
|
||||
|
|
@ -239,7 +240,7 @@ func parseSectionNumeral(s string) (int, bool) {
|
|||
case r >= '0' && r <= '9':
|
||||
num = num*10 + int(r-'0')
|
||||
any = true
|
||||
case r == '〇' || r == '零':
|
||||
case lang.DefaultCJKSection().Zero[r]:
|
||||
num = num * 10
|
||||
any = true
|
||||
default:
|
||||
|
|
@ -267,40 +268,16 @@ func parseSectionNumeral(s string) (int, bool) {
|
|||
return v, true
|
||||
}
|
||||
|
||||
// cjkSectionDigit / cjkSectionUnit read the shared lang.CJKSection (pair-14 §4): the SAME digit/unit value
|
||||
// tables the ingest chapter-numeral inventory derives from, so there is ONE source of truth.
|
||||
func cjkSectionDigit(r rune) (int, bool) {
|
||||
switch r {
|
||||
case '一':
|
||||
return 1, true
|
||||
case '二', '两', '兩':
|
||||
return 2, true
|
||||
case '三':
|
||||
return 3, true
|
||||
case '四':
|
||||
return 4, true
|
||||
case '五':
|
||||
return 5, true
|
||||
case '六':
|
||||
return 6, true
|
||||
case '七':
|
||||
return 7, true
|
||||
case '八':
|
||||
return 8, true
|
||||
case '九':
|
||||
return 9, true
|
||||
}
|
||||
return 0, false
|
||||
v, ok := lang.DefaultCJKSection().Digit[r]
|
||||
return v, ok
|
||||
}
|
||||
|
||||
func cjkSectionUnit(r rune) (int, bool) {
|
||||
switch r {
|
||||
case '十':
|
||||
return 10, true
|
||||
case '百':
|
||||
return 100, true
|
||||
case '千':
|
||||
return 1000, true
|
||||
}
|
||||
return 0, false
|
||||
v, ok := lang.DefaultCJKSection().Unit[r]
|
||||
return v, ok
|
||||
}
|
||||
|
||||
// chapterDraftChunks packs one chapter's paragraphs into DRAFT chunks (the fine tiling). Rule:
|
||||
|
|
@ -621,17 +598,13 @@ func isASCIISpace(r rune) bool {
|
|||
return false
|
||||
}
|
||||
|
||||
// sourceAbbrevs are trailing tokens after which a lone "." is treated as an
|
||||
// abbreviation, not a sentence end (case-insensitive). A pragmatic set for prose;
|
||||
// the acceptance path is ja→ru where CJK terminators dominate, so this only guards
|
||||
// the en source. Single-letter initials are handled separately (isAbbrevBefore).
|
||||
var sourceAbbrevs = map[string]bool{
|
||||
"mr": true, "mrs": true, "ms": true, "dr": true, "prof": true, "st": true,
|
||||
"jr": true, "sr": true, "vs": true, "no": true, "vol": true, "ch": true,
|
||||
"fig": true, "col": true, "gen": true, "sgt": true, "capt": true, "lt": true,
|
||||
"rev": true, "gov": true, "sen": true, "rep": true, "etc": true, "inc": true,
|
||||
"ltd": true, "co": true, "mt": true, "ave": true, "rd": true,
|
||||
}
|
||||
// sourceAbbrevs are trailing tokens after which a lone "." is treated as an abbreviation, not a sentence
|
||||
// end (case-insensitive). DATA lives in internal/lang (embedded, sectioned per SOURCE language, pair-14 §4);
|
||||
// the splitter consumes the "en" section — the one source whose ASCII period needs the guard (the acceptance
|
||||
// path is zh/ja→ru where 。 terminators dominate, so those sources ship no abbreviations). Making the splitter
|
||||
// consume the BOOK's source section is a shallow follow-up (thread sourceLang into SplitChunks). Single-
|
||||
// letter initials are handled separately (isAbbrevBefore).
|
||||
var sourceAbbrevs = lang.SentenceAbbrev("en")
|
||||
|
||||
// isAbbrevBefore reports whether the text immediately before a lone "." ends in a
|
||||
// known abbreviation or a single-letter initial (so the "." is not a boundary).
|
||||
|
|
|
|||
|
|
@ -5,6 +5,8 @@ import (
|
|||
"maps"
|
||||
"slices"
|
||||
"strings"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
"textmachine/backend/internal/store"
|
||||
)
|
||||
|
||||
|
|
@ -77,8 +79,9 @@ func (r *Runner) classifyOutput(role, source, output, finish string, isFinal boo
|
|||
return cls, ""
|
||||
}
|
||||
|
||||
// injectionRenderer serializes a chunk's selected memory records into a role's injection message.
|
||||
type injectionRenderer func(injected []pickedEntry) string
|
||||
// injectionRenderer serializes a chunk's selected memory records into a role's injection message, using the
|
||||
// target-language wire-text (lang.InjectionTexts, pair-14 §2) — a target with no texts renders nothing.
|
||||
type injectionRenderer func(injected []pickedEntry, tx lang.InjectionTexts) string
|
||||
|
||||
// roleInjectionRenderers maps a stage ROLE to the memory-injection renderer it consumes — a registry
|
||||
// instead of a hand-edited switch (D39 layer 7, L8-role-injection-hardcoded-switch), so a new Phase-2
|
||||
|
|
@ -171,5 +174,6 @@ func (r *Runner) cheapGateConfig() cheapGateConfig {
|
|||
yoPolicy: r.Book.YoPolicy,
|
||||
allowlist: allow,
|
||||
regressionEnabled: r.Pipeline.Gates.RegressionGuard.Enabled,
|
||||
checkers: r.checkers, // compiled once in openRunner (pair-14 data-out); nil-inert for a no-pack book
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -7,6 +7,7 @@ import (
|
|||
"unicode"
|
||||
"unicode/utf8"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
"textmachine/backend/internal/llm"
|
||||
)
|
||||
|
||||
|
|
@ -231,23 +232,12 @@ func classify(in classifyInput) classification {
|
|||
}
|
||||
|
||||
// --- refusal blacklist (en/ru/zh/ja) — ported 1:1 from eval/refusal_bench.py ---
|
||||
|
||||
var refusalPatterns = []string{
|
||||
`I (?:can(?:'|no)t|cannot|won'?t|am (?:not able|unable)) (?:to )?(?:assist|help|translate|continue|provide|fulfill)`,
|
||||
`I'?m sorry,? but`,
|
||||
`against (?:my|our) (?:content )?polic`,
|
||||
`content polic(?:y|ies)`,
|
||||
`не могу (?:помочь|перевести|продолжить|выполнить)`,
|
||||
`я не буду переводить`,
|
||||
`противоречит (?:политике|правилам)`,
|
||||
`としてお答えできません`,
|
||||
`申し訳ありませんが`,
|
||||
`我(?:不能|无法)(?:协助|翻译|继续)`,
|
||||
`抱歉[,,]`,
|
||||
`^\s*\[?(?:REFUSED|ОТКАЗ)\]?\s*$`,
|
||||
}
|
||||
|
||||
var refusalRE = regexp.MustCompile("(?im)" + strings.Join(refusalPatterns, "|"))
|
||||
//
|
||||
// The patterns are UNIVERSAL engine safety data (a model can refuse in any language regardless of the
|
||||
// book's pair), so they live in internal/lang as an embedded, per-language-sectioned file (pair-14 §3), NOT
|
||||
// a book pack — a no-langpack book (the ja→ru golden) still flags a refusal. Joined here into one
|
||||
// case-insensitive regex, byte-identically to the ported reference.
|
||||
var refusalRE = regexp.MustCompile("(?im)" + strings.Join(lang.RefusalPatterns(), "|"))
|
||||
|
||||
var thinkRE = regexp.MustCompile(`(?s)<think>.*?</think>\s*`)
|
||||
|
||||
|
|
|
|||
|
|
@ -64,7 +64,7 @@ func TestFewShotEnabledDefault(t *testing.T) {
|
|||
// the 物是人非 chengyu-atom), and few_shot OFF keeps the discourse core (incl. the chengyu-atom)
|
||||
// but drops the examples — never losing the meaning-preservation instructions.
|
||||
func TestEditorFewShotToggleRealFile(t *testing.T) {
|
||||
tpl, err := LoadPromptTemplate(filepath.Join("..", "..", "prompts", "editor.md"))
|
||||
tpl, err := LoadPromptTemplate(filepath.Join("..", "..", "prompts", "zh-ru", "editor.md"))
|
||||
if err != nil {
|
||||
t.Fatalf("load editor.md: %v", err)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -11,9 +11,12 @@ import (
|
|||
"path"
|
||||
"regexp"
|
||||
"strings"
|
||||
"sync"
|
||||
"unicode"
|
||||
"unicode/utf8"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
|
||||
"golang.org/x/text/encoding/simplifiedchinese"
|
||||
xunicode "golang.org/x/text/encoding/unicode"
|
||||
"golang.org/x/text/transform"
|
||||
|
|
@ -112,12 +115,20 @@ func ingestTXT(p, encoding, sourceLang string) (*Document, error) {
|
|||
|
||||
// --- CJK chapter-header splitting (D18: real zh/ja txt mark chapters as «第N章/节/回») ---------
|
||||
|
||||
// chapterUnitRunes are the section-level chapter markers auto-detected in a txt. 卷 (volume) is
|
||||
// deliberately EXCLUDED — it is coarser than a chapter and would carve a tiny title-only "chapter".
|
||||
var chapterUnitRunes = []rune{'章', '节', '節', '回'}
|
||||
// chapterNumeralRE matches the numeral run of a chapter header: Arabic (half/fullwidth) or CJK. Built ONCE
|
||||
// from the shared lang.CJKSection (pair-14 §4 + addendum-A): the section markers (章节節回) and the CJK
|
||||
// numeral class are the SAME data the chunker heading rule reads — ONE source, no ingest↔chunker byte-drift.
|
||||
// A CONSTANT (the CJK numeral system is language-invariant), so it is available even to a no-langpack book
|
||||
// (the ja→ru golden splits 第X章 here). Arabic ranges stay in code (not language data).
|
||||
var chapterNumeralReOnce sync.Once
|
||||
var chapterNumeralReVal *regexp.Regexp
|
||||
|
||||
// chapterNumeralRE matches the numeral run of a chapter header: Arabic (half/fullwidth) or CJK.
|
||||
var chapterNumeralRE = regexp.MustCompile(`^\s*第[0-90-9〇零一二三四五六七八九十百千两兩]+`)
|
||||
func chapterNumeralRE() *regexp.Regexp {
|
||||
chapterNumeralReOnce.Do(func() {
|
||||
chapterNumeralReVal = regexp.MustCompile(`^\s*第[0-90-9` + lang.DefaultCJKSection().HeadingNumeralClass() + `]+`)
|
||||
})
|
||||
return chapterNumeralReVal
|
||||
}
|
||||
|
||||
// chapterHeaderMaxRunes bounds a header line so a prose sentence that merely opens with «第三节…»
|
||||
// (a longer line) is not mistaken for a header. The 蛊真人 headers are ≤23 runes; 60 leaves room for
|
||||
|
|
@ -138,7 +149,7 @@ func isCJKChapterHeader(line string, unit rune) bool {
|
|||
if utf8.RuneCountInString(t) == 0 || utf8.RuneCountInString(t) > chapterHeaderMaxRunes {
|
||||
return false
|
||||
}
|
||||
loc := chapterNumeralRE.FindStringIndex(t)
|
||||
loc := chapterNumeralRE().FindStringIndex(t)
|
||||
if loc == nil {
|
||||
return false
|
||||
}
|
||||
|
|
@ -170,7 +181,9 @@ func isHeaderSeparator(r rune) bool {
|
|||
// single stray match). Returns 0 when no unit qualifies (→ the part stays a single chapter).
|
||||
func detectChapterUnit(lines []string) rune {
|
||||
best, bestN := rune(0), 0
|
||||
for _, unit := range chapterUnitRunes {
|
||||
// The section markers (章节節回) are shared lang.CJKSection data (pair-14 addendum-A); iterated in AUTHORED
|
||||
// order so a tie between two units breaks deterministically (first-seen), as the fixed slice once did.
|
||||
for _, unit := range lang.DefaultCJKSection().ChapterUnitOrdered {
|
||||
n := 0
|
||||
for _, ln := range lines {
|
||||
if isCJKChapterHeader(ln, unit) {
|
||||
|
|
|
|||
|
|
@ -11,6 +11,7 @@ import (
|
|||
"strings"
|
||||
"unicode"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
"textmachine/backend/internal/store"
|
||||
)
|
||||
|
||||
|
|
@ -498,17 +499,17 @@ func priorityRank(p pickedEntry) int {
|
|||
return rank
|
||||
}
|
||||
|
||||
// glossaryBlockHeader introduces the injected glossary block. The layout mirrors the
|
||||
// SakuraLLM/GalTransl convention "src → dst", with unverified (ambiguous) records
|
||||
// tagged so the model — and a human reviewer — see they are not authoritative (A2).
|
||||
const glossaryBlockHeader = "ГЛОССАРИЙ (используй эти утверждённые переводы имён и терминов последовательно; строки с пометкой ⟨проверить⟩ — неподтверждённые кандидаты):"
|
||||
|
||||
// renderGlossaryBlock serializes the selected records into the injection message
|
||||
// (§C). Records with no dst yet (ruby candidates) are skipped — a "src → " line
|
||||
// carries nothing. Returns "" when nothing renders, so an empty selection injects NO
|
||||
// message at all ("better nothing than garbage"). Deterministic: the injected order is the
|
||||
// budget priority order fixed by Select.
|
||||
func renderGlossaryBlock(injected []pickedEntry) string {
|
||||
// budget priority order fixed by Select. tx is the TARGET-language wire-text (pair-14 §2): a
|
||||
// target with no injection texts (HasData()==false) injects nothing (a non-ru book gets no
|
||||
// Russian block), so the whole render is gated on it.
|
||||
func renderGlossaryBlock(injected []pickedEntry, tx lang.InjectionTexts) string {
|
||||
if !tx.HasData() {
|
||||
return ""
|
||||
}
|
||||
var lines []string
|
||||
for _, p := range injected {
|
||||
if strings.TrimSpace(p.entry.dst) == "" {
|
||||
|
|
@ -516,21 +517,21 @@ func renderGlossaryBlock(injected []pickedEntry) string {
|
|||
}
|
||||
line := p.entry.src + " → " + p.entry.dst
|
||||
if p.disp != memConfirmed {
|
||||
line += " ⟨проверить⟩"
|
||||
line += tx.UnverifiedMarker
|
||||
} else {
|
||||
// pack-13 injection-completeness fix (D39.21 owner directive: «род должен доезжать»): the
|
||||
// gender of a CONFIRMED named term (蛊 Надежды = ж.р., a cicada gu, …) now reaches the DRAFT
|
||||
// wire too, not only the editor — so the TRANSLATOR renders the right родовые формы FIRST,
|
||||
// instead of leaving the editor to repair a wrong gender. "" for a genderless term (the common
|
||||
// case → byte-identical). Mirrors the editor block's confirmed-only gender (renderEditorConstraintBlock).
|
||||
line += genderConstraintNote(p.entry.gender)
|
||||
line += genderConstraintNote(p.entry.gender, tx)
|
||||
}
|
||||
lines = append(lines, line)
|
||||
}
|
||||
if len(lines) == 0 {
|
||||
return ""
|
||||
}
|
||||
return glossaryBlockHeader + "\n" + strings.Join(lines, "\n")
|
||||
return tx.GlossaryHeader + "\n" + strings.Join(lines, "\n")
|
||||
}
|
||||
|
||||
// renderFormatVersion versions the FORMAT of the role-injection renderers whose output is NOT
|
||||
|
|
@ -547,11 +548,10 @@ func renderGlossaryBlock(injected []pickedEntry) string {
|
|||
// draft injected bytes → the wire, so a loud --resnapshot; a genderless bank is byte-identical to v2.
|
||||
const renderFormatVersion = "renderfmt-v3-draft-gender+editor-src2dst+dc3-gender"
|
||||
|
||||
// editorConstraintHeader introduces the editor's canonical-constraint block. It gives the BILINGUAL
|
||||
// editor (D30.1) the approved src→dst bindings as consistency constraints — it must render each
|
||||
// listed source term with EXACTLY the paired canonical form (inflecting for context) and touch
|
||||
// nothing else.
|
||||
const editorConstraintHeader = "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имён и терминов (в черновике термин исходника слева ДОЛЖЕН быть передан именно указанной формой справа — приводи к ней любые расхождения, склоняя по контексту; не вводи иных вариантов и не меняй ничего другого):"
|
||||
// The editor's canonical-constraint block header and the glossary header are TARGET-language wire-text
|
||||
// (lang.InjectionTexts, pair-14 §2) — they gave the BILINGUAL editor (D30.1) the approved src→dst bindings
|
||||
// as consistency constraints. Relocated out of the pipeline (a Russian header rendered for a →en book was a
|
||||
// leak); now gated on the target and sourced from embedded per-target data.
|
||||
|
||||
// renderEditorConstraintBlock serializes the selected records' CONFIRMED renderings into the
|
||||
// editor's injection as a src→dst MAPPING (WS2 §2в resolves the D30.1-open question): "源термин →
|
||||
|
|
@ -568,7 +568,10 @@ const editorConstraintHeader = "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ имё
|
|||
// separate token budget. The src→dst FORMAT is snapshot-folded via renderFormatVersion (a format
|
||||
// edit is a loud --resnapshot); the injection is a message, so it also enters request_hash
|
||||
// directly (no silent false-hit).
|
||||
func renderEditorConstraintBlock(injected []pickedEntry) string {
|
||||
func renderEditorConstraintBlock(injected []pickedEntry, tx lang.InjectionTexts) string {
|
||||
if !tx.HasData() {
|
||||
return ""
|
||||
}
|
||||
var lines []string
|
||||
seen := map[[2]string]bool{}
|
||||
for _, p := range injected {
|
||||
|
|
@ -585,12 +588,12 @@ func renderEditorConstraintBlock(injected []pickedEntry) string {
|
|||
continue
|
||||
}
|
||||
seen[key] = true
|
||||
lines = append(lines, "- "+src+" → «"+dst+"»"+genderConstraintNote(p.entry.gender))
|
||||
lines = append(lines, "- "+src+" → «"+dst+"»"+genderConstraintNote(p.entry.gender, tx))
|
||||
}
|
||||
if len(lines) == 0 {
|
||||
return ""
|
||||
}
|
||||
return editorConstraintHeader + "\n" + strings.Join(lines, "\n")
|
||||
return tx.EditorHeader + "\n" + strings.Join(lines, "\n")
|
||||
}
|
||||
|
||||
// genderConstraintNote is the DC3 gender directive appended to a CONFIRMED editor-constraint line (WS5
|
||||
|
|
@ -598,14 +601,14 @@ func renderEditorConstraintBlock(injected []pickedEntry) string {
|
|||
// the Bai Ninbing class — no coreference needed. male/female ⇒ hard gender forms; hidden ⇒ a mandate to
|
||||
// AVOID gender-marking constructions until the reveal (a masculine default when unavoidable, D19.3). ""
|
||||
// for a term with no gender datum (the common case → the line is unchanged, byte-identical to before).
|
||||
func genderConstraintNote(gender string) string {
|
||||
func genderConstraintNote(gender string, tx lang.InjectionTexts) string {
|
||||
switch gender {
|
||||
case "male", "m":
|
||||
return " (муж. — мужские родовые формы)"
|
||||
return tx.GenderMale
|
||||
case "female", "f":
|
||||
return " (жен. — женские родовые формы)"
|
||||
return tx.GenderFemale
|
||||
case "hidden":
|
||||
return " (пол СКРЫТ до раскрытия — избегай родовых форм; при неизбежности — мужские)"
|
||||
return tx.GenderHidden
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
|
|
|||
|
|
@ -23,7 +23,9 @@ func declJSON(invariant bool, forms ...string) string {
|
|||
return string(b)
|
||||
}
|
||||
|
||||
func alias(a string) store.GlossaryAlias { return store.GlossaryAlias{Alias: a, AliasType: "прозвище"} }
|
||||
func alias(a string) store.GlossaryAlias {
|
||||
return store.GlossaryAlias{Alias: a, AliasType: "прозвище"}
|
||||
}
|
||||
|
||||
// specGlossary mirrors eval/memory_hotpath.py's G (with decl forms replacing the
|
||||
// toy accept-regexps).
|
||||
|
|
@ -237,7 +239,7 @@ func TestSingleKeyBanAndAllowShort(t *testing.T) {
|
|||
func TestBudgetEviction(t *testing.T) {
|
||||
entries := []store.GlossaryEntry{
|
||||
gl("甲甲", "Альфа", "", "approved"),
|
||||
gl("乙乙", "Бета", "", "auto"), // ambiguous → lower priority
|
||||
gl("乙乙", "Бета", "", "auto"), // ambiguous → lower priority
|
||||
gl("丙丙", "Гамма", "", "approved"),
|
||||
}
|
||||
b := bankFrom(entries)
|
||||
|
|
@ -295,7 +297,7 @@ func TestEmptyDstNotMatchable(t *testing.T) {
|
|||
if len(sel.injected) != 1 || sel.injected[0].entry.src != "乙乙" {
|
||||
t.Fatalf("empty-dst phantom polluted selection: injected=%v", injMap(sel))
|
||||
}
|
||||
if renderGlossaryBlock(sel.injected) == "" {
|
||||
if renderGlossaryBlock(sel.injected, ruTX()) == "" {
|
||||
t.Error("a renderable line was evicted by an empty-dst phantom")
|
||||
}
|
||||
// Unbounded: the phantom never appears as an exact hit.
|
||||
|
|
@ -368,8 +370,8 @@ func TestPerLanguageMinKeyAndCollisionDisposition(t *testing.T) {
|
|||
gl("鈴木", "Судзуки", "", "approved"), // 2 Han → fires, CONFIRMED (ideographic anchor)
|
||||
gl("すずき", "Судзуки-х", "s2", "approved"), // 3 kana → fires but AMBIGUOUS (collision-prone)
|
||||
gl("ながいなまえ", "Длинное имя", "", "approved"), // 6 kana → fires, CONFIRMED (long enough)
|
||||
gl("リン", "Рин", "", "approved"), // 2 kana → BANNED (phonetic min 3)
|
||||
gl("ai", "ИИ", "", "approved"), // 2 latin → BANNED (phonetic min 3)
|
||||
gl("リン", "Рин", "", "approved"), // 2 kana → BANNED (phonetic min 3)
|
||||
gl("ai", "ИИ", "", "approved"), // 2 latin → BANNED (phonetic min 3)
|
||||
}
|
||||
b := bankFrom(entries)
|
||||
sel := b.Select("鈴木とすずきとながいなまえが会った。", 1, nil, 0)
|
||||
|
|
@ -548,7 +550,7 @@ func TestRenderEditorConstraintBlock(t *testing.T) {
|
|||
{entry: tanakaDup, via: "田中", disp: memConfirmed},
|
||||
{entry: homonym, via: "済", disp: memConfirmed},
|
||||
}
|
||||
block := renderEditorConstraintBlock(sel)
|
||||
block := renderEditorConstraintBlock(sel, ruTX())
|
||||
|
||||
// src→dst mapping present under the editor's own header.
|
||||
if !strings.Contains(block, "КАНОНИЧЕСКИЕ ПЕРЕВОДЫ") {
|
||||
|
|
@ -570,7 +572,7 @@ func TestRenderEditorConstraintBlock(t *testing.T) {
|
|||
t.Errorf("homonymic dst must stay distinguishable by source term: %q", block)
|
||||
}
|
||||
// An all-AMBIGUOUS (or empty) selection yields no block at all.
|
||||
if got := renderEditorConstraintBlock([]pickedEntry{{entry: cand, disp: memAmbiguous}}); got != "" {
|
||||
if got := renderEditorConstraintBlock([]pickedEntry{{entry: cand, disp: memAmbiguous}}, ruTX()); got != "" {
|
||||
t.Errorf("an all-AMBIGUOUS selection must yield no editor block, got %q", got)
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -43,7 +43,7 @@ func TestSuppressorTrustGateFires(t *testing.T) {
|
|||
t.Errorf("draft 四代族长 must still inject AMBIGUOUS alongside, got %q", m["四代族长"])
|
||||
}
|
||||
// Editor constraint block: the approved term is now present (was empty before the fix).
|
||||
if blk := renderEditorConstraintBlock(sel.injected); !strings.Contains(blk, "глава клана") {
|
||||
if blk := renderEditorConstraintBlock(sel.injected, ruTX()); !strings.Contains(blk, "глава клана") {
|
||||
t.Errorf("editor constraint block must carry the approved «глава клана», got %q", blk)
|
||||
}
|
||||
// Loud record of the refused suppression.
|
||||
|
|
|
|||
|
|
@ -47,10 +47,10 @@ func isFragment(src string, seedSurfaces map[string]bool, p *lang.Pack) bool {
|
|||
}
|
||||
|
||||
// cooccurSameSentence reports whether a and b appear in one source sentence anywhere (alias.
|
||||
// cooccur_same_sentence: split on 。!?\n).
|
||||
func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string) bool {
|
||||
// cooccur_same_sentence: split on the pack's sentence terminators + \n).
|
||||
func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string, p *lang.Pack) bool {
|
||||
for _, c := range chunks {
|
||||
for _, sent := range splitMinerSentences(c.NSource) {
|
||||
for _, sent := range splitMinerSentences(c.NSource, p) {
|
||||
if strings.Contains(sent, aNorm) && strings.Contains(sent, bNorm) {
|
||||
return true
|
||||
}
|
||||
|
|
@ -59,9 +59,11 @@ func cooccurSameSentence(chunks []MinerChunk, aNorm, bNorm string) bool {
|
|||
return false
|
||||
}
|
||||
|
||||
func splitMinerSentences(s string) []string {
|
||||
// splitMinerSentences splits on the pair's source sentence terminators (langpack DATA, pair-14) plus the
|
||||
// structural newline (a layout mark, not a language one, so it stays in code).
|
||||
func splitMinerSentences(s string, p *lang.Pack) []string {
|
||||
return strings.FieldsFunc(s, func(r rune) bool {
|
||||
return r == '。' || r == '!' || r == '?' || r == '\n'
|
||||
return r == '\n' || p.SentenceTerminator[r]
|
||||
})
|
||||
}
|
||||
|
||||
|
|
@ -112,7 +114,7 @@ func proposeAliasEdges(surfaces map[string]aliasSurface, chunks []MinerChunk, se
|
|||
givenA := trimRunePrefix(a.src, surA)
|
||||
givenB := trimRunePrefix(b.src, surB)
|
||||
if givenA != givenB && givenA != "" && givenB != "" { // R4-i different given names → family only
|
||||
if cooccurSameSentence(chunks, a.src, b.src) {
|
||||
if cooccurSameSentence(chunks, a.src, b.src, p) {
|
||||
family = append(family, aliasEdge{a.src, b.src, "R2+R4iv", "family_copresent"})
|
||||
} else {
|
||||
family = append(family, aliasEdge{a.src, b.src, "R2", "family"})
|
||||
|
|
|
|||
|
|
@ -21,32 +21,28 @@ import (
|
|||
// name-shape signal a future faithful-A variant (with a Go morphology backend) would consume, and the
|
||||
// pattern P4 transliteration-by-pair seam (§B5).
|
||||
|
||||
var palladiusJQX = map[string]bool{"j": true, "q": true, "x": true}
|
||||
|
||||
// buildPalladiusCyrSyllables derives the distinct Cyrillic syllable set, longest-first, for greedy
|
||||
// segmentation (palladius._CYR_SYL) from the pack's pinyin→Cyrillic table. Ties in length never both match
|
||||
// a position (distinct strings), so the order among equal-length forms is immaterial — sorted (len desc,
|
||||
// then value) for a stable artifact. The table (initials/finals/Y_W/SPECIAL_I) is langpack DATA.
|
||||
// then value) for a stable artifact. The table (initials/finals/Y_W/SPECIAL_I) AND the phonotactic
|
||||
// constraints (retroflex / ü-finals / their legal initials) are langpack DATA (pair-14).
|
||||
func buildPalladiusCyrSyllables(p *lang.Pack) []string {
|
||||
set := map[string]bool{}
|
||||
for _, v := range p.PalladiusYW {
|
||||
pal := p.Palladius
|
||||
for _, v := range pal.YW {
|
||||
set[v] = true
|
||||
}
|
||||
for _, v := range p.PalladiusSpecialI {
|
||||
for _, v := range pal.SpecialI {
|
||||
set[v] = true
|
||||
}
|
||||
retroflex := map[string]bool{"zh": true, "ch": true, "sh": true, "r": true, "z": true, "c": true, "s": true}
|
||||
for pi, ci := range p.PalladiusInitials {
|
||||
for pf, cf := range p.PalladiusFinals {
|
||||
// ü finals (written v) are valid only after j/q/x (pinyin writes ü as plain u) or l/n.
|
||||
switch pf {
|
||||
case "v", "ve", "van", "vn":
|
||||
if !palladiusJQX[pi] && pi != "l" && pi != "n" {
|
||||
continue
|
||||
}
|
||||
for pi, ci := range pal.Initials {
|
||||
for pf, cf := range pal.Finals {
|
||||
// ü finals (written v) are valid only after a vfinal_initial (pinyin writes ü as plain u).
|
||||
if pal.VFinal[pf] && !pal.VFinalInitial[pi] {
|
||||
continue
|
||||
}
|
||||
// skip the retroflex/sibilant + bare i (handled by SPECIAL_I).
|
||||
if pf == "i" && retroflex[pi] {
|
||||
if pf == "i" && pal.Retroflex[pi] {
|
||||
continue
|
||||
}
|
||||
set[ci+cf] = true
|
||||
|
|
|
|||
|
|
@ -56,13 +56,14 @@ func runeHasPrefixAt(run []rune, i int, p []rune) bool {
|
|||
return true
|
||||
}
|
||||
|
||||
// formantType maps a formant char to a candidate type (patterns.formant_type): topo→place, 等/转/阶→
|
||||
// title, else term.
|
||||
// formantType maps a formant char to a candidate type (patterns.formant_type): topo→place, title-formant
|
||||
// (等/转/阶)→title, else term. Both the place-side (TopoSuffix) and title-side (TitleFormant) inventories are
|
||||
// langpack DATA (pair-14: the title-formant set moved out of this switch, beside its TopoSuffix neighbour).
|
||||
func formantType(c rune, p *lang.Pack) string {
|
||||
if p.TopoSuffix[c] {
|
||||
return "place"
|
||||
}
|
||||
if c == '等' || c == '转' || c == '阶' {
|
||||
if p.TitleFormant[c] {
|
||||
return "title"
|
||||
}
|
||||
return "term"
|
||||
|
|
|
|||
|
|
@ -14,7 +14,10 @@ import (
|
|||
// exercised over the exact same data.
|
||||
func testLangPack(t *testing.T) *lang.Pack {
|
||||
t.Helper()
|
||||
p, err := lang.Load("../../configs/langpacks", "zh", "ru")
|
||||
// Load the shared zh-ru pack PLUS the 蛊真人 book-scoped overlay (pair-14 §1): 古月 is the book's clan,
|
||||
// carried in the overlay, not the shared pair langpack. Every miner fixture (and the stand parity test)
|
||||
// runs on this book's effective pack, so the surname-anchor cases (古月方源) and {古月:22} stay EXACT.
|
||||
p, err := lang.LoadWithOverlay("../../configs/langpacks", "zh", "ru", "testdata/langpack-overlay-guzhenren")
|
||||
if err != nil {
|
||||
t.Fatalf("load langpack: %v", err)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -3,8 +3,14 @@ package pipeline
|
|||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
)
|
||||
|
||||
// ruTX is the ru target injection wire-text (pair-14 §2), for the renderer fixtures that assert the
|
||||
// Russian glossary/editor block bytes (the values that moved to lang/data/injection.txt).
|
||||
func ruTX() lang.InjectionTexts { return lang.InjectionTextsFor("ru") }
|
||||
|
||||
func TestMessagesInjectionLayout(t *testing.T) {
|
||||
tpl := writeTemplate(t, "Стабильный системный префикс.\n---USER---\nПереведи: {{text}}")
|
||||
v := RenderVars{Book: testBook(), Text: "исходный чанк"}
|
||||
|
|
@ -73,7 +79,7 @@ func TestRenderGlossaryBlock(t *testing.T) {
|
|||
ambiguous := pickedEntry{entry: &memoryEntry{src: "小D", dst: "Малыш Дэ", status: "auto"}, disp: memAmbiguous}
|
||||
emptyDst := pickedEntry{entry: &memoryEntry{src: "鈴木", dst: "", status: "auto"}, disp: memAmbiguous}
|
||||
|
||||
block := renderGlossaryBlock([]pickedEntry{confirmed, ambiguous, emptyDst})
|
||||
block := renderGlossaryBlock([]pickedEntry{confirmed, ambiguous, emptyDst}, ruTX())
|
||||
if !strings.Contains(block, "阿Q → А-кью") {
|
||||
t.Errorf("confirmed line missing: %q", block)
|
||||
}
|
||||
|
|
@ -84,10 +90,10 @@ func TestRenderGlossaryBlock(t *testing.T) {
|
|||
t.Errorf("empty-dst candidate should be skipped (nothing to inject): %q", block)
|
||||
}
|
||||
// Nothing to inject → empty block (better nothing).
|
||||
if renderGlossaryBlock(nil) != "" {
|
||||
if renderGlossaryBlock(nil, ruTX()) != "" {
|
||||
t.Error("empty selection must render an empty block")
|
||||
}
|
||||
if renderGlossaryBlock([]pickedEntry{emptyDst}) != "" {
|
||||
if renderGlossaryBlock([]pickedEntry{emptyDst}, ruTX()) != "" {
|
||||
t.Error("only-empty-dst selection must render an empty block")
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -69,6 +69,11 @@ type Runner struct {
|
|||
// enriched `memory`. nil ⇔ memory is nil (materialized together in seedGlossary).
|
||||
baseMemory *MemoryBank
|
||||
|
||||
// checkers is the compiled WS5/pack-13 observability checker spec (pair-14 data-out): the pair's
|
||||
// DETECTION patterns + tables (from pack) + the target-general lists (from embedded target data),
|
||||
// compiled once in loadLangPack. nil-safe: a no-pack / no-target-data book runs the checkers inert.
|
||||
checkers *dcCheckers
|
||||
|
||||
// pack is the book's language-data pack (internal/lang), loaded once in openRunner from
|
||||
// book.LangpackRoot (D39.15/16). The the bank-mining stop bank-miner reads its tables; pack.Version() is folded into
|
||||
// the snapshot (a pack edit is a loud --resnapshot). nil when the book declares no langpack_root, or
|
||||
|
|
@ -191,6 +196,10 @@ func openRunner(bookPath string, logger *slog.Logger, forWrite bool) (*Runner, e
|
|||
// catalog — or a book with no langpack_root — runs with a nil pack (the miner is inert, the bank-mining stop auto-continues,
|
||||
// and the snapshot fold is omitted). Presence of the pair directory is the "catalog exists" signal.
|
||||
func (r *Runner) loadLangPack() error {
|
||||
// The observability checkers compile from the pair pack (DETECTION patterns + tables, may be nil) AND the
|
||||
// embedded target data (target-general lists, keyed by target lang). Built here so it exists even for a
|
||||
// no-langpack book (a nil pack → inert pair checkers; the target lists still load). r.pack is set below.
|
||||
defer func() { r.checkers = compileCheckers(dcCheckerData(r.pack), lang.TargetChecksFor(r.Book.TargetLang)) }()
|
||||
if r.Book.LangpackRoot == "" {
|
||||
return nil // no langpack declared → nil pack, miner inert
|
||||
}
|
||||
|
|
@ -199,7 +208,9 @@ func (r *Runner) loadLangPack() error {
|
|||
// No catalog for this pair → nil-and-run (a ja book against a zh-only root just runs without mining).
|
||||
return nil
|
||||
}
|
||||
pack, err := lang.Load(r.Book.LangpackRoot, r.Book.SourceLang, r.Book.TargetLang)
|
||||
// A book-scoped overlay (LangpackExtend) unions the book's PRIVATE canon onto the shared pair pack
|
||||
// (pair-14 §1: 古月 is a 蛊真人 clan, not a shared 百家姓 surname). "" ⇒ plain Load. Folds into Version().
|
||||
pack, err := lang.LoadWithOverlay(r.Book.LangpackRoot, r.Book.SourceLang, r.Book.TargetLang, r.Book.LangpackExtend)
|
||||
if err != nil {
|
||||
return fmt.Errorf("pipeline: load langpack for %s: %w", r.Book.LangPair(), err)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -316,7 +316,7 @@ func TestRunnerMemoryResnapshotOnApprovedChange(t *testing.T) {
|
|||
// ({{draft}}). The monolingual variant is PRESERVED as prompts/editor-mono.md (the D13.1
|
||||
// confirming pilot arm) — it must NOT reference {{text}}.
|
||||
func TestEditorPromptIsBilingual(t *testing.T) {
|
||||
raw, err := os.ReadFile(filepath.Join("..", "..", "prompts", "editor.md"))
|
||||
raw, err := os.ReadFile(filepath.Join("..", "..", "prompts", "zh-ru", "editor.md"))
|
||||
if err != nil {
|
||||
t.Fatalf("read editor.md: %v", err)
|
||||
}
|
||||
|
|
@ -326,7 +326,7 @@ func TestEditorPromptIsBilingual(t *testing.T) {
|
|||
if !strings.Contains(string(raw), "{{draft}}") {
|
||||
t.Error("editor.md must reference {{draft}} (it edits the draft)")
|
||||
}
|
||||
mono, err := os.ReadFile(filepath.Join("..", "..", "prompts", "editor-mono.md"))
|
||||
mono, err := os.ReadFile(filepath.Join("..", "..", "prompts", "zh-ru", "editor-mono.md"))
|
||||
if err != nil {
|
||||
t.Fatalf("read editor-mono.md: %v (D30.1: the monolingual variant must be preserved for the D13.1 pilot arm)", err)
|
||||
}
|
||||
|
|
|
|||
|
|
@ -7,10 +7,39 @@ import (
|
|||
"strings"
|
||||
"unicode"
|
||||
|
||||
"textmachine/backend/internal/lang"
|
||||
|
||||
"golang.org/x/text/unicode/norm"
|
||||
"golang.org/x/text/width"
|
||||
)
|
||||
|
||||
// ruSanitizer holds the ru-target sanitizer DETECTION patterns as DATA (pair-14 data-out): the sanitizer is
|
||||
// isRuTarget-gated (chunkrun/waverun), so its preamble / trailing-note / edit-meta / invalid-sign patterns
|
||||
// are ru-target data (lang.TargetChecks), loaded once at package init. The strip/detect ALGORITHM stays in
|
||||
// this file. The "ru" key mirrors the isRuTarget call-site gate; threading the target lang so a future non-ru
|
||||
// sanitizer loads its own patterns is a follow-up (like sourceAbbrevs).
|
||||
var ruSanitizer = lang.TargetChecksFor("ru")
|
||||
|
||||
// compileSanitizerREs compiles an ORDERED list of sanitizer patterns by key (e.g. the 4 preamble shapes). A
|
||||
// malformed pattern panics (a corrupt embed, caught by the sanitizer/golden tests).
|
||||
func compileSanitizerREs(key string) []*regexp.Regexp {
|
||||
vals := ruSanitizer.List(key)
|
||||
out := make([]*regexp.Regexp, len(vals))
|
||||
for i, v := range vals {
|
||||
out[i] = regexp.MustCompile(v)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// compileSanitizerRE compiles a single sanitizer pattern by key; the key must carry exactly one value.
|
||||
func compileSanitizerRE(key string) *regexp.Regexp {
|
||||
vals := ruSanitizer.List(key)
|
||||
if len(vals) != 1 {
|
||||
panic(fmt.Sprintf("pipeline: sanitizer pattern %q wants exactly 1 value, got %d", key, len(vals)))
|
||||
}
|
||||
return regexp.MustCompile(vals[0])
|
||||
}
|
||||
|
||||
// sanitizer.go: the output-sanitizer (D30.3) — a deterministic verdict-axis gate on the
|
||||
// FINAL chunk text that catches the "instant unreadability" defect classes NO existing
|
||||
// gate detects (exp12 root-cause + flagman §5). It is OPT-IN (Gates.Sanitizer.Enabled),
|
||||
|
|
@ -206,18 +235,11 @@ func sanitizeOutput(text string) sanitizerResult {
|
|||
// вариант noun, or «перевод фрагмента/текста/…», or «ниже приведён … перевод/текст».
|
||||
//
|
||||
// Go RE2 \w/\b are ASCII-only, so Cyrillic uses [а-яё].
|
||||
var leadingPreamblePatterns = []*regexp.Regexp{
|
||||
// «Вот/Представляю/Привожу … отредактированный/исправленный перевод/текст/вариант …:»
|
||||
// (the edit-adjective front-gate holds precision, so the tail-to-colon may be long — the
|
||||
// gemini leak carries «… с соблюдением всех терминов из глоссария:»).
|
||||
regexp.MustCompile(`(?is)^\s*(?:вот|представляю|привожу|держите)\s+[^\n:]{0,40}?(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:`),
|
||||
// «Отредактированный/Исправленный перевод/текст/вариант …:» (label at the very start)
|
||||
regexp.MustCompile(`(?is)^\s*(?:отредактированн|исправленн|улучшенн)[а-яё]+\s+(?:перевод|текст|вариант)[а-яё]*[^\n:]{0,120}?:`),
|
||||
// «(Вот) перевод фрагмента/текста/отрывка/главы/черновика:»
|
||||
regexp.MustCompile(`(?is)^\s*(?:вот\s+)?перевод\s+(?:фрагмента|текста|отрывка|главы|черновика)\s*:`),
|
||||
// «Ниже приведён/представлен … перевод/отредактированный текст:»
|
||||
regexp.MustCompile(`(?is)^\s*ниже\s+(?:привед[её]н|представлен|дан|след)[а-яё]*\s+[^\n:]{0,40}?(?:перевод|отредактированн|текст)[а-яё]*[^\n:]{0,20}?:`),
|
||||
}
|
||||
// Pair-14 data-out: the actual ru patterns are DATA (lang.TargetChecks "sanitizer_preamble", ORDERED — the
|
||||
// preview returns the FIRST match). The 4 shapes: (1) «Вот… отредактированный перевод…:» edit-adjective front-
|
||||
// gate; (2) «Отредактированный перевод…:» label at the start; (3) «(Вот) перевод фрагмента/…:»; (4) «Ниже
|
||||
// приведён … перевод/текст:».
|
||||
var leadingPreamblePatterns = compileSanitizerREs("sanitizer_preamble")
|
||||
|
||||
func detectLeadingPreamble(text string) string {
|
||||
t := strings.TrimSpace(text)
|
||||
|
|
@ -237,7 +259,7 @@ func detectLeadingPreamble(text string) string {
|
|||
// «Сноска объясняла обычай.» are ordinary narrative nouns/gerunds — NOT headers). So the header
|
||||
// word MUST be followed by a COLON, or be «примечание/заметки/комментарий + переводчика/редактора»,
|
||||
// or be the «(прим. перев./ред.)» marker. Only counted inside the trailing region (below).
|
||||
var trailingNoteRE = regexp.MustCompile(`(?i)^[*_>\s-]{0,4}(?:(?:примечани[ея]|заметк[аи]|комментари[йяю]|пояснени[ея]|сноск[аи])\s*:|(?:примечани[ея]|заметк[аи]|комментари[йяю])\s+(?:переводчик|редактор)[а-яё]*|прим\.\s*(?:перев|ред)\.?)`)
|
||||
var trailingNoteRE = compileSanitizerRE("sanitizer_trailing_note")
|
||||
|
||||
// editMetaAnywhereRE matches UNAMBIGUOUS editor meta-commentary that is a defect wherever it
|
||||
// appears (not only trailing): «Внесённые правки:», «Что изменено:», «Список правок:», and the
|
||||
|
|
@ -248,7 +270,7 @@ var trailingNoteRE = regexp.MustCompile(`(?i)^[*_>\s-]{0,4}(?:(?:примеча
|
|||
// are common narrative openers), keeping only editor-specific compounds that never occur in prose.
|
||||
// «основн…правк» is gated to «… для справки» OR a colon so a bare narrative «Основные правки внёс
|
||||
// редактор» (no summary label) does not fire.
|
||||
var editMetaAnywhereRE = regexp.MustCompile(`(?im)^[*_>\s-]{0,4}(?:(?:внесённ|внесен)[а-яё]+\s+правк|основн[а-яё]+\s+правк[а-яё]*(?:\s+для\s+справк[а-яё]+|\s*:)|что\s+(?:было\s+)?(?:изменен|исправлен)|список\s+(?:правок|изменений)|список\s+внесённых)`)
|
||||
var editMetaAnywhereRE = compileSanitizerRE("sanitizer_edit_meta")
|
||||
|
||||
func detectTrailingNote(text string) string {
|
||||
if loc := editMetaAnywhereRE.FindStringIndex(text); loc != nil {
|
||||
|
|
@ -359,7 +381,7 @@ func hasUpperLatin(s string) bool {
|
|||
// metalinguistic mention of the letter itself, «Буква ъ называлась ером.» (adversarial
|
||||
// review nit). High precision otherwise: word-initial-sign-then-letter and doubled signs are
|
||||
// impossible in well-formed text. \P{L} (non-letter), not RE2's ASCII-only \b.
|
||||
var invalidSignRE = regexp.MustCompile(`(?i)(?:^|\P{L})[ъь][а-яё]|[ъь][ъь]`)
|
||||
var invalidSignRE = compileSanitizerRE("sanitizer_invalid_sign")
|
||||
|
||||
func detectBrokenWords(text string) (int, []string) {
|
||||
var det []string
|
||||
|
|
|
|||
4
backend/internal/pipeline/testdata/langpack-overlay-guzhenren/zh/surnames-compound.txt
vendored
Normal file
4
backend/internal/pipeline/testdata/langpack-overlay-guzhenren/zh/surnames-compound.txt
vendored
Normal file
|
|
@ -0,0 +1,4 @@
|
|||
# 蛊真人 (Reverend Insanity) BOOK-SCOPED langpack overlay (pair-14 §1). 古月 is the protagonist's CLAN,
|
||||
# not a 百家姓 compound surname — it lives here, unioned onto the shared pack FOR THIS BOOK, so the
|
||||
# shared zh langpack stays a clean pair layer while the miner still anchors 古月方源 for 蛊真人.
|
||||
古月
|
||||
|
|
@ -8,6 +8,7 @@ import (
|
|||
"sync"
|
||||
|
||||
"textmachine/backend/internal/config"
|
||||
"textmachine/backend/internal/lang"
|
||||
"textmachine/backend/internal/store"
|
||||
)
|
||||
|
||||
|
|
@ -331,8 +332,9 @@ func (r *Runner) runDraftChunk(ctx context.Context, draftSnapshot string, ch Chu
|
|||
// Serialize the per-role injection blocks via the role→renderer registry (D39 layer 7) from the
|
||||
// precomputed base-bank selection: the translator gets its src→dst glossary block; any other
|
||||
// role's block is rendered too but consumed only if a stage of that role runs in this wave.
|
||||
tx := lang.InjectionTextsFor(r.Book.TargetLang)
|
||||
for role, render := range roleInjectionRenderers {
|
||||
injectionByRole[role] = render(memSel.injected) // pure → map-order-independent
|
||||
injectionByRole[role] = render(memSel.injected, tx) // pure → map-order-independent
|
||||
}
|
||||
if len(memSel.trustGated) > 0 {
|
||||
r.Log.WarnContext(ctx, "memory: lower-trust longer key refused from suppressing a higher-trust nested key (approved term preserved; reconcile the seed)",
|
||||
|
|
@ -442,8 +444,9 @@ func (r *Runner) runEditUnit(ctx context.Context, editSnapshot string, unit edit
|
|||
// Render the per-role injection from the ENRICHED unit selection via the registry (D39 layer 7): the
|
||||
// editor gets its CONFIRMED-dst constraint block; other roles' blocks are rendered but consumed
|
||||
// only if a stage of that role runs in the edit wave.
|
||||
tx := lang.InjectionTextsFor(r.Book.TargetLang)
|
||||
for role, render := range roleInjectionRenderers {
|
||||
injectionByRole[role] = render(editSel.injected)
|
||||
injectionByRole[role] = render(editSel.injected, tx)
|
||||
}
|
||||
}
|
||||
seq, err := r.runStageSequence(ctx, editStages, editSnapshot, leader, unitDraft, injectionByRole)
|
||||
|
|
|
|||
|
|
@ -9,6 +9,10 @@
|
|||
> - **Ждём от владельца:** тачпойнты exp16 (карта подписи/пол · мини-голд алиасов · precision@30, `books/gu-zhenren/exp16/`, ~30–40 мин — для сид-дельты к пере-прогону) · развилки плана §10 (W1.5-UX · DC5-стих-политика · Edit-ceiling · и др.) · реплика ja→ru (тест общности §B5) · **ре-чек прайса DeepSeek у катовера слагов 24.07** (до него платных deepseek-прогонов нет) · чтение пере-прогона (после R1+resnapshot) · старые висящие: планка запуска (лучше-фана/гибрид/издательский) · билингв-якорь пилота (D25.9-Q1) · контаминация пилот-корпуса (D27.4) · publishable/waiver (D25.1) · FN-bound L3 (D25.4) · юр-пакет · провенанс 12-*-доков.
|
||||
> - Архивы хроники: `archive/PROGRESS-2026-07-04-10.md` (D31) · `archive/PROGRESS-2026-07-10-13.md` (D39.6-гигиена). Записи ниже — живая эра D39.
|
||||
|
||||
## Оркестратор №7 — ПАК-14 «генеральность-пасс» ПРИНЯТ (воркфлоу 11 линз + доприёмка 4) и залендён; пак-15 «структура+форма» готовится, 24.07
|
||||
|
||||
Сессия исполнила 9 пунктов + оба аддендума + **расширение по директиве владельца «full data-out»** (детект-регексы чекеров/гейтов/санитайзера → данные, не только lookup-таблицы; `checkers_zh_ru.go` → пар-агностичный `checkers.go`). Два дома данных: `configs/langpacks/`+книго-overlay (`book.yaml: langpack_extend`, аддитивный union, fail-loud на чужой файл) для пар/книго-данных · `internal/lang/data/` go:embed для того, что нужно движку БЕЗ пака (golden ja→ru без langpack это доказал: refusal/CJK-нарезка/→ru-инъекция). **Приёмка исполнением:** 0 блокеров; байт-сверка каждого переноса против HEAD; парити EXACT (`{方源:0 蛊:1 蛊师:2 古月:22}`, 古月 теперь из base∪overlay — общий zh-пак чист); golden байт-нетронут; тесты грузят фикстуры из реального пака. Блокер доприёмки (аддендум parsePalladius не был исполнен) закрыт дельтой: generic-парсер категорий + один типизированный `Pack.Palladius`, required валидирует потребитель; + c2-segmentation явно, packAlgoVersion v2. **Поправка отчёту (ревью-шапка):** Приложение A «distinct/matches 8» — греп-артефакт (0 кодовых ссылок; 35→32/14), хойст-вывод не задет. **Гигиена-инцидент мой:** `77dd9b8` утащил staged `git mv` промптов сессии (git commit коммитит весь индекс) — одна неконсистентная bisect-точка, вылечена этим лендингом; правило: перед лендинг-коммитом `git diff --cached`. Прод-overlay 古月 подключён в оба стендовых book.yaml. **Пак-15 (готовится):** хойст `internal/text`+value-типов → сплит miner/checks/membank/chunk (карта связности+экспорт-поверхности = приложения A/B отчёта пака-14, с моей поправкой) + форм-фиксы Export/RequestHash/resolveChunkState + конфиг-слой пары (пар-калибровки из ран-конфигов, промпт-резолв конвенцией) + перенесённый остаток (тесты Палладия · fail-loud оверлея · unionStringMap-семантика · empty-probe гард · терминаторы chunker/coverage · `第`).
|
||||
|
||||
## Оркестратор №7 — research/22 (доменные харнессы) ПРИНЯТ и залендён; вердикт: калибрующий, не переворачивающий, 24.07
|
||||
|
||||
Сессия сдала ревизию-2 (после СОБСТВЕННОГО 9-агентного адверсариала синтеза, поймавшего её факт-ошибки — прецедент мандата 12.07 в лучшей форме). Мой спот-чек 5/5 (detectCJKLeak/cheapgate-v4/ledger наши · ProofreadTask/GenDic их). **Оценка:** ядровые ставки ПОДТВЕРЖДЕНЫ рынком+академией (fused-майнинг>bolt-on доказан смертью KeywordGacha; DelTA-роадмап = наш пройденный путь; TransAgents назвали нашу D39.20-проблему и не решили; north-star «измеренная редактура» НЕ роадмапится никем — окно открыто). **Мы безоговорочно впереди:** COGS-телеметрия · 18+ · общность · измеренность (recall 0.932, вклад ролей). **Позади:** Q4 выпуск/epub (разрыв РАСТЁТ — bbm/Immersive активно куют) · gate+repair-петля QA (они чинят typed-error→адресный ре-перевод, мы флагаем). **В планы:** слой-3 diff-редактор получает прецедент AiNiee/LinguaGacha (приоритет↑, пак-15/16) · epub-tag-rewrite не держать до Ф3 бесконечно (кандидат после пилота) · Q7-леджер дополнить (глагольный вид/addition/voice-flattening) · Problem.py-классы reflow-aware (char-freq, control-code) · LinguaGacha = бенчмарк-кандидат №1 · SakuraLLM мониторить (канал B/ja→ru). «D39.7-коллизия» разрешена: не коллизия, ссылка на стандарт. Промт → архив.
|
||||
|
|
|
|||
|
|
@ -11,17 +11,17 @@
|
|||
|
||||
## Структура
|
||||
|
||||
- `architecture/` — синтез. **Источник истины по решениям — [`05-decisions-log.md`](architecture/05-decisions-log.md) (D1–D39.16); при конфликте с любым доком он выше.**
|
||||
- `architecture/` — синтез. **Источник истины по решениям — [`05-decisions-log.md`](architecture/05-decisions-log.md) (D1–D39.22+); при конфликте с любым доком он выше.**
|
||||
- `01-decisions.md` — принципы Р1–Р10; `02-mvp-plan.md` — фазы и приёмка (v3, 09.07); `03-implementation-notes.md` — контракты Фазы 0; `04-unhappy-paths.md` — ~70 режимов отказа → механизм; `06-memory-risk-registry.md` — реестр рисков банка памяти; **[`07-strategic-review.md`](architecture/07-strategic-review.md) — стратегический аудит (09.07): вердикт end-to-end, топ-риски, курс-коррекции**; **[`08-sync-audit-ledger.md`](architecture/08-sync-audit-ledger.md) — верифицированный ледджер синк-аудита «сломанного телефона» (65 находок, D39)** · **[`09-target-architecture.md`](architecture/09-target-architecture.md) — целевая 7-слойная архитектура + фазовый план арх-ресета (D39; инвариант общности §0.1)** · [`10-prompt-architecture.md`](architecture/10-prompt-architecture.md) — консолидированная промпт-заметка (концерн 4); **[`11-implementation-plan.md`](architecture/11-implementation-plan.md) — РАТИФИЦИРОВАННЫЙ план стройки пере-прогонного стека (D39.12; дизайн-оф-рекорд бэкенд-пака, три рубежа верификации)**; `components.puml`/`pipeline.puml` — диаграммы v3 (владелец смотрит PlantUML-расширением VS Code; вручную НЕ рендерить).
|
||||
- `experiments/` — эмпирика «Полигона»: `00-provider-quirks` (читать перед любым вызовом провайдера), `01-token-calibration`, `02-refusal-benchmark`, `03-local-stand`, `04-editor-quality`, `06-local-extraction`, `07-coverage-precision`, `08-cost-model-v2` (актуальная денежная модель), `09-pilot-protocol` (пилот Ф2.5 + поправки D13), `10-explicit-benchmark` (18+ violence-рука канала B), `11-erotica-benchmark` (erotica по трём парам/регистрам — закрытие D14.4, D22).
|
||||
- `research/` — фактура исследований 04–05.07: `01–10` базовые, `11-gap-*` добор критиком, `12-*` режимы отказа/отзывы/таксономии (+ два внешних материала с провенанс-шапками), `13` валидация памяти, `14` адаптивная память, `15` голос и состояние (принят, D21), **`16` ридер-IDE (принят с ревью-шапкой, D29)**, **`17` внешняя критика GPT-5.6 (принят с ревью-шапкой, D25)** — у 16/17 читать шапку прежде тела. **`18` рычаги качества (два отчёта, D36)** · **`19` нарезка+когезия+контракт t/e (D39.1)** · **`20` банк-майнинг W1.5 (D39.6)** · **`21` обзор LLM-транспорта чужих харнессов (23.07: наш транспорт опережает/вровень со всеми 11)** · **`22` доменные харнессы перевода (24.07: калибрующий — ядро подтверждено, впереди COGS/18+/общность/измеренность, позади Q4-выпуск и gate+repair-петля; сиквел `02`/`05`)** — у всех ревью-шапки. ⚠ Часть под superseded-баннерами (01/02/03/04/05/09 и gap-1/2/5) — **читай баннер прежде содержимого**.
|
||||
- `PROGRESS.md` — **журнал** (CURRENT-STATE сверху, ниже хронология; НЕ источник решений).
|
||||
- **Активные хендофф-промты (пост-паки-12/13, 24.07):** [`BACKEND_GENERALITY_PASS_SESSION_PROMPT.md`](BACKEND_GENERALITY_PASS_SESSION_PROMPT.md) (**пак-14 генеральность-пасс — ТЕКУЩИЙ**: подтверждённые утечки пара/книга-данных из Go в langpack/сид; байт-точность+парити EXACT; норматив — [`architecture/12-go-style-notes.md`](architecture/12-go-style-notes.md)) · [`ORCHESTRATOR_SESSION_PROMPT.md`](ORCHESTRATOR_SESSION_PROMPT.md) (хендофф №5) · [`POLYGON_PACKAGE4_SESSION_PROMPT.md`](POLYGON_PACKAGE4_SESSION_PROMPT.md) (residual пилот/18+/echo). **Паки 12/13 ЗАЛЕНДЕНЫ 24.07** (`7fe0a5b`/`2f91b04`, приёмка 25-агентным воркфлоу; их промты + research/22-промт → архив после ревью research/22). Дальше: пак-15 = канал B вживую + хвосты слоя 7 → пак-16 = слой 3 diff-based редактор → голос/состояние D21 + C2/C3. Закрытые — в `archive/prompts/`.
|
||||
- **Активные хендофф-промты (пост-пак-14, 24.07):** [`ORCHESTRATOR_SESSION_PROMPT.md`](ORCHESTRATOR_SESSION_PROMPT.md) (хендофф №5) · [`POLYGON_PACKAGE4_SESSION_PROMPT.md`](POLYGON_PACKAGE4_SESSION_PROMPT.md) (residual пилот/18+/echo). **Паки 12/13/14 ЗАЛЕНДЕНЫ 24.07** (`7fe0a5b`/`2f91b04`/пак-14 — приёмка воркфлоу-линзами; отчёт пака-14 с ревью-шапкой — [`archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md`](archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md); норматив общности — [`architecture/12-go-style-notes.md`](architecture/12-go-style-notes.md)). Дальше: **пак-15 «структура+форма»** (хойст text/model-базы → сплит miner/checks/membank/chunk + Export/RequestHash/resolveChunkState + конфиг-слой пары; промт готовится) → пак-16 = канал B вживую / слой 3 diff-based редактор (прецедент research/22) → голос/состояние D21 + C2/C3. Закрытые — в `archive/prompts/`.
|
||||
- `archive/` — закрытые сессионные промты (только история, инструкции оттуда не исполнять).
|
||||
|
||||
## Статус (2026-07-23, пост-D39.20 — механизация выпускного качества)
|
||||
|
||||
Фаза 0 ✅; Ф1-инфра ✅ (D20–D28). **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ:** 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1–D39.10) → план (D39.12, `11-*`) → стройка пака-11 (D39.13–D39.19: волновой исполнитель · чанкер output-бюджет · src→dst-редактор · Go-майнер паритет-EXACT · банкнота · `internal/lang`+langpacks · reject-set · echo-сплит). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07):** операционка вся подтверждена живьём (ОДИН resnapshot · **$0-резюм драфта ×3 арма** · echo 0 · $6.63<$15 · майнинг-стоп→подпись→reject-set); **планка ≤2 НЕ пройдена** — но дефект-классы полностью механизируемы. Три сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232, floor-шум 0 · два слепых чтения): **анти-корреляция гладкость↔верность доказана трижды**; **mistral исключён как несущий редактор** (голос = north-star промпта); dspro-vs-glm → итерация №2 после механизации (драфт $0). research/21: транспорт наш опережает/вровень со всеми 11, рефакторинг отвергнут. **Эксп ЗАКРЫТ (D39.22), итерация №2 отложена; редактор = deepseek-v4-pro ИНТЕРИМ (glm-резерв); стайл-каноны D39.21 ратифицированы.** **Очередь (курс — разработка бэкенда):** пак-12 транспорт-гигиена (лендится первым) → пак-13 «выпускной QA» (wire-двигающий; вкл. дефолт-редактор dspro) → сид-дельта (元/赤城/学堂家老/春秋蝉) → масштаб целой книги + канал B вживую + Ф2-механизмы (D21 голос/состояние, native-Gemini судья) + ja→ru §B5 → пилот Ф2.5 (блокеры прежние: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Судья 18+ = grok-фолбэк (квирк D22.6 подтверждён живьём). Ключ xAI: единый, data-sharing, off перед продом (D27).
|
||||
Фаза 0 ✅; Ф1-инфра ✅ (D20–D28). **АРХ-РЕСЕТ D39 ИСПОЛНЕН ЦЕЛИКОМ:** 7-слойная архитектура (`09-*`) → исследовательская программа (D39.1–D39.10) → план (D39.12, `11-*`) → стройка пака-11 (D39.13–D39.19: волновой исполнитель · чанкер output-бюджет · src→dst-редактор · Go-майнер паритет-EXACT · банкнота · `internal/lang`+langpacks · reject-set · echo-сплит). **ПЕРЕ-ПРОГОН rerun2 ИСПОЛНЕН И ПРОЧИТАН (D39.20, 23.07):** операционка вся подтверждена живьём (ОДИН resnapshot · **$0-резюм драфта ×3 арма** · echo 0 · $6.63<$15 · майнинг-стоп→подпись→reject-set); **планка ≤2 НЕ пройдена** — но дефект-классы полностью механизируемы. Три сигнала (судья dspro 0.679>glm 0.589>>mistral 0.232, floor-шум 0 · два слепых чтения): **анти-корреляция гладкость↔верность доказана трижды**; **mistral исключён как несущий редактор** (голос = north-star промпта); dspro-vs-glm → итерация №2 после механизации (драфт $0). research/21: транспорт наш опережает/вровень со всеми 11, рефакторинг отвергнут. **Эксп ЗАКРЫТ (D39.22), итерация №2 отложена; редактор = deepseek-v4-pro ИНТЕРИМ (glm-резерв); стайл-каноны D39.21 ратифицированы.** **Очередь (курс — разработка бэкенда):** паки 12/13/14 залендены (транспорт-гигиена · выпускной QA · генеральность: данные→langpack/embedded, книго-overlay 古月, пар-агностичные чекеры) → пак-15 «структура+форма» → сид-дельта (元/赤城/学堂家老/春秋蝉) → канал B вживую + слой-3 diff-редактор + Ф2-механизмы (D21 голос/состояние, native-Gemini судья) + ja→ru §B5 → пилот Ф2.5 (блокеры прежние: билингв-якорь D25.9-Q1, корпус, судья-дублёр D22.6). Судья 18+ = grok-фолбэк (квирк D22.6 подтверждён живьём). Ключ xAI: единый, data-sharing, off перед продом (D27).
|
||||
|
||||
## Доступные ключи от моделей
|
||||
DEEPSEEK_API_KEY, ZAI_API_KEY, KIMI_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, XAI_API_KEY, MISTRAL_API_KEY
|
||||
|
|
|
|||
253
docs/archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md
Normal file
253
docs/archive/reports/PACK14_GENERALITY_REPORT_2026-07-24.md
Normal file
|
|
@ -0,0 +1,253 @@
|
|||
# PACK-14 «генеральность-пасс» — приёмочный отчёт (2026-07-24)
|
||||
|
||||
> **Ревью-шапка оркестратора №7 (24.07, приёмка исполнением, два воркфлоу: 11 линз ~970k ток. + доприёмка дельты 4 линзы; 0 блокеров).** Байт-точность КАЖДОГО переноса независимо сверена против git HEAD (refusal 12/12 · injection 6/6 с ведущими пробелами · dc-регексы · санитайзер · CJK-числительные · фонотактика · форманты — ноль расхождений); парити EXACT свежими `-count=1` прогонами; golden байт-нетронут; тесты не ослаблены (2 удалённых ассерта — законные замены). **Поправки к телу отчёта:** (а) Приложение A — ведро «distinct/matches (8 refs)» = греп-артефакт: все 8 в комментариях, кодовых ссылок 0; заголовок «miner→membank 35» реально 32 полных вхождений / 14 кодовых; остальные вёдра точны, вывод «хойст первым» НЕ задет; (б) ёфикатор = 38 слов (строка «40 ru-слов» — обсчёт в прозе); (в) свип-47 воспроизводится как 41 (точный класс) + 6 строк ё/кана (класс сессии был шире заявленного регекса), скрытых пар-данных нет; класс «пунктуация-сепараторы» в 47 не представлен (перенос из классификации-100). **Дельта доприёмки принята** (generic-Палладий по аддендуму, c2-segmentation, README-пути, want-list magnitude, packAlgoVersion v2). **Остаток в пак-15 (зафиксирован):** тесты unknown-category-толерантности и required-fail Палладия · fail-loud оверлея уровнем выше (пропавший overlayRoot/чужой подкаталог/категория-опечатка в overlay-файле — тихие) · семантика override unionStringMap · isRuTarget/'ru'-швы · empty-probe гард lintMagnitudeScale · дуп терминаторов chunker/coverage · `第`. Прод-оверлей 古月 подключён оркестратором в оба стендовых book.yaml (rerun2).
|
||||
|
||||
Бэкенд-сессия. Источник: `docs/BACKEND_GENERALITY_PASS_SESSION_PROMPT.md` (9 пунктов) + два аддендума
|
||||
оркестратора (A: ingest `章节節回`+CJK-числительные → та же ед. точки истины, что у чанкера; B: форманты
|
||||
`等/转/阶`). Норматив: `docs/architecture/12-go-style-notes.md` §0. **Перенос ДАННЫХ, не редизайн алгоритмов.**
|
||||
**Сессия НЕ коммитит.**
|
||||
|
||||
## Вердикт: ✅ жёсткий инвариант выполнен
|
||||
|
||||
- **golden БАЙТ-ИДЕНТИЧЕН** — `testdata/golden/capture.golden` не тронут (`git diff` пуст). Сильнее допущенного
|
||||
«version-only»: у golden-книги нет langpack (ja→ru, `langpack_root` не задан → `pack==nil`), поэтому
|
||||
`packVersion()==""` и её снапшот НЕ двигается ни от одного НОВОГО langpack-файла; `TestGolden` сверяет
|
||||
capture побайтно и зелёный.
|
||||
- **майнер-парити EXACT** — `TM_MINER_PARITY=1`: `n=13618 catastrophe{方源:0 蛊:1 蛊师:2 古月:22}
|
||||
recall@proposed=0.9655` — тождественно дореформенному (вкл. 古月-механику через книго-расширение).
|
||||
- **`go build/vet/test -race` зелёные** (весь модуль, вкл. новый `internal/lang`).
|
||||
|
||||
## Ключевой факт, определивший архитектуру
|
||||
|
||||
golden — ja→ru книга БЕЗ langpack (`pack==nil`), но она живьём: (1) флагает `hard_refusal` (RU «Не могу
|
||||
помочь»), (2) инъектит русские заголовки глоссария в wire (16×), (3) режет главы `第X章`, (4) style_flags=0.
|
||||
Значит рефузал-паттерны, CJK-числительные глав и →ru-инъекцию НЕЛЬЗЯ держать в книжном паке (nil для golden).
|
||||
`request_hash` фолдит `snapshotID` (render.go) — любое движение снапшота golden-книги сдвинуло бы wire. →
|
||||
|
||||
## Два дома данных (правило, выдержавшее «100 языков» линзу владельца)
|
||||
|
||||
1. **`configs/langpacks/<src>|<pair>/` + книжный overlay** — пер-исходник / пер-пара / пер-книга данные.
|
||||
Добавить язык = **положить каталог, без перекомпиляции**; убрать = удалить каталог; загрузчик **падает
|
||||
ГРОМКО, называя файл** (никогда тихо-пусто). Сюда: item 1 (surnames/古月), 5 (терминаторы, Палладий-
|
||||
фонотактика), 6 (DC-таблицы), addB (форманты). Overlay: добавить/убрать КНИГУ = ±один каталог;
|
||||
DC-таблицы OPT-IN (убрать файл → чекеры инертны, не болтаются).
|
||||
2. **`internal/lang/data/` (go:embed, секционировано по языку/таргету)** — данные, нужные движку БЕЗ книжного
|
||||
пака (golden доказал: рефузал, CJK-нарезка, →ru-инъекция — все без langpack). Добавить язык = добавить
|
||||
секцию в один data-файл; убрать = удалить секцию. Перекомпиляция — приемлемо: это инженерные КОНСТАНТЫ
|
||||
(версионируются с бинарём, как код), а байт-точность no-pack golden запрещает фолдить их в книжный пак.
|
||||
Сюда: item 3 (refusal.txt), 4+addA (cjk-section.txt), 2 (injection.txt), sourceAbbrevs (sentence-abbrev.txt).
|
||||
|
||||
## По пунктам (что сделано, чем доказана нейтральность)
|
||||
|
||||
**1. 古月 ВОН из общего zh-langpack.** Удалён из `configs/langpacks/zh/surnames-compound.txt`. Механизм —
|
||||
книго-скоупный overlay: новое поле `book.yaml: langpack_extend` (каталог той же раскладки), `lang.LoadWithOverlay`
|
||||
делает АДДИТИВНЫЙ union (множества дополняются, слайсы append; overlay ничего не удаляет), байты overlay
|
||||
фолдятся в `Version()` (правка книжного канона = громкий --resnapshot ТОЛЬКО этой книги). Парити: `testLangPack`
|
||||
грузит `testdata/langpack-overlay-guzhenren` → эффективный пак = старый (base ∪ {古月}) → {古月:22} EXACT. Стенд-
|
||||
overlay создан `/home/ubuntu/books/gu-zhenren/langpack-extend/zh/surnames-compound.txt`. Прод-book.yaml —
|
||||
см. «Хвосты владельцу».
|
||||
|
||||
**2. Инъекция-тексты → таргет-данные, с гейтом.** `glossaryBlockHeader`/`editorConstraintHeader`/маркер
|
||||
`⟨проверить⟩`/3 gender-ноты вынесены в `internal/lang/data/injection.txt` (target-keyed, `lang.InjectionTextsFor`),
|
||||
рендереры гейтятся на `tx.HasData()` — →en книга получит ПУСТО, не русский блок (это и есть фикс утечки).
|
||||
Значимые ведущие пробелы (` (муж.…`, ` ⟨проверить⟩`) сохранены дословно (парсер НЕ триммит value). `renderFormatVersion`
|
||||
НЕ тронут → снапшот golden не двигается → wire байт-идентичен (проверено golden). Langpack-ФАЙЛ (а не embedded)
|
||||
отложен: golden без пака потребовал бы target-loader вне book-langpack_root, что сдвинуло бы снапшот golden —
|
||||
запрещено инвариантом.
|
||||
|
||||
**3. Рефузал-блэклист → универсальные embedded-данные.** `internal/lang/data/refusal.txt` (секции en/ru/ja/zh
|
||||
+ маркер), `disposition.go` строит `refusalRE` из `lang.RefusalPatterns()`. Байт-в-байт 12 паттернов в том же
|
||||
порядке (сверено скриптом). golden флагает `hard_refusal` тождественно.
|
||||
|
||||
**4 + addendum-A. CJK-числительные — ОДНА точка истины.** `internal/lang/data/cjk-section.txt` (digit/unit
|
||||
значения, zero, chapter_unit) — читают И чанкер (`isHeadingNumeral`/`cjkSectionDigit`/`cjkSectionUnit`, путь
|
||||
заголовков, nil для golden), И ingest (`chapterNumeralRE` строится из `HeadingNumeralClass()`, `detectChapterUnit`
|
||||
итерит `ChapterUnitOrdered` — авторский порядок сохранён для детерминированного тай-брейка). Устранён
|
||||
ingest↔chunker байт-дубль. golden режет `第X章` тождественно (ingest живёт для golden — верифицировано TestGolden).
|
||||
`sourceAbbrevs` → `internal/lang/data/sentence-abbrev.txt` (секции по ИСХОДНИКУ, en=29 токенов); чанкер потребляет
|
||||
секцию `en` (как и раньше — универсально; исходник-параметризация сплиттера = мелкий отложенный шаг, чтобы не
|
||||
трогать 18 колл-сайтов SplitChunks на не-приёмочном пути).
|
||||
|
||||
**5. Терминаторы `。!?` + Палладий-фонотактика.** `zh/sentence-terminator.txt` (`p.SentenceTerminator`,
|
||||
`splitMinerSentences`); фонотактика jqx/ретрофлексы/ü-финалы → `zh-ru/palladius-phonotactics.txt`
|
||||
(`PalladiusRetroflex/VFinal/VFinalInitial`), l/n-исключение стало данными `vfinal_initial`. Майнер-only,
|
||||
парити EXACT.
|
||||
|
||||
**6. DC-таблицы → пар-данные, регексы = алгоритм.** `dcCNNum`+`dcRuHourWord`+`dcRegisterNegList` → OPTIONAL
|
||||
`zh-ru/dc-checkers.txt` (`lang.DCCheckerData`), прокинуто через `cheapGateConfig` (nil пак → пусто → чекеры
|
||||
0 → golden style_flags=0). ДЕТЕКТ-регексы (`时辰`/`千万`/`成`/`терем`-паттерны) остались как пар-скоупный
|
||||
алгоритм чекера (§12.2 — «данные наружу, алгоритм остаётся»).
|
||||
|
||||
**7. Калибровка (config/pipeline.go).** Только пере-документирование (числа НЕ тронуты): дефолты 1797/3200/
|
||||
1.1978/0.3852 помечены как «zh-ru пар-калибровка = последний generic-fallback», шиппинг-конфиги (c1/армы) ставят
|
||||
блок ЯВНО (пар-конфиг = источник истины). Релокация в langpack ОТВЕРГНУТА (была бы мёртвой для всех живых путей:
|
||||
конфиги ставят явно, у golden нет langpack; least-mechanism §12.1). Числа держатся EXACT: no-langpack книга
|
||||
(golden) чанкается на этом fallback — пере-выведенное число сдвинуло бы её границы→wire.
|
||||
|
||||
**addendum-B. Форманты `等/转/阶`.** `zh/title-formant.txt` (`p.TitleFormant`, `formantType`). Рядом с TopoSuffix
|
||||
(сосед уже data-driven). Майнер-only, парити EXACT.
|
||||
|
||||
**9. Промпты → `prompts/zh-ru/`.** `backend/prompts/{translator,editor,editor-mono,judge-selector}.md` →
|
||||
`prompts/zh-ru/` (git mv, БАЙТЫ не тронуты — снапшот фолдит `PromptSHA256` от СОДЕРЖИМОГО, не пути). Обновлены 5
|
||||
шиппинг-конфигов (`../prompts/zh-ru/…`) + 4 тест-ссылки на реальные файлы. Тесты, грузящие реальные c1/арм-конфиги
|
||||
(prompt_pack_test, echo_mine_test), зелёные.
|
||||
|
||||
## Item 8 — свип-добор (полный grep non-test `.go`, построчно)
|
||||
|
||||
Метод: изолирован **non-comment** grep `[\x{4e00}-\x{9fff}]|[а-яА-Я]` (строчные комментарии срезаны). Было 133
|
||||
дата-литерала → **перенесено 33** (items 1–6+addA/B) → **осталось 100**, классификация:
|
||||
|
||||
- **SQL-DDL комментарии (migrate.go, 26):** русская документация ВНУТРИ raw-string схемы. Не пар-данные, 0
|
||||
поведения. OK / вне скоупа (это перевод комментариев, отдельная забота).
|
||||
- **Диагностика/сообщения-флагов (~11):** `book.go:188`(Р7), `models.go:184`(Р4)/`231`, `pipeline.go:411`(Р2),
|
||||
`checkers_zh_ru.go:63/119/124/200`, `cheapgates.go:451/294/319`, `checkers:285` — operator-диагностика /
|
||||
observability-строки (англ.-первичные, цитируют матч для человека). НЕ wire, НЕ пар-данные. OK.
|
||||
- **Пар-скоупный ДЕТЕКТ-алгоритм (регексы), §12.2 (данные уже вынесены, регекс остаётся):**
|
||||
`checkers_zh_ru.go:34/38/101/102/103/117/118/122/175/178/187/283` (DC1/DC2/percent/broken-word детект);
|
||||
`sanitizer.go:213/215/217/219/240/251/362` (ru-таргет editorial-preamble/note/invalid-sign детект). OK.
|
||||
- **Алгоритм-инвариантные литералы (OK-generic):** `chunker.go:222` (класс СЕПАРАТОРОВ пунктуации :、,。…);
|
||||
`memnorm.go:180-181` (ё→е фолд ключа памяти, 1-символьное правило); `memseed.go:570` (ー/・ в кана-ридингах,
|
||||
ja-норм); `miner_palladius.go:71` (ъ-drop перед сегментацией). `ingest.go:128` — мой новый регекс, `第`
|
||||
остался литералом (кандидат — см. ниже).
|
||||
- **Ратифицированный DEFER (НЕ тронут):** `cheapgates.go:361` translit-interj блоклист (ара-ара/маа/…) —
|
||||
D39.16 явно (ja-филлеры, мис-кей-ловушка). ✅ оставлен.
|
||||
|
||||
### Кандидаты СЛЕДУЮЩЕГО пака (тот же класс, НЕ в подтверждённом фикс-листе — решение оркестратора)
|
||||
Одного класса с перенесёнными, но не названы в 9 пунктах/аддендумах; байт-нейтральны (golden фаерит 0), но
|
||||
переносить их сейчас = расширять скоуп сверх санкционированного + повтор D39.16-концерна «живой пара-агностик».
|
||||
Флагаю для след. пака:
|
||||
- **`cheapgates.go:265-276` yofikator homograph-whitelist (40 ru-слов)** — ru-таргет readability-данные (класс
|
||||
DC6). Нужен ru-TARGET langpack-слот (сейчас нет).
|
||||
- **`cheapgates.go:459-619` 万/億 magnitude-gate** — zh-исходник числ-значения (万萬億亿兆 — ДРУГОЙ инвентарь, не
|
||||
section) + ru-стемы (тысяч/миллион/…). Кандидат: расширить `lang.CJKSection` категорией `magnitude` + ru-слот.
|
||||
- **`sanitizer.go` ru-фрагменты** — data-экстракция лексем-детекторов в ru-таргет-слот (регексы = алгоритм).
|
||||
- **`ingest.go:128` `第`** — маркер главы; можно добавить в `cjk-section.txt` рядом с chapter_unit (1 руна).
|
||||
|
||||
## «100 языков» линза (ответ на вопрос владельца)
|
||||
|
||||
- **Добавить пару:** положить `configs/langpacks/<src>/`(11 required src-файлов)+`<src>-<tgt>/`(2 pair + опц.
|
||||
heading/dc-checkers) + book.yaml. НИ строки под `internal/pipeline`. Загрузчик fail-loud называет отсутствующий
|
||||
файл. Смелл: `title-formant`/`surnames-*` осмысленны только для CJK-исходника (весь минер CJK-центричен —
|
||||
ПРЕ-существующее свойство, не ухудшено; для не-CJK исходника миннер сам подлежит переосмыслению).
|
||||
- **Добавить таргет:** секция в `injection.txt`/`refusal.txt` — правка ДАННЫХ + перекомпиляция (go:embed).
|
||||
Оправдано: инженерные универсалии/таргет-генерики.
|
||||
- **Убрать язык:** удалить каталог пака / секцию embedded-файла — ничего не болтается (DC opt-in инертен;
|
||||
overlay опционален; загрузчик громкий). Тесты не ломаются на удалении пары (fail-loud по требованию).
|
||||
|
||||
## Хвосты владельцу (операционка, вне зоны backend/)
|
||||
|
||||
- **Прод-`book.yaml` gu-zhenren** (стенд `rerun2/book.yaml`, `book-mistral.yaml`): добавить
|
||||
`langpack_extend: /home/ubuntu/books/gu-zhenren/langpack-extend` — иначе прод-майнинг потеряет 古月 (overlay-
|
||||
каталог уже создан на стенде). Не тронул сам: зона книжных данных владельца + закрытый эксперимент rerun2.
|
||||
Любой ре-ран и так --resnapshot (langpack Version() сдвинут новыми файлами).
|
||||
- **Стенд pipeline-конфиги** (`rerun2/pipeline-*.yaml`) и **docs**, ссылающиеся на старый путь `backend/prompts/…`
|
||||
(истор. записи в archive/experiments) — обновить путь на `prompts/zh-ru/…` при следующем касании (зона
|
||||
оркестратора/владельца).
|
||||
|
||||
## Само-проверка (мандат 12.07) + адверсариальный 4-линзовый ревью
|
||||
|
||||
Ревью ИСПОЛНЕНИЕМ: каждый item верифицирован golden+parity сразу после кода (не в конце). Линзы: байт-точность /
|
||||
парити / общность-ja-ru / **масштаб-100-языков (доп. линза владельца)**.
|
||||
|
||||
### Итог адверсариального 4-линзового ревью (воркфлоу, 4 независимых агента, 289k токенов)
|
||||
|
||||
- **БАЙТ-ТОЧНОСТЬ: CLEAN** — не опровергнуто. Агент независимо байт-сверил КАЖДЫЙ перенесённый литерал против
|
||||
git HEAD: injection (6 значений вкл. ведущие пробелы: unverified_marker len=13, gender len=31), refusal
|
||||
(12 паттернов, joined_equal=True), cjk-section (digit/unit/zero MATCH, chapter_unit ПОРЯДОК [章 节 節 回],
|
||||
heading-класс symmetric-diff пуст), sentence-abbrev (en=29 exact), dc-checkers (11+9+6 rows match), Палладий-
|
||||
фонотактика/форманты/терминаторы MATCH, промпты = чистые rename (0 байт). golden capture пуст в git, TestGolden
|
||||
байт-сверка PASS.
|
||||
- **МАЙНЕР-ПАРИТИ: CLEAN** — не опровергнуто (свежий -count=1). Overlay-union доказан аддитивным (base−古月 ∪
|
||||
{古月} = старый 16-сет); l/n-исключение фонотактики логически тождественно старому switch.
|
||||
- **ОБЩНОСТЬ: CONCERNS** — 1 MINOR + 3 NOTE (ниже, все либо пофикшены, либо документированы).
|
||||
- **МАСШТАБ-100-ЯЗЫКОВ: CONCERNS** — 1 MAJOR (пофикшен) + 3 MINOR/NOTE (документированы).
|
||||
|
||||
**Пофикшено по ревью (пост-ревью, верифицировано golden+parity):**
|
||||
- **MAJOR (масштаб):** `LoadWithOverlay` ТИХО игнорировал overlay-файл вне манифеста (опечатка `surname-compound.txt`
|
||||
или книжный `heading.txt` в overlay → приватный канон не доезжает до майнера, recall падает БЕЗ сигнала). Фикс:
|
||||
`overlayDirFiles` сканирует overlay-каталоги и **падает громко**, называя неожиданный файл. + тест
|
||||
(misnamed-overlay fails loud). Восстанавливает fail-loud-гарантию на масштабе десятков книг-overlay.
|
||||
- **NOTE (масштаб):** `refusal.txt` — комментарий-заголовки `# --- xx ---` вводили в заблуждение («удали секцию =
|
||||
отключи детект для пары»). На деле `RefusalPatterns()` читает плоско, ВСЕ паттерны фаерят для ВСЕХ книг. Комментарий
|
||||
переписан: секции ОРГАНИЗАЦИОННЫЕ; удаление триммит УНИВЕРСАЛЬНЫЙ сет (только `#`-строки — байты паттернов, golden не сдвинут).
|
||||
- **MINOR (масштаб):** header `srcFiles` — уточнён скоуп: «drop a directory» верно ВНУТРИ zh-family name-miner
|
||||
морфо-схемы; другое семейство исходников требует новых Pack-полей (не просто каталог).
|
||||
|
||||
**Документированные остатки (не фикшу — конфликт с инвариантом/мандатом/пре-существующее):**
|
||||
- **MINOR (общность): калибровка = ТИХИЙ zh-ru дефолт** при пропуске блока (ломает fail-loud). НЕ фиксится:
|
||||
item 7 предписал «оставить механизм», а golden (ja→ru) ПРОПУСКАЕТ блок и опирается на этот fallback — падение
|
||||
громко на пропуске сломало бы golden. Тайтен возможен лишь когда golden сам явно проставит сегментацию (будущий
|
||||
пак, который и так пере-снимет golden).
|
||||
- **NOTE (общность+масштаб): `sourceAbbrevs` захардкожен на `en`** — исходник-параметризация сплиттера отложена (18
|
||||
колл-сайтов SplitChunks, не-приёмочный путь). Секция не-en = мёртвые данные до треда `sourceLang`.
|
||||
- **NOTE (общность+масштаб): DC2 `lintMagnitudeScale` держит НЕЗАВИСИМЫЙ CJK-числ-парсер** (万/億/兆), дублирующий
|
||||
`lang.CJKSection` — пре-существующий, self-gating (инертен на не-zh), не тронут этим диффом. Кандидат след. пака
|
||||
(в списке item-8).
|
||||
- **NOTE (общность): DC ДЕТЕКТ-регексы пар-скоупны в Go** (§12.2 by-design: данные вынесены, детект-алгоритм остаётся;
|
||||
self-gating → добавление ПАРЫ не требует Go, только СВОИ чекеры пары — да).
|
||||
- **NOTE (масштаб): двух-домовая когезия — шов.** →ru target-wire-text в ДВУХ домах: injection (embedded, recompile)
|
||||
vs heading.txt «Глава {n}» (pack, data-drop). Причина — байт-точность no-pack golden: то, что golden НУЖНО без пака
|
||||
→ embedded; пар-гейтнутое, что golden НЕ трогает → pack. Честный шов, задокументирован; унификация потребовала бы
|
||||
target-loader вне book-langpack_root (сдвиг снапшота golden — запрещён). + сиблинг target/source-данные ещё в Go
|
||||
(cheapgates ё-политика/оному/magnitude) — граница data/algorithm проведена неравномерно (кандидаты item-8).
|
||||
|
||||
**Финальная верификация (пост-фиксы):** `go build/vet` clean; `TM_MINER_PARITY=1 go test -race ./...` весь зелёный;
|
||||
парити `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}` EXACT; `git status` golden capture пуст (байт-идентичен).
|
||||
|
||||
---
|
||||
|
||||
# ДОБАВКА: полный data-out чекеров/гейтов/санитайзера (директива владельца + оркестратора)
|
||||
|
||||
**Триггер:** владелец (после core-пака) — «магические названия остались в `checkers_zh_ru.go`/`checkers_pack13_test.go`; что при +100 языках? это мелкие примеры, ты всё отловил?». Честный ответ: НЕТ — core-пак вынес LOOKUP-таблицы, но оставил DETECTION-регексы + сиблинг-гейты как «алгоритм/DEFER» (слишком мягко для планки 100 языков). Решение оркестратора: **доводи data-out, структуру НЕ трогай** (сплит — след. пак с форм-фиксами). Выбор владельца: «Full: все паттерны → данные, Go generic» + «тесты грузят фикстуры из пака».
|
||||
|
||||
## Внешняя сверка (харнесы `/home/ubuntu/projects/tmp`, 5-агентный воркфлоу)
|
||||
- **crush** (Charm, Go AI-агент, ~300 .go) — **~71 фокус-пакет**, один каталог = один концерн; провайдеры = ДАННЫЕ (внешний SDK + JSON-каталог, 0 per-provider Go); промпты = embedded `.md`; пост-обработка РАЗНЕСЕНА по фокус-пакетам, НЕ «checkers-куча». Сильное свидетельство против нашей 38-файловой кучи И за «volatile-ось = данные».
|
||||
- **anthropic/openai-go** — плоский root = кодоген-артефакт (не образец); но ручные части = МАЛЫЕ single-purpose пакеты (`apijson`/`apiquery`/…).
|
||||
- **GalTransl/AiNiee** (домен, что мы зеркалим) — валидируют data-мандат (глоссарий TSV, промпты per-язык, WHICH-чеки = config-список, control-codes `regex.json` = ДАННЫЕ). НО обе ТЕКУТ языком в ТЕЛА чекеров (残留日文/独白男他 gated on target_lang; per-lang Unicode-регексы) — ровно наша утечка; AiNiee `regex.json` = чистый контр-пример = наш фикс.
|
||||
|
||||
## Что вынесено (Full data-out)
|
||||
**`checkers_zh_ru.go` → `checkers.go`, ПОЛНОСТЬЮ пар-агностичен.** DC1/DC2/percent DETECTION-паттерны (`shichen_re`/`ru_hours_re`/`qianwan_*`/`shushiwan_*`/`cheng_re`/`decimal_fraction_re`/`percent_word`) → пар-пак `dc-checkers.txt` (`pattern<TAB>key<TAB>value`, verbatim); компилируются ОДИН раз в `compileCheckers` → `*dcCheckers` спек, живёт в `r.checkers` (openRunner) → `cheapGateConfig`. Target-general (broken-word `-йть`) → embedded `internal/lang/data/target-ru.txt`. dc*-символы стали ключами данных; чекеры — методы generic-спека, nil-инертны для no-pack.
|
||||
**`cheapgates.go`:** yofikator homograph-whitelist (38 слов) → target-ru `yo_homograph`; 万億 magnitude — CJK-парсер (digit/unit/**magnitude 万萬億亿兆**) **ДЕДУП в `lang.CJKSection`** (закрыл reviewer-NOTE о дубле), ru-стемы (тысяч/миллион/…) → target-ru `magnitude_stem`. `lintYofikation`/`lintNumberMagnitude` — методы спека. **Оставлено (документировано):** ё↔е фолд (1-символьная орфография = алгоритм), dialogue-dash (пунктуация), **translit-interj `ара-ара`/… = РАТИФИЦИРОВАННЫЙ DEFER** (D39.16, prompt NOT-do — единственный ru-output-блоклист, придержан).
|
||||
**`sanitizer.go`:** 4 preamble + trailing-note + edit-meta + invalid-sign регексы → target-ru (извлечены из исходника СКРИПТОМ, round-trip byte-match — не перепечатка). Санитайзер уже `isRuTarget`-гейтнут → паттерны = ru-target данные, загружаются `ruSanitizer = lang.TargetChecksFor("ru")`; strip/detect-АЛГОРИТМ остаётся. «ru»-ключ зеркалит гейт (тред таргета = follow-up, как sourceAbbrevs).
|
||||
**Тесты грузят из пака:** `testCheckers(t) = compileCheckers(pack.DCCheckers, TargetChecksFor("ru"))`; checkers/cheapgates-фикстуры прогоняют АЛГОРИТМ над паттернами пака — 0 дублирования пар-значений (входные примеры-сценарии остаются, §0.3 легитимно).
|
||||
|
||||
**Дома по ГЕЙТИНГУ:** source-gated пара-чекеры (DC1/DC2/percent, нужен zh-исходник) → пар-пак (nil для golden→инертны); target-general (register/broken/yofikator/magnitude-стемы/sanitizer, бегут на любом →ru) → embedded target-ru (нужны no-pack golden'у); source-CJK (magnitude-руны) → embedded cjk-section.
|
||||
|
||||
**Байт-точность (пост-каждый-инкремент):** golden БАЙТ-ИДЕНТИЧЕН (`capture.golden` не тронут — вкл. ch6, где санитайзер СРЕЗАЕТ ru-преамбулу → доказано, что паттерны из данных = те же байты); парити EXACT `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}`; `-race` весь зелёный. **Свип: 100→47** остаток (26 SQL-коммент migrate.go · flag-МESSAGE-диагностики checkers/cheapgates · 1-символьные фолды ё→е/ъ/kana · пунктуация-сепараторы · translit-DEFER · `第`-маркер ingest — кандидат). Все классифицированы/легитимны.
|
||||
|
||||
## Приложение A — карта связности подсистем (измерено grep'ом, вкл. тесты) — ДЛЯ ПРОМТА СЛЕД. ПАКА
|
||||
|
||||
Кандидаты-пакеты и их файлы: **miner** (miner*.go, mining.go — 8) · **membank** (memory/memnorm/mempostcheck/memseed/seeding — 5) · **chunk** (chunker/ingest — 2) · **checks** (cheapgates/checkers/sanitizer/coverage/quality/regressionguard — 6). Драйвер (runner/bookrun/waverun/wave/stagerun/chunkrun/resume) + инфра (snapshot/render/export/status/banknote/disposition/escalation) = композит-корень.
|
||||
|
||||
**Кто ссылается ВНУТРЬ каждого (non-test · TEST):**
|
||||
- **miner** ← non-test: runner:2, waverun:2, seeding:2 (узко!) · test: miner_test:63, miner_parity:7, seedlint:4, waverun:2.
|
||||
- **membank** ← non-test: waverun:18, wave:14, runner:13, snapshot:13, +многие (горячий путь) · test: memory:78, memseed:54, memnorm:38.
|
||||
- **chunk** ← non-test: waverun:12, wave:8, runner:5, status:5, stagerun:5, memseed:4 · test: chunker:35, ingest_encoding:28, ingest:23.
|
||||
- **checks** ← non-test: chunkrun:17, waverun:16, status:16, export:9, disposition:9 · test: sanitizer:90, coverage:29, regressionguard:19, cheapgates:15.
|
||||
|
||||
**Латеральная связность — НЕ «ноль» (уточнение к survey «star topology»):** есть ОБЩИЙ СУБСТРАТ, который надо ХОЙСТНУТЬ ПЕРВЫМ, иначе сплит потечёт:
|
||||
- **miner → membank: 35 refs**, но это ШАРЕД-утилиты, не membank-API: `normalizeSourceKey`@memnorm (12 — текст-нормализация, общая для майнера И банка), `loadGlossarySeed`/`seedTerm`/`seedAlias`/`seedFile`@memseed (9 — майнер читает СИД как GT/entity-guard), `runesEqual`@mempostcheck (3), `distinct`/`matches`@memory (8 — мелкие дженерик-хелперы).
|
||||
- **membank → chunk: 7** (`Chunk`-тип), **checks → membank: 19** (`memoryEntry`/`pickedEntry`/`tokenizeCyrillic`), **chunk/miner → checks: 6+6**.
|
||||
|
||||
**Вывод для сплита (порядок):** (1) ХОЙСТ шаред-базы ПЕРЕД подсистемами — `Chunk`/`MinerChunk` value-типы + текст-нормализация (`normalizeSourceKey`/`runesEqual`/`tokenizeCyrillic`) + сид-типы (`SeedTerm`/`loadGlossarySeed`) в базовый пакет (напр. `internal/pipeline/model` + `internal/text`); (2) затем miner (самый узкий: 6 внешних не-тест ссылок, импортит только lang+store+base) → checks (чистые ф-ии над строками, импортят config+lang+base) → membank → chunk. Драйвер импортит всё (композит-корень, как crush `internal/agent`). Форм-фиксы оркестратора (Export/RequestHash/resolveChunkState) едут в тот же пак — golden ревьюится ОДНИМ слоем.
|
||||
|
||||
## Приложение B — предлагаемая экспорт-поверхность (что капитализируется / что прячется)
|
||||
|
||||
- **`internal/text` (база):** `NormalizeSourceKey`, `RunesEqual`, `TokenizeCyrillic`. (сейчас memnorm/mempostcheck; общие для miner+membank+checks).
|
||||
- **`internal/pipeline/model` (база):** `Chunk`, `MinerChunk`, `SegBudget`, `Document` (value-типы, вокабуляр).
|
||||
- **`miner`:** ЭКСПОРТ `MineBank`, `MinerConfig`, `MinedTerm`, `Contrast`, `LoadContrast`; ПРЯЧЕТ mineDetect/proposeAliasEdges/buildPalladiusCyrSyllables/patternCandidates/emissionEligible. Импорт: lang, store, text, model.
|
||||
- **`membank`:** ЭКСПОРТ `MemoryBank`, `Select`, `Version`, `BaseVersion`, `MemoryEntry`, инъекц-рендереры; ПРЯЧЕТ pickedEntry/Aho-Corasick/suppressContained. Импорт: lang, store, text.
|
||||
- **`chunk`:** ЭКСПОРТ `SplitChunks`, `Ingest`, `IngestEncoded`; ПРЯЧЕТ chapterDraftChunks/packSentences/splitSourceSentences. Импорт: lang, model.
|
||||
- **`checks`:** ЭКСПОРТ `RunCheapGates`(+`CheapGateConfig`/`CheapGateResult`), `SanitizeOutput`(+`SanitizerResult`), `CompileCheckers`(+`Checkers`); ПРЯЧЕТ lint*-методы/dcCheckers-внутренности. Импорт: config, lang, text.
|
||||
|
||||
---
|
||||
|
||||
# ДОПРИЁМОЧНАЯ ДЕЛЬТА (лендинг-блокер оркестратора, 24.07)
|
||||
|
||||
**(1) ОБЯЗАТЕЛЬНО — аддендум `parsePalladius` ВЫПОЛНЕН.** (а) generic-парсер `parseCategoryRows` → `map[категория]map[key]value` (3-field key→cyr, 2-field key→"" set-member); НЕТ per-category switch, незнакомая категория ≠ ошибка парсера. (б) ОДИН типизированный `Palladius`-struct (Initials/Finals/YW/SpecialI + Retroflex/VFinal/VFinalInitial) вместо 4-картежа + 7 параллельных Pack-полей → `Pack.Palladius Palladius`. (в) required-категории валидирует ПОТРЕБИТЕЛЬ (`validate()`), не парсер. Удалены `parsePalladius`+`parsePalladiusPhonotactics`; `assignPair`/`mergePair` схлопнуты в parse+merge (оба пар-файла через generic-путь, union нил-сейф через `newPalladius()`); потребители (miner_palladius.go, langpack_test.go) обновлены. **Hash-нейтрально** (байты .txt не тронуты); `packAlgoVersion` v1→v2 (схемо-изменение по политике константы — выбрал БАМП; сдвиг Version() = version-only, golden без пака не тронут).
|
||||
|
||||
**(2) Хвосты:** `pipeline-c2.yaml` — явный segmentation-блок (1797/3200/1.1978/0.3852, байт-нейтрально, закрыл обещание коммента pipeline.go «все шиппинг ставят явно») · `backend/README.md:59` пути промптов → `prompts/zh-ru/` · `embedded.go` parseCJKSection unknown-category want-list → `+magnitude` · `packAlgoVersion` бамп (см. выше).
|
||||
|
||||
**(3) НЕ тронуто (кандидаты пака-15, по указанию оркестратора):** fail-loud оверлея уровнем выше · override-семантика unionStringMap · isRuTarget/'ru'-швы · empty-probe гард lintMagnitudeScale · дуп терминаторов chunker/coverage · 第.
|
||||
|
||||
**Верификация дельты:** `go build/vet/test -race` зелёные · golden БАЙТ-ИДЕНТИЧЕН (`capture.golden` не тронут) · парити EXACT `n=13618 {方源:0 蛊:1 蛊师:2 古月:22}`. Сессия НЕ коммитит.
|
||||
Loading…
Add table
Reference in a new issue