From 225abc09302f41aaa8f8dd1cad1a715ef79cf854 Mon Sep 17 00:00:00 2001 From: "Claude (backend session)" Date: Mon, 13 Jul 2026 01:45:42 +0300 Subject: [PATCH] Ratify D39: land the broken-telephone sync-audit ledger and the target seven-layer architecture with the phased build-now and research tracks, folding in the polygon sync --- docs/PROGRESS.md | 10 +- docs/architecture/05-decisions-log.md | 27 +- docs/architecture/08-sync-audit-ledger.md | 607 ++++++++++++++++++++ docs/architecture/09-target-architecture.md | 184 ++++++ 4 files changed, 825 insertions(+), 3 deletions(-) create mode 100644 docs/architecture/08-sync-audit-ledger.md create mode 100644 docs/architecture/09-target-architecture.md diff --git a/docs/PROGRESS.md b/docs/PROGRESS.md index dd03874..4e2a129 100644 --- a/docs/PROGRESS.md +++ b/docs/PROGRESS.md @@ -1,7 +1,7 @@ # Журнал прогресса > **⟶ ТЕКУЩЕЕ СОСТОЯНИЕ** (обновляемый блок; на 2026-07-12). **Источник истины по РЕШЕНИЯМ — `architecture/05-decisions-log.md` (D1–D38); этот файл — ЖУРНАЛ, не спека.** -> - **Фаза:** 0 ✅; Ф1-машинерия ✅ (D21–D28); приёмка этап A ✅ (D24). **Путь D26→D35: качество-первым.** Флип D1 моно→билингв (D30, supersede D17), reflow+output-санитайзер (D30/D33), сид v2 подписан владельцем (D34.1), переводчик flash сохранён (exp13/D32 — pro отклонён). **Пере-прогон 25 глав связкой ВЫПОЛНЕН; вердикт владельца: «лучше (из-за сида), но крайне слабо художественно» (2 претензии: абзацы/конвенции + понимание связей). ПИВОТ D35:** цикл приёмочных прогонов ЗАКРЫТ (инфра ✅ подтверждена исполнением / читаемость ❌ = блок на непостроенной Ф2 — НЕ провал; диагноз «Ф2-недострой + near-term промпт/нарезка-рычаги, не баг» верифицирован; флор D24.3 валидирован в проде). Планка впредь = 2 претензии владельца (ИНТЕРИМ пре-скрин, НЕ пилот Ф2.5). **Очередь «мерить→строить»: research/18 ЗАЛЕНДЕН (D36) → полигон exp14 ПРОГНАН+ратифицирован (D37, ACCEPT_WITH_FIXES): претензия-1 (абзацы) = рычаг ПРОМПТА/активации (F-base фронтир+старый промпт рубит хуже P0), дёшево — дискурс-промпт + крупный чанк; A-2pass отклонён (2-пасс дрейфует); разряды 甲乙丙 → кириллица «класс А/Б/В/Г» (владелец); re-gate хвост ~5-6 реальных редактор-инверсий (не 13 — kimi оверфлаг); претензия-2 хвост → exp14b ПРОГНАН+ратифицирован (D38): **«только gpt-5.4» ОПРОВЕРГНУТО** (дешёвые mistral/deepseek-pro чинят двойн.отриц, что glm-5+grok инвертируют; провалы разбросаны; фронтир в дефолт НЕ нужен — P1a≈F-disc по абзацам); **корень терм-дрейфа = статус-draft сида** (draft-термы не в editor-констрейнте + longest-match съедает верный approved) → promote грейды(А/Б/В/Г)+四代族长 approved; **кандидат glm-5→mistral/deepseek-pro** (дешевле+чинят) на бейк-офф ПРОЗЫ/omission → бэкенд под доказанный конфиг.** Эмпирика сошлась на ОРГАНИЗАЦИИ (не моделях) — D30.7-трек подтверждён. **h2h залендён (D38.1): ChatGPT↔Claude расходятся (фронтир-сильнее ↔ near-parity), сходятся — дефекты=инфра, никто не publishable без человека → разворот к СТРОИТЬ инфра-пак (`BACKEND_INFRA_PACK_SESSION_PROMPT.md`), не ещё-один-editor-эксп.** Бэкенд-фиксы (D35.4): экспорт-баг (ch5/ch20 пустые), ведущий `###` из draft, глоссарий разрядов 甲乙丙丁. Трек: дешёвые модели сейчас, фронтир потом (D30.7). Решающая точка — пилот Ф2.5 (блокеры: билингв-якорь/аннотаторы ≥3 — D25.9-Q1, корпус, судья-дублёр — D22.6). *(Детальная хроника D30–D34 — в записях ниже + D-лог.)* **Консолидация D38.2 (док-гигиена): eval спент-скрипты сгруппированы в `eval/exp12|13|14|14b/` (живой сид/оракулы в `eval/`), статус-доки + `.puml` актуализированы до D38.1, ссылки проверены (1799) — контракт-решения не тронуты.** **Инфра-пак D38.3 залендён** (санитайзер strip+export + число/omission-гард, верифиц. исполнением: build/vet/test зелёные, 3 span-дефекта запинены); дискурс+чэнъюй editor (v3, `0ac5f7a`) + reseed (`eb409f2`) ЗАЛЕНДЕНЫ (D38.5). **⚠ ПЕРЕ-ПРОГОН ПРИОСТАНОВЛЕН арх-ресетом (2026-07-13):** владелец остановил раунд — тактическое латание вместо архитектуры + «сломанный телефон» полигон→оркестратор→бэкенд. Хендофф новой оркестратор-сессии: `ORCHESTRATOR_ARCH_RESET_PROMPT.md` (5 концернов + синк-план). +> - **Фаза:** 0 ✅; Ф1-машинерия ✅ (D21–D28); приёмка этап A ✅ (D24). **Путь D26→D35: качество-первым.** Флип D1 моно→билингв (D30, supersede D17), reflow+output-санитайзер (D30/D33), сид v2 подписан владельцем (D34.1), переводчик flash сохранён (exp13/D32 — pro отклонён). **Пере-прогон 25 глав связкой ВЫПОЛНЕН; вердикт владельца: «лучше (из-за сида), но крайне слабо художественно» (2 претензии: абзацы/конвенции + понимание связей). ПИВОТ D35:** цикл приёмочных прогонов ЗАКРЫТ (инфра ✅ подтверждена исполнением / читаемость ❌ = блок на непостроенной Ф2 — НЕ провал; диагноз «Ф2-недострой + near-term промпт/нарезка-рычаги, не баг» верифицирован; флор D24.3 валидирован в проде). Планка впредь = 2 претензии владельца (ИНТЕРИМ пре-скрин, НЕ пилот Ф2.5). **Очередь «мерить→строить»: research/18 ЗАЛЕНДЕН (D36) → полигон exp14 ПРОГНАН+ратифицирован (D37, ACCEPT_WITH_FIXES): претензия-1 (абзацы) = рычаг ПРОМПТА/активации (F-base фронтир+старый промпт рубит хуже P0), дёшево — дискурс-промпт + крупный чанк; A-2pass отклонён (2-пасс дрейфует); разряды 甲乙丙 → кириллица «класс А/Б/В/Г» (владелец); re-gate хвост ~5-6 реальных редактор-инверсий (не 13 — kimi оверфлаг); претензия-2 хвост → exp14b ПРОГНАН+ратифицирован (D38): **«только gpt-5.4» ОПРОВЕРГНУТО** (дешёвые mistral/deepseek-pro чинят двойн.отриц, что glm-5+grok инвертируют; провалы разбросаны; фронтир в дефолт НЕ нужен — P1a≈F-disc по абзацам); **корень терм-дрейфа = статус-draft сида** (draft-термы не в editor-констрейнте + longest-match съедает верный approved) → promote грейды(А/Б/В/Г)+四代族长 approved; **кандидат glm-5→mistral/deepseek-pro** (дешевле+чинят) на бейк-офф ПРОЗЫ/omission → бэкенд под доказанный конфиг.** Эмпирика сошлась на ОРГАНИЗАЦИИ (не моделях) — D30.7-трек подтверждён. **h2h залендён (D38.1): ChatGPT↔Claude расходятся (фронтир-сильнее ↔ near-parity), сходятся — дефекты=инфра, никто не publishable без человека → разворот к СТРОИТЬ инфра-пак (`BACKEND_INFRA_PACK_SESSION_PROMPT.md`), не ещё-один-editor-эксп.** Бэкенд-фиксы (D35.4): экспорт-баг (ch5/ch20 пустые), ведущий `###` из draft, глоссарий разрядов 甲乙丙丁. Трек: дешёвые модели сейчас, фронтир потом (D30.7). Решающая точка — пилот Ф2.5 (блокеры: билингв-якорь/аннотаторы ≥3 — D25.9-Q1, корпус, судья-дублёр — D22.6). *(Детальная хроника D30–D34 — в записях ниже + D-лог.)* **Консолидация D38.2 (док-гигиена): eval спент-скрипты сгруппированы в `eval/exp12|13|14|14b/` (живой сид/оракулы в `eval/`), статус-доки + `.puml` актуализированы до D38.1, ссылки проверены (1799) — контракт-решения не тронуты.** **Инфра-пак D38.3 залендён** (санитайзер strip+export + число/omission-гард, верифиц. исполнением: build/vet/test зелёные, 3 span-дефекта запинены); дискурс+чэнъюй editor (v3, `0ac5f7a`) + reseed (`eb409f2`) ЗАЛЕНДЕНЫ (D38.5). **⚠ ПЕРЕ-ПРОГОН ПРИОСТАНОВЛЕН арх-ресетом (2026-07-13):** владелец остановил раунд — тактическое латание вместо архитектуры + «сломанный телефон» полигон→оркестратор→бэкенд. Хендофф `ORCHESTRATOR_ARCH_RESET_PROMPT.md` (5 концернов + синк-план). **→ РАТИФИЦИРОВАНО D39 (13.07):** синк-аудит (42 агента, $0, 65 находок / 0 refuted — `architecture/08-sync-audit-ledger.md`) → целевая 7-слойная архитектура + фазовый план (`architecture/09-target-architecture.md`). Вывод: потолок в НАШЕЙ организации (нарезка/проходы/контракт t-e/когезия), не в моделях. Владелец: **фазово** (трек A build-now ∥ трек B ресёрч, пере-прогон после трека B) + **слой пар = сеам сейчас, zh→ru контент**. Несущее: редактор=4-мандата-full-regen (причина инверсий); верности нет владельца (inversion-guard не построен); нарезка=потерянный рычаг (const в неверной единице, границы бюджетные, edit=draft-единица, когезия=мёртвые заглушки); терм-дрейф код-корень ЖИВ (сид-воркэраунд); ноль quality-сигнала в контуре; «D35.7» фантом→формализован. Дальше: хендофф-промты трека A (бэкенд) + трека B (полигон+ресёрч). > - **Стек (пост-D32):** черновик — **deepseek-v4-flash СОХРАНЁН** (exp13/D32: смена на pro отклонена по данным — сфабрикованная сходимость + катастрофа 蛊→«гусеницы» + верность у flash не хуже; thinking ON; escalate_to=pro без изменений) · редактор — **БИЛИНГВ (D30.1), кандидат glm-5-билингв** (gemini — премиум-эскалация за санитайзером; grok reasoning-off из редакторов СНЯТ — no-op) · судья/апекс Gemini 3.1 Pro (**слаг `-preview`**; mandatory thinking; `additive_total` — D22.3) · эскалация A: DeepSeek V4 Pro → GLM-5.1 → Gemini-preview · **канал B (18+): Mistral · эскалация grok-4.3 reasoning-ON (⚠ не на архаичных zh — D22.5) · abliterated-скрининг · судьи: эротика — Grok (дублёр гейтится пробой D22.6), violence/SFW — Grok+Gemini** (DeepSeek в explicit запрещён ToS — D14.1). 7 ключей. > - **Экономика:** `experiments/08-cost-model-v2.md` — до-флиповый ~$0.69 ранобэ / ~$10.94 вебновелла-500; **после флипа D1 → ~$0.85/ранобэ (+15–25%, D30.4)**; «$1.5–2» = опция Б (селективный апекс+память, НЕ подписана); пере-съём = exp08 v3 (D21.8/D25.8). ⚠ Resume: любой `--resnapshot` = переоплата книги — онгоинг гейтится реализацией D15.2 (идёт, D30.9). Gemini-недоучёт 146×/вызов закрыт additive_total (D22.3). > - **Ждём от владельца:** ~~подпись спорных терминов сида v2~~ ✅ ПОДПИСАНО (D34.1; владелец переопределил: Жэнь Цзу, Монах Цветочного Вина, гу Жизни, деревня Гуюэ, первобытный камень; род «гу» муж) · **чтение 25 глав пере-прогона** (арбитр читаемости — после связки+re-gate) · **флагман-сверка exp13** (переводчики, monitor-паттерн) · DashScope-ключ для qwen-арма exp13 (опционально) · **планка запуска** (лучше-фана/гибрид/издательский + готовность к COGS опции Б ~$1.5–2 — мемо-эскалация exp12; интерим D30.7: «лучше фана» на дешёвом миксе) · решение по билингв-якорю пилота (D25.9-Q1) · «да/нет» митигации контаминации пилотного корпуса (D27.4) · publishable/waiver — при экспорте (D25.1) · FN-bound L3 — при спеке (D25.4) · провенанс 12-*-доков · юр-пакет (+rights manifest, D25.8). @@ -66,6 +66,14 @@ Владелец остановил раунд: тактическое латание вместо архитектуры с первых принципов + **«сломанный телефон»** полигон(ресёрч)→оркестратор(выводы)→бэкенд(реализация); триггер — чуть не потеряли «умное чанкование» (оркестратор о нём не вспомнил, пока владелец не вернул). **Пере-прогон ПРИОСТАНОВЛЕН.** Выдан `docs/ORCHESTRATOR_ARCH_RESET_PROMPT.md` — **5 концернов владельца:** (1) translator/editor-разделение (абзацы легли на редактора — а переводчик? редактор переписывает всю главу? тестировали?); (2) чанкование НЕТРИВИАЛЬНО (влияет на структуру+смысл; глава>чанк→логическая нарезка; отдельный ресёрч, статический код, интернет-best-practices, фоллбеки); (3) аудит лендингов «сломанный телефон» (чисто/расширяемо/правильно ли залендено); (4) промты ПЕР-ЯЗЫК + где держать правила пар zh→ru/en→ru + few-shot-релевантность + русско-центричность (V4-п5 START_PROMT); (5) дыра в передаче знаний (что ещё забыли — систематический ledger). **План:** новая оркестратор-сессия синкается с полигоном, полигон с бэкендом; пере-прогон — только после архитектурной проработки. Эта сессия — ТОЛЬКО передача дел. +## Оркестратор №6 — синк-аудит + целевая архитектура ратифицированы (D39), 2026-07-13 + +Онбординг + **грунтовка по коду** (chunker/chunkrun/stagerun/memory/промты — не по заголовкам). Вопросы к предыдущему оркестратору (через владельца) подтвердили: концерн 1 контрактно ОТКРЫТ (D30.2, `05:366` — «механизм реверстки открытый вопрос»; арм «переводчик даёт структуру» не тестировался); «D35.7» — фантом (бэклог-кандидат `05:471`, не ратифицирован). + +**Синк-аудит (многоагентный воркфлоу, read-only, $0):** 42 агента, 8 дорожек (по несущему рычагу: цепочка полигон→оркестратор→бэкенд), адверсариальная верификация refute-by-default. **65 находок: 25 CONFIRMED, 9 PLAUSIBLE, 0 REFUTED**; верификаторы отделили настоящий «телефон» от осознанных ROI-деферов. Ледджер — `architecture/08-sync-audit-ledger.md` (первый экземпляр процесса концерна 5). Спот-верифицировал сам корневой пример: терм-дрейф-фикс = сид-воркэраунд (`memory.go:489` CONFIRMED-only + `suppressContained:615` disposition-слепой, код не тронут — `git log` подтвердил). + +**Ратификация D39:** целевая 7-слойная архитектура + фазовый план — `architecture/09-target-architecture.md`. **Решения владельца (AskUserQuestion):** фазово (трек A ∥ трек B, пере-прогон после B) + слой пар = сеам сейчас/zh→ru. **Док-долг трека A:** формализовать/вычистить фантомный «D35.7»; диспозиции незакрытых владельческих идей (V2-п1/2/3, V4-п1/4). **Мультисессия:** дерево чистое (только `START_PROMT.MD` владельца); написаны `08`/`09` + D39-блок + актуальность-карта + этот журнал; НЕ коммичено (ждёт обзора владельца + git-status перед стейджем). Дальше — хендофф-промты трека A (бэкенд) и трека B (полигон+ресёрч). + ## Компоненты единого resnapshot залендены (D38.5), 2026-07-12 Оба параллельных deliverable исполнены сессиями + верифицированы оркестратором ИСПОЛНЕНИЕМ + залендены: diff --git a/docs/architecture/05-decisions-log.md b/docs/architecture/05-decisions-log.md index 418b887..e26f2b5 100644 --- a/docs/architecture/05-decisions-log.md +++ b/docs/architecture/05-decisions-log.md @@ -3,7 +3,7 @@ > **⟶ КАРТА АКТУАЛЬНОСТИ (ревизия D31, продлена до D38.2 [12.07]; исторические записи ниже НЕ переписываются — дисциплина D23.3).** Читая контракт целиком, держи под рукой, что чем перекрыто: > - **Полностью superseded:** **D1 (моно-редактор) → D17 → D30.1 (редактор БИЛИНГВ)** · D9 (ja-приёмка) → D18 (蛊真人 zh→ru; ja — второй прогон) · D17 → D30.1 · скобка D19.4 «ключ без data-sharing» и D20.3 «чистый ключ = блокер пилота» → **D27** (единый ключ С data-sharing, отключение перед продом) · exp04-дефолт «editor grok-4.3» (D3) → D30.1 (grok reasoning-off из редакторских ролей СНЯТ — no-op; кандидат glm-5-билингв; gemini — премиум-эскалация только за санитайзером D30.3). > - **Частично амендировано:** D3 (канал B → D14.1/D19.1 + оговорки D22.5/D22.6; апекс-слаг `-preview` — D22.4; draft под вопросом exp13 — D30.6) · D6 («off/minimal для draft/edit» эродирован: draft thinking-ON принудительно, editor-off снят D30.1; живы per-роль принцип и D6.2-буфер) · D10 (+D24.2 alias-слепое пятно, истинная консистентность ≈99.5%) · D11-числа → exp08 → **D30.4** (флип D1 = +15–25%, ~$0.85/ранобэ; «$1.5–2» = неподписанная опция Б) · D12 (+D24.3 флоры max_tokens; +D29.1в rollup — амендмент вердикт-правила) · D13.1-арм из решающего → ПОДТВЕРЖДАЮЩИЙ (D30.1); +D25.3 prefix-анализ судьи · D14.2 закрыт D19.1/D22.4; D14.4 закрыт D22.1 · D15.3 исполнен; спека D15.2: v2→v3→v3.1 (D22.2), реализация = текущий пакет (D30.9) · D24.4 +D28.1 (precision/recall-трейдофф; hard-gate требует пере-замера) · порядок D24.5 отложен D26.1 (этап B — после читаемых 25 глав). -> - **Живое ядро без изменений:** D2 (+новый гейт-класс D30.3), D4, D5, D7, D8, D10-механизм, D12-таксономия отказов, D16, D18, D19.1–19.2, D21 (Ф2-механизмы voice/address/reveal), D22.5–22.7, D23 (golden = инвариант №8), D24–D38.5 — действующая голова контракта (D35 = пивот качество-первым; D36 = приёмка research/18; D37 = приёмка exp14: претензия-1=промпт-рычаг; D38 = приёмка exp14b: «только gpt-5.4» ОПРОВЕРГНУТО [mistral/deepseek-pro чинят], фронтир в дефолт НЕ нужен, корень терм-дрейфа = статус-draft сида; **D38.1 = слепой h2h P1a↔F-disc залендён [ChatGPT «фронтир сильнее» / Claude «near-parity» расходятся, СХОДЯТСЯ: дефекты = инфра] → РАЗВОРОТ «мерить→строить», выдан BACKEND_INFRA_PACK; D38.2 = консолидация/док-гигиена: спент exp-скрипты в `eval/exp12|13|14|14b/`, статус-доки+.puml актуализированы; D38.3 = инфра-пак читаемости залендён (санитайзер-классы strip+export + число/omission-гард), верифицирован исполнением [build/vet/test зелёные, 3 span-дефекта запинены]; D38.4 = аудит полноты exp14/14b: чэнъюй-рычаг забыт near-term→дописан в дискурс-промпт; inversion-гард D37 §2г не строился (число/omission ≠ полярность)→editor-swap = защита от инверсий (приоритет↑); нарезка = re-run-арм; гард-тумблер вкл. при resnapshot; D38.5 = компоненты resnapshot залендены+верифицированы [reseed-сид eb409f2 · дискурс+чэнъюй editor v3 0ac5f7a с few_shot-тумблером] → следующий шаг единый resnapshot + пере-прогон**). +> - **Живое ядро без изменений:** D2 (+новый гейт-класс D30.3), D4, D5, D7, D8, D10-механизм, D12-таксономия отказов, D16, D18, D19.1–19.2, D21 (Ф2-механизмы voice/address/reveal), D22.5–22.7, D23 (golden = инвариант №8), D24–D39 — действующая голова контракта (**D39 = АРХ-РЕСЕТ: синк-аудит сломанного телефона [65 находок, 0 refuted, `08-sync-audit-ledger.md`] + целевая 7-слойная архитектура [`09-target-architecture.md`] + фазовый план [трек A build-now ∥ трек B ресёрч, пере-прогон после]; формализует фантомный «D35.7» как цель слоя 1; терм-дрейф код-корень жив [сид-воркэраунд]; редактор 4-мандата-full-regen = причина инверсий**; D35 = пивот качество-первым; D36 = приёмка research/18; D37 = приёмка exp14: претензия-1=промпт-рычаг; D38 = приёмка exp14b: «только gpt-5.4» ОПРОВЕРГНУТО [mistral/deepseek-pro чинят], фронтир в дефолт НЕ нужен, корень терм-дрейфа = статус-draft сида; **D38.1 = слепой h2h P1a↔F-disc залендён [ChatGPT «фронтир сильнее» / Claude «near-parity» расходятся, СХОДЯТСЯ: дефекты = инфра] → РАЗВОРОТ «мерить→строить», выдан BACKEND_INFRA_PACK; D38.2 = консолидация/док-гигиена: спент exp-скрипты в `eval/exp12|13|14|14b/`, статус-доки+.puml актуализированы; D38.3 = инфра-пак читаемости залендён (санитайзер-классы strip+export + число/omission-гард), верифицирован исполнением [build/vet/test зелёные, 3 span-дефекта запинены]; D38.4 = аудит полноты exp14/14b: чэнъюй-рычаг забыт near-term→дописан в дискурс-промпт; inversion-гард D37 §2г не строился (число/omission ≠ полярность)→editor-swap = защита от инверсий (приоритет↑); нарезка = re-run-арм; гард-тумблер вкл. при resnapshot; D38.5 = компоненты resnapshot залендены+верифицированы [reseed-сид eb409f2 · дискурс+чэнъюй editor v3 0ac5f7a с few_shot-тумблером] → следующий шаг единый resnapshot + пере-прогон**). Ответ на вопросы бэкенд-сессии (PROGRESS §«Вопросы от бэкенда», §«[НУЖНО РЕШЕНИЕ] перед Фазой 1» — с ревизии D31 в `archive/PROGRESS-2026-07-04-10.md`) и полигона (think-режим). Каждое решение прошло адверсариальную панель из 3 критиков (research / экономика / исполнимость в коде Фазы 0); ни одно не отклонено, все уточнены. Это **контракт Фазы 1** — бэкенд исполняет отсюда; при конфликте с буквой 01-decisions/02-mvp-plan — источник истины здесь (потом сольём в основные доки). @@ -548,4 +548,27 @@ exp14b (D37-мандат) прогнал батарею DET-смысл-трап 1. **Reseed (полигон) — `eb409f2`:** трансформ `eval/exp14b/exp14b_reseed_promote.py` промоутит РОВНО 5 направленных термов (грейды `甲乙丙丁` dst «разряд»→«класс А/Б/В/Г» + draft→approved; `四代族长`→approved), остальные draft (`南疆/沈嬷嬷`/гу-имена) не тронуты (approved-инвариант). Идемпотентен, байт-стабилен (D30.5); стейл-хвост грейд-заметок вычищен детерминированно. Enforce-проба: `四代族长` (approved-длинный) выигрывает longest-match над `族长`, `家老`=«старейшина» легитимно цел. Сид на диске (ВНЕ git). 2. **Дискурс-editor (бэкенд) — `0ac5f7a`:** P1a-дискурс-reflow (Ван Цуй, репараграфизация 1:N, single-pass) в боевой `editor.md` = **APPEND к baseline** (=победитель exp14, SHA `b3b4f604`; clean-replace не делали — непротестированный новый промпт). + **чэнъюй-атом D38.4** (`物是人非` оба компонента; полный SHA сдвинулся РОВНО на эту строку `b3b4f604`→`95b1085`). + **few_shot-тумблер** (`---FEWSHOT---` + per-stage `Stage.FewShot *bool`, absent=ON): ON=ядро+примеры, OFF=ядро без примеров (deepseek-pro swap-арм — thinking, few-shot мешает CoT, exp14 §2а). Детерминизм тройной (флип=громкий resnapshot); golden байт-идентичен (стаб без FEWSHOT). Верификация: `go build/vet/gofmt/test -race ./...` зелёные; тесты snapshot-fold/wire-core-only/absent-noop/golden — пройдены; editor.md-дифф = ровно дискурс+чэнъюй+FEWSHOT (baseline цел). `prompt_version=v3` (ещё не гонялся). -**ИТОГ: все 3 компонента resnapshot ГОТОВЫ** — дискурс+чэнъюй editor (v3) · reseed-сид · инфра-пак (D38.3). **Следующий шаг = единый resnapshot** (D30.9, одна переоплата книги) + пере-прогон 3–10 глав: конфиг = glm=P1a-few-shot (дефолт) + армы {mistral few-shot · deepseek-pro `few_shot:false` = инверсия-защита D38.4 · крупная edit-единица = нарезка-арм D38.4} + `RegressionGuard` ON (D38.4). ПАРНО с чтением владельца (интерим-планка 2 претензии, D35). Крупная edit-единица — НЕ реализована (нарезка-арм, не промпт-порт). +**ИТОГ: все 3 компонента resnapshot ГОТОВЫ** — дискурс+чэнъюй editor (v3) · reseed-сид · инфра-пак (D38.3). **Следующий шаг = единый resnapshot** (D30.9, одна переоплата книги) + пере-прогон 3–10 глав: конфиг = glm=P1a-few-shot (дефолт) + армы {mistral few-shot · deepseek-pro `few_shot:false` = инверсия-защита D38.4 · крупная edit-единица = нарезка-арм D38.4} + `RegressionGuard` ON (D38.4). ПАРНО с чтением владельца (интерим-планка 2 претензии, D35). Крупная edit-единица — НЕ реализована (нарезка-арм, не промпт-порт). **[⟶ ПРИОСТАНОВЛЕНО D39: владелец остановил раунд для арх-ресета; пере-прогон гейтится треком B, см. D39.]** + +## D39 — Арх-ресет: синк-аудит «сломанного телефона» + целевая 7-слойная архитектура + фазовый план (13.07, оркестратор №6). ✅ + +Владелец остановил build-раунд: тактическое латание вместо архитектуры с первых принципов + «сломанный телефон» полигон(ресёрч)→оркестратор(выводы)→бэкенд(реализация) (триггер — чуть не потеряли «умное чанкование»). Хендофф `ORCHESTRATOR_ARCH_RESET_PROMPT.md` (5 концернов). **Проведён синк-аудит** (многоагентный воркфлоу, 42 агента, $0/только-чтение, адверсариальная верификация refute-by-default): **65 находок, 0 REFUTED, 25 CONFIRMED** — `docs/architecture/08-sync-audit-ledger.md`. Целевая архитектура и дорожная карта — `docs/architecture/09-target-architecture.md`. + +**Стратегический вывод (grounded):** потолок художественности — в НАШЕЙ организации пайплайна (нарезка + структура редакторских проходов + контракт переводчик/редактор + отсутствие меж-чанковой когезии), НЕ в моделях. Подтверждает D30.7; переопределяет источник следующего прироста качества = архитектурный. Рычаги **сцеплены** (редактор репараграфизует ВНУТРИ чанка → нарезка ограничивает потолок дискурс-переверстки) → концерны 1/2/4 проектируются вместе. + +**Ратифицированные диспозиции несущих находок (все CONFIRMED по file:line):** +1. **Редактор = ОДНА full-regen с 4 мандатами** (верность/стиль/глоссарий/reflow) — DISTORTION research/07 («узкие мандаты + диффы»); причина хвоста инверсий D34.3. → слой 3 (узкие проходы / diff-editing). +2. **Верности нет владельца в дешёвом пути** — inversion-guard D37 §2г НЕ построен; до Ф2 смысл в проде не контролирует ничто. → слой 5 (детерм. inversion/omission-бэкстоп). +3. **Нарезка — потерянный рычаг:** размер = const в неверной единице (символы исходника, не выходные токены); границы бюджетные, не логические; edit-единица=draft-единица; меж-чанковой когезии нет (мёртвые заглушки `STMDepth`/`OverlapTokens`). → слой 1. +4. **«D35.7» — фантомный номер** (декаплинг draft=чанк/edit=глава цитируется как решение в PROGRESS/exp14/research18, записи в D-логе НЕТ; в контракте — лишь un-ratified кандидат D35 п.7, `05:471`). **Формализуется здесь:** декаплинг = ратифицированная ЦЕЛЬ слоя 1, реализация гейтится треком B; **фантомные цитаты `D35.7` подлежат вычистке/переуказанию на D39** (док-долг трека A). +5. **Терм-дрейф: код-корень жив** — D38.5-фикс = сид-воркэраунд (промоут 5 термов), `memory.go` не тронут; `suppressContained` disposition-слепой + editor-блок CONFIRMED-only → любой draft-длиннее-approved снова тихо (и НЕлогируемо) уронит верный термин. → слой 4 (disposition-gate suppressor + громкий лог). +6. **В контуре ноль сигнала качества** — запрос владельца (п.25/п.31) + research/18 §C1 #10 «СЕЙЧАС» молча спущен в Ф2. → слой 5. +7. **Слоя пар языков нет** — промты zh→ru-baked, `Book.LangPair()` мёртв, ja-пример едет через «китайский паратаксис»; язык промпта всегда русский (V4-п5 не закрыт). Сам версионируемый слой был осознанно отложен под ROI (`05:471`, НЕ телефон), но латентный баг + V4-п5 открыты. → слой 2. +8. **Чистота бьёт по грядущим расширениям:** readability-гейты target-слепы; role→инъекция = рукоправимый switch; strip-and-export размазан по 5 файлам; сироты-промпты; мёртвый `gatesEnabled()`. → слой 7. +Плюс подтверждено ВЕРНОЕ (не трогать): разделение переводчик/редактор осознанно (D30.2); few-shot редактора структурны (не «правильно-переводные» — претензия владельца тут мимо); техники Ван Цуй перенесены дословно; чэнъюй-атом корректен. + +**Решения владельца по подходу (AskUserQuestion 13.07):** (1) **фазово** — трек A build-now (память-корень → детерм. quality-сигнал+inversion-бэкстоп → export-contract → гигиена-сеамы → слой-2-сеам → док-долг) ∥ трек B ресёрч (единая матрица `граница×edit-единица×когезия×где-reflow` + интернет-best-practices нарезки; ⚠ exp14 vs research/18 §A конфликтуют по размеру чанка — решает эмпирика); **пере-прогон — только после трека B** + приёмки ключевых build-now. (2) **Слой пар = СЕАМ сейчас, контент только zh→ru** (убирает латентный баг, разблокирует расширяемость; другие пары — под реальную книгу). + +**Процесс (концерн 5):** постоянный findings-ledger как повторяющийся шаг оркестратора (этот аудит = экземпляр 1, `08-*`); вопросы research/18 §D «к бэкенду» маршрутизируются обратно; completeness-critic на границе фаз. Незакрытые владельческие идеи (input-limit V2-п3, abuse-prescreen V2-п1, user-stop V2-п2, parallelism/speed V4-п1, glossary-auto-build V4-п4, self-improvement V1/V3) — получают явную диспозицию в треке A (док-долг), даже если «Ф2/Ф3». + +**Курс/столпы целы** (деньги/durability, snapshot, детерм. память, echo-гейт, 18+, дешёвый трек, конфиг-первое ядро); целевая арх = декомпозиция+достройка по сеамам, не переписывание. Планка пере-прогона — 2 претензии (интерим, не пилот Ф2.5). Следующее от оркестратора: хендофф-промты трека A (бэкенд) и трека B (полигон+ресёрч). diff --git a/docs/architecture/08-sync-audit-ledger.md b/docs/architecture/08-sync-audit-ledger.md new file mode 100644 index 0000000..47ace79 --- /dev/null +++ b/docs/architecture/08-sync-audit-ledger.md @@ -0,0 +1,607 @@ +# Синк-аудит «сломанного телефона»: полигон → оркестратор → бэкенд — ЛЕДЖЕР находок + +> **Статус:** ВЫВОД АУДИТА (2026-07-13), НЕ ратифицированный контракт. Это доказательная база для арх-ресета +> (`ORCHESTRATOR_ARCH_RESET_PROMPT.md`). Диспозиции и правки в D-лог/код рождаются из него ОТДЕЛЬНО, после решения владельца. +> Первый экземпляр процесса, которого требует **концерн 5** (сквозной findings-ledger). + +## Провенанс +- Многоагентный воркфлоу `sync-broken-telephone-audit`: **42 агента, 0 ошибок**, ~2.9M токенов, $0 (только чтение репо, без вызовов моделей). +- 8 дорожек: по каждому несущему рычагу восстановлена цепочка *полигон-ресёрч → вывод оркестратора → код бэкенда*. +- Каждая не-`FAITHFUL` находка средней/высокой тяжести прошла **независимую адверсариальную верификацию** (refute-by-default, повторное чтение первоисточников). +- Итог верификации: **25 CONFIRMED, 9 PLAUSIBLE, 0 REFUTED** — аудиторы грунтованы; но верификаторы уточнили ряд диспозиций/тяжести (см. колонку «верификатор»). + +## Калибровка диспозиций (важно — не путать) +- **DISTORTED / LOST / WORKAROUND_NOT_ROOT / UNTESTED_ASSUMPTION** — настоящий «сломанный телефон» или незакрытая дыра. +- **NEVER_CLOSED** — часть из них верификатор переквалифицировал в *осознанный ROI-гейтед бэклог-дефер* (запись в D-логе ЕСТЬ) — это НЕ потеря, это отложенное решение. Помечено в колонке верификатора. +- **CODE_SMELL** — чистота/расширяемость, не корректность. +- **FAITHFUL** — рычаг залендён верно (в т.ч. как эталон-сеам, который стоит копировать). + +**Сводка (65 находок):** NEVER_CLOSED=18, CODE_SMELL=15, FAITHFUL=8, UNTESTED_ASSUMPTION=8, DISTORTED=7, LOST=5, WORKAROUND_NOT_ROOT=4 + +--- +## L1-translator-editor-split — Концерн 1 — разделение труда переводчик↔редактор + +**Итог дорожки:** The draft=literal / edit=reflow+fidelity split is intentional and research-grounded (подстрочник+поэт), NOT an accident, and the strong discourse-reflow lever correctly sits on the editor (answers a, d). BUT that editor performs FOUR competing mandates (fidelity, style, glossary, discourse-reflow) in ONE full-chunk regeneration, contradicting research/07's explicit "narrow mandate + diffs, not full regeneration" rule — and this is exactly what produced the ~10-chunk meaning-inversion tail the re-gate found (answers c). The polygon NEVER tested "translator produces a structured near-correct draft" — every arm held the draft constant and varied only the editor (answers b). Deepest issue: fidelity has no owner in the cheap path — the draft model was never quality-measured and the editor is explicitly relieved of the duty to restore meaning, so nothing makes the translation "near-correct from the start," which is the owner's exact complaint. + +### `L1-editor-four-mandates-one-fullregen-pass` — **DISTORTED** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** The production editor performs FOUR competing mandates (bilingual fidelity check, RU style de-kalka, glossary enforcement, discourse-reflow) in a SINGLE full-regeneration pass over the whole chunk, contradicting research/07's core rule that each pass must have a narrow mandate and use diffs rather than full regeneration. +- **Интент ресёрча:** research/07 splits the work into narrow single-mandate passes (Self-reviser semantic → Self-reviser style → Line Editor → Continuity Editor → Proofreader) and states plainly that a broad mandate + full regeneration lets late passes destroy the work of early ones; it prescribes 'жёсткие инструкции что менять нельзя + диффы вместо полной перегенерации'. + *источник:* docs/research/07-translation-methodology.md §a l.32, §h l.105, §Выводы l.117 +- **Что в коде:** editor.md loads all four mandates at once (fidelity sverka :5, style :6, glossary :7, verstka :8, discourse-reflow :14-18) and is instructed to emit the FULL edited chunk text (:12). The runner feeds the whole prior draft in (Draft: prev), sizes max_tokens from the whole draft, then propagates prev = the full editor text — a per-chunk full rewrite in isolation. exp14 re-gate found this combined editor inverting/damaging meaning on ~10 chunks (D34.3 class 'editor bought smoothness at cost of meaning'); the A-2pass arm, which split semantic-repair from a separate reflow pass, confirmed the mandate competition is real (it improved paragraphs but the 2nd pass drifted meaning, e.g. 人上之人 inversion). + *evidence:* `backend/prompts/editor.md:5-18; backend/internal/pipeline/stagerun.go:61,87-90; backend/internal/pipeline/chunkrun.go:153` +- **Архитектурный вывод:** Split into narrow-mandate passes (or a diff-apply editor) with an omission/coverage gate between them so aggressive reflow cannot silently drop or invert meaning; exp14 §2Б.4 separately measured a larger edit unit as BOTH cheaper and better on paragraphs — per-chunk isolation is a choice, not a requirement — so the right architecture likely edits a larger unit through narrower passes. +- **Уточнение верификатора:** Core claim stands unchanged. One peripheral number is overstated: the finding's "~10 chunks" of editor meaning damage is the polygon's pre-verification figure ("13 катастроф"); the orchestrator's span-level verification corrected it DOWN to ~5–6 confirmed editor meaning-inversions (05-decisions-log.md line 488а). Also note the disposition could equally be read as WORKAROUND_NOT_ROOT: the ratified response to the diagnosed mandate-competition root (D34.3) was to prescribe an omission/inversion GATE for the reflow arms rather than restructure into narrow passes — and even that gate is not yet implemented in production code, so the root recurs. + +### `L1-fidelity-has-no-owner-in-cheap-path` — **UNTESTED_ASSUMPTION** / HIGH · верификатор: PLAUSIBLE (high) +- **Находка:** Fidelity is architecturally assigned to a cheap draft model that was never quality-measured, while the editor is explicitly relieved of the duty to restore distorted/omitted meaning — so nothing owns 'near-correct from the start' in the cheap path, which is precisely the owner's complaint. +- **Интент ресёрча:** research/07's подстрочник+поэт assumes the cheap draft is semantically dense/annotated and that a downstream pass restores and verifies meaning; TextMachine's own D1 уточнение #1 states plainly that after the draft NOTHING controls fidelity until the (unbuilt) Phase-2 bilingual judge. + *источник:* docs/research/07-translation-methodology.md §h l.109 and §Выводы l.122; D-log D1 уточнение #1 (docs/architecture/05-decisions-log.md:78); D-log D30.6/D1.1 (docs/architecture/05-decisions-log.md:395) +- **Что в коде:** Config pins draft = deepseek-v4-flash with the note 'draft качеством не мерен, D30.6'; D-log:395 asserts 'по верности flash ≥ pro' yet the draft fidelity was never gated, and the editor is told to preserve completeness but is NOT mandated to restore meaning ('редактор задаёт прозу, но не обязан восстанавливать искажённый/опущенный смысл'). exp14 T1 shows the flash draft inverting polarity ('ни один/никто' for 'все знали') and the glm editor NOT fixing it (only gpt-5.4 did); exp14b a1 reproduces this and shows cheap mistral/deepseek-pro editors CAN fix it. Every proposed remediation is editor-side (swap editor model, or selective gpt-5.4 escalation) — the draft's fidelity is never given a gate or a floor. + *evidence:* `backend/configs/pipeline-c1.yaml:37,49-64; backend/prompts/editor.md:9; docs/architecture/05-decisions-log.md:78,395` +- **Архитектурный вывод:** Give fidelity an explicit owner in the cheap path: a cheap deterministic draft→editor fidelity check (polarity / number / rank / omission — the exact classes exp14b enumerated), and/or a draft model floor, and/or make meaning-restoration an explicit editor mandate — rather than leaving fidelity uncontrolled until a Phase-2 bilingual judge that does not yet exist. +- **Уточнение верификатора:** Directionally correct and well-grounded on its CORE: the cheap path has no deterministic fidelity gate or draft floor, and the only landed/planned mitigations are editor-side (model swap glm→mistral/deepseek-pro) plus a scoring-only re-gate — the D-log itself states this at D38.4 §2 (05-decisions-log.md:537: "детерминированного прод-бэкстопа НЕТ... инверсии закрываются (а) fidelity re-gate = СКОРИНГ пере-прогона (не прод-гейт), (б) ВЫБОРОМ editor-модели"), and D1 уточнение #1 (:78) confirms nothing controls fidelity until the unbuilt Phase-2 judge. exp14 T1 / exp14b a1 reproduce as claimed. HOWEVER two specifics are wrong: (1) The editor is NOT "explicitly relieved of the duty to restore distorted/omitted meaning" — editor.md:5 explicitly TASKS it with correcting meaning distortions; the finding cited editor.md:9 (completeness) and missed line 5, and the "не обязан восстанавливать смысл" quote is D-log decision-reasoning (D32:395), not the deployed prompt. Consequently the finding's own architectural_implication ("make meaning-restoration an explicit editor mandate") is already satisfied by editor.md:5; the real gap is that the default glm-5 editor empirically FAILS to execute that mandate on inversions AND there is no verification gate — not that the mandate is absent. (2) "never quality-measured" is imprecise: draft fidelity WAS bench-measured in exp13/D32 (flash #1 on fidelity, "flash ≥ pro"); the accurate claim is "never quality-GATED / no fidelity FLOOR." Also the disposition is arguably mislabeled: the assumption was TESTED (exp14/14b) and refuted, and the root was explicitly diagnosed at D38.4 §2 but left live behind an editor-swap + scoring-only re-gate — which fits WORKAROUND_NOT_ROOT better than UNTESTED_ASSUMPTION. + +### `L1-diff-editing-never-built` — **NEVER_CLOSED** / MEDIUM · верификатор: PLAUSIBLE (high) +- **Находка:** Diff-based editing — research/07's explicit 'диффы вместо полной перегенерации' — was acknowledged by the orchestrator as a pilot arm gated on a search/replace apply machine, but was never built; production ships full-chunk regeneration. +- **Интент ресёрча:** research/07 §h: narrow mandate + diffs so a late/copyediting pass cannot rewrite (and wreck) earlier prose; the D-log parked a 'minimal-diff редактор ... гейтится машинерией search/replace-апплая' as a pilot arm with 'спека minimal-diff апплая — по готовности' owed to backend. + *источник:* docs/research/07-translation-methodology.md §h l.105; D-log D25.5 (docs/architecture/05-decisions-log.md:256,262) +- **Что в коде:** No diff/apply/search-replace machinery exists in the pipeline (grep for minimal-diff/applyDiff/SearchReplace over backend/ returns nothing); editor.md:12 emits full text and the runner treats that full text as the chunk output. The lever exists in neither a closed D-log decision nor in code — it is a live open loop parked behind a pilot that has not run. + *evidence:* `N/A — absence: no diff-apply in backend/internal/pipeline/*.go; backend/prompts/editor.md:12 mandates full-text output` +- **Архитектурный вывод:** Either build the promised diff-apply editor (bounded blast radius, cheaper output tokens, structurally omission-safe because it can only touch what it names) or explicitly retire the idea in the D-log so it stops being an unclosed research finding masquerading as a future pilot arm. +- **Уточнение верификатора:** Core claim is confirmed: diff/minimal-diff editing was researched, parked as a gated pilot arm, never built in code, never retired, and production ships full-chunk regeneration (editor.md:12 + stagerun.go runner) — disposition NEVER_CLOSED / MEDIUM stands (the item is parked-pending, never resolved to a closing disposition). Fix the citations: the D-log parking is D21.4 + D21.10 (05-decisions-log.md:256,262), NOT "D25.5"; the research/07 quote is at l.105 in section g), not §h; and the concrete search/replace-apply machinery lever originates in research/15 l.153 & l.245 (which D21 ratifies), with research/07 l.105 providing only the general narrow-mandate + diffs-vs-full-regeneration principle. One nuance on the enum's parenthetical "(not in D-log)": the arm IS present in the D-log as a parked pilot item — it is an open, unresolved loop rather than wholly absent. + +### `L1-translator-structure-arm-never-tested` — **UNTESTED_ASSUMPTION** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** No polygon arm ever tested 'translator produces a near-correct + structured draft'; every exp14/exp14b arm held the draft CONSTANT (flash v1-reflow) and varied only the editor, and the single frontier-as-translator probe was a 2-trap fidelity spot-check whose results were never even reported. +- **Интент ресёрча:** exp14 itself named the translate-stage-vs-edit-stage ceiling as a diagnostic question ('избегает ли сильный ЧЕРНОВИК инверсии в принципе — потолок на translate-стадии vs edit-стадии'), i.e. whether pushing structure/fidelity into the translator could beat the draft→editor split — the exact question the owner raises. + *источник:* docs/experiments/14-quality-empirics.md §1.1 arms table l.116-127, §1.3 l.148, §1.7 l.196 +- **Что в коде:** All exp14 arms (P0/P1a/P1b/P1c/A-layout/A-2pass/A-ref-*) reuse the same flash v1-reflow draft and vary editor × prompt × slicing; exp14b likewise holds one flash draft constant and swaps editor models. §1.3 planned a frontier-as-translator spot-check on T1/T2 for FIDELITY only (2 traps), and §2's result tables report editor arms exclusively — the translate-stage spot-check outcome is absent from the landed report. So 'structure/fidelity at the translator' is an untested cell that the current draft→editor architecture silently assumes is suboptimal. + *evidence:* `N/A — no arm/data exists; docs/experiments/14-quality-empirics.md §2.1-§2.6 tabulate editor arms only` +- **Архитектурный вывод:** Run the missing cell — a strong translator emitting structured, near-correct RU — head-to-head against draft→editor on BOTH span-fidelity and COGS before treating the rough-draft/editor-reflow division as optimal; it could collapse two model calls into one and remove the editor's need to reflow at all. +- **Уточнение верификатора:** Minor refinement to the one-line claim only: "held the draft CONSTANT (flash v1-reflow) and varied only the editor" is loosely worded — P1b/P1c reslice the draft to whole-chapter (exp14 l.120-121), A-layout deformats it (l.124), and A-2pass adds a second editor pass, so the draft's slicing/layout WAS varied. The load-bearing point is intact: the translator MODEL stayed the cheap deepseek-v4-flash in every reported arm — structure/fidelity was never pushed into a strong translator — and the finding's own code_reality already states "editor × prompt × slicing," so the substantive assertion (strong-translator cell is an untested, unreported cell) is fully supported. + +### `L1-split-intended-not-accident` — **FAITHFUL** / LOW · верификатор: — +- **Находка:** The draft=literal / edit=reflow+fidelity split is a research-grounded intentional design (подстрочник+поэт), not an accident of where the lever landed; the strong discourse-reflow lever (Ван Цуй 5 techniques + few-shot) correctly sits on the editor, but the translator keeps an empirically INERT reflow one-liner. +- **Интент ресёрча:** research/07 prescribes draft = rough literal подстрочник produced by a cheap model, with prose/reflow/rep-aragraphing as a downstream editorial task ('поэт'/stylist), and direct-speech dash punctuation as a proofreader task — reflow is legitimately editorial, not translator work. + *источник:* docs/research/07-translation-methodology.md §h l.109, §Выводы l.117 and l.122, §g l.103 +- **Что в коде:** The strong ДИСКУРС-ПЕРЕВЁРСТКА block (5 Ван-Цуй merge techniques + narrative/dialogue/focal-shift few-shot) lives entirely in the editor prompt; the translator carries only a weak one-liner that exp14 measured as practically inert (median 1 sentence/narrative-paragraph until the strong lever was appended to the editor as v3-discourse-reflow). D30.2 actively downgraded the translator's paragraph instruction to this one-liner. + *evidence:* `backend/prompts/editor.md:14-31; backend/prompts/translator.md:9; backend/configs/pipeline-c1.yaml:40-62` +- **Архитектурный вывод:** Keep the split, but delete the dead translator reflow one-liner or restate the translator mandate explicitly as 'literal подстрочник, do NOT reflow', so reflow has a single unambiguous owner and the inert duplicate stops implying a lever that does nothing. + +--- +## L2-chunking — Концерн 2 — чанкование нетривиально + +**Итог дорожки:** Chunking is the lever the chain most clearly dropped. The empirical curve (exp14: bigger/whole-chapter chunks were cheaper on output tokens AND gave better paragraphs, and whole-chapter uniquely linked a cross-boundary antecedent) and the "decouple draft-unit from edit-unit" idea were labeled "D35.7" and cited as ratified across four documents, but D35.7 does not exist in the D-log and nothing reached code: the editor still runs on the exact same 1500-token chunk as the translator; the size is a hardcoded a-priori const in the wrong unit (a source char-class estimate, not output tokens); and the cut is a pure token-budget boundary that can slice a scene mid-cohesion. No logical/semantic segmentation and no cross-chunk carryover were ever built — the only cross-chunk memory is the term glossary, which the research itself scoped to ~14% of discourse errors. The snapshot layer already has the exact folded-config pattern the chunker would need, so making size a first-class knob is cheap; it simply was never done. + +### `L2-d357-phantom-decision` — **NEVER_CLOSED** / HIGH · верификатор: PLAUSIBLE (high) +- **Находка:** The 'decouple draft=chunk / edit=chapter' lever is cited as ratified decision D35.7 in four documents, but no D35.7 entry exists in the decisions-log — it is closed nowhere (not in contract, not in code). +- **Интент ресёрча:** The polygon curve (exp14) and both frozen research/18 §A reports converged that the discourse/edit unit should be decoupled from the billing/chunk unit; the orchestrator gave this the tag 'D35.7 (draft=чанк/edit=глава, runner-уровень)' and marked it '✅ угадано' as if the team had already decided it. + *источник:* research/18 §B1 line 199 (table row: 'Декаплинг дискурс-единицы от биллинг-единицы | D35.7 (draft=чанк/edit=глава, runner-уровень) | ✅ угадано'); PROGRESS.md:94; docs/experiments/14-quality-empirics.md:350; docs/ORCHESTRATOR_ARCH_RESET_PROMPT.md:38 which itself admits 'записи в D-логе НЕТ' +- **Что в коде:** grep -rn 'D35.7' over docs/ returns only PROGRESS.md, exp14, research/18, and the reset prompt — 05-decisions-log.md has NO D35.7 heading. The D35 section (05-decisions-log.md:445+) ratifies D35, D35.1, D35.4a, D35.4c, D35.6, but never assigns a D35.7. The reset prompt (line 38) flags this as a to-do ('формализовать'). + *evidence:* `docs/architecture/05-decisions-log.md (no D35.7 present; grep 'D35\.[0-9]' yields only D35.1/.4a/.4c/.6); docs/ORCHESTRATOR_ARCH_RESET_PROMPT.md:38` +- **Архитектурный вывод:** A ratification-shaped label (a D-number) that exists in citing documents but not in the contract makes a live, unbuilt lever look closed — the exact broken-telephone failure. Either the decouple is a real decision (write D35.7 into the log with a concrete runner-level design and an implementation task) or it is not (strike the phantom citations from PROGRESS/exp14/research18). It must not stay half-cited as done. +- **Уточнение верификатора:** Core is confirmed and disposition NEVER_CLOSED (HIGH) stands: 'D35.7' is a ratification-shaped tag cited across four documents yet has no entry in the decisions-log, and the decouple lever is not implemented in code — the broken-telephone risk is real. Two specifics need correcting. (1) The finding's supporting enumeration is inaccurate: the log does NOT ratify 'D35.1' or 'D35.6' (no such tokens exist) and 'D35.4c' is actually 'D35.4в' (Cyrillic); the only D35 sub-tags in the log are inline body tags D35.4a/D35.4в mapping to D35's numbered points 4(а)/(в) — there is a single D35 heading (line 445), not a series of ratified sub-headings — so the parenthetical 'grep D35\.[0-9] yields only D35.1/.4a/.4c/.6' is itself false. (2) 'closed nowhere (not in contract)' is slightly too strong: the decouple DOES appear inside the contract at 05-decisions-log.md:471 (D35 point 7) as an explicit un-ratified backend candidate ('не сейчас, только под доказанный ROI'), never as a decided D35.7. This refinement strengthens rather than weakens the thesis — a 'not-now candidate' got re-tagged 'D35.7 ✅ угадано' as if decided. Remedy is unchanged: either write a real D35.7 (with runner-level design + impl task) or strike the phantom D35.7 citations from PROGRESS/exp14/research18. + +### `L2-edit-unit-equals-chunk` — **LOST** / HIGH · верификатор: PLAUSIBLE (high) +- **Находка:** The edit unit is the draft unit is the chunk: the editor runs over the same ~1500-token slice as the translator, so the empirically-supported 'large edit-unit' was never built and has been demoted to a 're-run arm'. +- **Интент ресёрча:** exp14 §2Б.4/§2Б.5 concluded the move is 'дискурс-промпт + крупная edit-единица' because bigger chunks were cheaper AND better on paragraphs; research/18 §A Ось-4 conclusion is to run the edit at a coarser discourse unit than the billing unit. The intent is that the EDITOR operates over a whole chapter (or coarser unit), not the raw 1500-token cut. + *источник:* docs/experiments/14-quality-empirics.md:350 and §2Б.5 line 356 ('катить дискурс-промпт + крупная edit-единица'); research/18 §A A3 Ось-4 line 155 +- **Что в коде:** translateChunk drives ONE Chunk through the full stage list in a single loop (translator + bilingual editor) — every stage sees the identical ch.Text; there is no chapter-level or coarser edit pass. bookrun.go iterates the flat chunk list. D38.5 in the log explicitly concedes the large edit-unit is unbuilt: 'Крупная edit-единица — НЕ реализована (нарезка-арм, не промпт-порт)'. + *evidence:* `backend/internal/pipeline/chunkrun.go:110 (for stageIdx, st := range r.Pipeline.Stages over one ch); backend/internal/pipeline/bookrun.go:142 (for i, ch := range chunks); admission at 05-decisions-log.md:551` +- **Архитектурный вывод:** The right architecture gives the runner a distinct edit granularity — e.g. buffer a chapter's chunks and run the editor/reflow pass over the assembled chapter text (the chapter already exists as doc.Chapters in ingest.go), while keeping the fine chunk as the draft/coverage/alignment atom. Instead the finding was reframed from 'supported by COGS AND quality' into a thing to measure again, and the code stayed single-unit. +- **Уточнение верификатора:** The disposition (LOST, HIGH) and the central claim hold: the edit unit == draft unit == ~1500-token chunk, the empirically-supported large edit-unit lever from exp14 (14-quality-empirics.md:350,356) was never built, and the D-log (05-decisions-log.md:539,551) explicitly demotes it to a re-run/backlog arm. However one of the two research citations is mischaracterized: research/18 §A Ось-4 (18-quality-levers.md:154-155) does NOT recommend running the editor over a coarser/whole-chapter unit — it recommends DECOUPLING discourse from billing by keeping the chunk SMALL and supplying cohesion via cached running-summary/carryover ("а НЕ размером единицы"), and explicitly warns against larger units. exp14 (bigger=better on 2 chapters, 0 retries, "retry-adjusted не активировался") and research/18 (predicted inverted-U, small-chunk+carryover) are in TENSION, so research/18 should not be enlisted as co-support for the bigger-edit-unit fix; the finding's suggested architecture (buffer a chapter, edit assembled text) follows exp14 but runs against research/18's recommendation. + +### `L2-1500-untested-apriori` — **UNTESTED_ASSUMPTION** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** targetChunkTokens=1500 is a hardcoded a-priori plan value that was never empirically derived, and the later measured curve (optimum ~3200c/whole) never reached it. +- **Интент ресёрча:** exp14 §2Б.4 measured, on the real tokenizer, that output tokens fell monotonically and mean sentences-per-paragraph rose as the chunk grew from 800c → 1600c(prod) → 3200c → whole, with 0 retries — i.e. the production 1600c/1500-token point is on the WRONG side of the curve; the optimum sits at ~3200c/whole. + *источник:* docs/experiments/14-quality-empirics.md §2Б.4 lines 344-350 (curve table + 'крупнее чанк = ДЕШЕВЛЕ И лучше абзацы'); research/18 §A A4.6 line 165 +- **Что в коде:** const targetChunkTokens = 1500 with a comment anchoring it to a plan corridor, not to any measurement ('~1500 sits in the plan's "1–2k токенов по границам абзацев" corridor (02-mvp Фаза 1)'). SplitChunks(chapters []string) takes no size parameter; the value is fixed at compile time. + *evidence:* `backend/internal/pipeline/chunker.go:47 (const); chunker.go:42-46 (comment anchoring to the 02-mvp plan corridor, not evidence); chunker.go:54 (SplitChunks signature has no size arg)` +- **Архитектурный вывод:** The const bakes an untested assumption the polygon's own curve later contradicted. Adversarial caveat: the empirics are thin (2 chapters, 1 text, 0 observed retries) and research/18 §A line 154-155 argues the OPPOSITE for prod (keep the chunk small, small retry blast-radius, supply cohesion via cached carryover), so the fix is NOT to hardcode 3200. The right move is to make size a swept, snapshot-folded knob so the pre-registered retry-adjusted curve (research/18 §C1 row 8) can actually run to a landed default — rather than freezing a value neither the curve nor the theory endorses. +- **Уточнение верификатора:** Every narrow factual clause reproduces exactly (hardcoded const, a-priori plan-corridor anchor with no measurement, no size param, curve optimum ~3200c/whole, prod point on the wrong side of the raw curve), so UNTESTED_ASSUMPTION is a fair disposition. Two adjustments to the framing/severity, not the facts: (1) This is NOT a broken-telephone distortion — the orchestrator faithfully handled the finding: D37 §2а (05-decisions-log.md:487,539) ratified the curve as DIRECTIONAL-ONLY ("2 главы/0 ретраев — направленно"), and D38.4/D38.5 (lines 539,551) explicitly hold the large-edit-unit as an open re-run ARM, "НЕ реализована… прод-изменение нарезки НЕ строить до подтверждения" — which is exactly the finding's own recommended fix (make it a swept knob, run the retry-adjusted curve first). (2) The const is snapshot-folded (chunkerVersion/render.go:33) and its tunability path is documented (chunker.go:15-22), and the curve contradicts 1500 only on the raw 2-chapter/0-retry cost/paragraph axes — not the retry-adjusted axis that governs prod (research/18 §A:154-155 argues the OPPOSITE for prod). So severity HIGH is defensible on per-chunk COGS/quality grounds but slightly overstated: the value is a consciously-held, actively-tracked open arm, not a silently-live untested assumption. + +### `L2-token-budget-not-semantic` — **DISTORTED** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** Chunking is purely a token-budget cut on paragraph/sentence boundaries; it has no scene/semantic-boundary logic and can slice a scene at an arbitrary 1500-token boundary, which is a demonstrated cohesion hazard for claim-2. +- **Интент ресёрча:** The owner requires cutting by LOGICAL chunks. research/18 §C1 row 8 pre-registered the chunking arm explicitly as '(жадная упаковка | сцен-граница)' — i.e. greedy packing versus a SCENE-BOUNDARY segmenter — as a factor to measure, precisely because where the boundary falls affects cohesion, not just cost. + *источник:* research/18 §C1 line 244 (row 8: '(0.75k/1.5k/3k/глава) × (жадная упаковка | сцен-граница)'); exp14 §2Б.5 line 358 ('P1c связал (1 датапоинт)' — whole-chapter, no mid-boundary cut, linked a cross-boundary antecedent); research/12 line 300 (chunking-error failure: a phrase translated twice or not at all) +- **Что в коде:** appendChapterChunks greedily packs whole paragraphs until the running token estimate would exceed targetChunkTokens, then flushes — the boundary falls wherever accumulation happens to cross 1500. An oversize paragraph descends to sentence boundaries. There is no discourse/scene detection: grep for 'scene' in the pipeline hits only memory sticky-inertia and the sanitizer's decorative scene-break, never the chunker's boundary logic. The scene-boundary arm was never built. + *evidence:* `backend/internal/pipeline/chunker.go:82-121 (greedy token-budget packing, sentence-descent); chunker.go:99,113,144 (boundary decided solely by estTokensFrom > targetChunkTokens); grep confirms no scene/semantic boundary code in chunker.go` +- **Архитектурный вывод:** Static logical/semantic segmentation (scene/discourse-move boundaries with a token-budget fallback and a never-split-a-sentence guarantee) is the owner's actual requirement and directly bears on claim-2 (a cross-boundary antecedent split into two isolated chunks cannot be resolved). The current cut minimizes only tokens; it should choose the LEAST-cohesive-cost boundary within the budget corridor. This boundary-quality dimension must become a first-class, testable strategy alongside size. +- **Уточнение верификатора:** Core confirmed; three precision refinements. (1) The research/12-common-failures.md:300 citation ('одна фраза переводится дважды или не переводится вообще') is MISAPPLIED — that is a LOSSY-chunking failure, and chunker.go is explicitly lossless-tiling (lines 123-124, 199-200: 'no character is lost or moved; concatenation reproduces the paragraph exactly'), so that specific failure mode is already guarded; it does not support the semantic-boundary point and should be dropped. Also that doc carries a PROVENANCE-NOT-FIXED banner (external LLM, 'не абсорбировать как факты'). (2) 'demonstrated cohesion hazard' is real (exp14 P1c, 1 datapoint) but exp14 §2Б.5 line 358 and §2.3 grade cross-boundary as NON-dominant/weak in this slice ('не доминирует'; 'не строить Ф2-память под слабый сигнал'); the within-chunk T1 signal is the dominant claim-2 failure per the empirics — so the hazard should be stated as demonstrated-but-graded-weak, with HIGH severity resting on the owner's architectural requirement (Concern 2) rather than on the empirical dominance of cross-boundary errors. (3) research/18 §C1 row 8 is a pre-registered MEASUREMENT arm (greedy vs scene-boundary to compare), not by itself a build mandate; the actual owner requirement lives in the arch-reset (ORCHESTRATOR_ARCH_RESET_PROMPT.md:37, PROGRESS.md:67). Relatedly, the decoupling tag 'D35.7' (draft=chunk/edit=chapter) is cited across PROGRESS/exp14/research18 but has NO D-log entry (ORCHESTRATOR_ARCH_RESET_PROMPT.md:38), and the larger/logical edit-unit is flagged only as a re-run arm 'НЕ реализована' (05-decisions-log.md:551) — reinforcing that logical/scene-boundary chunking is an un-landed, nearly-dropped lever, which is exactly the broken-telephone this finding names. + +### `L2-cross-chunk-only-glossary` — **LOST** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** The only cross-chunk cohesion mechanism is the term glossary; running-summary / prev-chunk carryover / reference-context — research/18's top unspent claim-2 lever — was never built, leaving the claim-2 root (chunk isolation + zh zero-anaphora) with no code mechanism. +- **Интент ресёрча:** research/18 §A A3 claim-2 ranks the best unspent lever as 'Running bilingual summary + previous-chunk carryover' and states the glossary is the WEAKEST cross-chunk memory, addressing only ~14% of Russian discourse errors and missing deixis (~37%) and ellipsis/morphology (~29%); §A Ось-4 conclusion is to keep the chunk small and supply cohesion via cached summary/carryover. research/12 and research/16 ground the same failure (book-distance inconsistency, loss of social distance/deixis). + *источник:* research/18 §A A1 line 88 ('Единственная кросс-чанковая память — термин-ГЛОССАРИЙ'), §A A3 claim-2 line 140 (glossary ≈14%, deixis ~37%, ellipsis ~29%), §A A3 line 143 (running summary #1 lever), §A Ось-4 line 155; research/12 lines 36/235/300; research/16 lines 128-131 +- **Что в коде:** translateChunk's only cross-chunk state is memory.Select → a glossary block injected per role (translator src→dst; editor CONFIRMED-dst constraints) plus sticky-scene entity ids reset at each chapter. There is no running summary, no previous-target/previous-source/next-source reference window, and no LLM carryover — each chunk is translated and edited in isolation. + *evidence:* `backend/internal/pipeline/chunkrun.go:104-108 (only memSel/glossary injections), chunkrun.go:134-140 (per-role injection is glossary only); bookrun.go:136-137 (sticky state is exact-matched ids, reset per chapter — not summary/carryover)` +- **Архитектурный вывод:** Claim-2 ('understanding of connections') is architecturally impossible under pure chunk isolation for zh zero-anaphora. The chunking design and the cohesion design are one problem: either the edit unit must be coarsened (chapter) so antecedents co-occur, or a cache-friendly carryover/summary must be threaded across chunk boundaries. Landing only the glossary and calling the discourse-unit decoupling 'guessed/ratified' leaves the load-bearing half of claim-2 unbuilt. +- **Уточнение верификатора:** Finding stands as written; one refinement that SHARPENS (not contradicts) it: the lever is not merely absent — it reached the config SCHEMA as inert stubs. ContextAssembly.STMDepth and .OverlapTokens exist (config/pipeline.go:46-47), are given real values in the shipped configs (pipeline-c1.yaml/pipeline-c2.yaml: stm_depth:2, overlap_tokens:200 with the comment 'перекрытие чанков = read-only контекст'), and are folded into the snapshot payload (snapshot.go:18-19, 225-226) — yet NO execution path (chunker.go, memory.go, chunkrun.go) ever reads them. So the cohesion lever is a schema/snapshot skeleton with zero runtime effect, which is if anything a worse form of LOST than 'never reached the code at all'. Also a minor doc imprecision (not a finding error): research/18 line 88 says the glossary is 'добавляемый редактору' (added to the editor), but the code injects glossary to BOTH translator (src→dst) and editor (CONFIRMED-dst) — the finding's code_reality correctly captures both. + +### `L2-size-const-vs-folded-knob` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** The most empirically load-bearing lever is frozen as a compile-time const even though the codebase already ships the exact snapshot-folded config pattern it would need, blocking the polygon's own pre-registered curve from ever landing as a setting. +- **Интент ресёрча:** research/18 §C1 row 8 pre-registers a sweep of (0.75k / 1.5k / 3k / whole) × (greedy | scene-boundary) on retry-adjusted output cost; §C3 row 15 frames the pair-conventions layer's load-bearing property as versioning + snapshot-fold, not the carrier. The size must be experimentable and, when chosen, snapshot-folded like any other gate config. + *источник:* research/18 §C1 line 244 (row 8 sweep); research/18 §C3 line 261 (row 15: 'несущее свойство — версионирование+снапшот-фолд, а не носитель'); chunker.go:20-22 self-comment describing the required fold +- **Что в коде:** The chunker comment concedes the size is config-worthy and states that IF made tunable it MUST be folded 'into the top-level snapshot like coverageSnap (§7d)'. That folded-knob pattern already exists and is proven in snapshot.go (coverageSnap, sanitizerSnap, contextSnap, each rendered {enabled/params} into the snapshotID). Yet the size remains a hardcoded const, so every point on the curve requires a source edit + chunkerVersion bump + full --resnapshot, making the lever practically un-experimentable and unavailable as a production/per-book setting. + *evidence:* `backend/internal/pipeline/chunker.go:20-22 & 45-46 (const-not-config rationale that already prescribes the coverageSnap fold); backend/internal/pipeline/snapshot.go:39/60/71 (coverageSnap/sanitizerSnap folded-knob pattern already implemented)` +- **Архитектурный вывод:** Right architecture: promote chunk size AND boundary-strategy (greedy|scene) to a snapshot-folded config knob mirroring coverageSnap — the mechanism is already there, so the const is a tactical shortcut, not a real safety constraint. The snapshot-loudness argument the comment makes is satisfied by folding, not by hardcoding. As-is, the single most consequential quality/cost lever is the one thing the config/polygon side cannot touch. + +### `L2-budget-wrong-unit` — **DISTORTED** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** The chunk budget is a source-side CJK/other char-class estimate, not output tokens on the real tokenizer — the exact unit research/18 said the budget MUST use. +- **Интент ресёрча:** research/18 §A A4.6 [CONFIRMED]: 'Бюджет чанка — в ВЫХОДНЫХ токенах на реальном токенизаторе (DeepSeek/GLM), а не в zh-символах' — recalibrated for Cyrillic fertility (~2–2.5× vs English, Petrov NeurIPS 2023) and the model's stable output ceiling; COGS is dominated by OUTPUT tokens. + *источник:* research/18 §A A4.6 line 165; §A Ось-4 line 152 (COGS output-dominated + Cyrillic fertility); exp14 §2Б.4 (curve was measured on the 'реальный токенайзер') +- **Что в коде:** estTokensFrom computes cjk + other/3, floored at 16 — an ad-hoc SOURCE char-class approximation of input tokens, unrelated to the target (Russian) output tokenizer or output-token count. The whole packing budget is expressed and enforced in this source proxy. + *evidence:* `backend/internal/pipeline/chunker.go:173-181 (estTokensFrom: cjk + other/3, floor 16); chunker.go:38-47 (budget defined in these source char-classes)` +- **Архитектурный вывод:** Even if the size const were re-tuned, it is measured in the wrong quantity: the budget that governs cost and the model's output ceiling should be output-token-based (or at least fertility-adjusted), while a source estimate is fine only for the coverage-floor safety check. The two concerns (source-side never-split segmentation vs output-token budgeting) are conflated into one source proxy. +- **Уточнение верификатора:** Claim stands as written; one refinement of framing only: the ~1500 budget originates from the mvp-plan corridor '1–2k токенов' (chunker.go:41) and predates research/18's A4.6, so this is simultaneously a DISTORTION (budget measured in the wrong unit) AND an un-actioned later lever — A4.6's output-token recalibration never reached the code. Additionally, exp14 §2Б.4 empirically found bigger chunk = cheaper AND better paragraphs (optimum ~3200c/whole), so the const is likely mis-tuned too, but the finding's scope is correctly the UNIT, not the tuning. + +### `L2-no-boundary-algo-research` — **LOST** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** No internet-best-practice research was ever done on logical/semantic CHUNKING ALGORITHMS (boundary detection); only chunk SIZE was studied, and the scene-boundary alternative was proposed but never investigated or built. +- **Интент ресёрча:** The owner requires logical chunking 'informed by internet best-practices' with fallbacks. That calls for a survey of segmentation algorithms — recursive/hierarchical splitting, semantic/discourse-based chunking, scene-boundary detection — not just a size curve. + *источник:* research/18 §A A3 Ось-4 lines 149-155 (studies SIZE only: Karpinska WMT2023, Lost-in-Middle TACL2024, fertility Petrov); research/16 line 144 (only name-drops 'рекурсивная ре-сегментация' as a detect-and-recover pattern, GalTransl/SakuraLLM); research/18 §C1 line 244 (scene-boundary appears only as a one-word proposed arm) +- **Что в коде:** The research corpus contains a well-cited chunk-SIZE curve and a document-level-context literature review, but no study of boundary-detection algorithms; research/16's single mention of recursive re-segmentation is about output-side detect-and-recover, not proactive source segmentation. The proposed 'сцен-граница' arm was never researched or implemented — the chunker remains greedy token-packing. + *evidence:* `N/A — no boundary-detection algorithm exists in code (backend/internal/pipeline/chunker.go implements only greedy token-budget packing) and no research doc surveys such algorithms` +- **Архитектурный вывод:** The boundary-QUALITY dimension the owner cares about most (cut by logical units, with fallbacks) was never given the same research rigor as size. A short best-practices survey (recursive splitting with semantic/scene-aware boundary selection and a never-split-sentence + token-budget fallback) should precede any implementation, so the static segmenter is grounded rather than inheriting today's cost-only greedy cut by default. +- **Уточнение верификатора:** Verdict stands; two precision refinements that sharpen but do not overturn it: (1) the code is greedy PARAGRAPH packing with a never-split-sentence fallback (chunker.go L82-152), not raw token-packing — it respects surface paragraph/sentence structure but has no LOGICAL/scene/discourse unit detection, which is exactly the missing dimension. (2) The boundary-algorithm research was never done AND is now freshly re-commissioned by the 2026-07-13 arch-reset (ORCHESTRATOR_ARCH_RESET_PROMPT.md), so it is an actively re-opened item rather than permanently abandoned; disposition could equally read NEVER_CLOSED, but LOST is defensible since the owner-required logical-chunking lever never reached code. + +--- +## L3-memory-glossary — Концерн 3 — сломанный телефон (память/глоссарий) + +**Итог дорожки:** D38 §3 diagnosed the term-drift root squarely IN CODE — the editor constraint block is CONFIRMED-only and suppressContained lets a longer DRAFT key eat a nested APPROVED term, so the reader gets nothing — but the landed D38.5 "fix" (eb409f2) is a pure seed-data promote of exactly 5 terms; backend/internal/pipeline/memory.go was untouched (last edit 296419c, 2026-07-10, two days before the reseed and unmodified by every D38.x commit). The code root is therefore live: any future draft-status longer key nested over an approved shorter key will silently drop the approved term from the editor again, and — worse — the drop is invisible to the post-check and retrieval-state because the suppressed entry never enters memSel.injected. The right fix is a disposition-gated suppressor (a lower-trust longer key must not suppress a higher-trust shorter key); an editor-block fallback is strictly weaker because the suppressed entry is destroyed upstream in Select. A secondary landing distortion: the editor block and its wiring still assume the D1 monolingual editor ("editor never sees the source") although D30.1 flipped the editor to bilingual and editor.md now feeds it the source. + +### `L3-seed-workaround-not-code-root` — **WORKAROUND_NOT_ROOT** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** The term-drift fix (D38.5, commit eb409f2) was a seed-data promote of 5 terms, not the code fix D38 §3 itself said was the root; memory.go is untouched, so the disposition-blind suppressor + CONFIRMED-only editor block remain live. +- **Интент ресёрча:** 14b-meaning-battery.md §2.1 note and §2.4 (and the D38 ratification) diagnose the ROOT in code by file:line: renderEditorConstraintBlock (memory.go:489, CONFIRMED-only) excludes the draft 四代族长, and suppressContained (longest-match A3) EATS the correct approved 族长→«глава клана» because it is nested inside the draft-longer 四代族长, leaving only the approved neighbour 家老→«старейшина». The chosen fix was 'promote грейды + 四代族长 draft→approved + reseed'. + *источник:* docs/experiments/14b-meaning-battery.md §2.1 (line 147) and §2.4 (line 175); review header lines 20-26; docs/architecture/05-decisions-log.md D38 §3 (line 499) and D38.5 (line 548) +- **Что в коде:** git shows eb409f2 touched only eval/exp14b/exp14b_reseed_promote.py (+228) and eval/README.md; the ratify commit 183c8ad touched only docs; the D38.5 editor commit 0ac5f7a touched editor.md/render.go/config but NOT memory.go. memory.go's last commit is 296419c (2026-07-10), before the reseed (2026-07-12), and no eb409f2~1..HEAD commit modifies it. The reseed script itself (exp14b_reseed_promote.py:13-14, 89-90) states its whole purpose is that the approved-long 四代族长 wins longest-match so 族长 'больше не съедается' — i.e. it fixes the ONE instance by data, leaving suppressContained's rule intact. + *evidence:* `backend/internal/pipeline/memory.go:489 (CONFIRMED-only), memory.go:615-641 (suppressContained); git show --stat eb409f2 / 183c8ad / 0ac5f7a; git log -1 -- memory.go = 296419c` +- **Архитектурный вывод:** A code root diagnosed at file:line must be closed in code; a per-term seed promote clears the current symptom but rebuilds the exact hole the moment any other draft term nests over an approved one. The correct architecture fixes memory.go (see L3-suppressor-disposition-blind) so no future seed edit is load-bearing for correctness. +- **Уточнение верификатора:** Minor framing precision (does not change disposition): D38 §3 does not prescribe a code fix that was then skipped — it explicitly names the glossary draft-STATUS as the root and deliberately chose the owner-authorized deterministic seed promote as THE fix, while diagnosing the code MECHANISM (memory.go:489 CONFIRMED-only + suppressContained longest-match) at file:line. So the workaround was a conscious design choice, not an accidental miss. The finding's architectural implication nonetheless holds exactly: the disposition-blind suppressContained (validSuppressor gates only on spoiler validity, memory.go:616-623/631) and the CONFIRMED-only editor block (memory.go:489) are untouched, so any future spoiler-valid DRAFT term nesting over an APPROVED shorter term rebuilds the identical hole — the seed promote fixes the one 四代族长/族长 instance by data, leaving the code rule live to recur. WORKAROUND_NOT_ROOT, severity HIGH, stands. + +### `L3-recurrence-draft-longer-eats-approved-shorter` — **WORKAROUND_NOT_ROOT** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** With memory.go unchanged, any draft/auto longer key nested over an approved shorter key silently drops the approved term from the editor constraint block again — and the drop is invisible to the post-check and retrieval-state. +- **Интент ресёрча:** research/13 wants longest-match/whole-entity replacement to raise PRECISION (a short term must not fire inside a longer entity, Q3 verifier note; Q4 wrong-sense row = longest-match+whole-entity DET), on the assumption the longer match is the correct, TRUSTED whole-entity; and it wants silent degradation converted to LOUD (retrieval-state). Neither sanctioned a lower-trust longer key deleting a higher-trust shorter one, nor a drop absent from the observability record. + *источник:* docs/research/13-memory-bank-validation.md Q3 (line 96), Q4 (line 108), TL;DR §7 (line 27) and Вердикт (line 220); docs/architecture/05-decisions-log.md D38 §3 (line 499) +- **Что в коде:** Exact recurrence: for a chunk read at chapter C, if the normalized chunk contains an occurrence of key K_long owned by a draft/auto entry E_long (status != approved, non-empty dst so it enters the automaton per memory.go:179, spoiler-valid at C), and strictly nested inside it an occurrence of key K_short owned by an APPROVED entry E_short (spoiler-valid at C) that fires nowhere else in the chunk, then: (1) suppressContained drops K_short — validSuppressor (memory.go:616-623) returns true because it checks ONLY spoilerBlocked, never status/disposition, and the containment test at memory.go:631 requires only strictly-longer + valid; (2) E_short never becomes a pickedEntry; (3) dispositionFor (memory.go:415-423) maps E_long to memAmbiguous; (4) renderEditorConstraintBlock skips E_long at memory.go:489 (memConfirmed-only) -> the editor gets NO canonical form; a nearby approved neighbour is normalised toward instead. Because the suppressed E_short is absent from memSel.injected, r.memory.postcheck (chunkrun.go:170) and persistRetrievalState (chunkrun.go:202) never see it — the degradation is not converted to loud. + *evidence:* `backend/internal/pipeline/memory.go:616-623 and :631 (validSuppressor / containment), :415-423 (dispositionFor), :489 (CONFIRMED-only); backend/internal/pipeline/chunkrun.go:104-106 (Select once, both blocks consume memSel.injected), :170 (postcheck over memSel.injected)` +- **Архитектурный вывод:** The seed still holds untouched draft terms (per exp14b_reseed_promote.py:10 — 南疆/沈嬷嬷/gu-names left draft); I cannot enumerate a live current-seed collision without reading the seed (forbidden path /home/ubuntu/books/), so a live instance is data-dependent, but the code guarantees the class recurs. The recurrence is doubly silent: the approved term is dropped AND unlogged, which also violates research/13's core promise to convert silent degradation into loud. + +### `L3-suppressor-disposition-blind-right-fix` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** The root fix is to gate the SUPPRESSOR on disposition (a longer draft/ambiguous key must not suppress a shorter approved/confirmed key), not an editor-block fallback; suppressContained is disposition-blind while its two consumers apply opposite disposition policies to the same suppressed set. +- **Интент ресёрча:** research/13 Q1 Находка 3 makes the three-way disposition (CONFIRMED/AMBIGUOUS/REJECT) the machinery that governs how much to TRUST an injection; longest-match (Q3 verifier note) is a precision device premised on the longer match being the trusted whole-entity. Trust should therefore also govern which match wins a containment contest. + *источник:* docs/research/13-memory-bank-validation.md Q1 Находка 3 (line 44) and Q3 verifier note (line 96) +- **Что в коде:** suppressContained runs ONCE in Select (memory.go:311) and its output feeds BOTH renderGlossaryBlock (which renders AMBIGUOUS lines with ⟨проверить⟩, memory.go:458-460) and renderEditorConstraintBlock (which DROPS all non-CONFIRMED, memory.go:489). The self-review #6 comment (memory.go:607-614) explicitly hardened validSuppressor against SPOILER-blocked suppressors but says nothing about DISPOSITION — the exact gap. An editor-block fallback cannot recover the term because suppressContained already deleted the approved occurrence upstream (not in memSel.injected), so a pure editor-side patch would need extra plumbing (carry suppressed-CONFIRMED forward) and would heal only one of the two consumers. + *evidence:* `backend/internal/pipeline/memory.go:311 (suppression call), :607-623 (validSuppressor + spoiler-only comment), :489 vs :458-460 (editor drops ambiguous / glossary keeps it)` +- **Архитектурный вывод:** Make validSuppressor require the longer match to be at least as trustworthy as the shorter it would suppress: a suppressor whose dispositionFor is AMBIGUOUS (draft/auto or collision-prone) must not suppress a nested match that dispositions CONFIRMED. This preserves correct whole-entity longest-match when the longer term is approved (approved 四代族长 still overrides approved 族长) and only un-suppresses when the longer term is lower-trust — closing the root where information is lost, for both consumers, instead of promoting seed data forever. + +### `L3-editor-block-stale-monolingual-D1` — **DISTORTED** / LOW · верификатор: — +- **Находка:** renderEditorConstraintBlock and its wiring still assume the D1 MONOLINGUAL editor ('editor never sees the source') as the rationale for target-forms-only and CONFIRMED-only, but D30.1 flipped the editor to BILINGUAL and editor.md now feeds it the source — the memory landing is stale w.r.t. the current pipeline. +- **Интент ресёрча:** The D-log map records D1 (моно-редактор) → D17 → D30.1 (редактор БИЛИНГВ) as FULLY superseded; the current editor is bilingual and does see the source. + *источник:* docs/architecture/05-decisions-log.md:4 (карта актуальности: D1→D30.1 билингв) +- **Что в коде:** backend/prompts/editor.md line 1 gives the editor BOTH the source and the draft ('Тебе даны ИСХОДНЫЙ текст и его черновой перевод'; template carries {{text}} source + {{draft}}). Yet memory.go:469-480 (editorConstraintHeader + renderEditorConstraintBlock doc) still say 'monolingual editor (D1)' and justify target-forms-only because 'the editor never sees the source: decisions-log D1'; chunkrun.go:129 repeats 'the monolingual EDITOR ... no source — decisions-log D1'. The CONFIRMED-only gate's second premise (editor can't check the source so don't push ambiguous candidates) is now weaker since the editor CAN verify against source — the code was not revisited after the flip. + *evidence:* `backend/internal/pipeline/memory.go:469-473 and :476-480; backend/internal/pipeline/chunkrun.go:128-133; backend/prompts/editor.md:1; docs/architecture/05-decisions-log.md:4` +- **Архитектурный вывод:** Correct the stale D1 rationale; substantively, revisit whether a bilingual editor should receive the src→dst mapping (like the translator) rather than bare Russian canonical forms, so it can bind each canonical form to its source term while cross-checking — an open design question the D30.1 flip left unclosed. Uncertainty: behavior is not proven wrong, only mis-rationalised, hence LOW. + +--- +## L4-per-language-prompts — Концерн 4 — пер-язык / слой правил пары / язык промпта + +**Итог дорожки:** Prompts are NOT architecturally per-language-pair: there is one global system-prompt file per stage (translator.md, editor.md), pinned by a fixed path in the pipeline yaml, with only surface {{source_lang}}/{{target_lang}}/{{genre}} value-substitution. The actual RULES (Russian dialogue dash, zh-parataxis→ru-hypotaxis reflow, chengyu policy) are hardcoded into a single de-facto zh→ru file — so much so that the shipped ja→ru example runs a prompt that talks about "китайский паратаксис" and 物是人非. The versioned per-pair "convention pack" layer that research/18 §B2/§C#15 proposed (and the owner named in concern 4) is a backlog note only — no code, no placeholder, no pair-keyed selection. The prompt instruction LANGUAGE is unconditionally Russian, an assumption the polygon only ever exercised on zh→ru; zh→en would drag Russian into the prompt with nothing that tested that. One point of good news: the editor few-shot is STRUCTURE-only reflow exemplars, not "correct-translation" cases — the owner's specific worry does not apply to what's there. + +### `L4-prompts-not-per-pair-zh-baked` — **DISTORTED** / HIGH · верификатор: PLAUSIBLE (high) +- **Находка:** Prompts are single zh→ru-baked files selected by a fixed path, not a per-language-pair architecture; only surface {{var}} values are templated while the rules themselves are hardcoded Chinese-specific prose. +- **Интент ресёрча:** research/18 §A4.4 + §B2.2 frame the correct zh→ru transform as a FOUR-operation, pair-SPECIFIC content asset (parataxis→hypotaxis, red-line, dash, zero-pronoun explicitation, 成语 policy) that must be keyed to (src→tgt[×genre]); the discourse content is inherently per-pair, not a genre knob. + *источник:* docs/research/18-quality-levers.md §A4.4 (line 162) and §B2.2 (line 209); owner START_PROMT V1 item 4 / V4 item 5 +- **Что в коде:** backend/internal/pipeline/runner.go:149 loads LoadPromptTemplate(st.Prompt) where st.Prompt is a fixed path from the pipeline yaml (pipeline-c1.yaml:38,57 → ../prompts/translator.md, ../prompts/editor.md). editor.md:14 hardcodes 'обязательная операция передачи КИТАЙСКОГО паратаксиса', editor.md:18 hardcodes the chengyu example 物是人非, editor.md:8/17 the Russian dash. Render() (render.go:135-173) only substitutes scalar values for {{source_lang}} etc. The book's language pair does NOT select the prompt: the shipped example/book.yaml (source_lang: ja, line 4) routes via pipeline-c1.yaml and would feed the Japanese book the Chinese-parataxis editor prompt. Book.LangPair() (book.go:193) is dead code — grep shows zero call sites; the only pair-aware logic is coverage.go:140-191 (len_ratio corridor), never prompt selection. + *evidence:* `backend/prompts/editor.md:14; backend/prompts/editor.md:18; backend/internal/pipeline/runner.go:149; backend/configs/pipeline-c1.yaml:57; backend/example/book.yaml:4` +- **Архитектурный вывод:** A stage should resolve its prompt from a pair-keyed layer (src→tgt[×genre]) so the zh-parataxis/chengyu discourse rules live in a zh→ru asset and a ja→ru or en→ru book picks its own; the language-specific rules must not be baked into one file that surface-templates the pair labels while keeping Chinese-only content. +- **Уточнение верификатора:** Factual claim is fully reproduced and correct: prompts ARE single zh→ru-baked files selected by a fixed path, only scalar {{vars}} are templated, and there is no per-pair prompt architecture (Book.LangPair is dead code; the ja example routes through the zh editor). But the disposition DISTORTED/severity HIGH is off. Two things are being conflated: (1) the per-pair VERSIONABLE LAYER (src→tgt[×genre]) — this was consciously DEFERRED to architecture backlog under ROI (D-log D36 §4г, line 471), and multi-language/prompt-language is an explicit stated FUTURE task (owner START_PROMT V4 item 5, line 90); the current zh discourse content in editor.md is a FAITHFUL near-term landing of the ratified discourse-prompt lever (D36 §4а). So this is a dispositioned backlog gap, not a broken-telephone distortion of a lever the backend was told to build per-pair. (2) The genuine live defect is a latent CODE_SMELL/hazard: zh-source-specific rules (parataxis, 物是人非) are baked into the DEFAULT editor.md selected by fixed path with NO guard, so a non-zh book silently inherits Chinese conventions (verified via the ja→ru example/book.yaml). This is real but currently INERT — the only production book is zh→ru (Gu Zhenren); it manifests only for a ja/en book, which is explicit future scope. Severity is better rated MEDIUM (latent, non-prod, consciously deferred) than HIGH. Minor specifics: "FOUR-operation" mislabels the §A4.4:162 FIVE-item pair-conventions list (parataxis→hypotaxis, red-line, dash, zero-pronoun, chengyu), conflated with §A4.2's separate 4-op reflow (a/b/c/d); and editor.md:8/17 dash is a target-side →ru convention, not zh-source-specific, so it is weaker evidence for the "Chinese-specific" framing. + +### `L4-convention-pack-layer-never-built` — **NEVER_CLOSED** / HIGH · верификатор: PLAUSIBLE (high) +- **Находка:** The versioned per-pair 'convention/rules pack' layer the owner asked for is proposed in research and parked as a backlog idea, but has zero code — no placeholder, no schema, no snapshot fold. +- **Интент ресёрча:** research/18 §C#15 recommends a 'слой пакет конвенций пары': content = 5 Ван-Цуй ops + Russian norms + Palladius, key (src→tgt[×genre]), snapshot-folded like chunkerVersion, carrying completeness-invariant limiters; §B2.2 calls its absence a 'структурный пробел'. This is the owner's own idea (concern 4). + *источник:* docs/research/18-quality-levers.md §C#15 (line 261) and §B2.2 (line 209); ratification backlog note D36 §4г, 05-decisions-log.md:471 +- **Что в коде:** No convention-pack layer exists. RenderVars (render.go:123-127) carries only Book/Text/Draft; the placeholder allow-map (render.go:137-150) has no pair-rules input and Render() hard-errors on any unknown {{marker}}. The D-log records only 'слой пакет конвенций пары — идея в бэклог архитектуры (зона оркестратора/бэкенда), стройка под ROI' (05-decisions-log.md:471) — an acknowledgment and deferral, never a build decision, so the layer never reached code. + *evidence:* `backend/internal/pipeline/render.go:123; backend/internal/pipeline/render.go:137; docs/architecture/05-decisions-log.md:471` +- **Архитектурный вывод:** Build a versioned convention-pack asset keyed by pair (×genre), loaded and injected as a snapshot-folded layer (so a pack edit is a loud --resnapshot, mirroring chunkerVersion), sitting between the pair-neutral task frame and the per-book memory/per-run knobs — exactly the three-level factorization research/18 §A4.4 draws. +- **Уточнение верификатора:** The factual core is CONFIRMED (zero convention-pack code — no RenderVars field, no allow-map input, no snapshot-fold version const, no schema; research proposed the layer; D-log parked it), but the disposition NEVER_CLOSED is incorrect. NEVER_CLOSED requires the item "never reached ANY disposition (not in D-log)"; here it DID — D36 §4г (05-decisions-log.md:471) explicitly gives it a documented deferral disposition ("идея в бэклог архитектуры, зона оркестратора/бэкенда, стройка под ROI"), which the finding's own code_reality text concedes ("an acknowledgment and deferral"). The research itself did not ask to build now: §C#15 (line 261) says "Дизайн сейчас, стройка под ROI" and §D-3 (line 271) flags the pack's size/form as unresolved pending the #9 2×2 and #1-2 injection arms. So the zero-code state faithfully reflects a deliberate ROI-gated deferral, not a broken-telephone loss — a more accurate disposition is a conscious backlog deferral (closer to FAITHFUL/deferred) with severity below HIGH, since building it prematurely would front-run its own unresolved shape question. (Minor: the Van-Cui convention CONTENT did land as prose in the polygon discourse-reflow prompt, snapshot-pinned via the template SHA256 at render.go:80 — but not as the factored pair-keyed LAYER, so the layer-absence claim itself stands.) + +### `L4-prompt-language-always-russian` — **UNTESTED_ASSUMPTION** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** The prompt instruction language is unconditionally Russian regardless of pair; a zh→en book would receive Russian instructions, and nothing in code or research ever tested that the Russian-instruction assumption generalizes past zh→ru. +- **Интент ресёрча:** Owner START_PROMT V4 item 5: requests currently go to the model IN RUSSIAN with target=ru; likely they should be in the TARGET language, and for zh→en 'русский вообще не должен участвовать чтоб не сбивать модель'. The owner marks this explicitly open ('вообще не точно') — no research or D-log decision resolved prompt language. + *источник:* START_PROMT.MD V4 item 5; no D-log entry exists on prompt-instruction language (grep of 05-decisions-log.md finds only the adjacent D19.2 language-CONSTRAINT-line note, line 258) +- **Что в коде:** translator.md:1 and editor.md:1 open with all-Russian system prose ('Ты — профессиональный литературный переводчик…') and only inject the pair labels as values; Render() (render.go:135-173) never varies the surrounding instruction language. The polygon has only ever run zh→ru (pipeline-c1.yaml, all rerun evidence in research/18 §A1), so 'Russian instructions work for any pair' is baked but never exercised. D19.2 (05-decisions-log.md:258) already flags that a hardcoded Russian language line (translator.md:10) 'сегодня едет только на thinking-ON DeepSeek' with untested effects on other paths — the same class of untested-language-assumption, unresolved for the pair dimension. + *evidence:* `backend/prompts/translator.md:1; backend/prompts/editor.md:1; backend/internal/pipeline/render.go:137; backend/internal/config/book.go:193` +- **Архитектурный вывод:** Prompt instruction language should be a property of the (pair/target-language) convention pack — the whole template body authored in the right language per pair (target-language or source-language, to be decided by an actual polygon measurement), so zh→en carries no Russian; today's single Russian file makes this impossible without a rewrite. +- **Уточнение верификатора:** Core claim stands; two supporting-citation imprecisions worth flagging (neither undermines the finding, disposition UNTESTED_ASSUMPTION correct). (1) The line-258 note that flags translator.md:10 as riding "только на thinking-ON DeepSeek" is under D21 §6, not a "D19.2" entry — D19.2 itself (context near line 274) is the echo-gate-on-all-configs decision; D21 §9 merely references a D19.2 polygon assignment. That note is also about a Russian STYLE-CONSTRAINT line on reasoning-off echo paths, a different axis the finding correctly uses only as an analogy, not as a resolution. (2) "polygon has only ever run zh→ru" is slightly off: the example/golden fixtures are ja→ru (example/book.yaml:4-5, testdata/golden/book.yaml:5-6) while target is always ru — which actually STRENGTHENS the finding, since editor.md:14's hardcoded "китайского паратаксиса" already mismatches the shipped ja→ru example, not just a hypothetical zh→en. + +### `L4-adding-a-pair-is-a-rewrite-not-config` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** Adding en→ru or zh→en today is a prompt/code authoring task (new files), not a config toggle: there is no pair-keyed prompt selection, and the {{source_lang}} variable is cosmetic against Chinese-specific instruction prose. +- **Интент ресёрча:** The owner wants prompts held per language pair with a place for per-pair rules 'correctly written in the RIGHT languages' (concern 4); a new pair should be an additive, config-selected asset, not a fork of the whole prompt+pipeline. + *источник:* START_PROMT.MD V4 item 5 and V1 item 4; research/18 §C#15 (line 261) versioned-pack intent +- **Что в коде:** To add a pair you must: (1) author a brand-new prompt file with rules and few-shot in the right language (the existing editor.md is unusable — its reflow theory is Chinese-parataxis-specific, editor.md:14/18), and (2) create a new pipeline yaml pinning stages to that file (config/pipeline.go:231-235 resolves st.Prompt relative to the yaml; runner.go:149 loads exactly that path). Nothing selects by book pair (Book.LangPair() is dead, book.go:193), so you maintain parallel pipeline configs by hand. It is a per-pair file+config fork, i.e. a tactical, non-extensible shape. + *evidence:* `backend/internal/pipeline/runner.go:149; backend/internal/config/pipeline.go:231; backend/prompts/editor.md:14; backend/internal/config/book.go:193` +- **Архитектурный вывод:** With a pair-keyed convention-pack layer, adding a pair becomes dropping in one versioned pack asset + a book.yaml pair value the runner uses to select it — a config/data addition, not a prompt+pipeline rewrite; the current design forces the rewrite because the pair is not a selection axis anywhere in the runner. + +### `L4-fewshot-is-structure-not-translation-exemplars` — **FAITHFUL** / LOW · верификатор: — +- **Находка:** The editor few-shot block is STRUCTURE-only reflow exemplars ('показана структура вывода, не стиль'), not correct-translation cases — so the owner's specific worry ('can't stock enough correct-translation examples') does not apply to what is actually in the prompt. +- **Интент ресёрча:** research/18 §A3/§C#2 recommends 'discourse-move few-shot' (chopped source → merged multi-sentence target paragraph; structural demonstrations of the reflow lever), NOT a bank of exemplary translations; the owner separately worried that few-shot examples are correct-translation cases you can never stock enough of. + *источник:* docs/research/18-quality-levers.md §A3 (lines 130-133) and §C1 #2 (line 238); owner START_PROMT V4 item 5 (few-shot worry) +- **Что в коде:** editor.md:21 states the block shows 'структура вывода, не стиль'; the three examples (editor.md:23-31) are [narrative N:1] line-merge, [dialogue 1:N] dash conversion, and [focal-shift] paragraphing — reflow patterns, not model translations to imitate. The block is toggleable per stage via few_shot (config/pipeline.go:92; render.go:108 SystemFor, render.go:117 fewShotEnabled). Caveat: the exemplars are themselves written in Russian and use a Chinese-specific chengyu (物是人非), so they are correct as STRUCTURE exemplars but also baked into the single zh→ru file, inheriting finding L4-prompts-not-per-pair. + *evidence:* `backend/prompts/editor.md:21; backend/prompts/editor.md:23; backend/internal/config/pipeline.go:92; backend/internal/pipeline/render.go:108` +- **Архитектурный вывод:** This lever is implemented as intended — keep it structure-only; when the per-pair pack lands, the few-shot exemplars should move into the pack (authored in the pair's language) so each pair carries its own reflow demonstrations rather than the current zh→ru-only block. + +--- +## L5-sanitizer-infra-pack — Концерн 3 — сломанный телефон (санитайзер/инфра-пак) + +**Итог дорожки:** The sanitizer core is honest and mostly faithful: preamble/trailing-note/markdown/CJK-leak detection map cleanly to the exp12 flagman classes, and the D35.4a strip-and-export mechanism is correctly guarded and resume-deterministic. But the meta-worry holds in three concrete ways. (1) Normalization is fused into the FLAG path instead of an export-contract layer, so identical cosmetic defects (U+3000 indents) survive to export on clean chunks — the reactive D35.4a patch never became the architectural fix (D29.1a export-normalization) it implies. (2) The landed "broken word" class and the diacritics class deliver on none of their named exp12/D30.3 examples, and stripCosmetic hand-rolls an NFKC fold that already exists 300 lines away. (3) One owner-flagged question (legit source-term glosses stripped) sits un-disposed and is in tension with a research lever. The regression guard's 40% threshold is an asserted, uncalibrated default and only observability. + +### `L5-normalization-fused-to-flag-path` — **WORKAROUND_NOT_ROOT** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** Text normalization lives inside the sanitizer FLAG path (stripCosmetic), not in a uniform export-contract layer, so an identical cosmetic defect is normalized on a flagged chunk but shipped raw on a clean one. +- **Интент ресёрча:** exp12 flagman asked for an output-sanitizer that CATCHES instant-unreadability, and research/18 §C#4 framed the CJK fix as a deterministic reflow POST-PROCESSOR ('CJK-глиф=0') applied to the final. D35.4a is the reactive discovery that flag-only silently dropped good content (ch5/ch20 exported empty). The implied clean fix — an export-normalization layer (D29.1a) — was named as backlog in D35 §7 and never built. + *источник:* docs/experiments/12-quality-diagnosis.md §5 (line 248); docs/architecture/05-decisions-log.md D35 §4a (line 452), D35 §7 (line 455 'экспорт-###-нормализация D29.1а'), D36 §4б (line 471) +- **Что в коде:** sanitizeOutput/stripCosmetic run only when the opt-in sanitizer fires and classifies cosmeticOnly; the export writer emits ch.FinalText verbatim. A chunk carrying only a U+3000 ideographic indent does NOT fire the CJK-leak class (U+3000 is deliberately excluded) so it stays disposition=ok and exports the U+3000 unchanged — even though stripCosmetic KNOWS to fold U+3000→space and would do so if any other leak co-occurred. Normalization and detection are entangled behind cosmeticOnly instead of being an unconditional pass over every final text. + *evidence:* `backend/internal/pipeline/sanitizer.go:375 (U+3000 excluded from firing) + :437 (U+3000→space inside stripCosmetic only); backend/internal/pipeline/chunkrun.go:41-56,213-215 (strip only on flag); backend/cmd/tmctl/render.go:31 (FinalText written verbatim, no normalization)` +- **Архитектурный вывод:** Split the concern: an export-contract normalization layer (NFKC fold + whitespace/CJK-punct/U+3000 normalize) applied to EVERY final text unconditionally, with the sanitizer reduced to a pure DETECTOR that only flags the genuinely-unrecoverable. That is the D29.1a layer the D-log itself keeps deferring. +- **Уточнение верификатора:** Claim stands as written; no correction needed. + +### `L5-legit-cjk-gloss-stripped-undisposed` — **NEVER_CLOSED** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** An intentional source-term gloss in parentheses (e.g. «Кулак Ло Хань (羅漢拳)») is stripped and its bracket emptied by stripCosmetic; whether this is desired is an owner question that was raised and never disposed, and it directly contradicts a research lever. +- **Интент ресёрча:** The exp12 flagman named source-term glosses and footnotes-on-allusions as one of the three concrete levers that beat the fan translation (things the fan has and we lack). D38.3 §4б then explicitly parked 'нужны ли издательские глоссы source-термов? да→сузить класс' as an open question for the owner. No later D-entry resolves it. + *источник:* docs/experiments/12-quality-diagnosis.md §5.4 point 2b / §5 point 2 (lines 247-250); docs/architecture/05-decisions-log.md D38.3 §4б (line 529) +- **Что в коде:** detectCJKLeak fires on ANY Han run including an intentional parenthetical gloss; stripCosmetic removes the Han and emptyBracketRE then deletes the now-empty «()», so «Кулак Ло Хань (羅漢拳)» exports as «Кулак Ло Хань». The chunk stays flagged sanitizer_stripped (a human can see it), so it is not fully silent — but the gloss content is gone from the exported text and the design question is unanswered. + *evidence:* `backend/internal/pipeline/sanitizer.go:384-407 (detectCJKLeak fires on any Han run), :459 emptyBracketRE, :445-446 (Han→space); D38.3 §4б` +- **Архитектурный вывод:** The offline detector cannot separate a leaked 特产 from a deliberate parenthetical gloss; a narrow whitelist (balanced Han-in-parens directly after a Latin/Cyrillic head term) or an explicit editorial-gloss convention would resolve it — but a decision is still owed. +- **Уточнение верификатора:** Two secondary imprecisions, neither changes the disposition: (1) the exp12 lever is precisely «footnotes on allusions» (§5 point 2в, examples Su Shi / 采花) plus «enforced glossary» (а); an inline parenthetical source-term gloss is a related-but-distinct mechanism, so the strip is «in tension with» the glossing/footnoting levers rather than «directly contradicts» a lever named «source-term glosses». (2) The citation «§5.4 point 2b» is a mislabel — the content is §5 (ТРЕТИЙ РАУНД) point 2, lines 245-250; the line numbers are accurate. Core finding stands: stripCosmetic strips the intentional gloss and empties the bracket, the owner question was raised in D38.3 §4б and never disposed (NEVER_CLOSED, MEDIUM). + +### `L5-broken-word-class-delivers-no-named-example` — **DISTORTED** / MEDIUM · верификатор: PLAUSIBLE (high) +- **Находка:** The 'broken word forms' class that D30.3 §3 promised (with examples «Первок предок»/«валуне») catches NONE of its named examples in the landed code; only two zero-FP signatures survive that were never the exp12 target. +- **Интент ресёрча:** exp12 flagman and D30.3 §3 named broken word forms («Первок предок», «Первопред предок», «Силённый», «клёцок», «валуне») as a sanitizer defect class to catch. + *источник:* docs/architecture/05-decisions-log.md D30.3 §3 (line 367 'битые словоформы («Первок предок», «валуне»)'); docs/experiments/12-quality-diagnosis.md §5 point 3 (line 248) +- **Что в коде:** The junction-duplication split detector was removed (D33.1б) because it false-flagged ordinary polyptoton; the surviving BrokenWord class catches only invalid Cyrillic sign bigrams («ьзал») and homoglyph-mixed tokens — neither of which is any of the named examples. The code's own comments concede «Первок предок» and grammatical-but-wrong «валуне» are not offline-detectable (need a Ф2 morphology pass). The real root was correctly identified and honestly deferred, but the D30.3 decision still advertises a capability the code does not have. + *evidence:* `backend/internal/pipeline/sanitizer.go:49-51 (recall-gap admission), :316-354 (BrokenWord = invalid-sign + homoglyph only), :340-347 (split detector removed)` +- **Архитектурный вывод:** Either land the Ф2 morphology/dictionary pass that the class actually needs, or retract the class label in the contract; the D-log should carry the retraction rather than leaving a named lever that silently delivers nothing. +- **Уточнение верификатора:** Core distortion is confirmed: the D30.3 §3 / exp12 named broken-word examples («Первок предок», «валуне», «клёцок», «Силённый») are all uncaught by the landed code, which explicitly admits the recall gap (sanitizer.go:46-51, :340-347) and defers them to a Ф2 morphology pass, while the D-log entry still advertises the class without a retraction — so DISTORTED (MEDIUM) and the fix-or-retract implication both stand. But the finding's specific claim that the two surviving signatures "were never the exp12 target" and "neither is any of the named examples" is inaccurate: the homoglyph-mixed-token signature (sanitizer.go:348-351) DOES catch exp12's «vалуне» (Latin-'v' + Cyrillic) named at 12-quality-diagnosis.md:248 — an exp12-named example, categorized there under «латиница в кириллице» (Latin-in-Cyrillic), not broken-word. The finding conflates D30.3's «валуне» (all-Cyrillic, genuinely uncaught) with exp12's «vалуне» (Latin-v, caught). Only the invalid-sign signature («ьзал») is truly not an exp12-named example. + +### `L5-stripcosmetic-duplicates-nfkc` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** stripCosmetic hand-rolls a fullwidth-ASCII fold (r-0xFEE0) plus per-char CJK-punct cases, duplicating the NFKC normalizer already used in the same package by the memory layer. +- **Интент ресёрча:** N/A — implementation cleanliness / reuse. + *источник:* none +- **Что в коде:** sanitizer.go:439-440 folds U+FF01–FF5E via arithmetic (r-0xFEE0) and adds individual cases for «、»/«。»; the comment even says it is 'the same NFKC fold the memory normaliser uses' — yet memnorm.go:104-105 already calls norm.NFKC.String, which folds exactly that fullwidth-ASCII range. The fold is reimplemented instead of reused, and only the genuinely-custom semantic maps (、→, and 。→. beyond NFKC) actually need bespoke cases. + *evidence:* `backend/internal/pipeline/sanitizer.go:439-444 vs backend/internal/pipeline/memnorm.go:104-105 (norm.NFKC.String)` +- **Архитектурный вывод:** Reuse norm.NFKC (already a dependency) for the width fold and keep only 、/。 as explicit semantic maps; this removes a divergent partial reimplementation of a Unicode operation the project already owns. + +### `L5-regressionguard-40pct-uncalibrated` — **UNTESTED_ASSUMPTION** / LOW · верификатор: — +- **Находка:** The regression guard's 40% length-collapse threshold is an asserted, uncalibrated default, and its number-drift flagger only covers arabic-side drift while the actually-observed defect is word↔word. +- **Интент ресёрча:** research/18 §C#5 and D38.4 wanted an omission + number-drift guard; D38.4 §3 explicitly records b1 (四成四=44%→«четыре десятых», word↔word) as only PARTIALLY covered and NOT to be treated as closed, and D38.3 §4а labels the 40% a revisable observability default. + *источник:* docs/architecture/05-decisions-log.md D38.4 §3 (line 538), D38.3 §4а (line 529) +- **Что в коде:** regressionMaxShrinkPct=40 is justified by an asserted 'Russian is ~fertility-neutral draft→final' with no measured draft→final length distribution behind it (D35 measured similarity 0.643, not a length ratio). arabicNumberTokens compares only multi-digit ASCII-digit runs, so a draft that spells numbers as words (the glm case) is invisible. It is observability-only (never a disposition), so risk is bounded. + *evidence:* `backend/internal/pipeline/regressionguard.go:30-40 (const + rationale), :24-28 (recall-gap admission), :94-130 (arabic-only tokens); folded observability-only via cheapgates.go:90-95` +- **Архитектурный вывод:** Calibrate the shrink threshold against the real pere-progon draft→final length distribution instead of asserting fertility-neutrality; keep treating b1 as open (D38.4 §3) rather than as delivered by this guard. + +### `L5-diacritics-class-lost` — **LOST** / LOW · верификатор: — +- **Находка:** The diacritics defect class the exp12 flagman explicitly named («источа́ло») was dropped at the D30.3 decision step and never reached any code — a clean lost-in-transit example. +- **Интент ресёрча:** exp12 flagman listed диакритики («источа́ло») as one of the concrete sanitizer defect families to catch (combining acute on a Cyrillic base). + *источник:* docs/experiments/12-quality-diagnosis.md §5 point 3 (line 248) +- **Что в коде:** D30.3's class list (05-decisions-log:367) already omits diacritics, and there is no combining-mark / unicode.Mn handling anywhere in sanitizer.go — the class never became one of the six. The defect is rare, but its disappearance between polygon finding and ratified decision is exactly the broken-telephone pattern being hunted. + *evidence:* `backend/internal/pipeline/sanitizer.go (no unicode.Mn / U+0300–036F detection; grep confirms absent); docs/architecture/05-decisions-log.md:367 (class list omits diacritics)` +- **Архитектурный вывод:** A combining-diacritic-on-Cyrillic detector is a zero-false-positive offline signal; restoring it (or a normalize-to-NFC pass in the export layer of L5-1) would close the named lever cheaply. + +### `L5-sanitizer-version-churn-patch-pile` — **CODE_SMELL** / LOW · верификатор: — +- **Находка:** sanitizerVersion has churned v1→v4, each bump bolting a targeted regex/case onto stripCosmetic in response to an adversarial-review edge, leaving a pile of ad-hoc tidy regexes rather than a structural cleaner. +- **Интент ресёрча:** N/A — the meta-worry that the project drifted into tactical patching. + *источник:* docs/architecture/05-decisions-log.md D33.1 (line 421), D38.3 §2-3 (lines 527-528) +- **Что в коде:** Four version bumps layer fixes: D33.1 (add tail edit-summary, REMOVE broken-word split, narrow markdown/latin), D38 v3 (add CJK class + cosmetic/substantive split), D38 v4 (fold fullwidth, ideographic comma/period, empty-bracket removal, the special-case '2 : 1' colon exclusion). The whitespace tidy is five separate regexes (multiSpaceRE, emptyBracketRE, spaceBeforePunctRE, spaceAfterOpenRE, trailingHSpaceRE) applied in sequence. Each patch is individually well-tested, but the accretion is the tactical-patch fingerprint the owner worries about. + *evidence:* `backend/internal/pipeline/sanitizer.go:53-69 (v1→v4 history), :458-472 (five-regex tidy pile + '2 : 1' carve-out)` +- **Архитектурный вывод:** A structural cleaner (NFKC pass + a single whitespace/punctuation-spacing normalizer shared with the export layer of L5-1) replaces the per-edge regex accretion and stops the version-churn treadmill. + +### `L5-strip-export-mechanism-sound` — **FAITHFUL** / LOW · верификатор: — +- **Находка:** Setting the above aside, the D35.4a strip-and-export mechanism itself is correct, well-guarded, and resume-deterministic — the answer to 'correct or patch with edge cases' is: correct, with the gloss case (L5-6) the one live edge. +- **Интент ресёрча:** D35.4a: do not drop a whole chunk to an empty placeholder for a removable cosmetic leak (ch5/ch20 lost ~44% of their openers to a leading «###»); export the cleaned remainder, still flagged for a human. + *источник:* docs/architecture/05-decisions-log.md D35 §4a (line 452), D38.3 §2a (line 527) +- **Что в коде:** cosmeticOnly() gates the strip tier and a substantive class always wins (drops); classifyOutput re-validates sanitizeOutput(stripped).total()==0 && non-empty BEFORE tagging FlagSanitizerStripped, so the derived export is guaranteed usable; commitSanitizedExport writes a content-addressed $0 derived checkpoint (tm-sanitized-v1:) that resume re-serves identically; the escalation-fallback strip path is handled (a stripped fallback is authoritative). Backed by real tests (sanitizer_test.go, regressionguard_test.go) and golden re-capture, per D38.3 verification. + *evidence:* `backend/internal/pipeline/chunkrun.go:49-54 (re-validation guard); backend/internal/pipeline/stagerun.go:154,170-183,228-239 (derived checkpoint + escalation path); backend/internal/pipeline/resume.go:54-70 (resume re-serve)` +- **Архитектурный вывод:** Keep the mechanism; the residual work is architectural placement (L5-1) and the gloss decision (L5-6), not correctness of the strip. + +--- +## L6-discourse-reflow-fidelity — Концерн 3 — сломанный телефон (дискурс-reflow промпт) + +**Итог дорожки:** The Ван Цуй 5 techniques + participial/adverbial constructions are encoded FAITHFULLY (near-verbatim) in editor.md point 2, and the chengyu atom (物是人非 = «вещи те же, люди не те») is present and correct. The one real research→prompt distortion is the "technique not norm" caveat: the shipped header over-claims the parataxis→hypotaxis transform as «это КОНВЕНЦИЯ русского языка» — the exact over-elevation D36 popravka #4 ratified as a correction — because the corrected framing reached exp14's meta-prose but not the frozen arm STRING, which was APPENDED verbatim to production (D38.5). Two secondary landing issues: the chengyu atom is in prod v3 UNMEASURED (exp14b showed the flattening is model-universal; no arm tested that a one-line prompt fixes it), and it sits unreconciled against editor.md line 5's opposing de-literalization mandate, a tension the orchestrator itself flagged in D38.4 §1 yet landed anyway. + +### `L6-technique-not-norm-overclaim` — **DISTORTED** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** The discourse-reflow header over-claims the parataxis→hypotaxis transform as a «convention of the Russian language», reproducing the exact over-elevation D36's ratified correction warned against; the caveat reached exp14 meta-prose but not the shipped prompt string. +- **Интент ресёрча:** D36 popravka #4 ratified: Ван Цуй is a prescription of translation TECHNIQUE (parataxis→hypotaxis), NOT a «connecting norm of the language»; the norm/convention answer to the owner («convention, not author-style») is carried FIRST by Rosenthal/Rules-1956 (paragraph/dialogue), and Ван Цуй is HOW to execute it. The frame «обязательная дискурс-операция наравне с нормой» was explicitly called «слегка пере-возвышена». + *источник:* docs/architecture/05-decisions-log.md D36 §2 (line 467); docs/research/18-quality-levers.md review-header popravka #4 (lines 50-54); correct meta-framing exists at docs/experiments/14-quality-empirics.md:130 («прескрипция ТЕХНИКИ; норму держат Розенталь/Правила-1956») +- **Что в коде:** editor.md's ДИСКУРС-ПЕРЕВЁРСТКА header labels the whole operation «обязательная операция передачи китайского паратаксиса в русскую гипотактическую прозу — это КОНВЕНЦИЯ русского языка, а не вольность», i.e. it grants convention/norm status to the syntactic transform itself rather than to the Rosenthal paragraph/dialogue rules the transform serves. The string is byte-identical to the frozen measurement arm and was appended verbatim to prod (D38.5 «APPEND к baseline»). + *evidence:* `backend/prompts/editor.md:14 (verbatim source: eval/exp14/exp14_prompts.py:32 DISCOURSE_CORE; landed per docs/architecture/05-decisions-log.md D38.5 line 549)` +- **Архитектурный вывод:** The shipped prompt should mirror the ratified split that already exists in exp14 §1.2 meta-prose: present the merger as a REQUIRED translation technique (Ван Цуй) that serves the Russian paragraph/dialogue CONVENTION (Rosenthal), rather than labeling the parataxis transform itself «конвенция русского языка». Low functional harm (the over-claim reads as compliance-pressure on the model), but it is exactly the fidelity axis the owner flagged, and it entered prod because the correction lived only in commentary while the frozen arm string bypassed it. +- **Уточнение верификатора:** Verdict holds; two precision refinements. (1) Severity is at the LOW end of MEDIUM: popravka #4 itself calls the over-elevation «слегка» (slight) and «дёшев → строить», and the finding concedes low functional harm — the header reads as compliance-pressure so the model does not treat the merger as unfaithful liberty. (2) Strictly, this is not a landing-time distortion: the over-claim was baked into the frozen measurement arm (exp14_prompts.py:32) from inception, and the prompt shipped byte-identical to the empirically-winning arm (D38.5 APPEND). The gap is therefore that D36's ratified technique-vs-norm correction targeted the research report's framing and lived only in commentary (research/18 popravka #4, exp14 §1.2 meta-prose) — it never bound the prompt string, so the string faithfully carried its original over-elevated framing through measurement into prod. That is a real fidelity gap on the exact axis the owner flagged, but it is a commentary-never-reached-the-string gap, not a mis-port that weakened a correctly-framed original. + +### `L6-deliteralize-vs-chengyu-tension` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** The chengyu both-components atom is landed into a prompt that already, three lines earlier, mandates DE-literalizing idiom calques — a direct contradiction the orchestrator itself flagged in D38.4 §1 but landed without reconciling, leaving the model to arbitrate the boundary. +- **Интент ресёрча:** D38.4 §1 explicitly noted that editor.md line 5 carries the OPPOSITE instruction («там ПРОТИВОПОЛОЖНОЕ — де-буквализация калек») to the chengyu two-component atom, and that the P1a atom-list (exp14_prompts.py) had no idiom atom at all — the chengyu lever was a late addition. + *источник:* docs/architecture/05-decisions-log.md D38.4 §1 (line 536) +- **Что в коде:** Line 5 instructs the editor to fix «кальки жестов и идиом — переводчик мог передать выражение буквально (напр. поклон как ‹ударил головой›)», i.e. de-literalize idioms rendered word-for-word. Line 18 instructs the OPPOSITE for chengyu: «передавай ОБОИМИ смысловыми компонентами, не сворачивая в общий смысл». The distinction (de-literalize meaning-obscuring gesture/realia calques vs preserve the two-part contrast of a chengyu whose meaning IS the contrast) is never made explicit; a model reading both must guess which idioms fall under which rule. + *evidence:* `backend/prompts/editor.md:5 vs backend/prompts/editor.md:18` +- **Архитектурный вывод:** Make the boundary explicit in the prompt so the two rules stop competing: de-literalize gesture/realia calques that OBSCURE meaning (line 5), but preserve the semantic contrast of a chengyu whose meaning is that contrast (line 18). Left unreconciled, line 5's de-literalization mandate can push 物是人非 back toward the very «всё изменилось» collapse the atom exists to prevent. + +### `L6-chengyu-atom-untested` — **UNTESTED_ASSUMPTION** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** The chengyu prompt atom ships in production prompt v3 but was never measured; exp14b established the 物是人非 flattening is UNIVERSAL across all 6 model families, yet no arm ever tested whether a one-line prompt instruction fixes a model-universal semantic defect. +- **Интент ресёрча:** D38 §2г found 物是人非→«всё изменилось» universal across glm/gpt/grok/deepseek-pro/mistral/kimi and concluded it is «рычаг промпта/глоссария, не модель» — but that conclusion is a HYPOTHESIS about fixability, not a measured result; D38.4 §1 states «пере-прогон измерит». + *источник:* docs/architecture/05-decisions-log.md D38 §2г (line 498); D38.4 §1 (line 536); docs/experiments/14-quality-empirics.md:253 («T2 (物是人非 chengyu): P0/P1a сплющивают») +- **Что в коде:** editor.md:18 carries the atom in prod prompt v3; D38.5 confirms the whole v3 discourse editor «ещё не гонялся» (never run). The lever moved from empirical observation of a DEFECT straight to a prod prompt-side FIX without an arm demonstrating the prompt closes it. + *evidence:* `backend/prompts/editor.md:18; docs/architecture/05-decisions-log.md D38.5 line 549 («prompt_version=v3 (ещё не гонялся)»)` +- **Архитектурный вывод:** Acceptable as a disclosed near-term probe (cheap, and the re-run will score 物是人非 via fidelity re-gate), but it must not be treated as the CLOSED fix: exp14 §2 (line 253) even suggests the flattening tracked model capability on some traps. If the atom does not move 物是人非 in re-run, the ratified long-term home is the Ф2 chengyu annotator-mask (04-unhappy §124/§62), which is the root. + +### `L6-vancui-5tech-faithful` — **FAITHFUL** / LOW · верификатор: — +- **Находка:** editor.md faithfully encodes all 5 Ван Цуй techniques plus participial/adverbial constructions and subordinating conjunctions, near-verbatim from research/18. +- **Интент ресёрча:** research/18 §A4.4/§C1 prescribes the 5 techniques — (1) build hierarchy, (2) segment by meaning units, (3) mark aspect/tense of predicates per Russian norms, (4) vary lexicon, (5) add connectives — with participial/adverbial constructions + subordinating conjunctions as the Russian-specific merger instrument. + *источник:* docs/research/18-quality-levers.md §A4.4 line 162; §C1 row 1 line 237; measurement arm docs/experiments/14-quality-empirics.md:132 +- **Что в коде:** editor.md point 2 reads «Инструмент слияния — причастные и деепричастные обороты, подчинительные союзы и связки (построй иерархию, сегментируй по смыслу, промаркируй вид/время предикатов, варьируй лексику, добавляй связки)» — all five imperatives map 1:1 to the source, and the participial/adverbial + subordinating-conjunction instrument is present. Only cosmetic compression: technique 3 drops «по нормам РЯ». No lossy paraphrase. + *evidence:* `backend/prompts/editor.md:16` +- **Архитектурный вывод:** None — this is the load-bearing content and it landed cleanly. This is the strongest evidence the polygon→prompt port succeeded on the substance; the distortion is confined to the framing header, not the technique list. + +### `L6-chengyu-atom-correct` — **FAITHFUL** / LOW · верификатор: — +- **Находка:** The chengyu atom is present and rendered with BOTH components correctly: 物是人非 = «вещи те же, люди не те», with «всё изменилось» correctly forbidden as the flattening. +- **Интент ресёрча:** research/18 §A2 and the exp14 trap table define 物是人非 as the contrast 物是 («вещи те же») + 人非 («люди иные»), with the flattening «всё изменилось» losing half the idiom (a semantic-atom loss). + *источник:* docs/research/18-quality-levers.md §A2 line 100; docs/experiments/14-quality-empirics.md:87 (trap T2) +- **Что в коде:** editor.md point 4 appends «идиомы и чэнъюй: их передавай ОБОИМИ смысловыми компонентами, не сворачивая в общий смысл — 物是人非 = ‹вещи те же, люди не те›, а не ‹всё изменилось›». 物是 (things unchanged) and 人非 (people changed) are both rendered; the anti-example matches the exact flattening the empirics observed. + *evidence:* `backend/prompts/editor.md:18` +- **Архитектурный вывод:** Correct as written. (Fidelity of the atom itself is clean; its open risks are captured separately in L6-chengyu-atom-untested and L6-deliteralize-vs-chengyu-tension.) + +### `L6-deterministic-reflow-ops-partial` — **LOST** / LOW · верификатор: — +- **Находка:** research/18 prescribed reflow ops (b) 引号→dash/«», (c) glyph-normalize, (d)   -strip as DETERMINISTIC code; only op (c) landed, and as a precision-over-recall CJK-leak DEFECT gate (not a normalizer), while dialogue formatting (b) is carried in the model prompt. +- **Интент ресёрча:** research/18 §A3 (line 127) / §A4.2 / §C1#4 frame correct zh→ru reflow as a 4-operation transform where ops (b),(c),(d) are «ДЕТЕРМИНИРОВАННЫЕ (код, без модели)» — a deterministic reflow post-processor, separate from the model's discourse merger op (a). + *источник:* docs/research/18-quality-levers.md §A3 line 127; §C1 row 4 line 240 +- **Что в коде:** sanitizer.go v3 (D38 infra-pack) added a CJK-leak class that STRIPS stray Han/kana/fullwidth glyphs as a high-confidence DEFECT gate (≥15% threshold, precision over recall), not a general normalizer; there is no deterministic 引号→«»/dash converter or   -indent stripper. Dialogue-dash is instead a MODEL instruction (editor.md point 3). + *evidence:* `backend/prompts/editor.md:17 (dialogue-dash as model instruction); backend/internal/pipeline/sanitizer.go:30-40 (CJK-leak = defect gate, cosmetic-strip); no b/d normalizer found` +- **Архитектурный вывод:** Adjacent to Lane 4, and largely DEFENSIBLE for this architecture: the editor model regenerates the whole Russian text, so 引号/   never survive unless the model copies them, and the CJK-leak gate backstops the residual copy-through. But the research's «reflow is PARTIALLY FREE / do b,c,d in code» framing did not land as a normalizer — the pipeline relies on the model for b/d and on a narrow ≥15% defect gate for c, so sub-threshold glyph leaks and any dialogue-punctuation errors have no deterministic backstop. + +--- +## L7-hole-sweep — Концерн 5 — дыра в передаче знаний (незакрытое) + +**Итог дорожки:** The un-closed tail is large and clusters exactly where the owner's arch-reset already pointed: "smart chunking" (D35.7) was tagged in 3 docs but never written into the D-log or the code (chunker is still mechanical greedy packing to a hardcoded 1500-token const); the translator↔editor split was never tested (draft-as-translator's structural contribution is unmeasured per D30.6, and reflow was dumped entirely on the whole-chunk-rewriting editor); and the prompt layer is russo-centric single-template with no per-language-pair rule architecture (owner V4 п5). Beyond those three, a systematic sweep surfaces ~15 more open items: owner V-series ideas that never reached ANY disposition (input-size limit, cheap abuse pre-screen, parallelism/speed as owner's stated "main V4 idea", user-initiated stop), researcher recommendations parked by default (running-summary near-term re-weigh, in-loop automatic quality eval the owner explicitly requested, direct §D render questions to backend that were never answered), and the D38.4 audit's own four un-built tail items (inversion-guard, word↔word number-drift, chengyu long-term, large-edit-unit). Critically, the owner's actual concern-5 ask — a standing findings-ledger PROCESS, not one-off audits — itself has no home; D38.4 was a manual one-shot and nothing institutionalizes it. + +### `H1-smart-chunking-never-formalized` — **NEVER_CLOSED** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** "Smart/logical chunking" (D35.7) is cited as a decision in PROGRESS/exp14/research18 but has NO D-log entry, and the code is still mechanical greedy paragraph-packing to a hardcoded const with no logical/scene-boundary segmentation and no fallbacks. +- **Интент ресёрча:** research/18 §D-1 says the coherence-failure context locus is unmapped and primary for polygon; exp14 curve found whole-chapter better AND cheaper; the intent (D35.7) is to decouple draft=chunk/edit=chapter and cut chapters into LOGICAL units when a chapter exceeds the model window. Owner arch-reset concern 2 states chunking affects both paragraph structure and source meaning, so it is non-trivial and must be solved in static code with internet best-practice research and code fallbacks. + *источник:* research/18 §C row 8 + §D-1 (docs/research/18-quality-levers.md:244,269); exp14:350; ORCHESTRATOR_ARCH_RESET_PROMPT.md:38 ("тег D35.7 … записи в D-логе НЕТ") +- **Что в коде:** SplitChunks packs whole paragraphs to a token budget, descending to sentence boundaries only when one paragraph overflows; the budget is a non-tunable const; no scene/logical boundary detection, no per-book adaptation, no fallback ladder. + *evidence:* `backend/internal/pipeline/chunker.go:47 (const targetChunkTokens = 1500); SplitChunks/appendChapterChunks (chunker.go:52-80)` +- **Архитектурный вывод:** A dedicated chunker module with logical-segmentation strategies (scene/dialogue/topic boundaries), configurable per source language, snapshot-folded, with a documented fallback chain — plus a formal D-log decision — rather than an unwritten tag over a 1500-token const. This is the lever the owner nearly lost entirely (the trigger for the whole reset). +- **Уточнение верификатора:** Verdict stands; two minor labeling nuances (not defects in the disposition): (1) 'D35.7' strictly denotes the runner-level decoupling draft=chunk/edit=chapter (itself also un-implemented — editor runs per-chunk, ARCH_RESET:84), while logical/scene chunking is the related lever from research/18 §C row 8 + owner concern 2; the finding bundles both under 'the intent (D35.7)' but correctly attributes the logical-cutting to concern 2 in its own body. (2) The arch-reset path is docs/ORCHESTRATOR_ARCH_RESET_PROMPT.md:38 (the citation dropped the docs/ prefix). Neither affects the HIGH-severity NEVER_CLOSED disposition. + +### `H2-per-language-prompt-architecture` — **NEVER_CLOSED** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** There is no per-language-pair prompt architecture: every stage uses ONE template with {{source_lang}}/{{target_lang}} variable substitution, but the actual RULES baked in are zh→ru-specific and the prompt language is hardcoded Russian (russo-centric). +- **Интент ресёрча:** Owner V4 п5 (START_PROMT.MD:90) and arch-reset concern 4: prompts must be per source→target pair, language rules/conventions/academic findings must be correctly multiplied and written in the RIGHT languages; a zh→en job should not route through Russian at all. research/18 §C row 15 proposes a versionable "language-pair conventions pack" keyed (src→tgt[×genre]) as a design idea only. + *источник:* START_PROMT.MD:90 (V4 п5); ORCHESTRATOR_ARCH_RESET_PROMPT.md:46-50 (concern 4); research/18 §C row 15 (docs/research/18-quality-levers.md:261) +- **Что в коде:** translator.md/editor.md are single Russian-language templates; reflow rule "НЕ копируй построчную разбивку исходника … в оригинале часто одно предложение — одна строка" and Rosenthal dialogue-dash conventions are hardcoded zh→ru assumptions injected for ANY pair; the whole system prompt is Russian regardless of target. + *evidence:* `backend/prompts/editor.md:5,6 (zh-line-strip + dash reflow); backend/prompts/translator.md:1,9 (Russian system prompt, line-strip rule); no per-pair rule store anywhere under backend/prompts or internal/config` +- **Архитектурный вывод:** A pair-keyed rules layer (src→tgt[×genre]) whose content is snapshot-folded and whose few-shot examples and convention text are stored per language and written in the correct prompt language — not one Russian template that silently assumes the source is CJK and the target is Russian. +- **Уточнение верификатора:** Claim confirmed. Two precision refinements (both strengthen it): (1) editor.md is even more russo-centric than stated — it hardcodes "на русский язык" with NO {{target_lang}} variable, and lines 14-18 literally name "китайский паратаксис" with a chengyu example, so it is explicitly zh→ru, not merely "assumed CJK"; the citation should point at editor.md:8,14-18 rather than 5,6. (2) The narrower research/18 §C row 15 "conventions-pack" sub-idea did reach ONE backlog disposition (D35 §4(г), 05-decisions-log.md:471, "идея в бэклог … стройка под ROI"), but that entry still assumes zh→ru content and does not address prompt-language/russo-centrism — the core V4-п5 concern (per-pair prompt architecture + which language the prompt itself is written in) remains genuinely undispositioned in the D-log and absent from code, so NEVER_CLOSED (HIGH) stands. + +### `H3-translator-editor-split-untested` — **UNTESTED_ASSUMPTION** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** The architecture assumes the editor fixes structure while the draft/translator is a raw pass, but the translator's contribution to structure/artistry was NEVER measured, and the translator-gives-structure vs editor-rewrites-everything split was never A/B tested. +- **Интент ресёрча:** D30.6 premise flags that flash-as-TRANSLATOR was never measured for artistry; exp13/D32 admitted the quality lift comes from the BUNDLE (bilingual editor+reflow+sanitizer+seed), not the translator in isolation. Owner arch-reset concern 1 asks why the draft is not produced near-correct-and-structured, whether the editor rewrites the whole chapter, and whether the polygon tested this at all. + *источник:* PROGRESS.md:178 ("Черновик flash как ПЕРЕВОДЧИК на художественность не мерен НИКОГДА (D30.6)"); D32/D-log:396 (lift = связка, not translator); ORCHESTRATOR_ARCH_RESET_PROMPT.md:24-29 (concern 1) +- **Что в коде:** reflow instructions live in BOTH translator.md and editor.md, but every exp14/14b arm held the draft constant and varied only the EDITOR; the editor emits the full re-written text of each chunk (whole-unit rewrite), so structure de-facto rests on the editor and the split was never isolated. + *evidence:* `backend/prompts/editor.md:8,10 ("собирай повествование в естественные русские абзацы"; "Выведи ТОЛЬКО отредактированный текст перевода" = full rewrite); exp14 arms vary editor only (docs/experiments/14-quality-empirics.md:88-90)` +- **Архитектурный вывод:** A translator/editor contract that measures and assigns responsibility for structure explicitly (translator emits a structured near-correct draft; editor does targeted meaning+style repair), validated by a polygon A/B, rather than an untested assumption that dumps reflow onto a whole-chunk-rewriting editor. +- **Уточнение верификатора:** Substance fully holds; only fix imprecise code_evidence pointers: "Выведи ТОЛЬКО отредактированный текст перевода" is editor.md:12 (not :10 — line 10 is honorifics/venuti); the arms-vary-editor evidence is 14-quality-empirics.md:114 and :141 (not :88-90, which is the trap-table); the "lift=bundle" quote is D-log:398 (:396 is the exp13 section header). Nuance: reflow lives in BOTH prompts but asymmetrically — translator.md:9 is a light layout rule only, the heavy parataxis→hypotaxis merge is editor-only (editor.md:14-31); and a fidelity-only frontier-as-translator spot-check did exist (14:148, T1/T2), so "translator NEVER measured" is exact for STRUCTURE/ARTISTRY but the translator was measured for FIDELITY in exp13 and via that spot-check. + +### `H5-no-in-loop-quality-signal` — **NEVER_CLOSED** / HIGH · верификатор: CONFIRMED (high) +- **Находка:** The owner explicitly requested online/offline quality measurement, and research/18 C1 #10 lists automatic quality eval as a NOW lever, but the running pipeline has ZERO quality signal in the loop and no per-segment "where did meaning slip" signal. +- **Интент ресёрча:** START_PROMT п.31 (measure online/offline quality) and V2-outputs §2; research/18 §C1 row 10 marks automatic quality eval (deterministic KPI for claim-1; span-citing bilingual judge + comprehension-QA for claim-2) as phase=Сейчас, an explicit owner request. + *источник:* research/18 §C1 row 10 (docs/research/18-quality-levers.md:246, "Автоматическая оценка качества (запрос владельца)"); START_PROMT.MD:42 (п.31); PROGRESS.md:178 ("ноль оценки качества в контуре, ноль per-segment сигнала") +- **Что в коде:** role=judge is gated as non-executable in Phase 0/1; the pipeline is a linear draft→editor + deterministic gates; the fidelity re-gate exists only as an off-line polygon script over the book output, not in-product; no comprehension-QA or span-citing judge anywhere in backend. + *evidence:* `backend/internal/config/pipeline.go:190-191 (role=judge не исполним); pipeline.go:180-183 (C2/C3/fanout gated); no judge/quality-eval call path in internal/pipeline` +- **Архитектурный вывод:** An in-loop (or at-least per-run) automatic quality signal — deterministic structural KPI plus a cheap span-level meaning check with per-segment flags — so the system can tell WHERE quality degraded, which is the precondition for both self-improvement and the owner's quality-telemetry-from-day-one goal. +- **Уточнение верификатора:** Two precision refinements (do not change the finding's validity, and actually sharpen the broken-telephone thesis): (1) "ZERO quality signal" means zero SEMANTIC/quality-eval signal — per-chunk deterministic defect/consistency flaggers DO exist and persist to retrieval-state (coverage/excision, sanitizer, glossary post-check n_postcheck_miss, cheap style flags, regression-guard length/number drift), but none is the research/18 row-10 quality eval (neither the deterministic paragraph-vs-reference KPI nor the meaning-slip judge). This matches PROGRESS.md:178's own framing. (2) The disposition is slightly imprecise as blanket NEVER_CLOSED: the in-loop JUDGE-selector IS consciously dispositioned — deferred to Phase 2 (pipeline.go:190-191 gate, chunkrun.go:65, 05-decisions-log.md:21, judge-dubler pilot blocker D22.6). What truly never reached its own disposition is the research/18 §C1 row-10 NOW measurement lever (deterministic quality-KPI без судьи + online/offline quality telemetry per START_PROMT п.25/п.31): it got collapsed into the blanket "Ф2 quality machinery не построена" framing (PROGRESS.md:178) and never landed as the NOW deliverable it was flagged to be — a NOW lever silently downgraded to a Phase-2 deferral. + +### `H4-editor-whole-chunk-rewrite-vs-minimal-diff` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** The editor rewrites the entire chunk (owner's "редактор переписывает ВСЮ ГЛАВУ" fear is real), and research/15's minimal-diff/search-replace editor — which would cut the editor's ~90% COGS AND give provable edits — is parked as a Ф2 pilot arm and re-surfaced by the owner as if new. +- **Интент ресёрча:** research/15 found search/replace blocks are the best frontier patch format (0.94-0.95 EM), flexible apply mandatory, cutting the editor's ~90% output-token COGS and yielding "provable edits" for free; recommended as a pilot arm with a Go apply spec due before pilot. + *источник:* research/15-voice-and-state.md:153,245; D25.6/D-log:256 (minimal-diff редактор gated by search/replace-apply machinery, спека бэкенду до пилота) +- **Что в коде:** editor stage emits the full re-written chunk text; no diff/patch channel, no NO_CHANGE path, no search/replace apply; every editor call re-emits all output tokens (the dominant cost line, editor ~87% of run cost). + *evidence:* `backend/prompts/editor.md:12 ("Выведи ТОЛЬКО отредактированный текст перевода"); PROGRESS.md:157 (editor glm-5 = 87.1% of committed cost); no diff/apply code in internal/pipeline` +- **Архитектурный вывод:** A minimal-diff editor contract (editor returns search/replace spans over the draft + deterministic flexible Go apply) so the editor stops re-emitting whole chunks — cheaper, lower-risk (no full-rewrite drift), and auditable. It is dispositioned Ф2 but the owner re-raising it is itself a broken-telephone symptom worth surfacing. + +### `H6-running-summary-near-term-reweigh` — **NEVER_CLOSED** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** The researcher explicitly recommended re-weighing DelTA-class running bilingual summary as a NEAR-TERM lever (not burying it on COGS until measured), but it was never measured and stayed Ф2 by default; §D-5 confirms it is un-measured. +- **Интент ресёрча:** research/18 §B2.3 argues the team consciously does not build LLM summaries on a COGS argument, but DelTA's benefit is exactly on the same domain (webnovels) and the COGS objection weakens under cache-friendly ordering (~1 small call per ~20 sentences, near-free input at prefix cache) → measure it near-term, do not bury. + *источник:* research/18 §B2.3 (docs/research/18-quality-levers.md:210) + §C2 row 11 (:252) + §D-5 (:273, "COGS running-summary при кэше … не замерено") +- **Что в коде:** Only deterministic Aho-Corasick memory-slice injection exists; no running-summary or entity/antecedent register; exp14's cross-boundary arm tested ±1 reference (A-ref, null result), not a running summary, so the researcher's specific near-term probe was never run. + *evidence:* `backend/internal/pipeline/memory.go:18 ($0, no LLM, no embeddings hot-path matcher); no summary/register table in internal/store/migrate.go beyond deterministic glossary` +- **Архитектурный вывод:** A cache-friendly running bilingual summary + antecedent register measured as a near-term claim-2 lever before being defaulted to Ф2 — closing the explicit researcher recommendation instead of letting the COGS prior silently win. + +### `H7-input-size-limit-missing` — **NEVER_CLOSED** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** The owner asked for an input-size limit / input-token counter for the book (V2 п3); the backend has a MONEY ceiling (book_usd) and cost projection but NO preflight input-size guard or rejection, and it has no D-log disposition. +- **Интент ресёрча:** Owner wants a configurable input-size/token limit so an oversized book is caught before processing, distinct from the money ceiling. + *источник:* START_PROMT.MD:70 (V2 п3, "ограничения по входу … настроечка, которая будет считать входные токены для книги") +- **Что в коде:** No maxInput / input-token-cap / oversize-rejection code exists; the only guard is the money ceiling (book_usd, D4) which pauses on cost, not on input size; grep for input-limit patterns returns nothing. + *evidence:* `No matches for input-size/token-limit patterns in backend/internal; only escalation.go:42 money pause/resume; D-log has no input-limit entry` +- **Архитектурный вывод:** A preflight input-size estimate + configurable cap that rejects/flags oversized input before any paid call, complementing (not conflated with) the money ceiling. + +### `H8-abuse-injection-pre-screen` — **NEVER_CLOSED** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** The owner's cheap abuse/misuse pre-screen (V2 п1 — detect prompt-injection or non-translation misuse WITHOUT burning tokens, high precision) is only partially dispositioned: the trust-boundary + injection probe are queued for Ф3, but the cheap "someone is misusing this to write code" pre-screen has no disposition. +- **Интент ресёрча:** Owner wants a cheap, high-precision, low-false-positive pre-check that detects abusive prompts/attacks or non-translation misuse before spending tokens, especially to avoid burning a whole book's tokens on an attack. + *источник:* START_PROMT.MD:68 (V2 п1); D25.5/D-log:311 (trust boundary in contract; action-security gate = named Ф3 blocker; adversarial injection probe queued before Ф3, "не срочно") +- **Что в коде:** No abuse/injection/misuse detection code exists; the trust-boundary is enforced only by construction (translator/editor roles have no tools) and the injection probe is an unqueued eval item; the misuse-detection (non-translation input) angle is entirely absent. + *evidence:* `No abuse/injection/misuse/jailbreak guard in backend/internal (grep empty); trust boundary is a documented invariant only (D-log:311)` +- **Архитектурный вывод:** A cheap deterministic (or single-cheap-call) input classifier that flags injection/misuse before the paid pipeline runs, dispositioned explicitly rather than folded entirely into a distant Ф3 action-security gate. +- **Уточнение верификатора:** Directionally minor: code_reality calls the injection probe "an unqueued eval item," but D-log:311 actually places it in the eval-queue ("в eval-очередь перед Ф3, не срочно") — i.e. queued-but-not-urgent, not unqueued. This is a peripheral detail (the injection probe is the already-dispositioned half); it does not affect the NEVER_CLOSED disposition, which correctly applies to the cheap non-translation-misuse pre-screen half that has no D-log disposition and no code. + +### `H9-parallelism-speed-owner-main-V4` — **NEVER_CLOSED** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** Owner V4 п1 (declared his "main V4 idea": significantly speed up translation via parallel processing of cut chunks) was never dispositioned; only NAIVE chunk parallelism was rejected by an early verifier caution that predates V4. +- **Интент ресёрча:** Owner V4 п1: after cutting the book, pull the whole batch into models in parallel to get the book N times faster; he flags it as polish-phase but explicitly the main V4 idea and asks how to significantly speed up book translation. + *источник:* START_PROMT.MD:86 (V4 п1); D-log:45 (early caution: "наивный параллелизм чанков" not to build — race on Runner.clients + commit-terms-on-job-boundary); ORCHESTRATOR_ARCH_RESET_PROMPT.md:73 (do not lose V4 п1) +- **Что в коде:** Translation is fully sequential; no goroutine/errgroup/worker-pool over chapters/chunks in the pipeline; the only concurrency note is the rejection of naive parallelism, which does not answer the owner's actual safe-parallel goal. + *evidence:* `No errgroup/worker/goroutine over chunks in backend/internal/pipeline (grep empty); D-log:45 rejects only the naive form` +- **Архитектурный вывод:** A safe parallel-execution design (isolated per-chunk clients, term-commit ordering resolved) dispositioned as a real backlog arm — the owner's stated main V4 idea deserves a decision, not a stale rejection of only its naive form. + +### `H11-backend-render-questions-unanswered` — **NEVER_CLOSED** / MEDIUM · верификатор: PLAUSIBLE (high) +- **Находка:** research/18 §D asks direct questions TO the backend (notably §D-7: can the runner supply neighbor src/tgt as a non-output reference without a new persistent domain?) that were never answered — leaving whether the ±1 reference lever is prod-minor or Ф2 undecided. This is a literal broken-telephone gap: a question to backend never routed back. +- **Интент ресёрча:** research/18 §D-7 explicitly poses a backend render question; §D-1 needs a marked trap-set to choose between the ±1-reference / running-summary / chapter-card levers; §D-8 requires owner decisions (degree of syntactic merging; front-matter in acceptance). + *источник:* research/18 §D-7 (docs/research/18-quality-levers.md:275), §D-1 (:269), §D-8 (:276) +- **Что в коде:** No mechanism to inject a neighbor chunk as non-output reference exists; the runner injects only glossary memory slices; the question of whether this is a small render change vs a persistent-domain build was never answered, so the ±1-reference lever sits undecided. + *evidence:* `backend/internal/pipeline/chunkrun.go:16 (injection surface, glossary/Ф2-roles only); no reference-context render path in render.go` +- **Архитектурный вывод:** A closed loop where research questions addressed to backend get an explicit answer (here: whether neighbor-context reference is a render-only change) so downstream levers can be classified and scheduled rather than left dangling. +- **Уточнение верификатора:** The literal broken-telephone gap is real: §D-7 (research/18:275) is a backend-directed render question ("вопрос бэкенду") whose answer was never routed back / ratified in the D-log, PROGRESS, or backend docs, and no neighbor-context injection path exists in code (render.go RenderVars/injection surface carries only glossary+editor-constraint). But the finding's characterization that this leaves "the ±1 reference lever undecided/dangling" is off on two counts. (1) The lever's EMPIRICAL value WAS tested — exp14 ran the A-ref-prev/next/both arms and found NO lift on this slice ("±1 reference НЕ дал лифта, трапы локально самодостаточны", experiments/14-quality-empirics.md:290; "A-ref дал байт-идентичный выход, чинить нечего", PROGRESS.md:90), so the lever is empirically SHELVED pending next-supported/cataphora cases — not a lever left dangling with no disposition. This makes the render-cost question currently MOOT rather than an open blocker (severity is below MEDIUM). (2) The render answer is largely DERIVABLE from existing code: MessagesWithInjection (render.go:203-219) already accepts an arbitrary post-cache-boundary injection string with NO new persistent domain, and prev/next-source + prev-target are available to the runner in chapter order — so §D-7's answer is most likely "PROD-MINOR" (the surface exists), it was simply never written down. Net: a genuine low-severity un-closed backend question, not an undecided lever blocking a decision. Disposition NEVER_CLOSED fits the §D-7 question itself; the "sits undecided" impact framing overstates it. + +### `H12-cross-boundary-coherence-untested` — **UNTESTED_ASSUMPTION** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** The cross-boundary coherence lever (±1 reference / next-supported / cataphora) is an UNTESTED assumption on the real domain: exp14 could not find a next-supported/cataphora trap cheaply on the early slice, A-ref gave a byte-identical (null) result, and §D-1 says the actual context-locus of the owner's coherence failures was never annotated. +- **Интент ресёрча:** research/18/exp14 intend to locate whether claim-2 coherence failures live in-chunk / prev / next / whole-chapter / world-knowledge, and to test cross-boundary reference; the honest caveat is the trap-set for this was not built on the early PD-classic slice. + *источник:* exp14:259,290 (docs/experiments/14-quality-empirics.md); research/18 §D-1 (:269); exp14:93 (no next-supported/cataphora trap found cheaply) +- **Что в коде:** No cross-boundary reference or coherence handling exists; the single whole-chapter datapoint (P1c) that linked an antecedent is not a built feature; the claim that cross-boundary is 'not dominant' rests on 2 locally-self-contained prev-traps on one text. + *evidence:* `exp14 §2.4/§3 (docs/experiments/14-quality-empirics.md:258-291); no cross-boundary code path in internal/pipeline` +- **Архитектурный вывод:** A marked cross-boundary trap-set on the ACTUAL webnovel domain before deciding the ±1-reference / running-summary / chapter-card levers — the current 'cross-boundary not dominant' conclusion is preliminary and domain-mismatched (PD-classic, not webnovel cut). + +### `H14-D38.4-tail-inversion-guard` — **WORKAROUND_NOT_ROOT** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** D38.4's audit found the polarity/rank/emotion inversion-guard from D37 §2г was never built; the deterministic backstop is absent and the sole near-term defense is editor-model choice (editor-swap) — a model-selection workaround for a missing structural guard. +- **Интент ресёрча:** D37 §2г / D38 §2г: glm inverts polarity (a1 没有一个不知道) independently of the prompt; a deterministic inversion guard was implied but D38.3 only built number/omission (length/digits), which by design cannot catch polarity inversions. + *источник:* D38.4 items 2-4 (docs/architecture/05-decisions-log.md:537-539); PROGRESS.md:58-60 +- **Что в коде:** RegressionGuard checks number-drift only when the draft uses Arabic digits (its own comment admits the word↔word 四成四→'четыре десятых' case is blind); no polarity/rank/emotion inversion detection; inversions are 'closed' only by fidelity re-gate scoring + choosing a different editor model. + *evidence:* `backend/internal/pipeline/regressionguard.go:24-27 (Arabic-digit-only number-drift, self-documented hole); D-log:537-539 confirms inversion-guard not built` +- **Архитектурный вывод:** A real polarity/meaning-inversion signal (span-level, deterministic or cheap-judge) plus canon-glossary number-word coverage — rather than promoting editor-model swap to the ONLY inversion defense, which leaves the root (no inversion detection) live to recur on any new editor. +- **Уточнение верификатора:** Central claim fully holds. One precision note (does not change verdict/disposition): the specific cite regressionguard.go:24-27 is the self-documented NUMBER-WORD recall gap (the D38.4-item-3/b1 hole), not itself an inversion confession — the inversion-guard's absence is proven by the total absence of any polarity/rank/emotion detection in regressionguard.go and across the pipeline. The finding's code_reality states this correctly (number-only guard + no inversion detection); only the line-anchor illustrates the adjacent number-word half rather than the inversion half. + +### `H15-glossary-auto-construction` — **NEVER_CLOSED** / MEDIUM · верификатор: PLAUSIBLE (high) +- **Находка:** The glossary-BUILDING methodology the owner described (model queries that record the context terms appear in, pass genre, and use a web-fetch model for hard terms — V4 п4) is not built; the glossary is a static hand-curated, owner-signed seed, and only the web-fetch sub-part is dispositioned (Ф3). +- **Интент ресёрча:** Owner V4 п4: the model should build the glossary via translation queries that record surrounding context and genre, with web-fetch (e.g. Gemini-lite) for culturally-loaded terms; owner calls the glossary crucial. + *источник:* START_PROMT.MD:89 (V4 п4); D25.5/D-log:372 (web-fetch = Ф3+; near-term = засев глоссария фан-конвенциями — manual) +- **Что в коде:** The glossary is materialized from a hand-authored YAML seed (memseed.go); there is no automated glossary-construction query loop, no context-recording, no genre-conditioned term extraction on the hot path; local extraction (exp06) was a settled bench, not a product loop. + *evidence:* `backend/internal/pipeline/memseed.go (seed materialization); reseed is a manual Python transform (eval/exp14b/exp14b_reseed_promote.py); no auto-glossary-build path in internal/pipeline` +- **Архитектурный вывод:** An automated glossary-construction stage (context-recording term queries, genre-conditioned, with an optional web-fetch tier under the Ф3 trust boundary) — the current static hand-seed is a manual stopgap for the owner's central 'glossary is crucial' mechanism. +- **Уточнение верификатора:** Core fact CONFIRMED: the owner's glossary-building methodology is not built; the glossary is materialized from a hand-curated owner-signed YAML seed + deterministic ruby classification (memseed.go, seeding.go — explicitly "No LLM"), with no auto-construction query loop, no context-recording, no genre-conditioned extraction; reseed is a manual Python transform. Two specifics are off, so PLAUSIBLE not CONFIRMED: (1) It is NOT true that "only the web-fetch sub-part is dispositioned" — the broader glossary autopopulation is a recognized, dispositioned-as-DEFERRED gap (G1 in the memory-risk-registry with an exp06 design "спот=локаль/рендер=облако", called "our main product gap" in the strategic review, and milestone-scoped in the MVP plan; memseed.go/snapshot.go call it "a later/future milestone"). (2) Because the auto-build reached that deferred disposition, the disposition is better read as a deferred/known-gap (workaround-via-manual-seed), not a clean NEVER_CLOSED. What is genuinely never-dispositioned (and where NEVER_CLOSED holds) is V4 п4's DISTINCTIVE levers — query-time context-recording and genre-conditioning of glossary-build queries — which never reached a D-log entry or an experiment: the G1/exp06 design is NER-spot + cloud-render and includes neither, and genre is passed only as a translation brief-var (exp13:39), never to any glossary-build query since none exists. Severity MEDIUM stands. + +### `H18-no-standing-findings-ledger-process` — **NEVER_CLOSED** / MEDIUM · верификатор: CONFIRMED (high) +- **Находка:** The owner's actual concern-5 ask — a STANDING findings-ledger PROCESS that runs across polygon→orchestrator→backend so findings cannot silently die — has no home; D38.4 was a one-off manual audit and nothing institutionalizes the mechanism. +- **Интент ресёрча:** Owner arch-reset concern 5: a mechanism so findings do not get lost, a systematic ledger finding→disposition as a PROCESS (not a one-shot like D38.4), threaded end-to-end through the chain; the near-loss of smart chunking proves the hole. + *источник:* ORCHESTRATOR_ARCH_RESET_PROMPT.md:52-55 (concern 5); D38.4 (one-off audit, docs/architecture/05-decisions-log.md:532) +- **Что в коде:** N/A — no ledger artifact, template, or recurring audit step exists in docs/ or tooling; PROGRESS flag-lists ('Находки для оркестратора', 'Ждём от владельца') are ad-hoc prose that accrete un-closed items (e.g. PROGRESS:7 still lists ~8 open owner decisions) with no forcing function to disposition them. + *evidence:* `docs/PROGRESS.md:7 (long-standing 'Ждём от владельца' list) and :159,241 (ad-hoc 'Находки' blocks); no ledger process document` +- **Архитектурный вывод:** A durable findings-ledger (every polygon finding gets an explicit disposition FAITHFUL/DISTORTED/LOST/etc. and is tracked until closed), owned as a recurring orchestrator step — the process the owner is actually asking for, of which this audit is one instance. +- **Уточнение верификатора:** Substance confirmed; two minor refinements that do not change the verdict: (1) PROGRESS:7 lists ~10 (not ~8) non-struck open owner items. (2) An informal partial channel does exist — the recurring "Находки для оркестратора" section convention plus the CLAUDE.md self-review mandate — and some flagged items are in fact dispositioned (e.g. PROGRESS:160 glm `### Глава` finding → closed by the D38.3 export-fix). But this convention is ad-hoc prose with no per-finding disposition ledger, no tracked-to-closure forcing function, and no institutionalized recurring step, so it is not the standing end-to-end process the owner requests; the near-loss of smart chunking is the owner's own evidence that findings still fall through. Finding stands as CONFIRMED. + +### `H13-broken-draft-joins-sanitizer-hole` — **NEVER_CLOSED** / LOW · верификатор: — +- **Находка:** exp14 found broken draft joins (склейки) that leak even in the frontier output — 'a sanitizer hole even in the frontier' — but this defect class was never dispositioned into the D38.3 sanitizer or the D-log. +- **Интент ресёрча:** exp14 §2Б.7 records broken draft joins (F-disc ch17) as a sanitizer gap present even with a frontier editor. + *источник:* exp14:372 (docs/experiments/14-quality-empirics.md, 'Битые склейки черновика … дыра санитайзера даже во фронтире') +- **Что в коде:** The D38.3 sanitizer handles CJK-leak, markdown-###, cosmetic strip, and number-drift observation, but has no join/boundary-artifact detector; broken joins are unaddressed. + *evidence:* `backend/internal/pipeline/sanitizer.go (CJK/cosmetic classes only); D-log D38.3 (:6) lists the built classes — joins absent` +- **Архитектурный вывод:** A boundary/join-artifact check in the sanitizer or a chunk-stitch validation, so broken joins are at least flagged rather than silently shipped. + +### `H10-user-initiated-stop` — **NEVER_CLOSED** / LOW · верификатор: — +- **Находка:** Owner V2 п2 asked for both stop AND resume of the pipeline (user command or internal signal), with abusive prompts non-resumable; resume and ceiling/outage pause are built, but a user-initiated stop/cancel path has no code and no disposition. +- **Интент ресёрча:** Owner V2 п2: provide pipeline stop AND continue, triggered by user command or an internal pipeline signal. + *источник:* START_PROMT.MD:69 (V2 п2); D4/D-log:108,149 (provider-outage durable pause/resume built) +- **Что в коде:** Resume exists (resume.go) and ceiling/outage cause durable pause; there is no user-facing stop/cancel command in tmctl and no internal-signal-driven stop beyond ceiling/outage. + *evidence:* `backend/internal/pipeline/resume.go (resume); escalation.go:42 / stagerun.go:390 (ceiling/outage pause only); no stop/cancel command in cmd/tmctl` +- **Архитектурный вывод:** A user-initiated stop/cancel (and a non-resumable flag for abuse-flagged jobs) — likely a small tmctl + runner addition — dispositioned rather than assumed covered by the money-ceiling pause. + +### `H16-self-improvement-harness` — **NEVER_CLOSED** / LOW · верификатор: — +- **Находка:** The owner's self-improvement idea (V1/V3 п2: how to self-improve on huge books without expensive human review) has no built mechanism; its frontier-meta-reviewer sub-proposal was killed, but the underlying self-improvement/A-B loop was never dispositioned. +- **Интент ресёрча:** Owner V3 п2 / V1: self-improve at scale because human editorial review of huge books is expensive and error-prone; V1 also floats building the backend to support A/B tests eventually. + *источник:* START_PROMT.MD:75 (V3 п2), :64 (V1 A/B tests); D25 pilot arms are eval-only, not a product self-improvement loop +- **Что в коде:** No self-improvement or A/B harness in the product; A/B exists only as polygon experiments and planned pilot arms; there is no closed-loop 'extract fixes from run reports' mechanism (the frontier-meta-reviewer that would have done this was killed by the owner). + *evidence:* `No A/B or self-improvement path in backend/internal; D25.6/D-log:256 pilot arms are eval-scoped; owner killed frontier-meta-reviewer (START_PROMT.MD:87)` +- **Архитектурный вывод:** A decision on whether/how the product self-improves (e.g. structured per-run defect reports feeding prompt/glossary iteration) — the owner's concern survives the death of its meta-reviewer sub-idea and deserves its own disposition. + +### `H17-frontier-meta-reviewer-closed` — **FAITHFUL** / LOW · верификатор: — +- **Находка:** The frontier meta-reviewer / 'judge over the pipeline' (V3 п3) was correctly CLOSED — the owner killed it in V4 п2 — and it survives only as a diagnostic flagship in exp12/13, not as a product feature. Recorded for completeness; not an open hole. +- **Интент ресёрча:** Owner V3 п3 proposed a frontier judge-over-pipeline emitting a report of fixes; owner V4 п2 then judged it pointless and killed it. + *источник:* START_PROMT.MD:76 (V3 п3), :87 (V4 п2 kill); exp12/13 flagship use (PROGRESS.md:178) +- **Что в коде:** Not built as a product feature (correct per owner kill); appears only as a self-family-exclusion guard concept in experiment prompts and as one-off flagship reads. + *evidence:* `D-log:453-454 (frontier-meta-reviewer only as eval self-family-exclusion guard); no product code` +- **Архитектурный вывод:** None required — the closure is faithful to the owner's decision; included only to confirm the sweep checked it and to prevent it being silently re-litigated. + +### `H19-owner-decisions-pending-D8` — **NEVER_CLOSED** / LOW · верификатор: — +- **Находка:** research/18 §D-8 owner decisions — the permissible degree of syntactic merging/splitting (which collides with the completeness invariant) and whether front-matter counts in artistic acceptance — were never made, and the merge↔completeness collision (§D-4) is un-measured. +- **Интент ресёрча:** research/18 §D-4/§D-8: the merge/explicitation vs completeness-invariant collision needs a semantic-atom-coverage omission-guard measurement, and two owner decisions gate acceptance criteria. + *источник:* research/18 §D-4 (docs/research/18-quality-levers.md:272), §D-8 (:276) +- **Что в коде:** The editor prompt hard-forbids adding/removing sentences (completeness invariant) while also mandating paragraph merging — the exact collision §D-4 flags — with no omission-guard measurement of it; the front-matter acceptance question is undecided. + *evidence:* `backend/prompts/editor.md:8,9 (merge paragraphs AND preserve completeness — the collision); D-log has no D-8 disposition` +- **Архитектурный вывод:** An owner decision on allowable merge/split latitude plus a semantic-atom omission guard that measures whether reflow-merging violates completeness — so the two mandates in editor.md stop implicitly conflicting. + +--- +## L8-code-cleanliness — Концерн 3 п4 — чистота/расширяемость кода + +**Итог дорожки:** The recent readability-infra landings (D30.3 sanitizer, D35.4a cosmetic-strip export, D38.3 regression guard, D38.5 discourse editor) are functionally coherent but architecturally "кое как": the cosmetic-strip/export feature is smeared across five files with a cross-function invariant re-derived twice; the whole readability-gate layer (sanitizer + cheap gates) hardcodes target=Russian with no TargetLang seam while coverage/classify are pair-aware; roles and gates are wired by hand-edited switches and bespoke struct fields rather than a registry; and prompt sprawl (three orphaned prompt files, one duplicating editor.md) plus a dead/misleading gatesEnabled() helper accrete tactical shapes. The extension the project is about to attempt (Ф2 annotator/voice role, editor-swap, per-pair) lands squarely on these seams. Two seams (config-driven stage list, baseRequestLog constructor) are genuinely clean and worth keeping as the model. + +### `L8-sanitizer-strip-smeared-5-files` — **CODE_SMELL** / HIGH · верификатор: — +- **Находка:** The D35.4a/D38.3 'strip cosmetic leak and export flagged' feature is one logical concept scattered across five files with coordinated special-case branches in each. +- **Интент ресёрча:** D38.3 landed disposition-by-class so a chunk with a strippable cosmetic leak (### header, CJK-leak) is cleaned and exported flagged instead of dropped to an empty placeholder — a single 'recoverable flag' behaviour. + *источник:* docs/architecture/05-decisions-log.md D38.3 (line 527, п2а) + D35.4a +- **Что в коде:** The one behaviour is spread: sanitizer.go owns sanitizeOutput/cosmeticOnly/stripCosmetic; chunkrun.go classifyOutput has the isFinal sanitizer branch and returns FlagSanitizerStripped; translateChunk threads a `recovered` string; stagerun.go runStage has a dedicated FlagSanitizerStripped arm (recovered = stripCosmetic + commitSanitizedExport) AND a second copy inside the escalation branch; commitSanitizedExport writes a namespaced derived checkpoint; bookrun.go's StageResult.RecoveredText/ChunkOutcome.FinalText carry it; disposition.go defines two parallel flag reasons. Adding a second recoverable class means editing all five. + *evidence:* `internal/pipeline/chunkrun.go:41-56,92,148-152,207-216; internal/pipeline/stagerun.go:154-160,170-183,223-239; internal/pipeline/disposition.go:92-102; internal/pipeline/bookrun.go:25-30,49` +- **Архитектурный вывод:** A single resolver that maps a classification to (disposition, exportText) and owns the strip+commit — or a RecoverableFlag strategy on FlagReason — so a new recoverable class is added in one place; runStage/escalation/resume would consume the resolved export text instead of each re-branching on FlagSanitizerStripped. + +### `L8-readability-gates-target-blind` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** The entire readability-gate layer (output sanitizer + 4 cheap gates) hardcodes target=Russian and is never passed TargetLang, so it fires unconditionally on the final text regardless of the book's target language — the opposite of the pair-aware coverage/classify gates right next to it. +- **Интент ресёрча:** Coverage was designed pair-parametric (len_ratio_bounds keyed 'zh-ru'/'ja-ru'/'en-ru') and classify() branches on isCJKTarget — the intended pattern is 'gate logic is a function of the language pair.' + *источник:* internal/config/pipeline.go:144-149 (CoverageGate LenRatio map) + disposition.go:294 isCJKTarget +- **Что в коде:** classifyOutput passes r.Book.TargetLang into classify() (chunkrun.go:28) and coverageCheck (chunkrun.go:69), but sanitizeOutput(output) (chunkrun.go:42) and runCheapGates(...) (chunkrun.go:188) receive NO target: sanitizer bakes in [а-яё] preamble patterns, Cyrillic sign-bigram broken-word detection, and a 'Latin insertion' class; cheap gates bake in Russian magnitude words (тысяч/миллион) and the Russian em-dash dialogue convention. Grep confirms TargetLang is absent from sanitizer.go and cheapgates.go entirely. + *evidence:* `internal/pipeline/chunkrun.go:42,188; internal/pipeline/sanitizer.go:171-181,290-301,318-324; internal/pipeline/cheapgates.go:77,564-573` +- **Архитектурный вывод:** Currently valid only because all targets are ru; but 'add a new language pair' (any non-ru target) is code surgery — a →en final would false-flag as Latin insertion. Gate the readability set behind TargetLang (a per-target gate registry, or an early no-op when target!=ru), mirroring the coverage/classify seam that already exists. + +### `L8-role-injection-hardcoded-switch` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** Role→glossary-injection is a hand-edited switch, and adding the Ф2 annotator/voice role the project is planning requires editing that switch plus writing a bespoke renderXBlock — code surgery, not config/registry. +- **Интент ресёрча:** README and chunkrun.go header state Ф2 roles (annotator, voice-layer) are the intended near-term extension and 'add here.' + *источник:* backend/README.md line 14 ('annotator/voice-инъекция → chunkrun.go') + docs/architecture/05-decisions-log.md D38.4 (editor-swap / new roles) +- **Что в коде:** translateChunk selects the per-role injection with `switch st.Role { case roleTranslator: ...; case roleEditor: ... }` and each arm points at a distinct hand-written renderer (renderGlossaryBlock vs renderEditorConstraintBlock, which differ in src→dst vs CONFIRMED-only dst). There is no role→renderer table; a new consumer role means a new switch arm, a new render func, and re-encoding the CONFIRMED-only knowledge by hand. + *evidence:* `internal/pipeline/chunkrun.go:134-140; internal/pipeline/memory.go:451,485` +- **Архитектурный вывод:** A map[role]injectionRenderer (or an injection-strategy per role resolved from config) so a new role is data + one renderer, not a switch edit — the seam the README claims exists is really an edit-the-switch seam. + +### `L8-stripcosmetic-recomputed-twice` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** classifyOutput validates the cosmetic strip by running stripCosmetic + a full second sanitizeOutput pass, then throws the stripped text away and returns only the flag; runStage/resume recompute stripCosmetic on the same text to actually produce the export — a cross-function invariant maintained by re-derivation. +- **Интент ресёрча:** D38.3 guarantees the stripped remainder is non-empty and clean before tagging sanitizer_stripped, and that fresh and resume derive the identical export. + *источник:* docs/architecture/05-decisions-log.md D38.3 (line 527) + sanitizer.go:418-423 (stripCosmetic doc) +- **Что в коде:** classifyOutput (chunkrun.go:50) computes `stripCosmetic(output)` and `sanitizeOutput(stripped).total()==0` purely to decide the flag, discarding `stripped`. The real export is then recomputed: runStage does `recovered = stripCosmetic(last.text)` (stagerun.go:177) and resume reads a derived checkpoint. So sanitizeOutput runs up to 3x and stripCosmetic up to 2x per final chunk, and the 'non-empty & clean' guarantee is asserted in one function but relied on in another; a change to stripCosmetic can silently break the guarantee the classifier claims. + *evidence:* `internal/pipeline/chunkrun.go:49-53; internal/pipeline/stagerun.go:170-183` +- **Архитектурный вывод:** classifyOutput should return the resolved export text alongside the verdict (compute once), threaded to commitSanitizedExport — eliminating the double strip and the two-place invariant. + +### `L8-prompt-sprawl-orphans` — **CODE_SMELL** / MEDIUM · верификатор: — +- **Находка:** Three of six prompt files are orphaned and one duplicates editor.md's instruction blocks; none of the orphans is snapshot-pinned or golden-tested, so they silently drift. +- **Интент ресёрча:** D30.1 flipped the editor mono→bilingual; editor-mono.md was meant to be kept only as a reference variant, and analyst/terminologist are Ф2 skeletons. + *источник:* docs/architecture/05-decisions-log.md D30.1 (mono→bilingual editor) + configs/pipeline-c1.yaml:61 ('Моно-вариант — prompts/editor-mono.md') +- **Что в коде:** No active stage `prompt:` references editor-mono.md, analyst.md, or terminologist.md — editor-mono.md appears only in c1/c2 YAML comments; analyst.md and terminologist.md have zero references anywhere in code or config. editor-mono.md re-states editor.md's style/вёрстка/glossary bullets almost verbatim; because it is unwired it is never SHA-pinned into a snapshot nor covered by golden, so it rots independently of the live editor.md. + *evidence:* `backend/prompts/editor-mono.md:4-11 (dup of editor.md:5-8); backend/prompts/analyst.md; backend/prompts/terminologist.md (both unreferenced)` +- **Архитектурный вывод:** Delete orphans (or move to docs/reference); if editor-mono is a live fallback, wire it as a config-selectable prompt so it is snapshot/golden-covered; factor the shared editor instruction block into one included fragment rather than copy-paste across editor.md/editor-mono.md. + +### `L8-gatesEnabled-dead-and-lying` — **CODE_SMELL** / LOW · верификатор: — +- **Находка:** Pipeline.gatesEnabled() is dead code that also lies: zero callers, and it reports 'any QA-gate on' by checking only Coverage while Sanitizer/Glossary/RegressionGuard gates now exist. +- **Интент ресёрча:** The helper predates the D30.3/D38.3 gate additions and was meant to answer 'is any QA gate reachable' for escalation. + *источник:* internal/config/pipeline.go:164-168 (its own doc comment) +- **Что в коде:** gatesEnabled() returns only p.Gates.Coverage.Enabled; grep finds no caller. Its doc says 'reports whether any QA-gate is on' — false since three more gate structs (Sanitizer, Glossary, RegressionGuard) were added to Gates without updating it. + *evidence:* `internal/config/pipeline.go:164-168; grep: 0 callers of gatesEnabled()` +- **Архитектурный вывод:** Delete it, or if a 'any gate on' predicate is genuinely needed make it enumerate every Gates field — a stale predicate that silently under-reports is a trap for the next gate author. + +### `L8-regressionguard-bolted-onto-cheapgates` — **CODE_SMELL** / LOW · верификатор: — +- **Находка:** The D38.3 reflow regression guard (a distinct draft→final concept) was appended as extra fields on cheapGateResult and gated by a bool inside cheapGateConfig, and its on/off toggle deliberately escapes the snapshot discipline that Sanitizer/Coverage/StyleCheck all follow. +- **Интент ресёрча:** D38.3 added an opt-in observability flagger (length collapse + number drift) over the reflow transform, never a disposition change. + *источник:* docs/architecture/05-decisions-log.md D38.3 (line 527, п2в: 'не фолдится в снапшот (тумблер без resnapshot)') +- **Что в коде:** regressionguard.go is its own file/concept but its output is jammed into cheapGateResult.LengthCollapse/NumberDrift (cheapgates.go:65-66) and folded via runCheapGates behind cfg.regressionEnabled (cheapgates.go:90-95). Sanitizer/Coverage fold a version into snapshot.go and StyleCheckVersion folds cheapGateVersion, but Gates.RegressionGuard.Enabled is intentionally NOT folded — so flipping it changes the recorded n_style_flags without a --resnapshot, an unprincipled exception to the snapshot rule the sibling gates obey. + *evidence:* `internal/pipeline/cheapgates.go:46-72,90-95; internal/pipeline/regressionguard.go:30-40; internal/pipeline/snapshot.go:202-207,232` +- **Архитектурный вывод:** Model observability flaggers as a named-signal list with their own detail/version, and treat the toggle like the other gates (fold it, since it changes reported output) — or explicitly document why observability toggles are snapshot-exempt as a rule, not a one-off. + +### `L8-snapshot-fold-copypaste-tripwire` — **CODE_SMELL** / LOW · верификатор: — +- **Находка:** The snapshot model-wire fold is knowingly duplicated (primary model + escalate_to) with an in-code TRIPWIRE saying 'extract foldModelWire() before adding a third model axis' — exactly the Ф2 channel-B/annotator extension now on the roadmap. +- **Интент ресёрча:** Snapshot must fold every wire-affecting input (provider triple, extra_body, capability) for both the primary and the escalation model so a change is a loud --resnapshot. + *источник:* backend/README.md invariant #2 (snapshot discipline) + docs/architecture/05-decisions-log.md D38.4 (new model axes for editor-swap) +- **Что в коде:** snapshot.go folds the prov-triple/extra/capability for st.Model, then mirrors the identical block for st.EscalateTo; a comment (snapshot.go:268-272) admits it is copy-paste 'below the extraction threshold of 3' and warns the next author to extract foldModelWire() first. The extension that trips it (channel B / annotator adding a third model axis) is the project's declared next step, so the deferred refactor is scheduled to bite. + *evidence:* `internal/pipeline/snapshot.go:235-280` +- **Архитектурный вывод:** Extract foldModelWire(model) now (it is 2x already and the 3rd axis is imminent), so a new model-bearing role can't silently forget a fold and produce a false snapshot hit/miss on escalated checkpoints. + +### `L8-config-stage-list-clean-seam` — **FAITHFUL** / LOW · верификатор: — +- **Находка:** The config→stage-plan seam is genuinely clean and extensible: stages/roles/models/prompts/gates/escalation are data in pipeline-*.yaml with a fail-fast validator, and the 'config vs runner' boundary is stated and enforced. +- **Интент ресёрча:** Р2 requires a hard config-vs-code boundary: config defines stage composition/order/role→model/prompt versions/gate thresholds; runner owns loops/branching/escalation mechanics. + *источник:* internal/config/pipeline.go:12-17 (boundary doc) + configs/pipeline-c1.yaml:1-8 +- **Что в коде:** Stage is a typed struct list; LoadPipeline validates names/roles/models/prompt paths/reasoning/channel/escalate_to structurally and CheckRunnable/CheckAdultChannel reject un-runnable or unsafe configs before any store side effect. Adding a stage or changing a model is a YAML edit, not a code change — the seam works as intended. + *evidence:* `internal/config/pipeline.go:58-93,176-215,260-334` +- **Архитектурный вывод:** This is the model to copy for the smells above: the role→injection switch, the readability-gate target hardcoding, and the gate wiring should be pulled toward this same config/registry-driven shape. + +### `L8-baseRequestLog-constructor-clean` — **FAITHFUL** / LOW · верификатор: — +- **Находка:** The baseRequestLog constructor and the setJobStatus/releaseReservation helpers are a clean, extension-safe seam that a new call path (annotator stage, channel-B hop) inherits for free. +- **Интент ресёрча:** Every request_log row must carry the join keys (book/chapter/chunk/stage/role/model/hash) so telemetry ties to spend and chunk_status, and money-state updates must never be silently swallowed. + *источник:* backend/README.md invariant #1 (money) + stagerun.go:465-500 doc +- **Что в коде:** baseRequestLog centralizes the mandatory attribution prefix so a new call path cannot forget a key and emit an unjoinable row, while deliberately leaving outcome-specific tails at each site; setJobStatus/releaseReservation wrap the 6 prior `_ =`-swallowed error sites into one warning/error path. This is the right level of abstraction — shared invariant centralized, real asymmetries left local. + *evidence:* `internal/pipeline/stagerun.go:465-500` +- **Архитектурный вывод:** Keep this pattern; the sanitizer-export smear (L8-sanitizer-strip-smeared-5-files) is the inverse anti-pattern and should be refactored toward this constructor-owns-the-invariant shape. diff --git a/docs/architecture/09-target-architecture.md b/docs/architecture/09-target-architecture.md new file mode 100644 index 0000000..f9a6d68 --- /dev/null +++ b/docs/architecture/09-target-architecture.md @@ -0,0 +1,184 @@ +# Целевая архитектура и дорожная карта (арх-ресет 2026-07-13) + +> **Статус:** проектный документ оркестратора, ратифицирован **D39** (`05-decisions-log.md`). Рождён из арх-ресета +> (`ORCHESTRATOR_ARCH_RESET_PROMPT.md`, 5 концернов владельца) и синк-аудита «сломанного телефона» +> (`08-sync-audit-ledger.md`, 65 находок, адверсариально верифицированы). Решения владельца по подходу: +> **фазовый** (build-now ∥ ресёрч, пере-прогон после) + **слой пар языков = сеам сейчас, контент только zh→ru**. +> +> Это карта «куда и почему», а не пошаговая спека. Спеки — в хендофф-промтах сессий; контракт — в D-логе. + +--- + +## 0. Стратегический вывод аудита (почему ресет) + +Потолок художественности сейчас — **в НАШЕЙ организации пайплайна, а не в моделях.** Эмпирика (exp12–14b) и +код-аудит сошлись: дефекты, которые владелец назвал «крайне слабо художественно» (рубленые абзацы + «непонимание +связей»), рождаются из **нарезки, структуры редакторских проходов, контракта переводчик/редактор и отсутствия +меж-чанковой когезии** — всё это архитектура, не выбор модели. Это подтверждает трек «дешёвые модели сейчас» +(D30.7) и переопределяет источник следующего прироста качества: **он архитектурный.** + +Ключевое свойство проблемы: **рычаги сцеплены.** Редактор репараграфизует ВНУТРИ ~1500-токенного чанка → он +физически не может чинить когезию через границу чанка → нарезка (концерн 2) жёстко ограничивает потолок +дискурс-переверстки, которую мы требуем от редактора (концерн 1). Поэтому чинить их порознь = снова получить +локальный оптимум на неверно нарезанном тексте (ровно «слабый тест» exp14). **Концерны 1, 2, 4 проектируются +вместе.** + +--- + +## 1. Текущая архитектура (как есть, grounded) + +Пайплайн = линейный проход по плоскому списку чанков; на каждый чанк — стадии по порядку из `pipeline-*.yaml`: + +``` +книга → ingest (главы) → SplitChunks (чанкер) → [ по чанку: translator(draft) → editor(edit) ] → export + ↑ память v2 (Select→inject) + детерм. гейты +``` + +- **Нарезка** (`chunker.go`): жадная упаковка целых абзацев в бюджет `targetChunkTokens=1500` (**const**, в единице + «оценка символов исходника», не выходные токены), спуск до предложений при оверсайзе. Граница — там, где накопление + пересекло бюджет; **логической/сценовой границы нет**. Единица одна и для черновика, и для редактуры. +- **Переводчик** (`translator.md`, deepseek-v4-flash): подстрочник + слабая однострочка про вёрстку; на верности + валидирован (exp13), на художественность/структуру — **не мерен никогда** (D30.6). +- **Редактор** (`editor.md`, glm-5, БИЛИНГВ): ОДНА полная перегенерация чанка с **4 мандатами сразу** — верность, + стиль, глоссарий, дискурс-reflow (техники Ван Цуй + few-shot). +- **Память v2** (`memory.go`): Aho-Corasick матч → disposition → инъекция per-role (переводчику src→dst, редактору + CONFIRMED-dst-констрейнты) → детерм. post-check. Единственная меж-чанковая память. Меж-чанковой когезии + (summary/carryover) НЕТ — есть мёртвые config-заглушки `STMDepth`/`OverlapTokens`, которые код не читает. +- **Гейты** (детерминированные): intrinsic-классификатор (echo/empty/refusal/length), coverage/excision (только + translator), output-санитайзер (только final), glossary post-check (флаггер), cheap style-флаггеры, regression-guard. + **Семантического/качественного сигнала в контуре нет.** +- **Промты**: единые zh→ru-файлы, выбор по фикс-пути из yaml; `{{source_lang}}` — поверхностный шаблон; правила + захардкожены; язык промпта всегда русский. `Book.LangPair()` — мёртвый код. + +Хорошее, что подтвердил аудит (не ломать): деньги/durability (reserve→settle+checkpoint, committed==SUM, kill-9), +snapshot-дисциплина, эталонные сеамы `config→stage-list` и `baseRequestLog`, дословный перенос техник Ван Цуй, +структурные (не «правильно-переводные») few-shot, осознанность разделения переводчик/редактор. + +--- + +## 2. Целевая архитектура — 7 слоёв с чистыми сеамами + +Вместо 65 заплаток — слои, каждый первоклассный, независимо конфигурируемый и снапшот-folded. Ровно они закрывают +5 концернов. + +### Слой 1 — Сегментация (чанкер-модуль) · концерны 1+2 +- Размер бюджета — **снапшот-folded config-кноб в ВЫХОДНЫХ токенах** (не const, не в símволах исходника), чтобы + пре-регистрированная кривая полигона могла приземлиться как настройка. Механизм fold уже есть (`coverageSnap`). +- **Стратегия границы** — first-class ось: `greedy` (как сейчас) | `logical` (сцена/диалог/тема) — с фоллбек-лестницей + и инвариантом «никогда не резать предложение». Наполнение стратегии `logical` — по интернет-best-practices + ресёрч. +- **Декаплинг edit-единицы от draft-единицы** (бывший фантомный «D35.7»): черновик — мелкий чанк (COGS/coverage/ + выравнивание); reflow/edit — крупная дискурс-единица (глава/сцена), где антецеденты со-встречаются. +- **Меж-единичная когезия**: оживить `STMDepth`/`OverlapTokens` — running bilingual summary + carryover соседних + единиц как read-only контекст. (⚠ exp14 «крупнее лучше» и research/18 §A «мелкий чанк + carryover» **конфликтуют** — + решает ресёрч-матрица, см. §4.) + +### Слой 2 — Конвенции пары языков (pair-keyed pack) · концерн 4 +- Версионируемый пакет, ключ `src→tgt[×genre]`, снапшот-folded (правка = громкий `--resnapshot`, как `chunkerVersion`). +- Держит: дискурс-правила, few-shot (в языке пары), транскрипцию/хонорифики, **и язык самого промпта** (свойство + пакета — не хардкод русского). +- Стадия резолвит промпт/правила из слоя по паре книги, а не из фикс-пути. `Book.LangPair()` оживает как ось выбора. +- **Высота сейчас (решение владельца):** построить СЕАМ, наполнить только zh→ru; другие пары — когда придёт книга. + Убирает латентный баг (ja-книга едет через «китайский паратаксис»-промпт) и разблокирует расширяемость дёшево. + +### Слой 3 — Контракт переводчик/редактор (роли = узкие проходы) · концерн 1 +- research/07 предписывает **узкие мандаты + диффы, а не полную перегенерацию** — сейчас ровно наоборот (4 мандата, + full-regen), что и породило хвост инверсий. +- Целевой контракт: переводчик отдаёт **структурный ~верный черновик**; редактор — узкая правка смысла+стиля, + в идеале **diff-based** (search/replace-спаны + детерм. Go-apply → ограниченный blast-radius, дешевле, аудируемо, + структурно omission-safe). Между проходами — гейт. +- **⚠ Арм «переводчик отдаёт структуру» никогда не тестировался** — решает ресёрч-матрица (§4), а не догадка. + +### Слой 4 — Память (код-корень) · концерн 3 +- Починить **disposition-слепой** `suppressContained`: более длинный НИЗКО-доверенный ключ (draft/ambiguous) не должен + подавлять вложенный ВЫСОКО-доверенный (approved/confirmed). Это закрывает терм-дрейф в КОДЕ (сейчас — сид-воркэраунд). +- **Тихие дропы → в громкие**: подавленный approved-термин обязан попадать в retrieval-state (обещание research/13). +- Убрать устаревшую D1-моно-рационализацию из `renderEditorConstraintBlock` (редактор теперь БИЛИНГВ, D30.1); открытый + вопрос — давать ли билингв-редактору src→dst-мэппинг, а не голые dst-формы. + +### Слой 5 — Измерение качества в контуре · главная цель + концерн 5 +- Per-run/per-segment авто-сигнал: **детерминированный структурный KPI** (предл./абзац, диалог-тире, CJK-утечка, + глоссарий-консистентность) — дёшево, build-now — + **дешёвая span-проверка смысла** (полярность/число/ранг/пропуск → + тот самый inversion/omission-бэкстоп; семантический span-судья — ресёрч-зависим, позже). +- Даёт то, что владелец просил (телеметрия качества с 1-го дня, п.25/п.31) и делает КАЖДУЮ правку ниже измеримой; + предпосылка самоулучшения. + +### Слой 6 — Export-contract (нормализация) · концерн 3 +- Единая нормализация КАЖДОГО финального текста (NFKC-fold + пробелы/CJK-пунктуация/U+3000) — тот самый D29.1a-слой, + который D-лог откладывает. Санитайзер сводится к **чистому детектору** (флагует только по-настоящему невосстановимое). +- Убирает: дефект чинится на флагнутом чанке, но уезжает сырым на чистом; дубль NFKC; churn регексов v1→v4. + +### Слой 7 — Гигиена расширяемости · концерн 3 п4 +- **Реестр** `role→инъекция` (map/стратегия из конфига) вместо рукоправимого switch — чтобы Ф2-annotator/voice + добавлялся данными, не хирургией. +- **Target-aware гейты**: readability-набор за `TargetLang` (сейчас слеп — любой не-ru target даст ложные флаги). +- Де-размазать strip-and-export (5 файлов → один resolver `classification→(disposition, exportText)`). +- Снести сироты-промпты (`analyst.md`/`terminologist.md`/дубль `editor-mono.md`), мёртвый `gatesEnabled()`; + вынести `foldModelWire()` (tripwire на 3-ю модель-ось уже стоит). +- Эталоны для копирования: `config→stage-list`, `baseRequestLog`. + +--- + +## 3. Карта находок → слои (полный ледджер — `08-sync-audit-ledger.md`) + +| Слой | Несущие находки (HIGH, CONFIRMED) | +|---|---| +| 1 Сегментация | `L2-1500-untested-apriori`, `L2-budget-wrong-unit`, `L2-token-budget-not-semantic`, `L2-edit-unit-equals-chunk`, `L2-cross-chunk-only-glossary`, `L2-d357-phantom-decision`, `H1-smart-chunking-never-formalized`, `L2-no-boundary-algo-research` | +| 2 Слой пар | `L4-prompts-not-per-pair-zh-baked`, `L4-convention-pack-layer-never-built`, `L4-prompt-language-always-russian`, `H2`, `L4-adding-a-pair-is-a-rewrite-not-config` | +| 3 Контракт t/e | `L1-editor-four-mandates-one-fullregen-pass`, `L1-fidelity-has-no-owner-in-cheap-path`, `L1-diff-editing-never-built`, `H3/H4`, `L1-translator-structure-arm-never-tested` | +| 4 Память | `L3-seed-workaround-not-code-root`, `L3-recurrence-draft-longer-eats-approved-shorter`, `L3-suppressor-disposition-blind-right-fix`, `L3-editor-block-stale-monolingual-D1` | +| 5 Качество | `H5-no-in-loop-quality-signal`, `H14-inversion-guard` (D37 §2г), `H12-cross-boundary-coherence-untested` | +| 6 Export | `L5-normalization-fused-to-flag-path`, `L5-broken-word-class-*`, `L5-diacritics-class-lost`, `L5-legit-cjk-gloss-stripped`, `L6-technique-not-norm-overclaim` | +| 7 Гигиена | `L8-*` (target-blind gates, role-switch, strip-smear, prompt-orphans, dead gatesEnabled, foldModelWire) | +| Процесс (к.5) | `H18-no-standing-findings-ledger-process` + незакрытые владельческие: `H6` running-summary, `H7` input-size-limit, `H8` abuse-prescreen, `H9` parallelism/speed (V4-п1), `H15` glossary-auto-build (V4-п4), `H10` user-stop, `H16` self-improvement | + +--- + +## 4. Дорожная карта — две параллельные фазы (решение владельца: фазово) + +### Трек A — Build-now (бэкенд, ресёрч НЕ нужен, высокий ROI, низкий риск) +Порядок ≈ по ROI/риску; всё под мандат самопроверки исполнением (CLAUDE.md): +1. **Слой 4 память-корень** — disposition-gate на suppressor + громкий лог дропа. Мал, хирургичен, закрывает терм-дрейф + в коде (снимает вечную зависимость от сид-воркэраунда). +2. **Слой 5 (детерм. часть)** — per-run структурный quality-report + **inversion/omission-бэкстоп** (полярность/число/ + ранг/пропуск, span). Измеряет всё остальное. +3. **Слой 6 export-contract** — нормализация каждого финала; санитайзер → детектор. +4. **Слой 7 гигиена** — реестр role→инъекция, target-aware гейты, де-размазка, снос сирот. (Разблокирует слой 2 сеам.) +5. **Слой 2 сеам** (после гигиены) — pair-keyed резолв промпта/правил; наполнить zh→ru; язык промпта = свойство пакета. +6. **Док-долг оркестратора:** формализовать/вычистить фантомный «D35.7»; ратифицировать диспозиции незакрытых + владельческих идей (input-limit / abuse-prescreen / parallelism / glossary-auto-build — явное решение, даже если «Ф2/Ф3»). + +### Трек B — Ресёрч ПЕРЕД постройкой (концерны 1+2+4 — сцеплены) +- **Единая полигон-матрица:** `{граница: greedy|logical} × {edit-единица: чанк|глава} × {когезия: нет|summary/carryover} + × {где reflow: переводчик|редактор|отдельный-пасс}` на реальном вебновелл-срезе. ⚠ **exp14 vs research/18 §A прямо + конфликтуют** по размеру чанка — это ключевой вопрос, который матрица решает эмпирически, не догадкой. +- **Полигон-синк (13.07):** exp14 «пробелы» по сцен-границе/кросс-границе/translate-стадии — не отрицательные + результаты, а **непротестированные допущения на неподходящем домене** (ранняя арка 蛊真人 плоская: 1 предл./строка, + мало сцен-маркеров). → **Домен-пригодность = первый шаг эмпирики:** характеризовать структуру целевой книги; если + плоская — логическая/сцен-граница малополезна ДЛЯ НЕЁ (доминанта = внутри-чанк верность + размер edit-единицы), + и это сама по себе находка. Трапы Q1/Q3/Q4 — на срезе С сцен-структурой + кросс-граничными зависимостями. Реюз: + кривая + A-ref-prev/both как есть; дозаказать `A-ref-next` + translate-плечо. +- **Интернет-best-practices** по логической/семантической нарезке крупных текстов (владелец просил явно) + фоллбеки. +- **Дизайн слоя 2** (форма pair-pack, язык промпта) — под данные матрицы (§D-3 research/18: форма не разрешена). +- Методика-guard D32.4 (fidelity-first, leave-one-out, катастроф-скрин — полигон трижды собирал ложную сходимость). + +### Гейт пере-прогона +Пере-прогон 3–10 глав — **только после** трека B (иначе снова слабый тест на голом конфиге) и приёмки ключевых +build-now (память-корень + quality-сигнал + export-contract). Планка — 2 претензии владельца (интерим, не пилот Ф2.5). + +--- + +## 5. Процесс против «сломанного телефона» (концерн 5) + +Постоянный **findings-ledger** как повторяющийся шаг оркестратора (этот аудит — первый экземпляр, `08-*`): +- Каждая находка полигона/ресёрча/аудита получает явную диспозицию (FAITHFUL/DISTORTED/LOST/WORKAROUND_NOT_ROOT/ + UNTESTED_ASSUMPTION/NEVER_CLOSED/CODE_SMELL) и трекается **до закрытия** (fixed / ratified-deferred / retracted). +- Вопросы «к бэкенду» из ресёрча (research/18 §D) обязаны получать ответ, маршрутизируемый обратно (сейчас висят). +- Периодический completeness-critic прогон (как этот) на границе фаз — не разовый. + +--- + +## 6. Что НЕ демонтируется + +Столпы (деньги/durability, snapshot-дисциплина, детерминированный банк памяти как ставка, echo-гейт, 18+ линия, +дешёвый трек, конфиг-первое ядро) — целы. Целевая архитектура — это **декомпозиция и до-стройка** существующей +машинерии по чистым сеамам, а не переписывание. Фронтир-тир остаётся «потом» (новый yaml, не переделка кода).