diff --git a/docs/experiments/15-segmentation-empirics.md b/docs/experiments/15-segmentation-empirics.md index 532ae11..7b4e667 100644 --- a/docs/experiments/15-segmentation-empirics.md +++ b/docs/experiments/15-segmentation-empirics.md @@ -536,3 +536,148 @@ mean(A0−do-nothing) **+0.041** верна, но **TOST ПРОВАЛЕН**: sd **Интерпретация V2-брака (несущая):** армы eval-рига идут БЕЗ прод-гейтов — санитайзер v6 (fold-first) поймал бы 待遇/一步登天 как cjk_leak; брак V2 = (рига-контекст без гейтов) + вероятный механизм «source-carryover подаёт больше исходника в контекст редактора → выше вероятность CJK-протечки» (фиксируется как гипотеза к пере-прогону, где carryover не строится). V2-брак НЕ говорит о carryover как рычаге когезии. **Следствия (ратификация D39.8):** (1) вывод D39.7 «сцен-детекцию/carryover-машинерию не строить» подтверждён и на читательском уровне; (2) битва за качество пере-прогона — в ДЕФЕКТ-КЛАССАХ, все банко/пак/гейт-уровня: 时辰-юниты (детерминированная конверсия в пак пары — НОВЫЙ класс), числа-масштабы (b1, свежая эмпирика), gender-enforce по полю сида (Бай Нинбин hidden→male-до-reveal) + диалог-род говорящего (Ф2-annotator), стих/аллюзии (пак-политика + сноски-на-аллюзии lever), регистр-лексикон (негатив-лист «терем»-класса в пак); (3) **инструмент-фиксы будущих слепых пакетов:** авто-QA до человека (CJK-в-ru/битые формы/род-консистентность), обрезка всех версий по общему финальному предложению, 2-way пакеты для тонких контрастов, тир-структура претензий (критические/смысловые/редакторские) вместо плоского «≤2». + +## §7-Q4a — ИСПОЛНЕНИЕ Q4a «сильный переводчик» (отложенный арм, мини-сессия, 2026-07-18) + +> **B-hybrid, санкция оркестратора/владельца 18.07.** Свой кап **$2, ВНЕ $15 exp15**. Скрипты `eval/exp15/q4a_*.py` +> (аддитивные; существующие риги НЕ редактировались), артефакты `/home/ubuntu/books/gu-zhenren/exp15/q4a/`, леджеры +> `q4a_*.jsonl` — ВСЁ ВНЕ git, лендит оркестратор. Pro-генерация — в DeepSeek-долине (все вызовы UTC≈14–17, вне пиков 01–04/06–10). + +> **⚠ РАУНД КОРРЕКЦИИ (2026-07-18, адверс-верификация результатов, author≠reviewer).** Первый DET-прогон нёс +> **ИНФЛИРОВАННЫЙ заголовок «pro омитит 6×»**, пойманный пост-хок адверс-верификатором (как HEAD_CHARS в §7.8): «6 омиссий +> pro» = **1 truncation** (pro_raw/7.0.s0 finish=length — reasoning съел max_tokens 8000, хвост с a1/a2/d2 потерян; НЕ +> «омиссия модели») **+ 2 локатор-ложных** (a1/d1 локаторы ловили лексикон flash, мимо валидных синонимов pro Сказание/незнакомо, +> будь…не) **+ 1 genuine** (b2 s0). Фиксы: max_tokens draft 8000→16000 + reject finish=length; truncated-ячейка перегенерена +> (finish=stop); a1/d1/d2-локаторы уточнены симметрично (d2 переанкорен на lament-структуру — generic 'если бы/а то' ловили ЧУЖИЕ +> контрфактивы чанка). Self-test 19/19. **Скорректированный итог — ниже; заголовок «pro омитит» ОТОЗВАН.** Ядро вывода (pro НЕ +> ратифицируется; редактор = слабое звено) НЕ изменилось и УСИЛИЛОСЬ. + +### §7-Q4a.0 Итог одной строкой (СКОРР.) +**Верность НЕ покупается на стадии перевода.** DET-скорер смысл-трапов (floor-immune): **и flash, и pro черновики верны на +6/9 трапах** (a1/a2 двойн-отриц, b1/b2 числа, c1 ранг, d1 контрфактив — все 1.0). flash **25 fix / 0 fail / 27 (покрытие +26/27)**; **pro НЕ точнее и слегка ХУЖЕ — роняет контрфактив d2** (0.333 vs flash 1.0; 2 fail), **покрытие РАВНОЕ (26/27 = 26/27; +differential genuine омиссия pro ≈ 1 — b2 s0, НЕ 6)**, при **COGS ×3.67**. → **Pro-арм НЕ ратифицируется** (§D5-знаменатель ПУСТ: flash не валит ни одного трапа; pro +чинить нечего, а сам слегка хуже). Реальные смысл-катастрофы рождает **glm-редактор** (инвертирует 丙→разряд **Б**, 族长→**старейшина**; +**4 fail на flash_glm, 4 на pro_glm**; editor-delta flash −0.185, pro −0.093) — прямое подтверждение D38 «слабое звено = редактор, +не черновик». Ось — DET-истина + чтение (все ключевые вердикты сверены с реальными рендерами). + +### §7-Q4a.1 Дизайн (B-hybrid; развилка §3-аддендума решена) +Концерн №1 «может ли ЧЕРНОВИК быть ~верным сразу». Развилка (A) S2′-когезия-трапы (судьи, floor-уязвимо) vs (B) exp14 +смысл-трапы (DET, floor-immune) → **владелец/оркестратор выбрали B-hybrid**: материал = 7 exp14-чанков с 9 DET-смысл-трапами +exp14b (a1/a2 двойн-отриц = катастроф-класс, b1/b2 числа, c1 ранг = кат, c2/c3 ранг/роль, d1/d2 контрфактив); **PRIMARY = +$0 DET-скорер** (иммунен к судейскому полу 0.126, что утопил все S2′-блоки); **SECONDARY = судейский мини-пасс grok+mistral** +(калибровка судейского пола НА смысл-трапах + design-комплаенс). Интент пина «те же чанки» = **ПАРНОСТЬ** сравнения — сохранена. +Полный **квадрат 2×2** {flash, pro} × {сырой черновик, +glm-редактор}, **×3 seed**, 84 ячейки. draft=`deepseek-v4-pro` +(mistral исключён — судейская семья); суб-арм do-nothing = сам pro-черновик; редактор `glm-5`. + +### §7-Q4a.2 Целостность пар — байт-верифицирована (несущее, пин оркестратора «парность превыше») +flash и pro черновики собраны **байт-идентичными** промптом/инъекцией (exp14 `bilingual_glossary`→черновик, +`editor_constraint`→редактор; injected_ids-реконструкция БЕЗ sticky — несовершенство, ОБЩЕЕ обеим сторонам, НЕ «улучшалось» +mem_select-портом). Подтверждено: `flash_raw` и `pro_raw` каждого чанка — `prompt_tokens`=идентично (напр. 6.0: 1605/1605). +Единственная переменная строки черновика = модель; редактор (glm-5 discourse) идентичен обеим сторонам → editor-delta модель-чист. +Флэш-сторона **перегенерена** (не reuse on-disk backend-flash) ради ГАРАНТИРОВАННОЙ парности с pro (backend-flash = иная, +sticky-инъекция; оставлен как провенанс-кросс-чек). Провенанс: свежий flash_raw DET-сверен с backend-flash + arm C — +совпадает, кроме a1/c1 (backend — единичный неудачный семпл с инверсией; по 3 seed flash их держит) → sampling-вариация, не баг. + +### §7-Q4a.3 Инструмент (DET) + честность пре-регистрации +Правила exp14b **переписаны как additive-оверрайды** в `q4a_traps.py` (риг exp14b НЕ редактировался): sentence-scoping +(raw≈editor сопоставимы), double-neg faithful=fix (a1/a2), **FAIL-ветка контрфактива** (d1/d2 — иначе класс DET-слеп к своему +провалу), b1 требует ОБА числа, b2 «сотых»+склонения, c1/a1 anchors сужены, stray 'мёртв'→fix снят. В раунде коррекции: a1/d1 +локаторы расширены на pro-синонимы (Сказание/незнакомо, будь…не), **d2 переанкорен на lament-структуру** (generic 'если бы/а то' +ловили ЧУЖИЕ контрфактивы чанка → ложные вердикты). **Self-test 19/19.** Калибровка на exp14b-армах + probe- И финал-рендерах +(та же книга) — **симметрична к модели** (уточнения ловят синонимы ОБЕИХ сторон), **не к исходу** (d2-фикс ВЫЯВИЛ pro-хуже, не +подгонял под pro-лучше). '?' исключается из ставок, репортится отдельно (addendum §2). **Кавета: контрфактив-класс (d1/d2) +DET-СЛАБ** (пред-ревью F2 предсказал) — предложение-феномен тонет среди других контрфактивов; вердикты d1/d2 ниже-доверие, чем +a/b/c. Инструмент-здоровье: flash_raw 25fix/**0fail**/2'?'; **a1/a2/b1/b2/c1/d1 = 1.0 во всех 4 ячейках** → устойчивы. + +### §7-Q4a.4 Результаты (СКОРР. после раунда коррекции) +**(i) 2×2 fix-rate (PAIRED, n=9 — трапы decided во ВСЕХ 4 ячейках; marginal совпал). Покрытие = доля из 27 (9×3), где +фраза-феномен ЛОКАЛИЗОВАНА (рендер):** + +| | сырой черновик (fix-rate / покрытие) | +glm-редактор (fix-rate / покрытие) | +|---|---|---| +| **flash** | **1.0** / **0.963** (25fix 0fail; 1 miss = d1-локатор, НЕ омиссия) | 0.778 / 0.889 | +| **pro** | **0.889** / **0.963** (23fix **2fail=d2**; 1 genuine омиссия b2 s0) | 0.889 / 1.0 | + +Per-trap: a1/a2/b1/b2/c1/d1 = **1.0 во ВСЕХ 4 ячейках** (оба черновика верны на двойн-отриц, числах, ранге, cicada-контрфактиве). +Единственные не-1.0: **c2** (flash_glm 0.0, pro_glm 0.5 — редактор), **c3** (flash_glm 0.333 — редактор), **d2** (pro_raw 0.333, +pro_glm 0.0 — pro роняет контрфактив). + +**(ii) COVERAGE / ОМИССИЯ — заголовок «pro омитит» ОТОЗВАН (раунд коррекции):** покрытие **РАВНОЕ: flash_raw 26/27 = pro_raw +26/27**. Первичный «pro 6 омиссий» распался: **3 = 1 truncation** (pro_raw/7.0.s0 finish=length — перегенерена), **2 = локатор- +ложные** на ВАЛИДНЫХ pro-рендерах (a1 s1 «Сказание…не было незнакомо ни одному», d1 s2 «будь цикада не так слаба…» — локаторы +уточнены симметрично), **1 = genuine** (b2 s0 — pro finish=stop, но пропустил предложение 44%). Стохастическая полн-омиссия +существует (sanity-семпл 6.0.s0 выкинул опенинг), но по probe РЕДКА (~1/11) и **НЕ даёт систематического преимущества flash**. + +**(iii) EDITOR-RESTORATION (СКОРР.):** «pro restored 5» инфлировано truncation-recovery; **genuine restored = 1** (b2 s0: +pro_raw без 44% → pro_glm вернул «сорока четырёх сотых высоты апертуры» — редактор видит исходник). flash: restored 0, **broke 2** +(c2 s0, d2 s2). Направление реально (билингв-редактор доб. потерянное), но масштаб = 1, не 5. + +**(iv) EDITOR-DELTA (T-inv/D38 «редактор ломает верное черновика») — ПОДТВЕРЖДЁН (несущее, read-verified):** editor-delta +**flash −0.185** (breaks **c2, c3**), **pro −0.093** (breaks **c2, d2**). Редактор ввёл **4 fail на flash_glm + 4 на pro_glm** vs +**0 fail на сырых черновиках flash / 2 на pro**. Верифицировано чтением: flash_glm c2 = «Фан Юань — разряд **Б**» (черновик держал +«разряд В»=丙 верно → редактор ИНВЕРТИРОВАЛ грейд + спутал имена «Фан Юань…Фан Юань»); c3 = «Четвёртый **старейшина** рода» +(черновик: «глава рода» верно). Seed-non-unanimity (DET-пол): сырые черновики стабильны (flash 0.0, pro 0.111), **редактор +добавляет вариативность (0.111/0.125)**. + +**(v) COGS-дельта (метрика iii, фриз-цены, долина):** flash-черновик **$0.02431** vs pro-черновик **$0.08926** = **×3.67** +(сходится с фриз-оценкой ×3.1). Pro-черновик дороже ×3.67 при НУЛЕВОМ приросте точности (слегка ХУЖЕ на d2), покрытие равное. + +**(vi) Судейский мини-пасс (secondary, grok-4.3 reasoning-ON + mistral-large-latest, оба порядка, 6→10, LOO, per-trap seed +где ОБЕ стороны рендерят; 496 голосов, 0 parse-fail, 0 truncation):** ключевой мета-вывод — **судейский пол НА смысл-трапах = +mean|Δ| 0.261** (floor-контраст flash-s0 vs flash-s1, ЭКВИВАЛЕНТНЫЕ входы) — **БОЛЬШЕ, чем S2′-пол когезии (0.126)**: судьи +ещё шумнее на этих трапах (a2 0.583, c2 0.792 — огромные Δ на DET-РАВНЫХ черновиках). *Это прямо ВАЛИДИРУЕТ выбор DET-primary: +судьи не смогли бы ответить на Q4a.* Отметка: floor тут = верхняя граница (судейский шум + seed-content-вариация, review F11). +- **main (pro vs flash): diff −0.001, mean|Δ| 0.268 ≈ пол** → преимущества pro НЕТ, и судья слишком шумен, чтобы что-либо + различить (a2 −0.958, c2 +1.0 — дикий разброс на DET-равных ячейках, mean washes out). Сходится с DET (оба верны при рендере). +- **editor (pro_raw vs pro_glm): diff −0.102 ≈ DET editor-delta −0.093** → судья И DET СОГЛАСНЫ, что редактор слегка деградирует + pro; **слом d2 подтверждён ОБОИМИ** (судья Δ −0.917, LOO-устойчиво: без-grok −1.0 / без-mistral −0.833, оба судьи arm_acc≈0; + DET=fail). Единственный сигнал выше шума на судейской стороне — editor-слом, совпадает с DET. + +### §7-Q4a.5 Дерево решений §D5 применённое (СКОРР.) +**§D5 [INSUFFICIENT-POWER / вакуумно]:** знаменатель «flash-черновик валит смысл-трап» = **0** (flash держит все 9 при рендере, +25fix/0fail) → «pro чинит ≥0.5 flash-провалов» неопределимо. **Вывод: пре-мисса «нужен сильный черновик для верности» +опровергнута** — flash-черновик уже верен на смысле; **pro НЕ точнее, а слегка ХУЖЕ (роняет d2-контрфактив), покрытие равное, +цена ×3.67.** → **арм «структура у переводчика» НЕ ратифицируется; подтверждается путь D38 «сильный редактор»** (свап +glm→mistral/deepseek-pro как редактор): реальные смысл-катастрофы — на **editor-стадии** (c2-грейд 丙→Б, c3-роль 族长→старейшина — +4 fail из чистого черновика; d2-контрфактив), там же и restoration-омиссий. **Каветы (честно):** (1) вакуумность §D5 ХРУПКА — +свежий 3-seed flash держит a1/c1, но САМ прод-бэкенд-черновик (records.json) ИХ ВАЛИТ (a1 инверсия «не знал ни один», c1 «среди +людей») → на шипнутом семпле знаменатель ≥2 и вопрос «чинит ли pro» был бы жив; not-ratify держится на **COGS+нет-прироста**, не +на «flash идеален». (2) **Контрфактив-класс (d1/d2) DET-СЛАБ** (пред-ревью F2 предсказал): локаторы уточнялись пост-хок под lament/ +subjunctive-формы — d2-сигнал «pro хуже» вторичен (n мал, класс хрупок); заголовок держат a/b/c + editor-breaks + равное покрытие, +НЕ зависящие от d1/d2. (3) n=9 трапов (honest-power); материал = exp14 in-претрейн канон — ПРЕДВАРИТЕЛЬНО до вебновелл-среза. +Гейтящий эндпоинт остаётся — чтение владельца на большем/трудном корпусе. + +### §7-Q4a.6 Деньги (свой кап $2) +Генерация (84 ячейки + 1 перегенер-ячейка + sanity + провенанс-диаг + omission-probe, ВСЁ персистировано, D30.10): **$0.602**. +Судейский мини-пасс (3 контраста × 9 трапов, 496 голосов): **$0.716**. **Итого $1.317 / $2** (маржа **$0.683**). Гейты: +per-call predicted-cost + свой Spender seed=gen-spend, hard-cap $2 (НЕ наследует exp15 seed $9.82/$14.5). Ошибки: генерация +**0/84**, судьи **0 parse-fail / 0 truncation**. Все pro-вызовы — в долине (UTC≈14–17). +**A-мини (сокращённый S2′-когезия судейский пасс) — ПРОПУЩЕН (решение по правилу оркестратора).** Формально остаток $0.683 ≥ +порог $0.5, НО: (1) «ответ его класса известен» (S2′-судейский null, §7.4); (2) потребовал бы НОВОЙ S2′-pro-генерации (не просто +судейства); (3) Q4a-судейский пол на смысл-трапах (0.261) уже демонстрирует судейскую-шумовую проблему на ЭТОМ материале +СИЛЬНЕЕ, чем S2′ (0.126) — A-мини добавил бы шум к чистому результату. «Пропустить без сожалений» (формулировка оркестратора). + +### §7-Q4a.7 Самопроверка + ОТКЛОНЕНИЯ (явно) +**Ревью ИСПОЛНЕНИЕМ, 3 рубежа (мандат владельца) — сработал, поймал ИНФЛИРОВАННЫЙ заголовок ДО лендинга:** +(1) **pre-spend 4-линзовый адверс-ревью** (workflow, author≠reviewer) → **11 находок**, все закрыты ДО первого платного вызова: +§D5 без power-гейта → min-denom≥3 + completeness; контрфактив без fail-ветки; span-scoping; polarity_envy ложный fail; b1 +or-логика; matrix не-парная → paired headline; budget_truncated контаминация → drop; judge gate-block/ретрай → задокументированы. +(2) **sanity-гейт** поймал стохастическую омиссию + editor-restoration + b2/a1 локатор-гэпы ДО полного прогона → 3 seed + +coverage-метрика. (3) **ПОСТ-ХОК адверс-верификация результатов** (independent agent, author≠reviewer) — **поймала, что заголовок +«pro омитит 6×» ЛОЖЕН**: 3/6 «омиссий» = 1 truncation (finish=length не отклонялся — БАГ рига; «0/84 ошибок» его пропустил), +2/6 = локатор-ложные на валидных pro-рендерах (a1/d1 локаторы под лексикон flash), + моя ручная «верификация чтением» ОШИБЛАСЬ +(спутал truncated s0 с валидным s1). **Фиксы раунда коррекции:** reject finish=length + max_tokens 16000, перегенер truncated- +ячейки, a1/d1/d2-локаторы уточнены симметрично (d2 переанкорен — generic-маркеры ловили чужие контрфактивы). Пере-скор → +покрытие РАВНОЕ, заголовок отозван, ядро вывода усилилось. **Это ровно тот случай (как HEAD_CHARS §7.8), ради которого 2-й +независимый рубеж существует.** +**Отклонения (санкционированы/задекларированы):** материал=exp14 не S2′ (санкция оркестратора); редактор=exp14b discourse не +A0 editor.md (матч материала; editor-delta = реальный glm-5); DET-правила = additive-оверрайды (exp14b-риг не тронут); +флэш-сторона перегенерена (не reuse) ради парности; judge seed выбирается per-trap где обе рендерят (омиссия — отдельная +метрика). exp14b DET-правила a1/a2/b1/b2/c1/d1/d2 признаны дефектными ревью и НЕ использованы as-is (пинг оркестратору: если +exp14/exp14b пере-читаются — брать оверрайды из q4a_traps.py, не исходные rule_*). + +### §7-Q4a.8 Артефакты +`q4a_det_results.json` (полн. per-trap × seed × cell), `q4a_omission_probe.json`, `q4a_judge_results_{floor,main,editor}.json`, +ледж `q4a_gen_costs.jsonl`/`q4a_judge_costs.jsonl`, ячейки `q4a/{flash,pro}_{raw,glm}/{chunk}.s{seed}.txt`. Скрипты +`eval/exp15/q4a_{traps,generate,score,judge,probe}.py`. diff --git a/eval/exp15/q4a_generate.py b/eval/exp15/q4a_generate.py new file mode 100644 index 0000000..99ec5af --- /dev/null +++ b/eval/exp15/q4a_generate.py @@ -0,0 +1,168 @@ +#!/usr/bin/env python3 +"""exp15/Q4a — paired 2x2 generator (B-hybrid, orchestrator-ratified 2026-07-18). + +Generates the full square {flash, pro} x {raw draft, +glm-editor}, x2 draft seeds, on the 7 exp14 +meaning-trap chunks. PAIR INTEGRITY IS SUPREME (orchestrator pin): both flash and pro drafts get +BYTE-IDENTICAL prompts + injection — the exp14 injected_ids reconstruction (bilingual_glossary for +the draft, editor_constraint for the editor), imperfect (no sticky) but SHARED by both sides; we do +NOT "improve" it with the S2' mem_select port. The ONLY variable in the draft row is the draft model. + +Assembly (exp14 harness, matches the saved flash-draft provenance): + draft : messages_with_injection(translator_system(), bilingual_glossary, translator_user(src)) + deepseek-v4-{flash|pro}, temp 0.3, thinking-ON (never disabled — echo-mine mandate) + editor : messages_with_injection(editor_system("discourse"), editor_constraint, editor_user(src,draft)) + glm-5, temp 0.4, thinking-OFF (same editor for both sides -> editor-delta is model-clean) + +Why regenerate the flash side too (vs reuse on-disk r["draft"]/arm C): the on-disk flash came from +the Go backend (real sticky injection); regenerating flash with the SAME reconstruction assembly as +pro GUARANTEES the flash<->pro pairing the orchestrator ranks above reuse. On-disk backend flash is +retained (records.json / exp14b arm C) as a provenance cross-check, DET-scored in q4a_score.py. + +Valley guard: deepseek-v4-pro calls are BLOCKED in DeepSeek peak windows (UTC 01-04 & 06-10, the +x2 surge) — abort with a defer message; flash allowed anytime. Per-call predicted-cost gate + own +ledger (every attempt incl. failures persisted, D30.10). Resumable (skip non-empty existing files). +""" +from __future__ import annotations +import argparse +import json +import sys +import time +from datetime import datetime, timezone +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14")) +import exp14_common as C +import exp14_prompts as P + +SD = Path("/home/ubuntu/books/gu-zhenren/exp15") +OUT = SD / "q4a" +LEDGER = SD / "q4a_gen_costs.jsonl" +RECS = {(r["chapter"], r["chunk_idx"]): r + for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))} +INJ = json.load(open("/home/ubuntu/books/gu-zhenren/exp14/injection_blocks.json")) + +TRAP_CHUNKS = ["6.0", "7.0", "9.1", "10.1", "16.0", "17.0", "19.0"] # 9 meaning-traps live here +SIDES = {"flash": "deepseek-v4-flash", "pro": "deepseek-v4-pro"} +SEEDS = [0, 1, 2] +DRAFT_TEMP, EDIT_TEMP = 0.3, 0.4 +# DRAFT_MAX raised 8000->16000: deepseek-v4-pro reasoning can consume the whole 8000 cap and TRUNCATE the +# translation (seen: pro_raw/7.0.s0 finish=length, reasoning starved content -> a false "omission" of the +# lost tail; adversarial verification 2026-07-18). finish=length is now REJECTED in _gen (never accepted). +DRAFT_MAX, EDIT_MAX = 16000, 12000 +PEAK_UTC_HOURS = {1, 2, 3, 6, 7, 8, 9} # DeepSeek surge windows (01-04 & 06-10 UTC) — pro forbidden +GEN_HARD_STOP = 1.10 # cumulative gen ceiling ($; leaves room for judging under $2) + +CAPS = {"deepseek-v4-flash": 0.03, "deepseek-v4-pro": 0.06, "glm-5": 0.08} + + +def _now_utc_hour(): + return datetime.now(timezone.utc).hour + + +def _src_inj(cid): + ch, ck = cid.split("."); k = (int(ch), int(ck)) + b = INJ.get(f"{ch}/{ck}", {}) + return RECS[k]["source"], b.get("bilingual_glossary", ""), b.get("editor_constraint", "") + + +def _ledger_total(): + if not LEDGER.exists(): + return 0.0 + t = 0.0 + for line in LEDGER.read_text(encoding="utf-8").splitlines(): + try: + t += json.loads(line).get("cost", 0.0) + except Exception: + pass + return t + + +def _gen(sp, model, messages, temp, max_out, meta): + """Gate -> call -> persist EVERY attempt (incl. gate-block/error). Returns text or None.""" + in_est = sum(len(m["content"]) for m in messages) // 2 + ok, pred = sp.guard(model, in_est, max_out) + if not ok: + sp.record({**meta, "model": model, "keyenv": model, "usage": {}, + "err": f"per-call gate pred ${pred:.4f} > cap ${CAPS.get(model)}", "gate_block": True}) + print(f" [GATE BLOCK] {meta} pred ${pred:.4f}") + return None + total = _ledger_total() + if total + pred > GEN_HARD_STOP: + print(f" [GEN HARD-STOP] {total:.3f}+{pred:.4f} > {GEN_HARD_STOP} — stopping generation") + return "__STOP__" + s = C.spec(model, temp=temp, max_tokens=max_out) + text, err, usage = C.call(s, messages) + fin = usage.get("finish_reason") + if not err and fin == "length": # truncated output (reasoning ate the cap) — NOT a complete draft + err = f"truncated|finish=length|completion={usage.get('completion_tokens')}" + sp.record({**meta, "model": model, "keyenv": model, "usage": usage, "err": err}) + if err or not text: + print(f" [BAD] {meta} err={err} finish={fin}") + return None + return text + + +def run(limit_chunks=0, sanity=False): + OUT.mkdir(parents=True, exist_ok=True) + for cell in ("flash_raw", "pro_raw", "flash_glm", "pro_glm"): + (OUT / cell).mkdir(exist_ok=True) + sp = C.Spender(LEDGER, CAPS) + # sanity gate: FIRST chunk, seed 0 only, but BOTH sides -> flash+pro draft AND editor (4 cells), so + # the pro/deepseek path (content non-empty, thinking-ON) is exercised before the full valley run. + chunks = TRAP_CHUNKS[:1] if sanity else (TRAP_CHUNKS[:limit_chunks] if limit_chunks else TRAP_CHUNKS) + seeds = [0] if sanity else SEEDS + ed_sys = P.editor_system("discourse") + tr_sys = P.translator_system() + made, skipped = 0, 0 + for cid in chunks: + src, gloss, edconstr = _src_inj(cid) + tr_user = P.translator_user(src) + for side, dmodel in SIDES.items(): + for seed in seeds: + # ---- draft ---- + dpath = OUT / f"{side}_raw" / f"{cid}.s{seed}.txt" + if dpath.exists() and dpath.read_text(encoding="utf-8").strip(): + draft = dpath.read_text(encoding="utf-8"); skipped += 1 + else: + if dmodel == "deepseek-v4-pro" and _now_utc_hour() in PEAK_UTC_HOURS: + print(f" [VALLEY DEFER] pro draft {cid} s{seed}: UTC hour {_now_utc_hour()} in peak " + f"{sorted(PEAK_UTC_HOURS)} — rerun in valley."); return + msgs = C.messages_with_injection(tr_sys, gloss, tr_user) + draft = _gen(sp, dmodel, msgs, DRAFT_TEMP, DRAFT_MAX, + {"cell": f"{side}_raw", "chunk": cid, "seed": seed, "stage": "draft"}) + if draft == "__STOP__": + return + if not draft: + continue + dpath.write_text(draft, encoding="utf-8"); made += 1 + print(f" {side}_raw {cid} s{seed}: {len(draft)}c gen-cum ${_ledger_total():.4f}") + # ---- editor over that draft ---- + epath = OUT / f"{side}_glm" / f"{cid}.s{seed}.txt" + if epath.exists() and epath.read_text(encoding="utf-8").strip(): + skipped += 1 + else: + emsgs = C.messages_with_injection(ed_sys, edconstr, P.editor_user(src, draft)) + final = _gen(sp, "glm-5", emsgs, EDIT_TEMP, EDIT_MAX, + {"cell": f"{side}_glm", "chunk": cid, "seed": seed, "stage": "edit"}) + if final == "__STOP__": + return + if not final: + continue + epath.write_text(final, encoding="utf-8"); made += 1 + print(f" {side}_glm {cid} s{seed}: {len(final)}c gen-cum ${_ledger_total():.4f}") + print(f"\ngen done: {made} made, {skipped} reused. gen total ${_ledger_total():.4f}") + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--limit-chunks", type=int, default=0) + ap.add_argument("--sanity", action="store_true", + help="gate: 1 chunk, seed 0, BOTH sides (flash+pro draft+editor = 4 cells)") + a = ap.parse_args() + print(f"[q4a-gen] UTC hour {_now_utc_hour()} (peaks {sorted(PEAK_UTC_HOURS)}); " + f"gen ledger total ${_ledger_total():.4f}; hard-stop ${GEN_HARD_STOP}") + run(a.limit_chunks, a.sanity) + + +if __name__ == "__main__": + main() diff --git a/eval/exp15/q4a_judge.py b/eval/exp15/q4a_judge.py new file mode 100644 index 0000000..767157d --- /dev/null +++ b/eval/exp15/q4a_judge.py @@ -0,0 +1,158 @@ +#!/usr/bin/env python3 +"""exp15/Q4a — SECONDARY grok+mistral judge mini-pass (design-compliance + judge-floor calibration +ON the meaning-traps). The $0 DET scorer (q4a_score.py) is PRIMARY and floor-immune; this pass exists +to (1) honour the frozen judge design (grok-4.3 reasoning-ON + mistral-large-latest, both orders, +6->10, LOO, catastrophe, per-vote persist) and (2) MEASURE how well judges even see these traps — +does the judge noise-floor (which swamped every S2' block at 0.126) also swamp DET-obvious meaning +errors? Reuses judges.judge_cell + exp15_llm verbatim (NO edits to those rigs). + +Contrasts (--contrast): + main : a0=flash-raw(s0), arm=pro-raw(s0) judge's pro-vs-flash draft signal (compare to DET) + floor : a0=flash-raw(s0), arm=flash-raw(s1) two seeds of the SAME model = judge-noise + same-model + seed-content-variance UPPER BOUND (review F11): it is + pure judge noise ONLY where the two seeds DET-agree + (cross-referenced against q4a_score seed_floor). + editor : a0=pro-raw(s0), arm=pro-glm(s0) judge's editor-delta (does editor break the draft) + +Judge inputs are WINDOWED around the trap's phenomenon sentence (+/-700 chars, symmetric — flash and +pro translate the SAME source so the sentence sits at ~the same position; addendum §4 authorises +windowing for same-chunk symmetric contrasts) with a full-text fallback when the sentence isn't found. +This keeps the phenomenon visible while holding per-cell cost far under the per-call cap (a full ~5k +draft would push predict near the cap and could truncate the run mid-way). + +Budget: OWN Spender, ledger q4a_judge_costs.jsonl, seed_total = current Q4a GEN spend (so the hard cap +is the GLOBAL $2 Q4a ceiling across gen+judge, NOT the exp15 $14.5). budget_truncated cells are DROPPED +from the analysis (review F3): a cell whose votes were cut off by the cap is partial/asymmetric and must +not enter the summary. per_call_cap kept high enough that ONLY the hard-cap governs (review F8). Per-call +gate + every completed vote persisted (judge gate-blocks are $0 and unledgered in the reused judges.py — +review F9, documented; the cutoff is recorded via budget_truncated in the results JSON).""" +from __future__ import annotations +import argparse +import json +import sys +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import exp15_llm as L +import q4a_traps as Q +from judges import judge_cell + +SD = Path("/home/ubuntu/books/gu-zhenren/exp15") +OUT = SD / "q4a" +GEN_LEDGER = SD / "q4a_gen_costs.jsonl" +JUDGE_LEDGER = SD / "q4a_judge_costs.jsonl" +Q4A_CAP = 2.0 # GLOBAL Q4a ceiling ($; own budget, OUTSIDE exp15 $15 — orchestrator) + +SEEDS = [0, 1, 2] +CONTRAST = { # name -> (a0 cell, arm cell); the SEED is chosen per-trap (see _pick_seed) + "main": ("flash_raw", "pro_raw"), # judge's pro-vs-flash draft signal + "floor": ("flash_raw", "flash_raw"), # same model, two DIFFERENT seeds -> noise+seed-variance floor + "editor": ("pro_raw", "pro_glm"), # judge's editor-delta +} + + +def _read(cell, cid, seed): + p = OUT / cell / f"{cid}.s{seed}.txt" + return p.read_text(encoding="utf-8") if p.exists() else None + + +def _located(cell, cid, seed, trap): + txt = _read(cell, cid, seed) + return bool(txt) and bool(Q.sentence_with(txt, trap["locator"])) + + +def _pick_seed(a0cell, armcell, trap, floor): + """Per-trap seed selection: judge the meaning GIVEN both sides RENDERED the phenomenon (omission is a + separate DET-coverage metric). Return (a0seed, armseed) or None. For 'floor', a0/arm are the same cell + -> pick two DISTINCT seeds both rendering the trap (else the floor conflates omission with judge noise).""" + cid = trap["chunk"] + if floor: + rs = [s for s in SEEDS if _located(a0cell, cid, s, trap)] + return (rs[0], rs[1]) if len(rs) >= 2 else None + for s in SEEDS: # same seed both sides, both rendering + if _located(a0cell, cid, s, trap) and _located(armcell, cid, s, trap): + return (s, s) + return None + + +def _window(text, trap, half=700): + """Symmetric +/-half window around the trap's phenomenon sentence; full text if not located.""" + sent = Q.sentence_with(text, trap["locator"]) + if not sent: + return text + probe = sent[:40] + i = text.find(probe) + if i < 0: + return text + lo, hi = max(0, i - half), min(len(text), i + len(sent) + half) + return text[lo:hi] + + +def _gen_spent(): + t = 0.0 + if GEN_LEDGER.exists(): + for line in GEN_LEDGER.read_text(encoding="utf-8").splitlines(): + try: + t += json.loads(line).get("cost", 0.0) + except Exception: + pass + return t + + +def run(contrast, per_call_cap, limit=0): + traps = Q.traps() + if limit: + traps = traps[:limit] + a0cell, armcell = CONTRAST[contrast] + gen = _gen_spent() + sp = L.Spender(str(JUDGE_LEDGER), per_call_cap=per_call_cap, hard_cap=Q4A_CAP, seed_total=gen) + print(f"[budget] Q4a gen spent ${gen:.4f}; judge ledger start ${sp.total:.4f}; global cap ${Q4A_CAP}; " + f"contrast={contrast} (a0={a0cell} arm={armcell}, seed chosen per-trap where both render)") + out, skipped, truncated = [], [], [] + for t in traps: + pick = _pick_seed(a0cell, armcell, t, floor=(contrast == "floor")) + if pick is None: + skipped.append((t["id"], "no-seed-both-render")); continue + a0seed, armseed = pick + a0_full = _read(a0cell, t["chunk"], a0seed) + arm_full = _read(armcell, t["chunk"], armseed) + det_a0, _ = Q.det_score(a0_full, t) + det_arm, _ = Q.det_score(arm_full, t) + agg = judge_cell(sp, Q.judge_trap(t), a0_text=_window(a0_full, t), arm_text=_window(arm_full, t)) + if agg.get("budget_truncated"): + # DROP a cap-truncated (partial/asymmetric) cell — do NOT let it enter the analysis (review F3) + truncated.append(t["id"]) + print(f" [HARD-CAP] Q4a ${Q4A_CAP} bound at {t['id']} — cell dropped, stopping."); break + agg.update(contrast=contrast, id=t["id"], cat_class=t["catastrophe_class"], + a0_seed=a0seed, arm_seed=armseed, det_a0=det_a0, det_arm=det_arm) + out.append(agg) + print(f" {contrast} {t['id']} (s{a0seed}/s{armseed}): judge diff(arm-a0) {agg['paired_diff_arm_minus_a0']} " + f"(a0_acc {agg['a0_accuracy']} arm_acc {agg['arm_accuracy']}; DET a0={det_a0} arm={det_arm})" + f" spend ${sp.total:.4f}") + resfile = SD / f"q4a_judge_results_{contrast}.json" + resfile.write_text(json.dumps({"contrast": contrast, "n": len(out), "skipped": skipped, + "truncated_dropped": truncated, "results": out}, + ensure_ascii=False, indent=2), encoding="utf-8") + # quick summary + decided = [a for a in out if a["paired_diff_arm_minus_a0"] is not None] + if decided: + import statistics + diffs = [a["paired_diff_arm_minus_a0"] for a in decided] + print(f"\n{contrast}: n={len(decided)} mean diff {round(statistics.mean(diffs),3)} " + f"mean|diff| {round(statistics.mean(abs(d) for d in diffs),3)} spend ${sp.total:.4f}") + print(f"{contrast}: {len(out)} scored, {len(skipped)} skipped {skipped}, " + f"{len(truncated)} cap-dropped {truncated}") + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--contrast", choices=list(CONTRAST), required=True) + # per-call cap high enough that ONLY the $2 hard-cap governs (windowed inputs predict ~$0.03; review F8) + ap.add_argument("--per-call-cap", type=float, default=0.09) + ap.add_argument("--limit", type=int, default=0) + a = ap.parse_args() + run(a.contrast, a.per_call_cap, a.limit) + + +if __name__ == "__main__": + main() diff --git a/eval/exp15/q4a_probe.py b/eval/exp15/q4a_probe.py new file mode 100644 index 0000000..5bf2fc8 --- /dev/null +++ b/eval/exp15/q4a_probe.py @@ -0,0 +1,69 @@ +#!/usr/bin/env python3 +"""exp15/Q4a — omission characterisation probe (sanity-gate follow-up). The first sanity pro-draft +OMITTED the opening ~13 source paragraphs (aperture/44%/丙 exposition); a fresh pro call was COMPLETE +-> stochastic omission at temp 0.3, NOT a capture artifact (translation is in content, not reasoning). +This probe samples K drafts of a chunk from BOTH models and measures, per meaning-trap, whether the +trap's phenomenon sentence is RENDERED (locatable) vs OMITTED — to decide seed count + whether omission +is a Q4a signal. $ persisted to the gen ledger (D30.10).""" +from __future__ import annotations +import argparse +import json +import sys +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14")) +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import exp14_common as C +import exp14_prompts as P +import q4a_traps as Q + +RECS = {(r["chapter"], r["chunk_idx"]): r + for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))} +INJ = json.load(open("/home/ubuntu/books/gu-zhenren/exp14/injection_blocks.json")) +LEDGER = Path("/home/ubuntu/books/gu-zhenren/exp15/q4a_gen_costs.jsonl") +OUTF = Path("/home/ubuntu/books/gu-zhenren/exp15/q4a_omission_probe.json") + + +def run(chunks, k): + traps = Q.traps() + results = {} + for cid in chunks: + ch, ck = cid.split("."); src = RECS[(int(ch), int(ck))]["source"] + gloss = INJ[f"{ch}/{ck}"]["bilingual_glossary"] + msgs = C.messages_with_injection(P.translator_system(), gloss, P.translator_user(src)) + chunk_traps = [t for t in traps if t["chunk"] == cid] + results[cid] = {"traps": [t["id"] for t in chunk_traps], "samples": {}} + for model in ("deepseek-v4-flash", "deepseek-v4-pro"): + rows = [] + for s in range(k): + spec = C.spec(model, temp=0.3, max_tokens=8000) + text, err, usage = C.call(spec, msgs) + with LEDGER.open("a", encoding="utf-8") as f: + f.write(json.dumps({"cell": f"PROBE_{model}", "chunk": cid, "seed": s, "model": model, + "usage": usage, "cost": C.cost_of(model, usage), "err": err}, + ensure_ascii=False) + "\n") + if err: + rows.append({"seed": s, "err": err}); continue + # per-trap: rendered (fix/fail) vs omitted ('?') + verds = {t["id"]: Q.det_score(text, t)[0] for t in chunk_traps} + rendered = sum(1 for v in verds.values() if v in ("fix", "fail")) + rows.append({"seed": s, "chars": len(text), "finish": usage.get("finish_reason"), + "rendered": rendered, "n_traps": len(chunk_traps), "verds": verds}) + print(f" {model:<18} {cid} s{s}: {len(text)}c finish={usage.get('finish_reason')} " + f"rendered {rendered}/{len(chunk_traps)} {verds}") + results[cid]["samples"][model] = rows + OUTF.write_text(json.dumps(results, ensure_ascii=False, indent=2), encoding="utf-8") + tot = sum(json.loads(l).get("cost", 0) for l in LEDGER.read_text().splitlines()) + print(f"\nprobe done -> {OUTF}; gen ledger total ${tot:.4f}") + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--chunks", nargs="+", default=["6.0", "7.0"]) + ap.add_argument("--k", type=int, default=5) + a = ap.parse_args() + run(a.chunks, a.k) + + +if __name__ == "__main__": + main() diff --git a/eval/exp15/q4a_score.py b/eval/exp15/q4a_score.py new file mode 100644 index 0000000..7e6aeaa --- /dev/null +++ b/eval/exp15/q4a_score.py @@ -0,0 +1,294 @@ +#!/usr/bin/env python3 +"""exp15/Q4a — DETERMINISTIC scorer (PRIMARY instrument, $0). Reads the paired 2x2 cells generated +by q4a_generate.py, DET-scores each meaning-trap with the exp14b rules (q4a_traps), and produces: + + * 2x2 fix-rate matrix {flash,pro} x {raw, +glm-editor} + * §D5 primary rule among traps FLASH-raw FAILS, fraction PRO-raw FIXES (threshold >=0.5) + * editor delta (pro_glm - pro_raw) and (flash_glm - flash_raw): does the editor + PRESERVE or BREAK the draft's fidelity (T-inv = editor property, D38) + * catastrophe class a1/a2/c1 (polarity/rank inversions) tracked separately + * seed variance flash-s0 vs flash-s1 (and pro) disagreement = the DET-level noise floor + (should be ~0 for polarity/number traps -> confirms instrument reliability) + * provenance x-check DET-score the on-disk BACKEND flash draft (records.json) + exp14b arm C, + to confirm the freshly-regenerated flash side agrees with the saved one + * COGS delta pro-draft vs flash-draft stage $ from the gen ledger (metric iii; pro ~x3.1) + +Scoring convention: fix=1, fail=0, '?'=undecidable (EXCLUDED from rates, reported separately — +addendum §2: no silent drops). Cell score over seeds = mean of decided seeds; cell 'fails' if +score<0.5, 'fixes' if score>=0.5, 'undecided' if all seeds '?'. A cell whose file(s) are MISSING +(not generated / valley-defer / GEN hard-stop) is tracked as 'missing' SEPARATELY from all-'?' +(review F1). §D5 ratify/not-ratify is gated on (i) a complete 2x2 square for the trap and (ii) a +pre-registered minimum decided denominator (>=3); otherwise the verdict is INSUFFICIENT-POWER, never +a clean pass/fail off a tiny partial set. $0 (no api calls).""" +from __future__ import annotations +import json +import sys +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import q4a_traps as Q + +SD = Path("/home/ubuntu/books/gu-zhenren/exp15") +OUT = SD / "q4a" +GEN_LEDGER = SD / "q4a_gen_costs.jsonl" +CELLS = ["flash_raw", "pro_raw", "flash_glm", "pro_glm"] +SEEDS = [0, 1, 2] +RECS = {(r["chapter"], r["chunk_idx"]): r + for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))} +EX14B = Path("/home/ubuntu/books/gu-zhenren/exp14b/arms") + + +def _read(cell, cid, seed): + p = OUT / cell / f"{cid}.s{seed}.txt" + return p.read_text(encoding="utf-8") if p.exists() else None + + +def _cell_verdict(scores): + """scores = list of 'fix'/'fail'/'?' over seeds -> (label, mean_of_decided|None, n_decided, n_q).""" + dec = [1.0 if s == "fix" else 0.0 for s in scores if s in ("fix", "fail")] + nq = sum(1 for s in scores if s == "?") + if not dec: + return "undecided", None, 0, nq + m = sum(dec) / len(dec) + return ("fix" if m >= 0.5 else "fail"), round(m, 3), len(dec), nq + + +def score_all(): + traps = Q.traps() + per_trap = [] + for t in traps: + row = {"id": t["id"], "chunk": t["chunk"], "cat_class": t["catastrophe_class"], "desc": t["desc"], + "cells": {}, "seed_raw": {}} + row["seed_located"] = {} + for cell in CELLS: + seed_scores, located = [], [] + for seed in SEEDS: + txt = _read(cell, t["chunk"], seed) + if txt is None: + seed_scores.append(None); located.append(None); continue + v, span = Q.det_score(txt, t) + seed_scores.append(v) + located.append(bool(span)) # sentence found = phenomenon RENDERED (else omitted/locator-miss) + row["seed_raw"][cell] = seed_scores + row["seed_located"][cell] = located + present = [s for s in seed_scores if s is not None] + label, mean, ndec, nq = _cell_verdict(present) + missing = len(present) < len(SEEDS) # some seed file absent (not generated / hard-stop) + if len(present) == 0: + label = "missing" + loc = [x for x in located if x is not None] + row["cells"][cell] = {"label": label, "mean": mean, "n_decided": ndec, "n_q": nq, + "n_present": len(present), "missing": missing, + "n_located": sum(1 for x in loc if x), "n_seen": len(loc)} + # a trap is COMPLETE only if every cell has all seeds present (review F1: completeness gate) + row["complete"] = all(not row["cells"][c]["missing"] for c in CELLS) + row["paired_decided"] = all(row["cells"][c]["mean"] is not None for c in CELLS) + per_trap.append(row) + return traps, per_trap + + +MIN_DENOM = 3 # pre-registered minimum decided flash-fail denominator for a §D5 ratify/not-ratify verdict + + +def cell_fixrate(per_trap, cell, cat_only=False, paired_only=False): + """MARGINAL fix-rate: over this cell's own decided traps (paired_only=True -> only traps decided in + ALL 4 cells, so the four cells are comparable — the paired headline, review F6).""" + rows = [r for r in per_trap if (r["cat_class"] if cat_only else True) + and (r["paired_decided"] if paired_only else True)] + dec = [r["cells"][cell]["mean"] for r in rows if r["cells"][cell]["mean"] is not None] + return (round(sum(1 for m in dec if m >= 0.5) / len(dec), 3), len(dec)) if dec else (None, 0) + + +def d5_rule(per_trap): + """§D5 (review F1): among COMPLETE traps FLASH-raw FAILS (mean<0.5), fraction PRO-raw FIXES + (mean>=0.5). ratified is a clean bool ONLY when the denominator >= MIN_DENOM AND the 2x2 square is + complete for every counted trap; otherwise ratified=None and status='INSUFFICIENT-POWER'.""" + flash_fail, pro_fixes, detail = [], 0, [] + incomplete = [r["id"] for r in per_trap if not r["complete"]] + for r in per_trap: + if not r["complete"]: + continue # never let a partial square enter the ratification metric + f = r["cells"]["flash_raw"]["mean"]; p = r["cells"]["pro_raw"]["mean"] + if f is None or p is None: + continue + if f < 0.5: # flash-raw fails this meaning-trap + flash_fail.append(r["id"]) + fixed = p >= 0.5 + pro_fixes += int(fixed) + detail.append({"id": r["id"], "cat": r["cat_class"], "flash": f, "pro": p, "pro_fixes": fixed}) + n = len(flash_fail) + rate = round(pro_fixes / n, 3) if n else None + powered = n >= MIN_DENOM + ratified = (rate >= 0.5) if (powered and rate is not None) else None + status = ("RATIFIED" if ratified else "NOT-RATIFIED") if powered else "INSUFFICIENT-POWER" + return {"flash_fail_traps": flash_fail, "n_denominator": n, "min_denominator": MIN_DENOM, + "pro_fixes": pro_fixes, "fix_rate": rate, "threshold": 0.5, "powered": powered, + "ratified": ratified, "status": status, "incomplete_traps": incomplete, "detail": detail} + + +def editor_delta(per_trap, side): + """(side_glm mean - side_raw mean) per trap; positive = editor improves, negative = editor breaks.""" + deltas, breaks = [], [] + for r in per_trap: + raw = r["cells"][f"{side}_raw"]["mean"]; glm = r["cells"][f"{side}_glm"]["mean"] + if raw is None or glm is None: + continue + d = glm - raw + deltas.append(d) + if d < 0: + breaks.append({"id": r["id"], "raw": raw, "glm": glm}) + mean = round(sum(deltas) / len(deltas), 3) if deltas else None + return {"mean_delta": mean, "n": len(deltas), "editor_breaks": breaks} + + +def seed_floor(per_trap): + """DET-level noise: per cell, fraction of traps whose decided seeds are NOT unanimous on fix/fail + (>=2 decided seeds required). ~0 expected for polarity/number traps -> confirms DET reliability.""" + out = {} + for cell in CELLS: + disagree, n = 0, 0 + for r in per_trap: + ss = [s for s in r["seed_raw"][cell] if s in ("fix", "fail")] + if len(ss) >= 2: + n += 1; disagree += int(len(set(ss)) > 1) + out[cell] = {"disagree": disagree, "n_multi_decided": n, + "rate": round(disagree / n, 3) if n else None} + return out + + +def coverage(per_trap): + """OMISSION metric (sanity-gate finding): per cell, fraction of traps whose phenomenon sentence is + LOCATED (rendered) vs not (omitted or locator-miss). flash_raw vs pro_raw = does the strong draft + omit MORE? Reported honestly with the caveat that 'not-located' conflates true omission with a + locator gap (symmetric-ish across models).""" + out = {} + for cell in CELLS: + loc = tot = 0 + for r in per_trap: + for x in r["seed_located"][cell]: + if x is not None: + tot += 1; loc += int(x) + out[cell] = {"located": loc, "seen": tot, "rate": round(loc / tot, 3) if tot else None} + return out + + +def editor_restoration(per_trap): + """Sanity-gate signal: per side, count (trap,seed) where the RAW draft did NOT render the phenomenon + (not located) but the glm editor DID (located) — the editor, seeing the source, RESTORED an omission. + Direct evidence the editor stays load-bearing (argues against strong-draft+do-nothing).""" + out = {} + for side in ("flash", "pro"): + restored, broke, both_absent = [], [], 0 + for r in per_trap: + rawloc = r["seed_located"][f"{side}_raw"]; glmloc = r["seed_located"][f"{side}_glm"] + for i in SEEDS: + rw, gl = rawloc[i], glmloc[i] + if rw is None or gl is None: + continue + if not rw and gl: + restored.append({"id": r["id"], "seed": i}) + elif rw and not gl: + broke.append({"id": r["id"], "seed": i}) + elif not rw and not gl: + both_absent += 1 + out[side] = {"restored": restored, "n_restored": len(restored), + "broke": broke, "n_broke": len(broke), "both_absent": both_absent} + return out + + +def provenance_xcheck(traps): + """DET-score the SAVED backend flash artifacts to confirm the fresh flash side agrees.""" + out = {"backend_flash_raw": {}, "exp14b_arm_C": {}} + for t in traps: + ch, ck = t["chunk"].split("."); r = RECS.get((int(ch), int(ck))) + v_bk, _ = Q.det_score(r["draft"] if r else "", t) + out["backend_flash_raw"][t["id"]] = v_bk + cp = EX14B / "C" / f"{t['chunk']}.txt" + v_c, _ = Q.det_score(cp.read_text(encoding="utf-8") if cp.exists() else "", t) + out["exp14b_arm_C"][t["id"]] = v_c + return out + + +def cogs(): + if not GEN_LEDGER.exists(): + return {} + agg = {} + for line in GEN_LEDGER.read_text(encoding="utf-8").splitlines(): + try: + d = json.loads(line) + except Exception: + continue + cell = d.get("cell", "?"); agg.setdefault(cell, {"cost": 0.0, "n": 0}) + agg[cell]["cost"] += d.get("cost", 0.0); agg[cell]["n"] += 1 + fr = agg.get("flash_raw", {}).get("cost", 0.0); pr = agg.get("pro_raw", {}).get("cost", 0.0) + return {"by_cell": {k: {"cost": round(v["cost"], 5), "n": v["n"]} for k, v in agg.items()}, + "draft_stage_flash": round(fr, 5), "draft_stage_pro": round(pr, 5), + "pro_over_flash_x": round(pr / fr, 2) if fr else None} + + +def main(): + traps, per_trap = score_all() + # MARGINAL matrix (each cell over its own decided traps) + PAIRED matrix (only traps decided in ALL + # 4 cells -> the comparable headline; review F6). Marginal cells can be over DIFFERENT trap subsets. + matrix = {cell: {"marginal": cell_fixrate(per_trap, cell), + "paired": cell_fixrate(per_trap, cell, paired_only=True), + "cat": cell_fixrate(per_trap, cell, cat_only=True)} for cell in CELLS} + n_paired = sum(1 for r in per_trap if r["paired_decided"]) + n_complete = sum(1 for r in per_trap if r["complete"]) + result = { + "n_traps": len(traps), "seeds": SEEDS, + "n_complete_square": n_complete, "n_paired_decided": n_paired, + "fixrate_matrix": matrix, + "d5_rule": d5_rule(per_trap), + "editor_delta_pro": editor_delta(per_trap, "pro"), + "editor_delta_flash": editor_delta(per_trap, "flash"), + "seed_floor": seed_floor(per_trap), + "coverage": coverage(per_trap), + "editor_restoration": editor_restoration(per_trap), + "provenance_xcheck": provenance_xcheck(traps), + "cogs": cogs(), + "per_trap": per_trap, + } + (SD / "q4a_det_results.json").write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8") + + print(f"=== Q4a DET 2x2 fix-rate matrix (n_traps={result['n_traps']}, complete-square={n_complete}, " + f"paired-decided={n_paired}) ===") + print("cell = PAIRED (all-4-cells-decided; comparable) | marginal (own decided subset):") + print(f"{'':<10}{'raw draft':<30}{'+glm editor':<30}") + for side in ("flash", "pro"): + raw, glm = matrix[f"{side}_raw"], matrix[f"{side}_glm"] + rc = f"{raw['paired']} | {raw['marginal']}" + gc = f"{glm['paired']} | {glm['marginal']}" + print(f" {side:<8}{rc:<30}{gc:<30}") + d5 = result["d5_rule"] + print(f"\n§D5 [{d5['status']}]: among COMPLETE traps flash-raw FAILS {d5['n_denominator']} " + f"{d5['flash_fail_traps']} (min for verdict {d5['min_denominator']}); pro-raw FIXES {d5['pro_fixes']} " + f"-> fix-rate {d5['fix_rate']} (thr 0.5); incomplete traps {d5['incomplete_traps']}") + print(f"editor-delta pro (glm-raw): {result['editor_delta_pro']['mean_delta']} " + f"breaks={[b['id'] for b in result['editor_delta_pro']['editor_breaks']]}") + print(f"editor-delta flash(glm-raw): {result['editor_delta_flash']['mean_delta']} " + f"breaks={[b['id'] for b in result['editor_delta_flash']['editor_breaks']]}") + print(f"seed-floor (DET non-unanimity rate): " + f"{ {c: result['seed_floor'][c]['rate'] for c in CELLS} }") + cov = result["coverage"] + print(f"coverage (phenomenon RENDERED rate; omission proxy): " + f"flash_raw {cov['flash_raw']['rate']} ({cov['flash_raw']['located']}/{cov['flash_raw']['seen']}) " + f"pro_raw {cov['pro_raw']['rate']} ({cov['pro_raw']['located']}/{cov['pro_raw']['seen']}) | " + f"flash_glm {cov['flash_glm']['rate']} pro_glm {cov['pro_glm']['rate']}") + er = result["editor_restoration"] + print(f"editor restoration (raw-omitted -> glm-rendered): " + f"flash restored {er['flash']['n_restored']} broke {er['flash']['n_broke']} | " + f"pro restored {er['pro']['n_restored']} broke {er['pro']['n_broke']}") + print(f"COGS: flash-draft ${result['cogs'].get('draft_stage_flash')} vs " + f"pro-draft ${result['cogs'].get('draft_stage_pro')} (x{result['cogs'].get('pro_over_flash_x')})") + print("\nprovenance x-check (fresh flash_raw vs backend saved):") + for t in traps: + fresh = [per for per in per_trap if per["id"] == t["id"]][0]["cells"]["flash_raw"]["label"] + bk = result["provenance_xcheck"]["backend_flash_raw"][t["id"]] + armc = result["provenance_xcheck"]["exp14b_arm_C"][t["id"]] + flag = "" if fresh == bk else " <-- DIFF" + print(f" {t['id']:<3} fresh={fresh:<10} backend_flash={bk:<10} armC(flash+glm)={armc:<10}{flag}") + + +if __name__ == "__main__": + main() diff --git a/eval/exp15/q4a_traps.py b/eval/exp15/q4a_traps.py new file mode 100644 index 0000000..bd5cfbf --- /dev/null +++ b/eval/exp15/q4a_traps.py @@ -0,0 +1,259 @@ +#!/usr/bin/env python3 +"""exp15/Q4a — meaning-trap definitions + CORRECTED deterministic scorer (B-hybrid, orchestrator- +ratified 2026-07-18). PRIMARY instrument; $0; floor-immune (unlike the S2' judge blocks). + +The Q4a concern (#1): can a strong draft (deepseek-v4-pro) be faithful enough that the editor is +less load-bearing? Sharpest instrument = the exp14b DET meaning-trap battery (double-negation / +number-scale / rank / counterfactual), which is $0-scorable by rule and immune to the judge noise- +floor (0.126) that swamped every S2' block (addendum §1). + +RULE CORRECTIONS (frozen here PRE-generation — no pro data exists yet, so this is legitimate pre- +registration refinement, driven by the pre-spend adversarial review, NOT by outcomes). We may NOT +edit the existing rig eval/exp14b/exp14b_score.py, so its flawed rules are OVERRIDDEN in this +additive layer; its sound rules (b2/c2/c3) are reused verbatim. Corrections: + * SENTENCE-scoping (review F4): score the single SENTENCE carrying the phenomenon, not the whole + matched line/paragraph — so a raw draft (short line) and a glm editor output (merged paragraph) + are scored over comparable spans; an unrelated 'никто не' elsewhere in a glm paragraph can no + longer flip a1/a2. + * a1/a2 double-negation (review F5): a FAITHFUL double-negation ('не было ни одного, кто НЕ знал' + / 'ни одного взгляда БЕЗ зависти' = everyone) must score FIX, not FAIL — checked before the + polarity-inversion fail pattern. + * a1/c1 anchors tightened; the stray 'мёртв'→fix branch of rule_polarity_all dropped (review F4). + * counterfactual d1/d2 (review F2): a FAIL branch added (anchor sentence present but the + counterfactual marker 否则→'иначе' dropped → fail) + 'а то'/'а не то' markers — else the class + could never enter the §D5 flash-fail denominator. + * b1 (review F7): FIX now requires BOTH numbers (1/12 AND 30%); a mangled fraction → fail. + +Scoring: fix / fail / '?' (undecidable: sentence not found or unclassifiable — EXCLUDED from rates, +never a silent pass; addendum §2). Catastrophe class a1/a2/c1 = gross polarity/rank inversions.""" +from __future__ import annotations +import re +import sys +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14b")) +import exp14b_score as S # reuse SOUND rules only: rule_num_b2 (b2), rule_grade_bing (c2), rule_patriarch (c3) + +CAT_CLASS = {"a1", "a2", "c1"} + + +# ── sentence scoping ──────────────────────────────────────────────────────────────── +_ENDERS = ".!?…\n" + + +def sentence_with(text, locator): + """Return the single sentence of `text` containing the FIRST match of `locator` (regex, re.I). + '' if no match. Bounds on Russian/CJK sentence enders — makes raw vs editor spans comparable.""" + m = re.search(locator, text or "", re.I) + if not m: + return "" + i = m.start() + lo = max((text.rfind(p, 0, i) for p in _ENDERS), default=-1) + hi_cands = [text.find(p, i) for p in _ENDERS if text.find(p, i) >= 0] + hi = min(hi_cands) if hi_cands else len(text) + return text[lo + 1:hi + 1].strip() + + +# ── corrected rules (operate on the scoped SENTENCE) ──────────────────────────────── +def rule_a1(s): + """没有一个不知道 = ВСЕ знали. fail = инверсия 'никто не знал'. Double-neg faithful form = fix.""" + l = s.lower() + # faithful double-negation FIRST (so it isn't caught by the inversion pattern below). + # covers: 'не было ни одного, кто не знал', 'никто не оставался в неведении', + # 'не было незнакомо / неизвестно ни одному / никому' (было НЕ-незнакомо = знали все). + if re.search(r"(не было|нет|не оставалось|не остал)\s+(ни одного|никого|ни единого|человека)[^.]*(кто|котор|не зна|не ведал|неведени)", l) \ + or re.search(r"(никто|ни один)[^.]*не\s+(оставал|остал|был)[^.]*неведени", l) \ + or re.search(r"не\s+был[оа]?\s+(не\s*)?(знаком|извест)[а-я]*[^.]*(ни одному|никому|ни для кого|ни один|ни одного)", l): + return "fix" + # polarity INVERSION (nobody knew) — bidirectional: 'ни один … не знал' OR 'не знал … ни один' + if re.search(r"(никто|ни один|ни одного|ни единого)[^.]*не\s*(знал|ведал|слышал|слыхал)", l) \ + or re.search(r"не\s*(знал|ведал|слышал|слыхал)[^.]*(ни один|никто|ни одного|ни единого)", l) \ + or re.search(r"не\s+был[аои]?\s+извест", l): + return "fail" + # positive faithful (everyone knew) + if re.search(r"(все|всякий|каждый|весь род|в роду|поголовно|любой|всем)[^.]*(знал|знали|ведал|известн)", l) \ + or re.search(r"(знал[аи]?|ведал[аи]?|известн[ао]?)[^.]*(все|всякий|каждый|весь|любой)", l): + return "fix" + return "?" + + +def rule_a2(s): + """无不…羡慕嫉妒 = ВСЕ полны зависти. fail = 'без зависти/никто не завидовал'. Double-neg = fix.""" + l = s.lower() + if not re.search(r"завист|зависти|зависть|ревнос|завидова|зави́д", l): + return "?" + # faithful double-negation ('не было ни одного … без зависти') + if re.search(r"(не было|нет|ни один|ни одного|никого)[^.]*без\s+завист", l): + return "fix" + # inversion (no envy) + if re.search(r"без\s+завист|не было завист|(никто|ни один)[^.]*не\s*завид", l): + return "fail" + return "fix" # envy present, affirmative context + + +def rule_b1(s): + """十二分之一 (1/12) AND 三成 (30%). FIX requires BOTH; mangled fraction or 30→3 = fail.""" + l = s.lower() + has12 = bool(re.search(r"одн[ауой]{1,2} двенадцат|двенадцат[ауой]{1,2} (часть|долю)|1\s*/\s*12|двенадцатой", l)) + has30 = bool(re.search(r"тридцать процент|30\s*%|три десят|три десятых|0[.,]3\b", l)) + frac_bad = bool(re.search(r"одн[ауой]{1,2} втор|1\s*/\s*2|половин|одн[ауой]{1,2} трет|1\s*/\s*3|одн[ауой]{1,2} десят", l)) + pct_bad = bool(re.search(r"\b3\s*%|\bтри процент|три с полов", l)) and not has30 + if pct_bad or frac_bad: + return "fail" + if has12 and has30: + return "fix" + return "?" + + +def rule_b2(s): + """四成四 = 44% = 44 сотых. FIX: 'сорок(а) четыр[её]х процент/сотых', '44%'. FAIL: 4,4% / 4% (10x low). + Overrides exp14b rule_num_b2, which knew only the 'процент' surface form and missed the 'сотых' + rendering + declined 'сорока' (review: repeated b2 '?' on valid pro renderings).""" + l = s.lower() + if re.search(r"сорок\w*\s+четыр\w*\s*(процент|сот|%)|\b44\s*(процент|%|сот)|0[.,]44", l): + return "fix" + if re.search(r"\b4[.,]4\s*(процент|%)|четыре\s+процент|\b4\s*%|четыре\s+сот\w*\s+высот", l): + return "fail" + return "?" + + +def rule_c1(s): + """人上之人 = человек НАД людьми. fail = 'среди людей/один из людей'.""" + l = s.lower() + if re.search(r"среди людей|одним из людей|как все|наравне с людьми|обычным человеком|таким же, как", l): + return "fail" + if re.search(r"над людьми|над (простыми|обычными|прочими|остальными)? ?людьми|выше (простых|обычных|других|прочих|остальных|всех)? ?людей|возвыс|человек[а-я]* выше|над смертными|вознес", l): + return "fix" + return "?" + + +def _counterfactual(s): + """幸亏…否则 / 否则 contrafactual. Marker present = fix; anchor sentence but marker dropped = fail. + Markers include the SUBJUNCTIVE forms that render 否则 without the literal 'иначе' (review of pro + renderings: 'будь цикада не так слаба…', 'не будь…', 'если бы не…' = faithful counterfactual).""" + l = s.lower() + if re.search(r"иначе|в противном случае|не то бы|а не то|\bа то\b|иным образом|в ином случае|" + r"если бы(?:\s+\w+){0,3}\s+не|\bне будь\b|\bбудь\b(?:\s+\w+){0,4}\s+не|в противном разе", l): + return "fix" + return "fail" # the locator already selected the counterfactual sentence; no marker => modality lost + + +# rule dispatch: corrected rules override the flawed exp14b ones; sound exp14b rules reused verbatim. +RULES = {"a1": rule_a1, "a2": rule_a2, "b1": rule_b1, "b2": rule_b2, + "c1": rule_c1, "c2": S.rule_grade_bing, "c3": S.rule_patriarch, + "d1": _counterfactual, "d2": _counterfactual} + +# locator regex to find the phenomenon SENTENCE (broader than the classifier; tightened for a1/c1). +LOCATOR = { + # a1: the LEGEND (典故 of the 4th clan head & Huajiu Xingzhe) that "everyone/no-one-didn't" knew. + # story-words incl. сказани/быль; knowledge-words incl. знаком (не было НЕзнакомо = знали все). + "a1": r"(притч|легенд|истори|предани|типичн|наследи|сказани|быль|молв)[^.]{0,90}(знал|извест|ведал|неведени|знаком)|" + r"(знал|извест|ведал|неведени|знаком)[^.]{0,90}(притч|легенд|Монах|Цветочн|Четвёрт|сказани|быль)", + "a2": r"завист|зависти|зависть|ревнос|завидова", + "b1": r"двенадцат|(истинн\w* ци|真元)[^.]{0,60}(потрач|истрач|消耗)", # 蛊虫祭炼1/12, 真元耗三成 + "b2": r"сорок\w*\s+четыр|\b44\b|четыр\w*\s+сот|четыре десят|четырёх десят|половин\w*\s+высот|высот\w*\s+апертур", # 四成四 = 44% + "c1": r"стать человеком|человеком (над|среди|выше)|над людьми|среди людей|одним из людей|" + r"выше .{0,15}людей|над смертными|наравне с людьми", + "c2": r"Фан Юань[^.]{0,60}(разряд|класс|ранг|степен|丙)|(третьего|перв|втор) разряд|разряда [вбаг]|丙", + "c3": r"族长|поколени|старейшин|глав[аеуы].{0,20}(рода|клана|поколени)|патриарх|засад|подстро|ловушк", + # d1: 幸亏春秋蝉虚弱…否则. locate via 'к счастью'/'幸亏' OR a subjunctive counterfactual. NOT the cicada + # entity (in the chapter TITLE) and NOT bare 'повезло' (matches success-context 'наконец повезло'). + "d1": r"счастью|благо, что|хорошо, что|幸亏|庆幸|\bне будь\b|\bбудь\b\s+\w+\s+не|если бы\s+\w+\s+не", + # d2: 只恨我没有早生几百年,否则…揭破他的丑恶嘴脸. Anchor on the SPECIFIC lament ('жаль/досадно … родился … + # раньше') — generic counterfactual markers (иначе/а то/если бы) collide with OTHER counterfactual and + # villain sentences in this chunk (verified: they grab the wrong span). Fallback: the expose-face content. + "d2": r"(жаль|досадно|обидно|горько|сожал)[^.!?]{0,55}(родил|рожд|появил|на свет)[^.!?]{0,45}" + r"(раньше|ране|столет|сотен лет|веками|прежде|давно)|揭破|丑恶嘴脸|разоблач\w* его|сорв\w* (маск|личин)", +} + +ANNOT = { + "a1": dict(src_zone="四代族长和花酒行者的典故,古月族人没有一个不知道的。", + phenomenon="Двойное отрицание 没有一个不知道 ('нет НИ ОДНОГО, кто НЕ знал бы' = знали ВСЕ). " + "Провал = инверсия полярности ('никто не знал').", + expected="Все/каждый в клане знали эту легенду; никто не был в неведении."), + "a2": dict(src_zone="许多人不由自主地看向第一排正襟危坐的古月方正,这可是甲等资质啊,目光中无不充满了羡慕嫉妒的情感。", + phenomenon="无不 (двойное отрицание) — 目光中无不充满羡慕嫉妒 = ВСЕ взгляды полны зависти. " + "Провал = 'без зависти / никто не завидовал'.", + expected="Все взгляды полны зависти и ревности (无不 = поголовно)."), + "b1": dict(src_zone="蛊虫只祭炼了十二分之一,而我的真元却消耗了整整三成。", + phenomenon="十二分之一 = 1/12 (очищено гу); 三成 = 30% (истрачено ци). Провал = искажение дробей.", + expected="Одна двенадцатая (1/12) очищено, истрачено целых тридцать процентов (30%)."), + "b2": dict(src_zone="海面不到空窍的一半高度,只有四成四。", + phenomenon="四成四 = 4,4 десятых = 44%. Провал = '4,4%' или '4%' (занижение ~10×).", + expected="Около сорока четырёх процентов (44%)."), + "c1": dict(src_zone="蛊师能拥有超越凡人的力量,成为人上之人,但是这其中付出的代价,也是高昂的。", + phenomenon="人上之人 = человек НАД людьми. Провал = 'один из людей / среди людей'.", + expected="Стать человеком ВЫШЕ прочих / возвыситься над обычными людьми."), + "c2": dict(src_zone="方源只是个丙等,方正可是甲等资质。", + phenomenon="丙等 = разряд В / третий (ранг Фан Юаня). Провал = разряд Б / второй / A.", + expected="Разряд В (третий, 丙), в противопоставление 甲 (А) у брата."), + "c3": dict(src_zone="当年,四代族长暗算花酒行者不成,战败后又偷袭,虽然击退了后者,但是他也因此身亡。", + phenomenon="四代族长 = Четвёртый ГЛАВА рода/клана (族长). Провал = 'старейшина' (家老).", + expected="Глава рода/клана (族长), а не старейшина."), + "d1": dict(src_zone="庆幸的是,幸亏春秋蝉虚弱到这种程度,否则自己麻烦就大了!", + phenomenon="幸亏…否则 = контрфактив ('к счастью …, ИНАЧЕ была бы беда'). Провал = потеря 'иначе'.", + expected="К счастью (幸亏) …, иначе / в противном случае (否则) была бы беда."), + "d2": dict(src_zone="“只恨我没有早生几百年,否则见到那个魔头,定要拼死揭破他的丑恶嘴脸。”", + phenomenon="否则 = контрфактив ('жаль, что не родился раньше, ИНАЧЕ разоблачил бы демона'). " + "Провал = потеря 'иначе/否则'.", + expected="… иначе / в противном случае (否则) я бы разоблачил его."), +} + +# chunk + short label per trap (from exp14b TRAPS) +CHUNK = {tid: cid for tid, cid, *_ in S.TRAPS} +DESC = {tid: desc for tid, _c, _a, _r, desc in S.TRAPS} + + +def traps(): + out = [] + for tid in ("a1", "a2", "b1", "b2", "c1", "c2", "c3", "d1", "d2"): + out.append(dict(id=tid, chunk=CHUNK[tid], locator=LOCATOR[tid], rule=RULES[tid], + desc=DESC[tid], catastrophe_class=tid in CAT_CLASS, **ANNOT[tid])) + return out + + +def det_score(text, trap): + """Sentence-scoped fix/fail/? . '?' = phenomenon sentence not found OR rule unclassifiable + (never a silent pass). Returns (verdict, scoped_sentence).""" + sent = sentence_with(text, trap["locator"]) + if not sent: + return "?", "" + return trap["rule"](sent), sent + + +def judge_trap(trap): + return {"id": trap["id"], "type": "T-inv" if trap["id"][0] in "acd" else "T-num", + "src_zone": trap["src_zone"], "phenomenon": trap["phenomenon"], "expected": trap["expected"]} + + +if __name__ == "__main__": + # rule self-test on HAND-WRITTEN fix/fail strings (validates the corrected rules pre-generation) + cases = [ + ("a1", "Эту легенду в роду Гуюэ знал каждый.", "fix"), + ("a1", "О легенде не было ни одного человека, кто не знал бы её.", "fix"), + ("a1", "О легенде про главу рода в роду Гуюэ не знал ни один человек.", "fail"), + ("a2", "Во всех взглядах читались зависть и ревность.", "fix"), + ("a2", "Не было ни одного взгляда без зависти.", "fix"), + ("a2", "В их взглядах не было зависти.", "fail"), + ("b1", "Гу-червь очищен лишь на одну двенадцатую, а истинная ци потрачена на целых тридцать процентов.", "fix"), + ("b1", "Гу-червь очищен наполовину, а истинная ци потрачена на три процента.", "fail"), + ("b2", "Вода в море достигает лишь сорока четырёх сотых высоты апертуры.", "fix"), + ("b2", "Уровень воды — только сорок четыре процента высоты апертуры.", "fix"), + ("b2", "Уровень воды — лишь четыре процента высоты апертуры.", "fail"), + ("c1", "Заклинатель гу становится человеком выше прочих людей.", "fix"), + ("c1", "Заклинатель гу остаётся одним из людей.", "fail"), + ("a1", "Сказание о Четвёртом главе рода не было незнакомо ни одному человеку из рода Гуюэ.", "fix"), + ("d1", "К счастью, Цикада так слаба, иначе была бы большая беда.", "fix"), + ("d1", "Будь цикада не так чудовищно слаба, у него самого были бы огромные проблемы.", "fix"), + ("d1", "К счастью, Цикада оказалась совсем слабой.", "fail"), + ("d2", "Жаль, что не родился раньше, а то разоблачил бы того демона.", "fix"), + ("d2", "Жаль, что не родился раньше; я бы разоблачил того демона.", "fail"), + ] + tby = {t["id"]: t for t in traps()} + ok = 0 + for tid, txt, want in cases: + got, sent = det_score(txt, tby[tid]) + flag = "OK" if got == want else "**FAIL**" + ok += got == want + print(f" {tid} want={want:<5} got={got:<5} {flag} [{sent[:50]}]") + print(f"\nrule self-test: {ok}/{len(cases)} pass")