Land Q4a: fidelity is not bought at the translate stage, the pro arm is not ratified, the glm editor is the confirmed weak link, with corrected trap rules and the truncation-accepting harness bug documented

This commit is contained in:
Claude (backend session) 2026-07-18 22:08:20 +03:00
parent 94aeeb55fd
commit cfcf0f08e3
6 changed files with 1093 additions and 0 deletions

View file

@ -536,3 +536,148 @@ mean(A0do-nothing) **+0.041** верна, но **TOST ПРОВАЛЕН**: sd
**Интерпретация V2-брака (несущая):** армы eval-рига идут БЕЗ прод-гейтов — санитайзер v6 (fold-first) поймал бы 待遇/一步登天 как cjk_leak; брак V2 = (рига-контекст без гейтов) + вероятный механизм «source-carryover подаёт больше исходника в контекст редактора → выше вероятность CJK-протечки» (фиксируется как гипотеза к пере-прогону, где carryover не строится). V2-брак НЕ говорит о carryover как рычаге когезии.
**Следствия (ратификация D39.8):** (1) вывод D39.7 «сцен-детекцию/carryover-машинерию не строить» подтверждён и на читательском уровне; (2) битва за качество пере-прогона — в ДЕФЕКТ-КЛАССАХ, все банко/пак/гейт-уровня: 时辰-юниты (детерминированная конверсия в пак пары — НОВЫЙ класс), числа-масштабы (b1, свежая эмпирика), gender-enforce по полю сида (Бай Нинбин hidden→male-до-reveal) + диалог-род говорящего (Ф2-annotator), стих/аллюзии (пак-политика + сноски-на-аллюзии lever), регистр-лексикон (негатив-лист «терем»-класса в пак); (3) **инструмент-фиксы будущих слепых пакетов:** авто-QA до человека (CJK-в-ru/битые формы/род-консистентность), обрезка всех версий по общему финальному предложению, 2-way пакеты для тонких контрастов, тир-структура претензий (критические/смысловые/редакторские) вместо плоского «≤2».
## §7-Q4a — ИСПОЛНЕНИЕ Q4a «сильный переводчик» (отложенный арм, мини-сессия, 2026-07-18)
> **B-hybrid, санкция оркестратора/владельца 18.07.** Свой кап **$2, ВНЕ $15 exp15**. Скрипты `eval/exp15/q4a_*.py`
> (аддитивные; существующие риги НЕ редактировались), артефакты `/home/ubuntu/books/gu-zhenren/exp15/q4a/`, леджеры
> `q4a_*.jsonl` — ВСЁ ВНЕ git, лендит оркестратор. Pro-генерация — в DeepSeek-долине (все вызовы UTC≈1417, вне пиков 0104/0610).
> **⚠ РАУНД КОРРЕКЦИИ (2026-07-18, адверс-верификация результатов, author≠reviewer).** Первый DET-прогон нёс
> **ИНФЛИРОВАННЫЙ заголовок «pro омитит 6×»**, пойманный пост-хок адверс-верификатором (как HEAD_CHARS в §7.8): «6 омиссий
> pro» = **1 truncation** (pro_raw/7.0.s0 finish=length — reasoning съел max_tokens 8000, хвост с a1/a2/d2 потерян; НЕ
> «омиссия модели») **+ 2 локатор-ложных** (a1/d1 локаторы ловили лексикон flash, мимо валидных синонимов pro Сказание/незнакомо,
> будь…не) **+ 1 genuine** (b2 s0). Фиксы: max_tokens draft 8000→16000 + reject finish=length; truncated-ячейка перегенерена
> (finish=stop); a1/d1/d2-локаторы уточнены симметрично (d2 переанкорен на lament-структуру — generic 'если бы/а то' ловили ЧУЖИЕ
> контрфактивы чанка). Self-test 19/19. **Скорректированный итог — ниже; заголовок «pro омитит» ОТОЗВАН.** Ядро вывода (pro НЕ
> ратифицируется; редактор = слабое звено) НЕ изменилось и УСИЛИЛОСЬ.
### §7-Q4a.0 Итог одной строкой (СКОРР.)
**Верность НЕ покупается на стадии перевода.** DET-скорер смысл-трапов (floor-immune): **и flash, и pro черновики верны на
6/9 трапах** (a1/a2 двойн-отриц, b1/b2 числа, c1 ранг, d1 контрфактив — все 1.0). flash **25 fix / 0 fail / 27 (покрытие
26/27)**; **pro НЕ точнее и слегка ХУЖЕ — роняет контрфактив d2** (0.333 vs flash 1.0; 2 fail), **покрытие РАВНОЕ (26/27 = 26/27;
differential genuine омиссия pro ≈ 1 — b2 s0, НЕ 6)**, при **COGS ×3.67**. → **Pro-арм НЕ ратифицируется** (§D5-знаменатель ПУСТ: flash не валит ни одного трапа; pro
чинить нечего, а сам слегка хуже). Реальные смысл-катастрофы рождает **glm-редактор** (инвертирует 丙→разряд **Б**, 族长→**старейшина**;
**4 fail на flash_glm, 4 на pro_glm**; editor-delta flash 0.185, pro 0.093) — прямое подтверждение D38 «слабое звено = редактор,
не черновик». Ось — DET-истина + чтение (все ключевые вердикты сверены с реальными рендерами).
### §7-Q4a.1 Дизайн (B-hybrid; развилка §3-аддендума решена)
Концерн №1 «может ли ЧЕРНОВИК быть ~верным сразу». Развилка (A) S2-когезия-трапы (судьи, floor-уязвимо) vs (B) exp14
смысл-трапы (DET, floor-immune) → **владелец/оркестратор выбрали B-hybrid**: материал = 7 exp14-чанков с 9 DET-смысл-трапами
exp14b (a1/a2 двойн-отриц = катастроф-класс, b1/b2 числа, c1 ранг = кат, c2/c3 ранг/роль, d1/d2 контрфактив); **PRIMARY =
$0 DET-скорер** (иммунен к судейскому полу 0.126, что утопил все S2-блоки); **SECONDARY = судейский мини-пасс grok+mistral**
(калибровка судейского пола НА смысл-трапах + design-комплаенс). Интент пина «те же чанки» = **ПАРНОСТЬ** сравнения — сохранена.
Полный **квадрат 2×2** {flash, pro} × {сырой черновик, +glm-редактор}, **×3 seed**, 84 ячейки. draft=`deepseek-v4-pro`
(mistral исключён — судейская семья); суб-арм do-nothing = сам pro-черновик; редактор `glm-5`.
### §7-Q4a.2 Целостность пар — байт-верифицирована (несущее, пин оркестратора «парность превыше»)
flash и pro черновики собраны **байт-идентичными** промптом/инъекцией (exp14 `bilingual_glossary`→черновик,
`editor_constraint`→редактор; injected_ids-реконструкция БЕЗ sticky — несовершенство, ОБЩЕЕ обеим сторонам, НЕ «улучшалось»
mem_select-портом). Подтверждено: `flash_raw` и `pro_raw` каждого чанка — `prompt_tokens`=идентично (напр. 6.0: 1605/1605).
Единственная переменная строки черновика = модель; редактор (glm-5 discourse) идентичен обеим сторонам → editor-delta модель-чист.
Флэш-сторона **перегенерена** (не reuse on-disk backend-flash) ради ГАРАНТИРОВАННОЙ парности с pro (backend-flash = иная,
sticky-инъекция; оставлен как провенанс-кросс-чек). Провенанс: свежий flash_raw DET-сверен с backend-flash + arm C —
совпадает, кроме a1/c1 (backend — единичный неудачный семпл с инверсией; по 3 seed flash их держит) → sampling-вариация, не баг.
### §7-Q4a.3 Инструмент (DET) + честность пре-регистрации
Правила exp14b **переписаны как additive-оверрайды** в `q4a_traps.py` (риг exp14b НЕ редактировался): sentence-scoping
(raw≈editor сопоставимы), double-neg faithful=fix (a1/a2), **FAIL-ветка контрфактива** (d1/d2 — иначе класс DET-слеп к своему
провалу), b1 требует ОБА числа, b2 «сотых»+склонения, c1/a1 anchors сужены, stray 'мёртв'→fix снят. В раунде коррекции: a1/d1
локаторы расширены на pro-синонимы (Сказание/незнакомо, будь…не), **d2 переанкорен на lament-структуру** (generic 'если бы/а то'
ловили ЧУЖИЕ контрфактивы чанка → ложные вердикты). **Self-test 19/19.** Калибровка на exp14b-армах + probe- И финал-рендерах
(та же книга) — **симметрична к модели** (уточнения ловят синонимы ОБЕИХ сторон), **не к исходу** (d2-фикс ВЫЯВИЛ pro-хуже, не
подгонял под pro-лучше). '?' исключается из ставок, репортится отдельно (addendum §2). **Кавета: контрфактив-класс (d1/d2)
DET-СЛАБ** (пред-ревью F2 предсказал) — предложение-феномен тонет среди других контрфактивов; вердикты d1/d2 ниже-доверие, чем
a/b/c. Инструмент-здоровье: flash_raw 25fix/**0fail**/2'?'; **a1/a2/b1/b2/c1/d1 = 1.0 во всех 4 ячейках** → устойчивы.
### §7-Q4a.4 Результаты (СКОРР. после раунда коррекции)
**(i) 2×2 fix-rate (PAIRED, n=9 — трапы decided во ВСЕХ 4 ячейках; marginal совпал). Покрытие = доля из 27 (9×3), где
фраза-феномен ЛОКАЛИЗОВАНА (рендер):**
| | сырой черновик (fix-rate / покрытие) | +glm-редактор (fix-rate / покрытие) |
|---|---|---|
| **flash** | **1.0** / **0.963** (25fix 0fail; 1 miss = d1-локатор, НЕ омиссия) | 0.778 / 0.889 |
| **pro** | **0.889** / **0.963** (23fix **2fail=d2**; 1 genuine омиссия b2 s0) | 0.889 / 1.0 |
Per-trap: a1/a2/b1/b2/c1/d1 = **1.0 во ВСЕХ 4 ячейках** (оба черновика верны на двойн-отриц, числах, ранге, cicada-контрфактиве).
Единственные не-1.0: **c2** (flash_glm 0.0, pro_glm 0.5 — редактор), **c3** (flash_glm 0.333 — редактор), **d2** (pro_raw 0.333,
pro_glm 0.0 — pro роняет контрфактив).
**(ii) COVERAGE / ОМИССИЯ — заголовок «pro омитит» ОТОЗВАН (раунд коррекции):** покрытие **РАВНОЕ: flash_raw 26/27 = pro_raw
26/27**. Первичный «pro 6 омиссий» распался: **3 = 1 truncation** (pro_raw/7.0.s0 finish=length — перегенерена), **2 = локатор-
ложные** на ВАЛИДНЫХ pro-рендерах (a1 s1 «Сказание…не было незнакомо ни одному», d1 s2 «будь цикада не так слаба…» — локаторы
уточнены симметрично), **1 = genuine** (b2 s0 — pro finish=stop, но пропустил предложение 44%). Стохастическая полн-омиссия
существует (sanity-семпл 6.0.s0 выкинул опенинг), но по probe РЕДКА (~1/11) и **НЕ даёт систематического преимущества flash**.
**(iii) EDITOR-RESTORATION (СКОРР.):** «pro restored 5» инфлировано truncation-recovery; **genuine restored = 1** (b2 s0:
pro_raw без 44% → pro_glm вернул «сорока четырёх сотых высоты апертуры» — редактор видит исходник). flash: restored 0, **broke 2**
(c2 s0, d2 s2). Направление реально (билингв-редактор доб. потерянное), но масштаб = 1, не 5.
**(iv) EDITOR-DELTA (T-inv/D38 «редактор ломает верное черновика») — ПОДТВЕРЖДЁН (несущее, read-verified):** editor-delta
**flash 0.185** (breaks **c2, c3**), **pro 0.093** (breaks **c2, d2**). Редактор ввёл **4 fail на flash_glm + 4 на pro_glm** vs
**0 fail на сырых черновиках flash / 2 на pro**. Верифицировано чтением: flash_glm c2 = «Фан Юань — разряд **Б**» (черновик держал
«разряд В»=丙 верно → редактор ИНВЕРТИРОВАЛ грейд + спутал имена «Фан Юань…Фан Юань»); c3 = «Четвёртый **старейшина** рода»
(черновик: «глава рода» верно). Seed-non-unanimity (DET-пол): сырые черновики стабильны (flash 0.0, pro 0.111), **редактор
добавляет вариативность (0.111/0.125)**.
**(v) COGS-дельта (метрика iii, фриз-цены, долина):** flash-черновик **$0.02431** vs pro-черновик **$0.08926** = **×3.67**
(сходится с фриз-оценкой ×3.1). Pro-черновик дороже ×3.67 при НУЛЕВОМ приросте точности (слегка ХУЖЕ на d2), покрытие равное.
**(vi) Судейский мини-пасс (secondary, grok-4.3 reasoning-ON + mistral-large-latest, оба порядка, 6→10, LOO, per-trap seed
где ОБЕ стороны рендерят; 496 голосов, 0 parse-fail, 0 truncation):** ключевой мета-вывод — **судейский пол НА смысл-трапах =
mean|Δ| 0.261** (floor-контраст flash-s0 vs flash-s1, ЭКВИВАЛЕНТНЫЕ входы) — **БОЛЬШЕ, чем S2-пол когезии (0.126)**: судьи
ещё шумнее на этих трапах (a2 0.583, c2 0.792 — огромные Δ на DET-РАВНЫХ черновиках). *Это прямо ВАЛИДИРУЕТ выбор DET-primary:
судьи не смогли бы ответить на Q4a.* Отметка: floor тут = верхняя граница (судейский шум + seed-content-вариация, review F11).
- **main (pro vs flash): diff 0.001, mean|Δ| 0.268 ≈ пол** → преимущества pro НЕТ, и судья слишком шумен, чтобы что-либо
различить (a2 0.958, c2 +1.0 — дикий разброс на DET-равных ячейках, mean washes out). Сходится с DET (оба верны при рендере).
- **editor (pro_raw vs pro_glm): diff 0.102 ≈ DET editor-delta 0.093** → судья И DET СОГЛАСНЫ, что редактор слегка деградирует
pro; **слом d2 подтверждён ОБОИМИ** (судья Δ 0.917, LOO-устойчиво: без-grok 1.0 / без-mistral 0.833, оба судьи arm_acc≈0;
DET=fail). Единственный сигнал выше шума на судейской стороне — editor-слом, совпадает с DET.
### §7-Q4a.5 Дерево решений §D5 применённое (СКОРР.)
**§D5 [INSUFFICIENT-POWER / вакуумно]:** знаменатель «flash-черновик валит смысл-трап» = **0** (flash держит все 9 при рендере,
25fix/0fail) → «pro чинит ≥0.5 flash-провалов» неопределимо. **Вывод: пре-мисса «нужен сильный черновик для верности»
опровергнута** — flash-черновик уже верен на смысле; **pro НЕ точнее, а слегка ХУЖЕ (роняет d2-контрфактив), покрытие равное,
цена ×3.67.** → **арм «структура у переводчика» НЕ ратифицируется; подтверждается путь D38 «сильный редактор»** (свап
glm→mistral/deepseek-pro как редактор): реальные смысл-катастрофы — на **editor-стадии** (c2-грейд 丙→Б, c3-роль 族长→старейшина —
4 fail из чистого черновика; d2-контрфактив), там же и restoration-омиссий. **Каветы (честно):** (1) вакуумность §D5 ХРУПКА —
свежий 3-seed flash держит a1/c1, но САМ прод-бэкенд-черновик (records.json) ИХ ВАЛИТ (a1 инверсия «не знал ни один», c1 «среди
людей») → на шипнутом семпле знаменатель ≥2 и вопрос «чинит ли pro» был бы жив; not-ratify держится на **COGS+нет-прироста**, не
на «flash идеален». (2) **Контрфактив-класс (d1/d2) DET-СЛАБ** (пред-ревью F2 предсказал): локаторы уточнялись пост-хок под lament/
subjunctive-формы — d2-сигнал «pro хуже» вторичен (n мал, класс хрупок); заголовок держат a/b/c + editor-breaks + равное покрытие,
НЕ зависящие от d1/d2. (3) n=9 трапов (honest-power); материал = exp14 in-претрейн канон — ПРЕДВАРИТЕЛЬНО до вебновелл-среза.
Гейтящий эндпоинт остаётся — чтение владельца на большем/трудном корпусе.
### §7-Q4a.6 Деньги (свой кап $2)
Генерация (84 ячейки + 1 перегенер-ячейка + sanity + провенанс-диаг + omission-probe, ВСЁ персистировано, D30.10): **$0.602**.
Судейский мини-пасс (3 контраста × 9 трапов, 496 голосов): **$0.716**. **Итого $1.317 / $2** (маржа **$0.683**). Гейты:
per-call predicted-cost + свой Spender seed=gen-spend, hard-cap $2 (НЕ наследует exp15 seed $9.82/$14.5). Ошибки: генерация
**0/84**, судьи **0 parse-fail / 0 truncation**. Все pro-вызовы — в долине (UTC≈1417).
**A-мини (сокращённый S2-когезия судейский пасс) — ПРОПУЩЕН (решение по правилу оркестратора).** Формально остаток $0.683 ≥
порог $0.5, НО: (1) «ответ его класса известен» (S2-судейский null, §7.4); (2) потребовал бы НОВОЙ S2-pro-генерации (не просто
судейства); (3) Q4a-судейский пол на смысл-трапах (0.261) уже демонстрирует судейскую-шумовую проблему на ЭТОМ материале
СИЛЬНЕЕ, чем S2 (0.126) — A-мини добавил бы шум к чистому результату. «Пропустить без сожалений» (формулировка оркестратора).
### §7-Q4a.7 Самопроверка + ОТКЛОНЕНИЯ (явно)
**Ревью ИСПОЛНЕНИЕМ, 3 рубежа (мандат владельца) — сработал, поймал ИНФЛИРОВАННЫЙ заголовок ДО лендинга:**
(1) **pre-spend 4-линзовый адверс-ревью** (workflow, author≠reviewer) → **11 находок**, все закрыты ДО первого платного вызова:
§D5 без power-гейта → min-denom≥3 + completeness; контрфактив без fail-ветки; span-scoping; polarity_envy ложный fail; b1
or-логика; matrix не-парная → paired headline; budget_truncated контаминация → drop; judge gate-block/ретрай → задокументированы.
(2) **sanity-гейт** поймал стохастическую омиссию + editor-restoration + b2/a1 локатор-гэпы ДО полного прогона → 3 seed +
coverage-метрика. (3) **ПОСТ-ХОК адверс-верификация результатов** (independent agent, author≠reviewer) — **поймала, что заголовок
«pro омитит 6×» ЛОЖЕН**: 3/6 «омиссий» = 1 truncation (finish=length не отклонялся — БАГ рига; «0/84 ошибок» его пропустил),
2/6 = локатор-ложные на валидных pro-рендерах (a1/d1 локаторы под лексикон flash), + моя ручная «верификация чтением» ОШИБЛАСЬ
(спутал truncated s0 с валидным s1). **Фиксы раунда коррекции:** reject finish=length + max_tokens 16000, перегенер truncated-
ячейки, a1/d1/d2-локаторы уточнены симметрично (d2 переанкорен — generic-маркеры ловили чужие контрфактивы). Пере-скор →
покрытие РАВНОЕ, заголовок отозван, ядро вывода усилилось. **Это ровно тот случай (как HEAD_CHARS §7.8), ради которого 2-й
независимый рубеж существует.**
**Отклонения (санкционированы/задекларированы):** материал=exp14 не S2 (санкция оркестратора); редактор=exp14b discourse не
A0 editor.md (матч материала; editor-delta = реальный glm-5); DET-правила = additive-оверрайды (exp14b-риг не тронут);
флэш-сторона перегенерена (не reuse) ради парности; judge seed выбирается per-trap где обе рендерят (омиссия — отдельная
метрика). exp14b DET-правила a1/a2/b1/b2/c1/d1/d2 признаны дефектными ревью и НЕ использованы as-is (пинг оркестратору: если
exp14/exp14b пере-читаются — брать оверрайды из q4a_traps.py, не исходные rule_*).
### §7-Q4a.8 Артефакты
`q4a_det_results.json` (полн. per-trap × seed × cell), `q4a_omission_probe.json`, `q4a_judge_results_{floor,main,editor}.json`,
ледж `q4a_gen_costs.jsonl`/`q4a_judge_costs.jsonl`, ячейки `q4a/{flash,pro}_{raw,glm}/{chunk}.s{seed}.txt`. Скрипты
`eval/exp15/q4a_{traps,generate,score,judge,probe}.py`.

168
eval/exp15/q4a_generate.py Normal file
View file

@ -0,0 +1,168 @@
#!/usr/bin/env python3
"""exp15/Q4a — paired 2x2 generator (B-hybrid, orchestrator-ratified 2026-07-18).
Generates the full square {flash, pro} x {raw draft, +glm-editor}, x2 draft seeds, on the 7 exp14
meaning-trap chunks. PAIR INTEGRITY IS SUPREME (orchestrator pin): both flash and pro drafts get
BYTE-IDENTICAL prompts + injection the exp14 injected_ids reconstruction (bilingual_glossary for
the draft, editor_constraint for the editor), imperfect (no sticky) but SHARED by both sides; we do
NOT "improve" it with the S2' mem_select port. The ONLY variable in the draft row is the draft model.
Assembly (exp14 harness, matches the saved flash-draft provenance):
draft : messages_with_injection(translator_system(), bilingual_glossary, translator_user(src))
deepseek-v4-{flash|pro}, temp 0.3, thinking-ON (never disabled echo-mine mandate)
editor : messages_with_injection(editor_system("discourse"), editor_constraint, editor_user(src,draft))
glm-5, temp 0.4, thinking-OFF (same editor for both sides -> editor-delta is model-clean)
Why regenerate the flash side too (vs reuse on-disk r["draft"]/arm C): the on-disk flash came from
the Go backend (real sticky injection); regenerating flash with the SAME reconstruction assembly as
pro GUARANTEES the flash<->pro pairing the orchestrator ranks above reuse. On-disk backend flash is
retained (records.json / exp14b arm C) as a provenance cross-check, DET-scored in q4a_score.py.
Valley guard: deepseek-v4-pro calls are BLOCKED in DeepSeek peak windows (UTC 01-04 & 06-10, the
x2 surge) abort with a defer message; flash allowed anytime. Per-call predicted-cost gate + own
ledger (every attempt incl. failures persisted, D30.10). Resumable (skip non-empty existing files).
"""
from __future__ import annotations
import argparse
import json
import sys
import time
from datetime import datetime, timezone
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14"))
import exp14_common as C
import exp14_prompts as P
SD = Path("/home/ubuntu/books/gu-zhenren/exp15")
OUT = SD / "q4a"
LEDGER = SD / "q4a_gen_costs.jsonl"
RECS = {(r["chapter"], r["chunk_idx"]): r
for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))}
INJ = json.load(open("/home/ubuntu/books/gu-zhenren/exp14/injection_blocks.json"))
TRAP_CHUNKS = ["6.0", "7.0", "9.1", "10.1", "16.0", "17.0", "19.0"] # 9 meaning-traps live here
SIDES = {"flash": "deepseek-v4-flash", "pro": "deepseek-v4-pro"}
SEEDS = [0, 1, 2]
DRAFT_TEMP, EDIT_TEMP = 0.3, 0.4
# DRAFT_MAX raised 8000->16000: deepseek-v4-pro reasoning can consume the whole 8000 cap and TRUNCATE the
# translation (seen: pro_raw/7.0.s0 finish=length, reasoning starved content -> a false "omission" of the
# lost tail; adversarial verification 2026-07-18). finish=length is now REJECTED in _gen (never accepted).
DRAFT_MAX, EDIT_MAX = 16000, 12000
PEAK_UTC_HOURS = {1, 2, 3, 6, 7, 8, 9} # DeepSeek surge windows (01-04 & 06-10 UTC) — pro forbidden
GEN_HARD_STOP = 1.10 # cumulative gen ceiling ($; leaves room for judging under $2)
CAPS = {"deepseek-v4-flash": 0.03, "deepseek-v4-pro": 0.06, "glm-5": 0.08}
def _now_utc_hour():
return datetime.now(timezone.utc).hour
def _src_inj(cid):
ch, ck = cid.split("."); k = (int(ch), int(ck))
b = INJ.get(f"{ch}/{ck}", {})
return RECS[k]["source"], b.get("bilingual_glossary", ""), b.get("editor_constraint", "")
def _ledger_total():
if not LEDGER.exists():
return 0.0
t = 0.0
for line in LEDGER.read_text(encoding="utf-8").splitlines():
try:
t += json.loads(line).get("cost", 0.0)
except Exception:
pass
return t
def _gen(sp, model, messages, temp, max_out, meta):
"""Gate -> call -> persist EVERY attempt (incl. gate-block/error). Returns text or None."""
in_est = sum(len(m["content"]) for m in messages) // 2
ok, pred = sp.guard(model, in_est, max_out)
if not ok:
sp.record({**meta, "model": model, "keyenv": model, "usage": {},
"err": f"per-call gate pred ${pred:.4f} > cap ${CAPS.get(model)}", "gate_block": True})
print(f" [GATE BLOCK] {meta} pred ${pred:.4f}")
return None
total = _ledger_total()
if total + pred > GEN_HARD_STOP:
print(f" [GEN HARD-STOP] {total:.3f}+{pred:.4f} > {GEN_HARD_STOP} — stopping generation")
return "__STOP__"
s = C.spec(model, temp=temp, max_tokens=max_out)
text, err, usage = C.call(s, messages)
fin = usage.get("finish_reason")
if not err and fin == "length": # truncated output (reasoning ate the cap) — NOT a complete draft
err = f"truncated|finish=length|completion={usage.get('completion_tokens')}"
sp.record({**meta, "model": model, "keyenv": model, "usage": usage, "err": err})
if err or not text:
print(f" [BAD] {meta} err={err} finish={fin}")
return None
return text
def run(limit_chunks=0, sanity=False):
OUT.mkdir(parents=True, exist_ok=True)
for cell in ("flash_raw", "pro_raw", "flash_glm", "pro_glm"):
(OUT / cell).mkdir(exist_ok=True)
sp = C.Spender(LEDGER, CAPS)
# sanity gate: FIRST chunk, seed 0 only, but BOTH sides -> flash+pro draft AND editor (4 cells), so
# the pro/deepseek path (content non-empty, thinking-ON) is exercised before the full valley run.
chunks = TRAP_CHUNKS[:1] if sanity else (TRAP_CHUNKS[:limit_chunks] if limit_chunks else TRAP_CHUNKS)
seeds = [0] if sanity else SEEDS
ed_sys = P.editor_system("discourse")
tr_sys = P.translator_system()
made, skipped = 0, 0
for cid in chunks:
src, gloss, edconstr = _src_inj(cid)
tr_user = P.translator_user(src)
for side, dmodel in SIDES.items():
for seed in seeds:
# ---- draft ----
dpath = OUT / f"{side}_raw" / f"{cid}.s{seed}.txt"
if dpath.exists() and dpath.read_text(encoding="utf-8").strip():
draft = dpath.read_text(encoding="utf-8"); skipped += 1
else:
if dmodel == "deepseek-v4-pro" and _now_utc_hour() in PEAK_UTC_HOURS:
print(f" [VALLEY DEFER] pro draft {cid} s{seed}: UTC hour {_now_utc_hour()} in peak "
f"{sorted(PEAK_UTC_HOURS)} — rerun in valley."); return
msgs = C.messages_with_injection(tr_sys, gloss, tr_user)
draft = _gen(sp, dmodel, msgs, DRAFT_TEMP, DRAFT_MAX,
{"cell": f"{side}_raw", "chunk": cid, "seed": seed, "stage": "draft"})
if draft == "__STOP__":
return
if not draft:
continue
dpath.write_text(draft, encoding="utf-8"); made += 1
print(f" {side}_raw {cid} s{seed}: {len(draft)}c gen-cum ${_ledger_total():.4f}")
# ---- editor over that draft ----
epath = OUT / f"{side}_glm" / f"{cid}.s{seed}.txt"
if epath.exists() and epath.read_text(encoding="utf-8").strip():
skipped += 1
else:
emsgs = C.messages_with_injection(ed_sys, edconstr, P.editor_user(src, draft))
final = _gen(sp, "glm-5", emsgs, EDIT_TEMP, EDIT_MAX,
{"cell": f"{side}_glm", "chunk": cid, "seed": seed, "stage": "edit"})
if final == "__STOP__":
return
if not final:
continue
epath.write_text(final, encoding="utf-8"); made += 1
print(f" {side}_glm {cid} s{seed}: {len(final)}c gen-cum ${_ledger_total():.4f}")
print(f"\ngen done: {made} made, {skipped} reused. gen total ${_ledger_total():.4f}")
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--limit-chunks", type=int, default=0)
ap.add_argument("--sanity", action="store_true",
help="gate: 1 chunk, seed 0, BOTH sides (flash+pro draft+editor = 4 cells)")
a = ap.parse_args()
print(f"[q4a-gen] UTC hour {_now_utc_hour()} (peaks {sorted(PEAK_UTC_HOURS)}); "
f"gen ledger total ${_ledger_total():.4f}; hard-stop ${GEN_HARD_STOP}")
run(a.limit_chunks, a.sanity)
if __name__ == "__main__":
main()

158
eval/exp15/q4a_judge.py Normal file
View file

@ -0,0 +1,158 @@
#!/usr/bin/env python3
"""exp15/Q4a — SECONDARY grok+mistral judge mini-pass (design-compliance + judge-floor calibration
ON the meaning-traps). The $0 DET scorer (q4a_score.py) is PRIMARY and floor-immune; this pass exists
to (1) honour the frozen judge design (grok-4.3 reasoning-ON + mistral-large-latest, both orders,
6->10, LOO, catastrophe, per-vote persist) and (2) MEASURE how well judges even see these traps
does the judge noise-floor (which swamped every S2' block at 0.126) also swamp DET-obvious meaning
errors? Reuses judges.judge_cell + exp15_llm verbatim (NO edits to those rigs).
Contrasts (--contrast):
main : a0=flash-raw(s0), arm=pro-raw(s0) judge's pro-vs-flash draft signal (compare to DET)
floor : a0=flash-raw(s0), arm=flash-raw(s1) two seeds of the SAME model = judge-noise + same-model
seed-content-variance UPPER BOUND (review F11): it is
pure judge noise ONLY where the two seeds DET-agree
(cross-referenced against q4a_score seed_floor).
editor : a0=pro-raw(s0), arm=pro-glm(s0) judge's editor-delta (does editor break the draft)
Judge inputs are WINDOWED around the trap's phenomenon sentence (+/-700 chars, symmetric — flash and
pro translate the SAME source so the sentence sits at ~the same position; addendum §4 authorises
windowing for same-chunk symmetric contrasts) with a full-text fallback when the sentence isn't found.
This keeps the phenomenon visible while holding per-cell cost far under the per-call cap (a full ~5k
draft would push predict near the cap and could truncate the run mid-way).
Budget: OWN Spender, ledger q4a_judge_costs.jsonl, seed_total = current Q4a GEN spend (so the hard cap
is the GLOBAL $2 Q4a ceiling across gen+judge, NOT the exp15 $14.5). budget_truncated cells are DROPPED
from the analysis (review F3): a cell whose votes were cut off by the cap is partial/asymmetric and must
not enter the summary. per_call_cap kept high enough that ONLY the hard-cap governs (review F8). Per-call
gate + every completed vote persisted (judge gate-blocks are $0 and unledgered in the reused judges.py
review F9, documented; the cutoff is recorded via budget_truncated in the results JSON)."""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import exp15_llm as L
import q4a_traps as Q
from judges import judge_cell
SD = Path("/home/ubuntu/books/gu-zhenren/exp15")
OUT = SD / "q4a"
GEN_LEDGER = SD / "q4a_gen_costs.jsonl"
JUDGE_LEDGER = SD / "q4a_judge_costs.jsonl"
Q4A_CAP = 2.0 # GLOBAL Q4a ceiling ($; own budget, OUTSIDE exp15 $15 — orchestrator)
SEEDS = [0, 1, 2]
CONTRAST = { # name -> (a0 cell, arm cell); the SEED is chosen per-trap (see _pick_seed)
"main": ("flash_raw", "pro_raw"), # judge's pro-vs-flash draft signal
"floor": ("flash_raw", "flash_raw"), # same model, two DIFFERENT seeds -> noise+seed-variance floor
"editor": ("pro_raw", "pro_glm"), # judge's editor-delta
}
def _read(cell, cid, seed):
p = OUT / cell / f"{cid}.s{seed}.txt"
return p.read_text(encoding="utf-8") if p.exists() else None
def _located(cell, cid, seed, trap):
txt = _read(cell, cid, seed)
return bool(txt) and bool(Q.sentence_with(txt, trap["locator"]))
def _pick_seed(a0cell, armcell, trap, floor):
"""Per-trap seed selection: judge the meaning GIVEN both sides RENDERED the phenomenon (omission is a
separate DET-coverage metric). Return (a0seed, armseed) or None. For 'floor', a0/arm are the same cell
-> pick two DISTINCT seeds both rendering the trap (else the floor conflates omission with judge noise)."""
cid = trap["chunk"]
if floor:
rs = [s for s in SEEDS if _located(a0cell, cid, s, trap)]
return (rs[0], rs[1]) if len(rs) >= 2 else None
for s in SEEDS: # same seed both sides, both rendering
if _located(a0cell, cid, s, trap) and _located(armcell, cid, s, trap):
return (s, s)
return None
def _window(text, trap, half=700):
"""Symmetric +/-half window around the trap's phenomenon sentence; full text if not located."""
sent = Q.sentence_with(text, trap["locator"])
if not sent:
return text
probe = sent[:40]
i = text.find(probe)
if i < 0:
return text
lo, hi = max(0, i - half), min(len(text), i + len(sent) + half)
return text[lo:hi]
def _gen_spent():
t = 0.0
if GEN_LEDGER.exists():
for line in GEN_LEDGER.read_text(encoding="utf-8").splitlines():
try:
t += json.loads(line).get("cost", 0.0)
except Exception:
pass
return t
def run(contrast, per_call_cap, limit=0):
traps = Q.traps()
if limit:
traps = traps[:limit]
a0cell, armcell = CONTRAST[contrast]
gen = _gen_spent()
sp = L.Spender(str(JUDGE_LEDGER), per_call_cap=per_call_cap, hard_cap=Q4A_CAP, seed_total=gen)
print(f"[budget] Q4a gen spent ${gen:.4f}; judge ledger start ${sp.total:.4f}; global cap ${Q4A_CAP}; "
f"contrast={contrast} (a0={a0cell} arm={armcell}, seed chosen per-trap where both render)")
out, skipped, truncated = [], [], []
for t in traps:
pick = _pick_seed(a0cell, armcell, t, floor=(contrast == "floor"))
if pick is None:
skipped.append((t["id"], "no-seed-both-render")); continue
a0seed, armseed = pick
a0_full = _read(a0cell, t["chunk"], a0seed)
arm_full = _read(armcell, t["chunk"], armseed)
det_a0, _ = Q.det_score(a0_full, t)
det_arm, _ = Q.det_score(arm_full, t)
agg = judge_cell(sp, Q.judge_trap(t), a0_text=_window(a0_full, t), arm_text=_window(arm_full, t))
if agg.get("budget_truncated"):
# DROP a cap-truncated (partial/asymmetric) cell — do NOT let it enter the analysis (review F3)
truncated.append(t["id"])
print(f" [HARD-CAP] Q4a ${Q4A_CAP} bound at {t['id']} — cell dropped, stopping."); break
agg.update(contrast=contrast, id=t["id"], cat_class=t["catastrophe_class"],
a0_seed=a0seed, arm_seed=armseed, det_a0=det_a0, det_arm=det_arm)
out.append(agg)
print(f" {contrast} {t['id']} (s{a0seed}/s{armseed}): judge diff(arm-a0) {agg['paired_diff_arm_minus_a0']} "
f"(a0_acc {agg['a0_accuracy']} arm_acc {agg['arm_accuracy']}; DET a0={det_a0} arm={det_arm})"
f" spend ${sp.total:.4f}")
resfile = SD / f"q4a_judge_results_{contrast}.json"
resfile.write_text(json.dumps({"contrast": contrast, "n": len(out), "skipped": skipped,
"truncated_dropped": truncated, "results": out},
ensure_ascii=False, indent=2), encoding="utf-8")
# quick summary
decided = [a for a in out if a["paired_diff_arm_minus_a0"] is not None]
if decided:
import statistics
diffs = [a["paired_diff_arm_minus_a0"] for a in decided]
print(f"\n{contrast}: n={len(decided)} mean diff {round(statistics.mean(diffs),3)} "
f"mean|diff| {round(statistics.mean(abs(d) for d in diffs),3)} spend ${sp.total:.4f}")
print(f"{contrast}: {len(out)} scored, {len(skipped)} skipped {skipped}, "
f"{len(truncated)} cap-dropped {truncated}")
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--contrast", choices=list(CONTRAST), required=True)
# per-call cap high enough that ONLY the $2 hard-cap governs (windowed inputs predict ~$0.03; review F8)
ap.add_argument("--per-call-cap", type=float, default=0.09)
ap.add_argument("--limit", type=int, default=0)
a = ap.parse_args()
run(a.contrast, a.per_call_cap, a.limit)
if __name__ == "__main__":
main()

69
eval/exp15/q4a_probe.py Normal file
View file

@ -0,0 +1,69 @@
#!/usr/bin/env python3
"""exp15/Q4a — omission characterisation probe (sanity-gate follow-up). The first sanity pro-draft
OMITTED the opening ~13 source paragraphs (aperture/44%/ exposition); a fresh pro call was COMPLETE
-> stochastic omission at temp 0.3, NOT a capture artifact (translation is in content, not reasoning).
This probe samples K drafts of a chunk from BOTH models and measures, per meaning-trap, whether the
trap's phenomenon sentence is RENDERED (locatable) vs OMITTED — to decide seed count + whether omission
is a Q4a signal. $ persisted to the gen ledger (D30.10)."""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14"))
sys.path.insert(0, str(Path(__file__).resolve().parent))
import exp14_common as C
import exp14_prompts as P
import q4a_traps as Q
RECS = {(r["chapter"], r["chunk_idx"]): r
for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))}
INJ = json.load(open("/home/ubuntu/books/gu-zhenren/exp14/injection_blocks.json"))
LEDGER = Path("/home/ubuntu/books/gu-zhenren/exp15/q4a_gen_costs.jsonl")
OUTF = Path("/home/ubuntu/books/gu-zhenren/exp15/q4a_omission_probe.json")
def run(chunks, k):
traps = Q.traps()
results = {}
for cid in chunks:
ch, ck = cid.split("."); src = RECS[(int(ch), int(ck))]["source"]
gloss = INJ[f"{ch}/{ck}"]["bilingual_glossary"]
msgs = C.messages_with_injection(P.translator_system(), gloss, P.translator_user(src))
chunk_traps = [t for t in traps if t["chunk"] == cid]
results[cid] = {"traps": [t["id"] for t in chunk_traps], "samples": {}}
for model in ("deepseek-v4-flash", "deepseek-v4-pro"):
rows = []
for s in range(k):
spec = C.spec(model, temp=0.3, max_tokens=8000)
text, err, usage = C.call(spec, msgs)
with LEDGER.open("a", encoding="utf-8") as f:
f.write(json.dumps({"cell": f"PROBE_{model}", "chunk": cid, "seed": s, "model": model,
"usage": usage, "cost": C.cost_of(model, usage), "err": err},
ensure_ascii=False) + "\n")
if err:
rows.append({"seed": s, "err": err}); continue
# per-trap: rendered (fix/fail) vs omitted ('?')
verds = {t["id"]: Q.det_score(text, t)[0] for t in chunk_traps}
rendered = sum(1 for v in verds.values() if v in ("fix", "fail"))
rows.append({"seed": s, "chars": len(text), "finish": usage.get("finish_reason"),
"rendered": rendered, "n_traps": len(chunk_traps), "verds": verds})
print(f" {model:<18} {cid} s{s}: {len(text)}c finish={usage.get('finish_reason')} "
f"rendered {rendered}/{len(chunk_traps)} {verds}")
results[cid]["samples"][model] = rows
OUTF.write_text(json.dumps(results, ensure_ascii=False, indent=2), encoding="utf-8")
tot = sum(json.loads(l).get("cost", 0) for l in LEDGER.read_text().splitlines())
print(f"\nprobe done -> {OUTF}; gen ledger total ${tot:.4f}")
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--chunks", nargs="+", default=["6.0", "7.0"])
ap.add_argument("--k", type=int, default=5)
a = ap.parse_args()
run(a.chunks, a.k)
if __name__ == "__main__":
main()

294
eval/exp15/q4a_score.py Normal file
View file

@ -0,0 +1,294 @@
#!/usr/bin/env python3
"""exp15/Q4a — DETERMINISTIC scorer (PRIMARY instrument, $0). Reads the paired 2x2 cells generated
by q4a_generate.py, DET-scores each meaning-trap with the exp14b rules (q4a_traps), and produces:
* 2x2 fix-rate matrix {flash,pro} x {raw, +glm-editor}
* §D5 primary rule among traps FLASH-raw FAILS, fraction PRO-raw FIXES (threshold >=0.5)
* editor delta (pro_glm - pro_raw) and (flash_glm - flash_raw): does the editor
PRESERVE or BREAK the draft's fidelity (T-inv = editor property, D38)
* catastrophe class a1/a2/c1 (polarity/rank inversions) tracked separately
* seed variance flash-s0 vs flash-s1 (and pro) disagreement = the DET-level noise floor
(should be ~0 for polarity/number traps -> confirms instrument reliability)
* provenance x-check DET-score the on-disk BACKEND flash draft (records.json) + exp14b arm C,
to confirm the freshly-regenerated flash side agrees with the saved one
* COGS delta pro-draft vs flash-draft stage $ from the gen ledger (metric iii; pro ~x3.1)
Scoring convention: fix=1, fail=0, '?'=undecidable (EXCLUDED from rates, reported separately
addendum §2: no silent drops). Cell score over seeds = mean of decided seeds; cell 'fails' if
score<0.5, 'fixes' if score>=0.5, 'undecided' if all seeds '?'. A cell whose file(s) are MISSING
(not generated / valley-defer / GEN hard-stop) is tracked as 'missing' SEPARATELY from all-'?'
(review F1). §D5 ratify/not-ratify is gated on (i) a complete 2x2 square for the trap and (ii) a
pre-registered minimum decided denominator (>=3); otherwise the verdict is INSUFFICIENT-POWER, never
a clean pass/fail off a tiny partial set. $0 (no api calls)."""
from __future__ import annotations
import json
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import q4a_traps as Q
SD = Path("/home/ubuntu/books/gu-zhenren/exp15")
OUT = SD / "q4a"
GEN_LEDGER = SD / "q4a_gen_costs.jsonl"
CELLS = ["flash_raw", "pro_raw", "flash_glm", "pro_glm"]
SEEDS = [0, 1, 2]
RECS = {(r["chapter"], r["chunk_idx"]): r
for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))}
EX14B = Path("/home/ubuntu/books/gu-zhenren/exp14b/arms")
def _read(cell, cid, seed):
p = OUT / cell / f"{cid}.s{seed}.txt"
return p.read_text(encoding="utf-8") if p.exists() else None
def _cell_verdict(scores):
"""scores = list of 'fix'/'fail'/'?' over seeds -> (label, mean_of_decided|None, n_decided, n_q)."""
dec = [1.0 if s == "fix" else 0.0 for s in scores if s in ("fix", "fail")]
nq = sum(1 for s in scores if s == "?")
if not dec:
return "undecided", None, 0, nq
m = sum(dec) / len(dec)
return ("fix" if m >= 0.5 else "fail"), round(m, 3), len(dec), nq
def score_all():
traps = Q.traps()
per_trap = []
for t in traps:
row = {"id": t["id"], "chunk": t["chunk"], "cat_class": t["catastrophe_class"], "desc": t["desc"],
"cells": {}, "seed_raw": {}}
row["seed_located"] = {}
for cell in CELLS:
seed_scores, located = [], []
for seed in SEEDS:
txt = _read(cell, t["chunk"], seed)
if txt is None:
seed_scores.append(None); located.append(None); continue
v, span = Q.det_score(txt, t)
seed_scores.append(v)
located.append(bool(span)) # sentence found = phenomenon RENDERED (else omitted/locator-miss)
row["seed_raw"][cell] = seed_scores
row["seed_located"][cell] = located
present = [s for s in seed_scores if s is not None]
label, mean, ndec, nq = _cell_verdict(present)
missing = len(present) < len(SEEDS) # some seed file absent (not generated / hard-stop)
if len(present) == 0:
label = "missing"
loc = [x for x in located if x is not None]
row["cells"][cell] = {"label": label, "mean": mean, "n_decided": ndec, "n_q": nq,
"n_present": len(present), "missing": missing,
"n_located": sum(1 for x in loc if x), "n_seen": len(loc)}
# a trap is COMPLETE only if every cell has all seeds present (review F1: completeness gate)
row["complete"] = all(not row["cells"][c]["missing"] for c in CELLS)
row["paired_decided"] = all(row["cells"][c]["mean"] is not None for c in CELLS)
per_trap.append(row)
return traps, per_trap
MIN_DENOM = 3 # pre-registered minimum decided flash-fail denominator for a §D5 ratify/not-ratify verdict
def cell_fixrate(per_trap, cell, cat_only=False, paired_only=False):
"""MARGINAL fix-rate: over this cell's own decided traps (paired_only=True -> only traps decided in
ALL 4 cells, so the four cells are comparable the paired headline, review F6)."""
rows = [r for r in per_trap if (r["cat_class"] if cat_only else True)
and (r["paired_decided"] if paired_only else True)]
dec = [r["cells"][cell]["mean"] for r in rows if r["cells"][cell]["mean"] is not None]
return (round(sum(1 for m in dec if m >= 0.5) / len(dec), 3), len(dec)) if dec else (None, 0)
def d5_rule(per_trap):
"""§D5 (review F1): among COMPLETE traps FLASH-raw FAILS (mean<0.5), fraction PRO-raw FIXES
(mean>=0.5). ratified is a clean bool ONLY when the denominator >= MIN_DENOM AND the 2x2 square is
complete for every counted trap; otherwise ratified=None and status='INSUFFICIENT-POWER'."""
flash_fail, pro_fixes, detail = [], 0, []
incomplete = [r["id"] for r in per_trap if not r["complete"]]
for r in per_trap:
if not r["complete"]:
continue # never let a partial square enter the ratification metric
f = r["cells"]["flash_raw"]["mean"]; p = r["cells"]["pro_raw"]["mean"]
if f is None or p is None:
continue
if f < 0.5: # flash-raw fails this meaning-trap
flash_fail.append(r["id"])
fixed = p >= 0.5
pro_fixes += int(fixed)
detail.append({"id": r["id"], "cat": r["cat_class"], "flash": f, "pro": p, "pro_fixes": fixed})
n = len(flash_fail)
rate = round(pro_fixes / n, 3) if n else None
powered = n >= MIN_DENOM
ratified = (rate >= 0.5) if (powered and rate is not None) else None
status = ("RATIFIED" if ratified else "NOT-RATIFIED") if powered else "INSUFFICIENT-POWER"
return {"flash_fail_traps": flash_fail, "n_denominator": n, "min_denominator": MIN_DENOM,
"pro_fixes": pro_fixes, "fix_rate": rate, "threshold": 0.5, "powered": powered,
"ratified": ratified, "status": status, "incomplete_traps": incomplete, "detail": detail}
def editor_delta(per_trap, side):
"""(side_glm mean - side_raw mean) per trap; positive = editor improves, negative = editor breaks."""
deltas, breaks = [], []
for r in per_trap:
raw = r["cells"][f"{side}_raw"]["mean"]; glm = r["cells"][f"{side}_glm"]["mean"]
if raw is None or glm is None:
continue
d = glm - raw
deltas.append(d)
if d < 0:
breaks.append({"id": r["id"], "raw": raw, "glm": glm})
mean = round(sum(deltas) / len(deltas), 3) if deltas else None
return {"mean_delta": mean, "n": len(deltas), "editor_breaks": breaks}
def seed_floor(per_trap):
"""DET-level noise: per cell, fraction of traps whose decided seeds are NOT unanimous on fix/fail
(>=2 decided seeds required). ~0 expected for polarity/number traps -> confirms DET reliability."""
out = {}
for cell in CELLS:
disagree, n = 0, 0
for r in per_trap:
ss = [s for s in r["seed_raw"][cell] if s in ("fix", "fail")]
if len(ss) >= 2:
n += 1; disagree += int(len(set(ss)) > 1)
out[cell] = {"disagree": disagree, "n_multi_decided": n,
"rate": round(disagree / n, 3) if n else None}
return out
def coverage(per_trap):
"""OMISSION metric (sanity-gate finding): per cell, fraction of traps whose phenomenon sentence is
LOCATED (rendered) vs not (omitted or locator-miss). flash_raw vs pro_raw = does the strong draft
omit MORE? Reported honestly with the caveat that 'not-located' conflates true omission with a
locator gap (symmetric-ish across models)."""
out = {}
for cell in CELLS:
loc = tot = 0
for r in per_trap:
for x in r["seed_located"][cell]:
if x is not None:
tot += 1; loc += int(x)
out[cell] = {"located": loc, "seen": tot, "rate": round(loc / tot, 3) if tot else None}
return out
def editor_restoration(per_trap):
"""Sanity-gate signal: per side, count (trap,seed) where the RAW draft did NOT render the phenomenon
(not located) but the glm editor DID (located) the editor, seeing the source, RESTORED an omission.
Direct evidence the editor stays load-bearing (argues against strong-draft+do-nothing)."""
out = {}
for side in ("flash", "pro"):
restored, broke, both_absent = [], [], 0
for r in per_trap:
rawloc = r["seed_located"][f"{side}_raw"]; glmloc = r["seed_located"][f"{side}_glm"]
for i in SEEDS:
rw, gl = rawloc[i], glmloc[i]
if rw is None or gl is None:
continue
if not rw and gl:
restored.append({"id": r["id"], "seed": i})
elif rw and not gl:
broke.append({"id": r["id"], "seed": i})
elif not rw and not gl:
both_absent += 1
out[side] = {"restored": restored, "n_restored": len(restored),
"broke": broke, "n_broke": len(broke), "both_absent": both_absent}
return out
def provenance_xcheck(traps):
"""DET-score the SAVED backend flash artifacts to confirm the fresh flash side agrees."""
out = {"backend_flash_raw": {}, "exp14b_arm_C": {}}
for t in traps:
ch, ck = t["chunk"].split("."); r = RECS.get((int(ch), int(ck)))
v_bk, _ = Q.det_score(r["draft"] if r else "", t)
out["backend_flash_raw"][t["id"]] = v_bk
cp = EX14B / "C" / f"{t['chunk']}.txt"
v_c, _ = Q.det_score(cp.read_text(encoding="utf-8") if cp.exists() else "", t)
out["exp14b_arm_C"][t["id"]] = v_c
return out
def cogs():
if not GEN_LEDGER.exists():
return {}
agg = {}
for line in GEN_LEDGER.read_text(encoding="utf-8").splitlines():
try:
d = json.loads(line)
except Exception:
continue
cell = d.get("cell", "?"); agg.setdefault(cell, {"cost": 0.0, "n": 0})
agg[cell]["cost"] += d.get("cost", 0.0); agg[cell]["n"] += 1
fr = agg.get("flash_raw", {}).get("cost", 0.0); pr = agg.get("pro_raw", {}).get("cost", 0.0)
return {"by_cell": {k: {"cost": round(v["cost"], 5), "n": v["n"]} for k, v in agg.items()},
"draft_stage_flash": round(fr, 5), "draft_stage_pro": round(pr, 5),
"pro_over_flash_x": round(pr / fr, 2) if fr else None}
def main():
traps, per_trap = score_all()
# MARGINAL matrix (each cell over its own decided traps) + PAIRED matrix (only traps decided in ALL
# 4 cells -> the comparable headline; review F6). Marginal cells can be over DIFFERENT trap subsets.
matrix = {cell: {"marginal": cell_fixrate(per_trap, cell),
"paired": cell_fixrate(per_trap, cell, paired_only=True),
"cat": cell_fixrate(per_trap, cell, cat_only=True)} for cell in CELLS}
n_paired = sum(1 for r in per_trap if r["paired_decided"])
n_complete = sum(1 for r in per_trap if r["complete"])
result = {
"n_traps": len(traps), "seeds": SEEDS,
"n_complete_square": n_complete, "n_paired_decided": n_paired,
"fixrate_matrix": matrix,
"d5_rule": d5_rule(per_trap),
"editor_delta_pro": editor_delta(per_trap, "pro"),
"editor_delta_flash": editor_delta(per_trap, "flash"),
"seed_floor": seed_floor(per_trap),
"coverage": coverage(per_trap),
"editor_restoration": editor_restoration(per_trap),
"provenance_xcheck": provenance_xcheck(traps),
"cogs": cogs(),
"per_trap": per_trap,
}
(SD / "q4a_det_results.json").write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
print(f"=== Q4a DET 2x2 fix-rate matrix (n_traps={result['n_traps']}, complete-square={n_complete}, "
f"paired-decided={n_paired}) ===")
print("cell = PAIRED (all-4-cells-decided; comparable) | marginal (own decided subset):")
print(f"{'':<10}{'raw draft':<30}{'+glm editor':<30}")
for side in ("flash", "pro"):
raw, glm = matrix[f"{side}_raw"], matrix[f"{side}_glm"]
rc = f"{raw['paired']} | {raw['marginal']}"
gc = f"{glm['paired']} | {glm['marginal']}"
print(f" {side:<8}{rc:<30}{gc:<30}")
d5 = result["d5_rule"]
print(f"\n§D5 [{d5['status']}]: among COMPLETE traps flash-raw FAILS {d5['n_denominator']} "
f"{d5['flash_fail_traps']} (min for verdict {d5['min_denominator']}); pro-raw FIXES {d5['pro_fixes']} "
f"-> fix-rate {d5['fix_rate']} (thr 0.5); incomplete traps {d5['incomplete_traps']}")
print(f"editor-delta pro (glm-raw): {result['editor_delta_pro']['mean_delta']} "
f"breaks={[b['id'] for b in result['editor_delta_pro']['editor_breaks']]}")
print(f"editor-delta flash(glm-raw): {result['editor_delta_flash']['mean_delta']} "
f"breaks={[b['id'] for b in result['editor_delta_flash']['editor_breaks']]}")
print(f"seed-floor (DET non-unanimity rate): "
f"{ {c: result['seed_floor'][c]['rate'] for c in CELLS} }")
cov = result["coverage"]
print(f"coverage (phenomenon RENDERED rate; omission proxy): "
f"flash_raw {cov['flash_raw']['rate']} ({cov['flash_raw']['located']}/{cov['flash_raw']['seen']}) "
f"pro_raw {cov['pro_raw']['rate']} ({cov['pro_raw']['located']}/{cov['pro_raw']['seen']}) | "
f"flash_glm {cov['flash_glm']['rate']} pro_glm {cov['pro_glm']['rate']}")
er = result["editor_restoration"]
print(f"editor restoration (raw-omitted -> glm-rendered): "
f"flash restored {er['flash']['n_restored']} broke {er['flash']['n_broke']} | "
f"pro restored {er['pro']['n_restored']} broke {er['pro']['n_broke']}")
print(f"COGS: flash-draft ${result['cogs'].get('draft_stage_flash')} vs "
f"pro-draft ${result['cogs'].get('draft_stage_pro')} (x{result['cogs'].get('pro_over_flash_x')})")
print("\nprovenance x-check (fresh flash_raw vs backend saved):")
for t in traps:
fresh = [per for per in per_trap if per["id"] == t["id"]][0]["cells"]["flash_raw"]["label"]
bk = result["provenance_xcheck"]["backend_flash_raw"][t["id"]]
armc = result["provenance_xcheck"]["exp14b_arm_C"][t["id"]]
flag = "" if fresh == bk else " <-- DIFF"
print(f" {t['id']:<3} fresh={fresh:<10} backend_flash={bk:<10} armC(flash+glm)={armc:<10}{flag}")
if __name__ == "__main__":
main()

259
eval/exp15/q4a_traps.py Normal file
View file

@ -0,0 +1,259 @@
#!/usr/bin/env python3
"""exp15/Q4a — meaning-trap definitions + CORRECTED deterministic scorer (B-hybrid, orchestrator-
ratified 2026-07-18). PRIMARY instrument; $0; floor-immune (unlike the S2' judge blocks).
The Q4a concern (#1): can a strong draft (deepseek-v4-pro) be faithful enough that the editor is
less load-bearing? Sharpest instrument = the exp14b DET meaning-trap battery (double-negation /
number-scale / rank / counterfactual), which is $0-scorable by rule and immune to the judge noise-
floor (0.126) that swamped every S2' block (addendum §1).
RULE CORRECTIONS (frozen here PRE-generation no pro data exists yet, so this is legitimate pre-
registration refinement, driven by the pre-spend adversarial review, NOT by outcomes). We may NOT
edit the existing rig eval/exp14b/exp14b_score.py, so its flawed rules are OVERRIDDEN in this
additive layer; its sound rules (b2/c2/c3) are reused verbatim. Corrections:
* SENTENCE-scoping (review F4): score the single SENTENCE carrying the phenomenon, not the whole
matched line/paragraph so a raw draft (short line) and a glm editor output (merged paragraph)
are scored over comparable spans; an unrelated 'никто не' elsewhere in a glm paragraph can no
longer flip a1/a2.
* a1/a2 double-negation (review F5): a FAITHFUL double-negation ('не было ни одного, кто НЕ знал'
/ 'ни одного взгляда БЕЗ зависти' = everyone) must score FIX, not FAIL checked before the
polarity-inversion fail pattern.
* a1/c1 anchors tightened; the stray 'мёртв'fix branch of rule_polarity_all dropped (review F4).
* counterfactual d1/d2 (review F2): a FAIL branch added (anchor sentence present but the
counterfactual marker 否则'иначе' dropped fail) + 'а то'/'а не то' markers else the class
could never enter the §D5 flash-fail denominator.
* b1 (review F7): FIX now requires BOTH numbers (1/12 AND 30%); a mangled fraction fail.
Scoring: fix / fail / '?' (undecidable: sentence not found or unclassifiable EXCLUDED from rates,
never a silent pass; addendum §2). Catastrophe class a1/a2/c1 = gross polarity/rank inversions."""
from __future__ import annotations
import re
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14b"))
import exp14b_score as S # reuse SOUND rules only: rule_num_b2 (b2), rule_grade_bing (c2), rule_patriarch (c3)
CAT_CLASS = {"a1", "a2", "c1"}
# ── sentence scoping ────────────────────────────────────────────────────────────────
_ENDERS = ".!?…\n"
def sentence_with(text, locator):
"""Return the single sentence of `text` containing the FIRST match of `locator` (regex, re.I).
'' if no match. Bounds on Russian/CJK sentence enders makes raw vs editor spans comparable."""
m = re.search(locator, text or "", re.I)
if not m:
return ""
i = m.start()
lo = max((text.rfind(p, 0, i) for p in _ENDERS), default=-1)
hi_cands = [text.find(p, i) for p in _ENDERS if text.find(p, i) >= 0]
hi = min(hi_cands) if hi_cands else len(text)
return text[lo + 1:hi + 1].strip()
# ── corrected rules (operate on the scoped SENTENCE) ────────────────────────────────
def rule_a1(s):
"""没有一个不知道 = ВСЕ знали. fail = инверсия 'никто не знал'. Double-neg faithful form = fix."""
l = s.lower()
# faithful double-negation FIRST (so it isn't caught by the inversion pattern below).
# covers: 'не было ни одного, кто не знал', 'никто не оставался в неведении',
# 'не было незнакомо / неизвестно ни одному / никому' (было НЕ-незнакомо = знали все).
if re.search(r"(не было|нет|не оставалось|не остал)\s+(ни одного|никого|ни единого|человека)[^.]*(кто|котор|не зна|не ведал|неведени)", l) \
or re.search(r"(никто|ни один)[^.]*не\s+(оставал|остал|был)[^.]*неведени", l) \
or re.search(r"не\s+был[оа]?\s+(не\s*)?(знаком|извест)[а-я]*[^.]*(ни одному|никому|ни для кого|ни один|ни одного)", l):
return "fix"
# polarity INVERSION (nobody knew) — bidirectional: 'ни один … не знал' OR 'не знал … ни один'
if re.search(r"(никто|ни один|ни одного|ни единого)[^.]*не\s*(знал|ведал|слышал|слыхал)", l) \
or re.search(r"не\s*(знал|ведал|слышал|слыхал)[^.]*(ни один|никто|ни одного|ни единого)", l) \
or re.search(r"не\s+был[аои]?\s+извест", l):
return "fail"
# positive faithful (everyone knew)
if re.search(r"(все|всякий|каждый|весь род|в роду|поголовно|любой|всем)[^.]*(знал|знали|ведал|известн)", l) \
or re.search(r"(знал[аи]?|ведал[аи]?|известн[ао]?)[^.]*(все|всякий|каждый|весь|любой)", l):
return "fix"
return "?"
def rule_a2(s):
"""无不…羡慕嫉妒 = ВСЕ полны зависти. fail = 'без зависти/никто не завидовал'. Double-neg = fix."""
l = s.lower()
if not re.search(r"завист|зависти|зависть|ревнос|завидова|зави́д", l):
return "?"
# faithful double-negation ('не было ни одного … без зависти')
if re.search(r"(не было|нет|ни один|ни одного|никого)[^.]*без\s+завист", l):
return "fix"
# inversion (no envy)
if re.search(r"без\s+завист|не было завист|(никто|ни один)[^.]*не\s*завид", l):
return "fail"
return "fix" # envy present, affirmative context
def rule_b1(s):
"""十二分之一 (1/12) AND 三成 (30%). FIX requires BOTH; mangled fraction or 30→3 = fail."""
l = s.lower()
has12 = bool(re.search(r"одн[ауой]{1,2} двенадцат|двенадцат[ауой]{1,2} (часть|долю)|1\s*/\s*12|двенадцатой", l))
has30 = bool(re.search(r"тридцать процент|30\s*%|три десят|три десятых|0[.,]3\b", l))
frac_bad = bool(re.search(r"одн[ауой]{1,2} втор|1\s*/\s*2|половин|одн[ауой]{1,2} трет|1\s*/\s*3|одн[ауой]{1,2} десят", l))
pct_bad = bool(re.search(r"\b3\s*%|\ри процент|три с полов", l)) and not has30
if pct_bad or frac_bad:
return "fail"
if has12 and has30:
return "fix"
return "?"
def rule_b2(s):
"""四成四 = 44% = 44 сотых. FIX: 'сорок(а) четыр[её]х процент/сотых', '44%'. FAIL: 4,4% / 4% (10x low).
Overrides exp14b rule_num_b2, which knew only the 'процент' surface form and missed the 'сотых'
rendering + declined 'сорока' (review: repeated b2 '?' on valid pro renderings)."""
l = s.lower()
if re.search(r"сорок\w*\s+четыр\w*\s*(процент|сот|%)|\b44\s*(процент|%|сот)|0[.,]44", l):
return "fix"
if re.search(r"\b4[.,]4\s*(процент|%)|четыре\s+процент|\b4\s*%|четыре\s+сот\w*\s+высот", l):
return "fail"
return "?"
def rule_c1(s):
"""人上之人 = человек НАД людьми. fail = 'среди людей/один из людей'."""
l = s.lower()
if re.search(r"среди людей|одним из людей|как все|наравне с людьми|обычным человеком|таким же, как", l):
return "fail"
if re.search(r"над людьми|над (простыми|обычными|прочими|остальными)? ?людьми|выше (простых|обычных|других|прочих|остальных|всех)? ?людей|возвыс|человек[а-я]* выше|над смертными|вознес", l):
return "fix"
return "?"
def _counterfactual(s):
"""幸亏…否则 / 否则 contrafactual. Marker present = fix; anchor sentence but marker dropped = fail.
Markers include the SUBJUNCTIVE forms that render 否则 without the literal 'иначе' (review of pro
renderings: 'будь цикада не так слаба…', 'не будь…', 'если бы не…' = faithful counterfactual)."""
l = s.lower()
if re.search(r"иначе|в противном случае|не то бы|а не то|\bа то\b|иным образом|в ином случае|"
r"если бы(?:\s+\w+){0,3}\s+не|\е будь\b|\bбудь\b(?:\s+\w+){0,4}\s+не|в противном разе", l):
return "fix"
return "fail" # the locator already selected the counterfactual sentence; no marker => modality lost
# rule dispatch: corrected rules override the flawed exp14b ones; sound exp14b rules reused verbatim.
RULES = {"a1": rule_a1, "a2": rule_a2, "b1": rule_b1, "b2": rule_b2,
"c1": rule_c1, "c2": S.rule_grade_bing, "c3": S.rule_patriarch,
"d1": _counterfactual, "d2": _counterfactual}
# locator regex to find the phenomenon SENTENCE (broader than the classifier; tightened for a1/c1).
LOCATOR = {
# a1: the LEGEND (典故 of the 4th clan head & Huajiu Xingzhe) that "everyone/no-one-didn't" knew.
# story-words incl. сказани/быль; knowledge-words incl. знаком (не было НЕзнакомо = знали все).
"a1": r"(притч|легенд|истори|предани|типичн|наследи|сказани|быль|молв)[^.]{0,90}(знал|извест|ведал|неведени|знаком)|"
r"(знал|извест|ведал|неведени|знаком)[^.]{0,90}(притч|легенд|Монах|Цветочн|Четвёрт|сказани|быль)",
"a2": r"завист|зависти|зависть|ревнос|завидова",
"b1": r"двенадцат|(истинн\w* ци|真元)[^.]{0,60}(потрач|истрач|消耗)", # 蛊虫祭炼1/12, 真元耗三成
"b2": r"сорок\w*\s+четыр|\b44\b|четыр\w*\s+сот|четыре десят|четырёх десят|половин\w*\s+высот|высот\w*\s+апертур", # 四成四 = 44%
"c1": r"стать человеком|человеком (над|среди|выше)|над людьми|среди людей|одним из людей|"
r"выше .{0,15}людей|над смертными|наравне с людьми",
"c2": r"Фан Юань[^.]{0,60}(разряд|класс|ранг|степен|丙)|(третьего|перв|втор) разряд|разряда [вбаг]|丙",
"c3": r"族长|поколени|старейшин|глав[аеуы].{0,20}(рода|клана|поколени)|патриарх|засад|подстро|ловушк",
# d1: 幸亏春秋蝉虚弱…否则. locate via 'к счастью'/'幸亏' OR a subjunctive counterfactual. NOT the cicada
# entity (in the chapter TITLE) and NOT bare 'повезло' (matches success-context 'наконец повезло').
"d1": r"счастью|благо, что|хорошо, что|幸亏|庆幸|\е будь\b|\bбудь\b\s+\w+\s+не|если бы\s+\w+\s+не",
# d2: 只恨我没有早生几百年,否则…揭破他的丑恶嘴脸. Anchor on the SPECIFIC lament ('жаль/досадно … родился …
# раньше') — generic counterfactual markers (иначе/а то/если бы) collide with OTHER counterfactual and
# villain sentences in this chunk (verified: they grab the wrong span). Fallback: the expose-face content.
"d2": r"(жаль|досадно|обидно|горько|сожал)[^.!?]{0,55}(родил|рожд|появил|на свет)[^.!?]{0,45}"
r"(раньше|ране|столет|сотен лет|веками|прежде|давно)|揭破|丑恶嘴脸|разоблач\w* его|сорв\w* (маск|личин)",
}
ANNOT = {
"a1": dict(src_zone="四代族长和花酒行者的典故,古月族人没有一个不知道的。",
phenomenon="Двойное отрицание 没有一个不知道 ('нет НИ ОДНОГО, кто НЕ знал бы' = знали ВСЕ). "
"Провал = инверсия полярности ('никто не знал').",
expected="Все/каждый в клане знали эту легенду; никто не был в неведении."),
"a2": dict(src_zone="许多人不由自主地看向第一排正襟危坐的古月方正,这可是甲等资质啊,目光中无不充满了羡慕嫉妒的情感。",
phenomenon="无不 (двойное отрицание) — 目光中无不充满羡慕嫉妒 = ВСЕ взгляды полны зависти. "
"Провал = 'без зависти / никто не завидовал'.",
expected="Все взгляды полны зависти и ревности (无不 = поголовно)."),
"b1": dict(src_zone="蛊虫只祭炼了十二分之一,而我的真元却消耗了整整三成。",
phenomenon="十二分之一 = 1/12 (очищено гу); 三成 = 30% (истрачено ци). Провал = искажение дробей.",
expected="Одна двенадцатая (1/12) очищено, истрачено целых тридцать процентов (30%)."),
"b2": dict(src_zone="海面不到空窍的一半高度,只有四成四。",
phenomenon="四成四 = 4,4 десятых = 44%. Провал = '4,4%' или '4%' (занижение ~10×).",
expected="Около сорока четырёх процентов (44%)."),
"c1": dict(src_zone="蛊师能拥有超越凡人的力量,成为人上之人,但是这其中付出的代价,也是高昂的。",
phenomenon="人上之人 = человек НАД людьми. Провал = 'один из людей / среди людей'.",
expected="Стать человеком ВЫШЕ прочих / возвыситься над обычными людьми."),
"c2": dict(src_zone="方源只是个丙等,方正可是甲等资质。",
phenomenon="丙等 = разряд В / третий (ранг Фан Юаня). Провал = разряд Б / второй / A.",
expected="Разряд В (третий, 丙), в противопоставление 甲 (А) у брата."),
"c3": dict(src_zone="当年,四代族长暗算花酒行者不成,战败后又偷袭,虽然击退了后者,但是他也因此身亡。",
phenomenon="四代族长 = Четвёртый ГЛАВА рода/клана (族长). Провал = 'старейшина' (家老).",
expected="Глава рода/клана (族长), а не старейшина."),
"d1": dict(src_zone="庆幸的是,幸亏春秋蝉虚弱到这种程度,否则自己麻烦就大了!",
phenomenon="幸亏…否则 = контрфактив ('к счастью …, ИНАЧЕ была бы беда'). Провал = потеря 'иначе'.",
expected="К счастью (幸亏) …, иначе / в противном случае (否则) была бы беда."),
"d2": dict(src_zone="“只恨我没有早生几百年,否则见到那个魔头,定要拼死揭破他的丑恶嘴脸。”",
phenomenon="否则 = контрфактив ('жаль, что не родился раньше, ИНАЧЕ разоблачил бы демона'). "
"Провал = потеря 'иначе/否则'.",
expected="… иначе / в противном случае (否则) я бы разоблачил его."),
}
# chunk + short label per trap (from exp14b TRAPS)
CHUNK = {tid: cid for tid, cid, *_ in S.TRAPS}
DESC = {tid: desc for tid, _c, _a, _r, desc in S.TRAPS}
def traps():
out = []
for tid in ("a1", "a2", "b1", "b2", "c1", "c2", "c3", "d1", "d2"):
out.append(dict(id=tid, chunk=CHUNK[tid], locator=LOCATOR[tid], rule=RULES[tid],
desc=DESC[tid], catastrophe_class=tid in CAT_CLASS, **ANNOT[tid]))
return out
def det_score(text, trap):
"""Sentence-scoped fix/fail/? . '?' = phenomenon sentence not found OR rule unclassifiable
(never a silent pass). Returns (verdict, scoped_sentence)."""
sent = sentence_with(text, trap["locator"])
if not sent:
return "?", ""
return trap["rule"](sent), sent
def judge_trap(trap):
return {"id": trap["id"], "type": "T-inv" if trap["id"][0] in "acd" else "T-num",
"src_zone": trap["src_zone"], "phenomenon": trap["phenomenon"], "expected": trap["expected"]}
if __name__ == "__main__":
# rule self-test on HAND-WRITTEN fix/fail strings (validates the corrected rules pre-generation)
cases = [
("a1", "Эту легенду в роду Гуюэ знал каждый.", "fix"),
("a1", "О легенде не было ни одного человека, кто не знал бы её.", "fix"),
("a1", "О легенде про главу рода в роду Гуюэ не знал ни один человек.", "fail"),
("a2", "Во всех взглядах читались зависть и ревность.", "fix"),
("a2", "Не было ни одного взгляда без зависти.", "fix"),
("a2", "В их взглядах не было зависти.", "fail"),
("b1", "Гу-червь очищен лишь на одну двенадцатую, а истинная ци потрачена на целых тридцать процентов.", "fix"),
("b1", "Гу-червь очищен наполовину, а истинная ци потрачена на три процента.", "fail"),
("b2", "Вода в море достигает лишь сорока четырёх сотых высоты апертуры.", "fix"),
("b2", "Уровень воды — только сорок четыре процента высоты апертуры.", "fix"),
("b2", "Уровень воды — лишь четыре процента высоты апертуры.", "fail"),
("c1", "Заклинатель гу становится человеком выше прочих людей.", "fix"),
("c1", "Заклинатель гу остаётся одним из людей.", "fail"),
("a1", "Сказание о Четвёртом главе рода не было незнакомо ни одному человеку из рода Гуюэ.", "fix"),
("d1", "К счастью, Цикада так слаба, иначе была бы большая беда.", "fix"),
("d1", "Будь цикада не так чудовищно слаба, у него самого были бы огромные проблемы.", "fix"),
("d1", "К счастью, Цикада оказалась совсем слабой.", "fail"),
("d2", "Жаль, что не родился раньше, а то разоблачил бы того демона.", "fix"),
("d2", "Жаль, что не родился раньше; я бы разоблачил того демона.", "fail"),
]
tby = {t["id"]: t for t in traps()}
ok = 0
for tid, txt, want in cases:
got, sent = det_score(txt, tby[tid])
flag = "OK" if got == want else "**FAIL**"
ok += got == want
print(f" {tid} want={want:<5} got={got:<5} {flag} [{sent[:50]}]")
print(f"\nrule self-test: {ok}/{len(cases)} pass")