Land Q4a: fidelity is not bought at the translate stage, the pro arm is not ratified, the glm editor is the confirmed weak link, with corrected trap rules and the truncation-accepting harness bug documented
This commit is contained in:
parent
94aeeb55fd
commit
cfcf0f08e3
6 changed files with 1093 additions and 0 deletions
|
|
@ -536,3 +536,148 @@ mean(A0−do-nothing) **+0.041** верна, но **TOST ПРОВАЛЕН**: sd
|
|||
**Интерпретация V2-брака (несущая):** армы eval-рига идут БЕЗ прод-гейтов — санитайзер v6 (fold-first) поймал бы 待遇/一步登天 как cjk_leak; брак V2 = (рига-контекст без гейтов) + вероятный механизм «source-carryover подаёт больше исходника в контекст редактора → выше вероятность CJK-протечки» (фиксируется как гипотеза к пере-прогону, где carryover не строится). V2-брак НЕ говорит о carryover как рычаге когезии.
|
||||
|
||||
**Следствия (ратификация D39.8):** (1) вывод D39.7 «сцен-детекцию/carryover-машинерию не строить» подтверждён и на читательском уровне; (2) битва за качество пере-прогона — в ДЕФЕКТ-КЛАССАХ, все банко/пак/гейт-уровня: 时辰-юниты (детерминированная конверсия в пак пары — НОВЫЙ класс), числа-масштабы (b1, свежая эмпирика), gender-enforce по полю сида (Бай Нинбин hidden→male-до-reveal) + диалог-род говорящего (Ф2-annotator), стих/аллюзии (пак-политика + сноски-на-аллюзии lever), регистр-лексикон (негатив-лист «терем»-класса в пак); (3) **инструмент-фиксы будущих слепых пакетов:** авто-QA до человека (CJK-в-ru/битые формы/род-консистентность), обрезка всех версий по общему финальному предложению, 2-way пакеты для тонких контрастов, тир-структура претензий (критические/смысловые/редакторские) вместо плоского «≤2».
|
||||
|
||||
## §7-Q4a — ИСПОЛНЕНИЕ Q4a «сильный переводчик» (отложенный арм, мини-сессия, 2026-07-18)
|
||||
|
||||
> **B-hybrid, санкция оркестратора/владельца 18.07.** Свой кап **$2, ВНЕ $15 exp15**. Скрипты `eval/exp15/q4a_*.py`
|
||||
> (аддитивные; существующие риги НЕ редактировались), артефакты `/home/ubuntu/books/gu-zhenren/exp15/q4a/`, леджеры
|
||||
> `q4a_*.jsonl` — ВСЁ ВНЕ git, лендит оркестратор. Pro-генерация — в DeepSeek-долине (все вызовы UTC≈14–17, вне пиков 01–04/06–10).
|
||||
|
||||
> **⚠ РАУНД КОРРЕКЦИИ (2026-07-18, адверс-верификация результатов, author≠reviewer).** Первый DET-прогон нёс
|
||||
> **ИНФЛИРОВАННЫЙ заголовок «pro омитит 6×»**, пойманный пост-хок адверс-верификатором (как HEAD_CHARS в §7.8): «6 омиссий
|
||||
> pro» = **1 truncation** (pro_raw/7.0.s0 finish=length — reasoning съел max_tokens 8000, хвост с a1/a2/d2 потерян; НЕ
|
||||
> «омиссия модели») **+ 2 локатор-ложных** (a1/d1 локаторы ловили лексикон flash, мимо валидных синонимов pro Сказание/незнакомо,
|
||||
> будь…не) **+ 1 genuine** (b2 s0). Фиксы: max_tokens draft 8000→16000 + reject finish=length; truncated-ячейка перегенерена
|
||||
> (finish=stop); a1/d1/d2-локаторы уточнены симметрично (d2 переанкорен на lament-структуру — generic 'если бы/а то' ловили ЧУЖИЕ
|
||||
> контрфактивы чанка). Self-test 19/19. **Скорректированный итог — ниже; заголовок «pro омитит» ОТОЗВАН.** Ядро вывода (pro НЕ
|
||||
> ратифицируется; редактор = слабое звено) НЕ изменилось и УСИЛИЛОСЬ.
|
||||
|
||||
### §7-Q4a.0 Итог одной строкой (СКОРР.)
|
||||
**Верность НЕ покупается на стадии перевода.** DET-скорер смысл-трапов (floor-immune): **и flash, и pro черновики верны на
|
||||
6/9 трапах** (a1/a2 двойн-отриц, b1/b2 числа, c1 ранг, d1 контрфактив — все 1.0). flash **25 fix / 0 fail / 27 (покрытие
|
||||
26/27)**; **pro НЕ точнее и слегка ХУЖЕ — роняет контрфактив d2** (0.333 vs flash 1.0; 2 fail), **покрытие РАВНОЕ (26/27 = 26/27;
|
||||
differential genuine омиссия pro ≈ 1 — b2 s0, НЕ 6)**, при **COGS ×3.67**. → **Pro-арм НЕ ратифицируется** (§D5-знаменатель ПУСТ: flash не валит ни одного трапа; pro
|
||||
чинить нечего, а сам слегка хуже). Реальные смысл-катастрофы рождает **glm-редактор** (инвертирует 丙→разряд **Б**, 族长→**старейшина**;
|
||||
**4 fail на flash_glm, 4 на pro_glm**; editor-delta flash −0.185, pro −0.093) — прямое подтверждение D38 «слабое звено = редактор,
|
||||
не черновик». Ось — DET-истина + чтение (все ключевые вердикты сверены с реальными рендерами).
|
||||
|
||||
### §7-Q4a.1 Дизайн (B-hybrid; развилка §3-аддендума решена)
|
||||
Концерн №1 «может ли ЧЕРНОВИК быть ~верным сразу». Развилка (A) S2′-когезия-трапы (судьи, floor-уязвимо) vs (B) exp14
|
||||
смысл-трапы (DET, floor-immune) → **владелец/оркестратор выбрали B-hybrid**: материал = 7 exp14-чанков с 9 DET-смысл-трапами
|
||||
exp14b (a1/a2 двойн-отриц = катастроф-класс, b1/b2 числа, c1 ранг = кат, c2/c3 ранг/роль, d1/d2 контрфактив); **PRIMARY =
|
||||
$0 DET-скорер** (иммунен к судейскому полу 0.126, что утопил все S2′-блоки); **SECONDARY = судейский мини-пасс grok+mistral**
|
||||
(калибровка судейского пола НА смысл-трапах + design-комплаенс). Интент пина «те же чанки» = **ПАРНОСТЬ** сравнения — сохранена.
|
||||
Полный **квадрат 2×2** {flash, pro} × {сырой черновик, +glm-редактор}, **×3 seed**, 84 ячейки. draft=`deepseek-v4-pro`
|
||||
(mistral исключён — судейская семья); суб-арм do-nothing = сам pro-черновик; редактор `glm-5`.
|
||||
|
||||
### §7-Q4a.2 Целостность пар — байт-верифицирована (несущее, пин оркестратора «парность превыше»)
|
||||
flash и pro черновики собраны **байт-идентичными** промптом/инъекцией (exp14 `bilingual_glossary`→черновик,
|
||||
`editor_constraint`→редактор; injected_ids-реконструкция БЕЗ sticky — несовершенство, ОБЩЕЕ обеим сторонам, НЕ «улучшалось»
|
||||
mem_select-портом). Подтверждено: `flash_raw` и `pro_raw` каждого чанка — `prompt_tokens`=идентично (напр. 6.0: 1605/1605).
|
||||
Единственная переменная строки черновика = модель; редактор (glm-5 discourse) идентичен обеим сторонам → editor-delta модель-чист.
|
||||
Флэш-сторона **перегенерена** (не reuse on-disk backend-flash) ради ГАРАНТИРОВАННОЙ парности с pro (backend-flash = иная,
|
||||
sticky-инъекция; оставлен как провенанс-кросс-чек). Провенанс: свежий flash_raw DET-сверен с backend-flash + arm C —
|
||||
совпадает, кроме a1/c1 (backend — единичный неудачный семпл с инверсией; по 3 seed flash их держит) → sampling-вариация, не баг.
|
||||
|
||||
### §7-Q4a.3 Инструмент (DET) + честность пре-регистрации
|
||||
Правила exp14b **переписаны как additive-оверрайды** в `q4a_traps.py` (риг exp14b НЕ редактировался): sentence-scoping
|
||||
(raw≈editor сопоставимы), double-neg faithful=fix (a1/a2), **FAIL-ветка контрфактива** (d1/d2 — иначе класс DET-слеп к своему
|
||||
провалу), b1 требует ОБА числа, b2 «сотых»+склонения, c1/a1 anchors сужены, stray 'мёртв'→fix снят. В раунде коррекции: a1/d1
|
||||
локаторы расширены на pro-синонимы (Сказание/незнакомо, будь…не), **d2 переанкорен на lament-структуру** (generic 'если бы/а то'
|
||||
ловили ЧУЖИЕ контрфактивы чанка → ложные вердикты). **Self-test 19/19.** Калибровка на exp14b-армах + probe- И финал-рендерах
|
||||
(та же книга) — **симметрична к модели** (уточнения ловят синонимы ОБЕИХ сторон), **не к исходу** (d2-фикс ВЫЯВИЛ pro-хуже, не
|
||||
подгонял под pro-лучше). '?' исключается из ставок, репортится отдельно (addendum §2). **Кавета: контрфактив-класс (d1/d2)
|
||||
DET-СЛАБ** (пред-ревью F2 предсказал) — предложение-феномен тонет среди других контрфактивов; вердикты d1/d2 ниже-доверие, чем
|
||||
a/b/c. Инструмент-здоровье: flash_raw 25fix/**0fail**/2'?'; **a1/a2/b1/b2/c1/d1 = 1.0 во всех 4 ячейках** → устойчивы.
|
||||
|
||||
### §7-Q4a.4 Результаты (СКОРР. после раунда коррекции)
|
||||
**(i) 2×2 fix-rate (PAIRED, n=9 — трапы decided во ВСЕХ 4 ячейках; marginal совпал). Покрытие = доля из 27 (9×3), где
|
||||
фраза-феномен ЛОКАЛИЗОВАНА (рендер):**
|
||||
|
||||
| | сырой черновик (fix-rate / покрытие) | +glm-редактор (fix-rate / покрытие) |
|
||||
|---|---|---|
|
||||
| **flash** | **1.0** / **0.963** (25fix 0fail; 1 miss = d1-локатор, НЕ омиссия) | 0.778 / 0.889 |
|
||||
| **pro** | **0.889** / **0.963** (23fix **2fail=d2**; 1 genuine омиссия b2 s0) | 0.889 / 1.0 |
|
||||
|
||||
Per-trap: a1/a2/b1/b2/c1/d1 = **1.0 во ВСЕХ 4 ячейках** (оба черновика верны на двойн-отриц, числах, ранге, cicada-контрфактиве).
|
||||
Единственные не-1.0: **c2** (flash_glm 0.0, pro_glm 0.5 — редактор), **c3** (flash_glm 0.333 — редактор), **d2** (pro_raw 0.333,
|
||||
pro_glm 0.0 — pro роняет контрфактив).
|
||||
|
||||
**(ii) COVERAGE / ОМИССИЯ — заголовок «pro омитит» ОТОЗВАН (раунд коррекции):** покрытие **РАВНОЕ: flash_raw 26/27 = pro_raw
|
||||
26/27**. Первичный «pro 6 омиссий» распался: **3 = 1 truncation** (pro_raw/7.0.s0 finish=length — перегенерена), **2 = локатор-
|
||||
ложные** на ВАЛИДНЫХ pro-рендерах (a1 s1 «Сказание…не было незнакомо ни одному», d1 s2 «будь цикада не так слаба…» — локаторы
|
||||
уточнены симметрично), **1 = genuine** (b2 s0 — pro finish=stop, но пропустил предложение 44%). Стохастическая полн-омиссия
|
||||
существует (sanity-семпл 6.0.s0 выкинул опенинг), но по probe РЕДКА (~1/11) и **НЕ даёт систематического преимущества flash**.
|
||||
|
||||
**(iii) EDITOR-RESTORATION (СКОРР.):** «pro restored 5» инфлировано truncation-recovery; **genuine restored = 1** (b2 s0:
|
||||
pro_raw без 44% → pro_glm вернул «сорока четырёх сотых высоты апертуры» — редактор видит исходник). flash: restored 0, **broke 2**
|
||||
(c2 s0, d2 s2). Направление реально (билингв-редактор доб. потерянное), но масштаб = 1, не 5.
|
||||
|
||||
**(iv) EDITOR-DELTA (T-inv/D38 «редактор ломает верное черновика») — ПОДТВЕРЖДЁН (несущее, read-verified):** editor-delta
|
||||
**flash −0.185** (breaks **c2, c3**), **pro −0.093** (breaks **c2, d2**). Редактор ввёл **4 fail на flash_glm + 4 на pro_glm** vs
|
||||
**0 fail на сырых черновиках flash / 2 на pro**. Верифицировано чтением: flash_glm c2 = «Фан Юань — разряд **Б**» (черновик держал
|
||||
«разряд В»=丙 верно → редактор ИНВЕРТИРОВАЛ грейд + спутал имена «Фан Юань…Фан Юань»); c3 = «Четвёртый **старейшина** рода»
|
||||
(черновик: «глава рода» верно). Seed-non-unanimity (DET-пол): сырые черновики стабильны (flash 0.0, pro 0.111), **редактор
|
||||
добавляет вариативность (0.111/0.125)**.
|
||||
|
||||
**(v) COGS-дельта (метрика iii, фриз-цены, долина):** flash-черновик **$0.02431** vs pro-черновик **$0.08926** = **×3.67**
|
||||
(сходится с фриз-оценкой ×3.1). Pro-черновик дороже ×3.67 при НУЛЕВОМ приросте точности (слегка ХУЖЕ на d2), покрытие равное.
|
||||
|
||||
**(vi) Судейский мини-пасс (secondary, grok-4.3 reasoning-ON + mistral-large-latest, оба порядка, 6→10, LOO, per-trap seed
|
||||
где ОБЕ стороны рендерят; 496 голосов, 0 parse-fail, 0 truncation):** ключевой мета-вывод — **судейский пол НА смысл-трапах =
|
||||
mean|Δ| 0.261** (floor-контраст flash-s0 vs flash-s1, ЭКВИВАЛЕНТНЫЕ входы) — **БОЛЬШЕ, чем S2′-пол когезии (0.126)**: судьи
|
||||
ещё шумнее на этих трапах (a2 0.583, c2 0.792 — огромные Δ на DET-РАВНЫХ черновиках). *Это прямо ВАЛИДИРУЕТ выбор DET-primary:
|
||||
судьи не смогли бы ответить на Q4a.* Отметка: floor тут = верхняя граница (судейский шум + seed-content-вариация, review F11).
|
||||
- **main (pro vs flash): diff −0.001, mean|Δ| 0.268 ≈ пол** → преимущества pro НЕТ, и судья слишком шумен, чтобы что-либо
|
||||
различить (a2 −0.958, c2 +1.0 — дикий разброс на DET-равных ячейках, mean washes out). Сходится с DET (оба верны при рендере).
|
||||
- **editor (pro_raw vs pro_glm): diff −0.102 ≈ DET editor-delta −0.093** → судья И DET СОГЛАСНЫ, что редактор слегка деградирует
|
||||
pro; **слом d2 подтверждён ОБОИМИ** (судья Δ −0.917, LOO-устойчиво: без-grok −1.0 / без-mistral −0.833, оба судьи arm_acc≈0;
|
||||
DET=fail). Единственный сигнал выше шума на судейской стороне — editor-слом, совпадает с DET.
|
||||
|
||||
### §7-Q4a.5 Дерево решений §D5 применённое (СКОРР.)
|
||||
**§D5 [INSUFFICIENT-POWER / вакуумно]:** знаменатель «flash-черновик валит смысл-трап» = **0** (flash держит все 9 при рендере,
|
||||
25fix/0fail) → «pro чинит ≥0.5 flash-провалов» неопределимо. **Вывод: пре-мисса «нужен сильный черновик для верности»
|
||||
опровергнута** — flash-черновик уже верен на смысле; **pro НЕ точнее, а слегка ХУЖЕ (роняет d2-контрфактив), покрытие равное,
|
||||
цена ×3.67.** → **арм «структура у переводчика» НЕ ратифицируется; подтверждается путь D38 «сильный редактор»** (свап
|
||||
glm→mistral/deepseek-pro как редактор): реальные смысл-катастрофы — на **editor-стадии** (c2-грейд 丙→Б, c3-роль 族长→старейшина —
|
||||
4 fail из чистого черновика; d2-контрфактив), там же и restoration-омиссий. **Каветы (честно):** (1) вакуумность §D5 ХРУПКА —
|
||||
свежий 3-seed flash держит a1/c1, но САМ прод-бэкенд-черновик (records.json) ИХ ВАЛИТ (a1 инверсия «не знал ни один», c1 «среди
|
||||
людей») → на шипнутом семпле знаменатель ≥2 и вопрос «чинит ли pro» был бы жив; not-ratify держится на **COGS+нет-прироста**, не
|
||||
на «flash идеален». (2) **Контрфактив-класс (d1/d2) DET-СЛАБ** (пред-ревью F2 предсказал): локаторы уточнялись пост-хок под lament/
|
||||
subjunctive-формы — d2-сигнал «pro хуже» вторичен (n мал, класс хрупок); заголовок держат a/b/c + editor-breaks + равное покрытие,
|
||||
НЕ зависящие от d1/d2. (3) n=9 трапов (honest-power); материал = exp14 in-претрейн канон — ПРЕДВАРИТЕЛЬНО до вебновелл-среза.
|
||||
Гейтящий эндпоинт остаётся — чтение владельца на большем/трудном корпусе.
|
||||
|
||||
### §7-Q4a.6 Деньги (свой кап $2)
|
||||
Генерация (84 ячейки + 1 перегенер-ячейка + sanity + провенанс-диаг + omission-probe, ВСЁ персистировано, D30.10): **$0.602**.
|
||||
Судейский мини-пасс (3 контраста × 9 трапов, 496 голосов): **$0.716**. **Итого $1.317 / $2** (маржа **$0.683**). Гейты:
|
||||
per-call predicted-cost + свой Spender seed=gen-spend, hard-cap $2 (НЕ наследует exp15 seed $9.82/$14.5). Ошибки: генерация
|
||||
**0/84**, судьи **0 parse-fail / 0 truncation**. Все pro-вызовы — в долине (UTC≈14–17).
|
||||
**A-мини (сокращённый S2′-когезия судейский пасс) — ПРОПУЩЕН (решение по правилу оркестратора).** Формально остаток $0.683 ≥
|
||||
порог $0.5, НО: (1) «ответ его класса известен» (S2′-судейский null, §7.4); (2) потребовал бы НОВОЙ S2′-pro-генерации (не просто
|
||||
судейства); (3) Q4a-судейский пол на смысл-трапах (0.261) уже демонстрирует судейскую-шумовую проблему на ЭТОМ материале
|
||||
СИЛЬНЕЕ, чем S2′ (0.126) — A-мини добавил бы шум к чистому результату. «Пропустить без сожалений» (формулировка оркестратора).
|
||||
|
||||
### §7-Q4a.7 Самопроверка + ОТКЛОНЕНИЯ (явно)
|
||||
**Ревью ИСПОЛНЕНИЕМ, 3 рубежа (мандат владельца) — сработал, поймал ИНФЛИРОВАННЫЙ заголовок ДО лендинга:**
|
||||
(1) **pre-spend 4-линзовый адверс-ревью** (workflow, author≠reviewer) → **11 находок**, все закрыты ДО первого платного вызова:
|
||||
§D5 без power-гейта → min-denom≥3 + completeness; контрфактив без fail-ветки; span-scoping; polarity_envy ложный fail; b1
|
||||
or-логика; matrix не-парная → paired headline; budget_truncated контаминация → drop; judge gate-block/ретрай → задокументированы.
|
||||
(2) **sanity-гейт** поймал стохастическую омиссию + editor-restoration + b2/a1 локатор-гэпы ДО полного прогона → 3 seed +
|
||||
coverage-метрика. (3) **ПОСТ-ХОК адверс-верификация результатов** (independent agent, author≠reviewer) — **поймала, что заголовок
|
||||
«pro омитит 6×» ЛОЖЕН**: 3/6 «омиссий» = 1 truncation (finish=length не отклонялся — БАГ рига; «0/84 ошибок» его пропустил),
|
||||
2/6 = локатор-ложные на валидных pro-рендерах (a1/d1 локаторы под лексикон flash), + моя ручная «верификация чтением» ОШИБЛАСЬ
|
||||
(спутал truncated s0 с валидным s1). **Фиксы раунда коррекции:** reject finish=length + max_tokens 16000, перегенер truncated-
|
||||
ячейки, a1/d1/d2-локаторы уточнены симметрично (d2 переанкорен — generic-маркеры ловили чужие контрфактивы). Пере-скор →
|
||||
покрытие РАВНОЕ, заголовок отозван, ядро вывода усилилось. **Это ровно тот случай (как HEAD_CHARS §7.8), ради которого 2-й
|
||||
независимый рубеж существует.**
|
||||
**Отклонения (санкционированы/задекларированы):** материал=exp14 не S2′ (санкция оркестратора); редактор=exp14b discourse не
|
||||
A0 editor.md (матч материала; editor-delta = реальный glm-5); DET-правила = additive-оверрайды (exp14b-риг не тронут);
|
||||
флэш-сторона перегенерена (не reuse) ради парности; judge seed выбирается per-trap где обе рендерят (омиссия — отдельная
|
||||
метрика). exp14b DET-правила a1/a2/b1/b2/c1/d1/d2 признаны дефектными ревью и НЕ использованы as-is (пинг оркестратору: если
|
||||
exp14/exp14b пере-читаются — брать оверрайды из q4a_traps.py, не исходные rule_*).
|
||||
|
||||
### §7-Q4a.8 Артефакты
|
||||
`q4a_det_results.json` (полн. per-trap × seed × cell), `q4a_omission_probe.json`, `q4a_judge_results_{floor,main,editor}.json`,
|
||||
ледж `q4a_gen_costs.jsonl`/`q4a_judge_costs.jsonl`, ячейки `q4a/{flash,pro}_{raw,glm}/{chunk}.s{seed}.txt`. Скрипты
|
||||
`eval/exp15/q4a_{traps,generate,score,judge,probe}.py`.
|
||||
|
|
|
|||
168
eval/exp15/q4a_generate.py
Normal file
168
eval/exp15/q4a_generate.py
Normal file
|
|
@ -0,0 +1,168 @@
|
|||
#!/usr/bin/env python3
|
||||
"""exp15/Q4a — paired 2x2 generator (B-hybrid, orchestrator-ratified 2026-07-18).
|
||||
|
||||
Generates the full square {flash, pro} x {raw draft, +glm-editor}, x2 draft seeds, on the 7 exp14
|
||||
meaning-trap chunks. PAIR INTEGRITY IS SUPREME (orchestrator pin): both flash and pro drafts get
|
||||
BYTE-IDENTICAL prompts + injection — the exp14 injected_ids reconstruction (bilingual_glossary for
|
||||
the draft, editor_constraint for the editor), imperfect (no sticky) but SHARED by both sides; we do
|
||||
NOT "improve" it with the S2' mem_select port. The ONLY variable in the draft row is the draft model.
|
||||
|
||||
Assembly (exp14 harness, matches the saved flash-draft provenance):
|
||||
draft : messages_with_injection(translator_system(), bilingual_glossary, translator_user(src))
|
||||
deepseek-v4-{flash|pro}, temp 0.3, thinking-ON (never disabled — echo-mine mandate)
|
||||
editor : messages_with_injection(editor_system("discourse"), editor_constraint, editor_user(src,draft))
|
||||
glm-5, temp 0.4, thinking-OFF (same editor for both sides -> editor-delta is model-clean)
|
||||
|
||||
Why regenerate the flash side too (vs reuse on-disk r["draft"]/arm C): the on-disk flash came from
|
||||
the Go backend (real sticky injection); regenerating flash with the SAME reconstruction assembly as
|
||||
pro GUARANTEES the flash<->pro pairing the orchestrator ranks above reuse. On-disk backend flash is
|
||||
retained (records.json / exp14b arm C) as a provenance cross-check, DET-scored in q4a_score.py.
|
||||
|
||||
Valley guard: deepseek-v4-pro calls are BLOCKED in DeepSeek peak windows (UTC 01-04 & 06-10, the
|
||||
x2 surge) — abort with a defer message; flash allowed anytime. Per-call predicted-cost gate + own
|
||||
ledger (every attempt incl. failures persisted, D30.10). Resumable (skip non-empty existing files).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14"))
|
||||
import exp14_common as C
|
||||
import exp14_prompts as P
|
||||
|
||||
SD = Path("/home/ubuntu/books/gu-zhenren/exp15")
|
||||
OUT = SD / "q4a"
|
||||
LEDGER = SD / "q4a_gen_costs.jsonl"
|
||||
RECS = {(r["chapter"], r["chunk_idx"]): r
|
||||
for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))}
|
||||
INJ = json.load(open("/home/ubuntu/books/gu-zhenren/exp14/injection_blocks.json"))
|
||||
|
||||
TRAP_CHUNKS = ["6.0", "7.0", "9.1", "10.1", "16.0", "17.0", "19.0"] # 9 meaning-traps live here
|
||||
SIDES = {"flash": "deepseek-v4-flash", "pro": "deepseek-v4-pro"}
|
||||
SEEDS = [0, 1, 2]
|
||||
DRAFT_TEMP, EDIT_TEMP = 0.3, 0.4
|
||||
# DRAFT_MAX raised 8000->16000: deepseek-v4-pro reasoning can consume the whole 8000 cap and TRUNCATE the
|
||||
# translation (seen: pro_raw/7.0.s0 finish=length, reasoning starved content -> a false "omission" of the
|
||||
# lost tail; adversarial verification 2026-07-18). finish=length is now REJECTED in _gen (never accepted).
|
||||
DRAFT_MAX, EDIT_MAX = 16000, 12000
|
||||
PEAK_UTC_HOURS = {1, 2, 3, 6, 7, 8, 9} # DeepSeek surge windows (01-04 & 06-10 UTC) — pro forbidden
|
||||
GEN_HARD_STOP = 1.10 # cumulative gen ceiling ($; leaves room for judging under $2)
|
||||
|
||||
CAPS = {"deepseek-v4-flash": 0.03, "deepseek-v4-pro": 0.06, "glm-5": 0.08}
|
||||
|
||||
|
||||
def _now_utc_hour():
|
||||
return datetime.now(timezone.utc).hour
|
||||
|
||||
|
||||
def _src_inj(cid):
|
||||
ch, ck = cid.split("."); k = (int(ch), int(ck))
|
||||
b = INJ.get(f"{ch}/{ck}", {})
|
||||
return RECS[k]["source"], b.get("bilingual_glossary", ""), b.get("editor_constraint", "")
|
||||
|
||||
|
||||
def _ledger_total():
|
||||
if not LEDGER.exists():
|
||||
return 0.0
|
||||
t = 0.0
|
||||
for line in LEDGER.read_text(encoding="utf-8").splitlines():
|
||||
try:
|
||||
t += json.loads(line).get("cost", 0.0)
|
||||
except Exception:
|
||||
pass
|
||||
return t
|
||||
|
||||
|
||||
def _gen(sp, model, messages, temp, max_out, meta):
|
||||
"""Gate -> call -> persist EVERY attempt (incl. gate-block/error). Returns text or None."""
|
||||
in_est = sum(len(m["content"]) for m in messages) // 2
|
||||
ok, pred = sp.guard(model, in_est, max_out)
|
||||
if not ok:
|
||||
sp.record({**meta, "model": model, "keyenv": model, "usage": {},
|
||||
"err": f"per-call gate pred ${pred:.4f} > cap ${CAPS.get(model)}", "gate_block": True})
|
||||
print(f" [GATE BLOCK] {meta} pred ${pred:.4f}")
|
||||
return None
|
||||
total = _ledger_total()
|
||||
if total + pred > GEN_HARD_STOP:
|
||||
print(f" [GEN HARD-STOP] {total:.3f}+{pred:.4f} > {GEN_HARD_STOP} — stopping generation")
|
||||
return "__STOP__"
|
||||
s = C.spec(model, temp=temp, max_tokens=max_out)
|
||||
text, err, usage = C.call(s, messages)
|
||||
fin = usage.get("finish_reason")
|
||||
if not err and fin == "length": # truncated output (reasoning ate the cap) — NOT a complete draft
|
||||
err = f"truncated|finish=length|completion={usage.get('completion_tokens')}"
|
||||
sp.record({**meta, "model": model, "keyenv": model, "usage": usage, "err": err})
|
||||
if err or not text:
|
||||
print(f" [BAD] {meta} err={err} finish={fin}")
|
||||
return None
|
||||
return text
|
||||
|
||||
|
||||
def run(limit_chunks=0, sanity=False):
|
||||
OUT.mkdir(parents=True, exist_ok=True)
|
||||
for cell in ("flash_raw", "pro_raw", "flash_glm", "pro_glm"):
|
||||
(OUT / cell).mkdir(exist_ok=True)
|
||||
sp = C.Spender(LEDGER, CAPS)
|
||||
# sanity gate: FIRST chunk, seed 0 only, but BOTH sides -> flash+pro draft AND editor (4 cells), so
|
||||
# the pro/deepseek path (content non-empty, thinking-ON) is exercised before the full valley run.
|
||||
chunks = TRAP_CHUNKS[:1] if sanity else (TRAP_CHUNKS[:limit_chunks] if limit_chunks else TRAP_CHUNKS)
|
||||
seeds = [0] if sanity else SEEDS
|
||||
ed_sys = P.editor_system("discourse")
|
||||
tr_sys = P.translator_system()
|
||||
made, skipped = 0, 0
|
||||
for cid in chunks:
|
||||
src, gloss, edconstr = _src_inj(cid)
|
||||
tr_user = P.translator_user(src)
|
||||
for side, dmodel in SIDES.items():
|
||||
for seed in seeds:
|
||||
# ---- draft ----
|
||||
dpath = OUT / f"{side}_raw" / f"{cid}.s{seed}.txt"
|
||||
if dpath.exists() and dpath.read_text(encoding="utf-8").strip():
|
||||
draft = dpath.read_text(encoding="utf-8"); skipped += 1
|
||||
else:
|
||||
if dmodel == "deepseek-v4-pro" and _now_utc_hour() in PEAK_UTC_HOURS:
|
||||
print(f" [VALLEY DEFER] pro draft {cid} s{seed}: UTC hour {_now_utc_hour()} in peak "
|
||||
f"{sorted(PEAK_UTC_HOURS)} — rerun in valley."); return
|
||||
msgs = C.messages_with_injection(tr_sys, gloss, tr_user)
|
||||
draft = _gen(sp, dmodel, msgs, DRAFT_TEMP, DRAFT_MAX,
|
||||
{"cell": f"{side}_raw", "chunk": cid, "seed": seed, "stage": "draft"})
|
||||
if draft == "__STOP__":
|
||||
return
|
||||
if not draft:
|
||||
continue
|
||||
dpath.write_text(draft, encoding="utf-8"); made += 1
|
||||
print(f" {side}_raw {cid} s{seed}: {len(draft)}c gen-cum ${_ledger_total():.4f}")
|
||||
# ---- editor over that draft ----
|
||||
epath = OUT / f"{side}_glm" / f"{cid}.s{seed}.txt"
|
||||
if epath.exists() and epath.read_text(encoding="utf-8").strip():
|
||||
skipped += 1
|
||||
else:
|
||||
emsgs = C.messages_with_injection(ed_sys, edconstr, P.editor_user(src, draft))
|
||||
final = _gen(sp, "glm-5", emsgs, EDIT_TEMP, EDIT_MAX,
|
||||
{"cell": f"{side}_glm", "chunk": cid, "seed": seed, "stage": "edit"})
|
||||
if final == "__STOP__":
|
||||
return
|
||||
if not final:
|
||||
continue
|
||||
epath.write_text(final, encoding="utf-8"); made += 1
|
||||
print(f" {side}_glm {cid} s{seed}: {len(final)}c gen-cum ${_ledger_total():.4f}")
|
||||
print(f"\ngen done: {made} made, {skipped} reused. gen total ${_ledger_total():.4f}")
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--limit-chunks", type=int, default=0)
|
||||
ap.add_argument("--sanity", action="store_true",
|
||||
help="gate: 1 chunk, seed 0, BOTH sides (flash+pro draft+editor = 4 cells)")
|
||||
a = ap.parse_args()
|
||||
print(f"[q4a-gen] UTC hour {_now_utc_hour()} (peaks {sorted(PEAK_UTC_HOURS)}); "
|
||||
f"gen ledger total ${_ledger_total():.4f}; hard-stop ${GEN_HARD_STOP}")
|
||||
run(a.limit_chunks, a.sanity)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
158
eval/exp15/q4a_judge.py
Normal file
158
eval/exp15/q4a_judge.py
Normal file
|
|
@ -0,0 +1,158 @@
|
|||
#!/usr/bin/env python3
|
||||
"""exp15/Q4a — SECONDARY grok+mistral judge mini-pass (design-compliance + judge-floor calibration
|
||||
ON the meaning-traps). The $0 DET scorer (q4a_score.py) is PRIMARY and floor-immune; this pass exists
|
||||
to (1) honour the frozen judge design (grok-4.3 reasoning-ON + mistral-large-latest, both orders,
|
||||
6->10, LOO, catastrophe, per-vote persist) and (2) MEASURE how well judges even see these traps —
|
||||
does the judge noise-floor (which swamped every S2' block at 0.126) also swamp DET-obvious meaning
|
||||
errors? Reuses judges.judge_cell + exp15_llm verbatim (NO edits to those rigs).
|
||||
|
||||
Contrasts (--contrast):
|
||||
main : a0=flash-raw(s0), arm=pro-raw(s0) judge's pro-vs-flash draft signal (compare to DET)
|
||||
floor : a0=flash-raw(s0), arm=flash-raw(s1) two seeds of the SAME model = judge-noise + same-model
|
||||
seed-content-variance UPPER BOUND (review F11): it is
|
||||
pure judge noise ONLY where the two seeds DET-agree
|
||||
(cross-referenced against q4a_score seed_floor).
|
||||
editor : a0=pro-raw(s0), arm=pro-glm(s0) judge's editor-delta (does editor break the draft)
|
||||
|
||||
Judge inputs are WINDOWED around the trap's phenomenon sentence (+/-700 chars, symmetric — flash and
|
||||
pro translate the SAME source so the sentence sits at ~the same position; addendum §4 authorises
|
||||
windowing for same-chunk symmetric contrasts) with a full-text fallback when the sentence isn't found.
|
||||
This keeps the phenomenon visible while holding per-cell cost far under the per-call cap (a full ~5k
|
||||
draft would push predict near the cap and could truncate the run mid-way).
|
||||
|
||||
Budget: OWN Spender, ledger q4a_judge_costs.jsonl, seed_total = current Q4a GEN spend (so the hard cap
|
||||
is the GLOBAL $2 Q4a ceiling across gen+judge, NOT the exp15 $14.5). budget_truncated cells are DROPPED
|
||||
from the analysis (review F3): a cell whose votes were cut off by the cap is partial/asymmetric and must
|
||||
not enter the summary. per_call_cap kept high enough that ONLY the hard-cap governs (review F8). Per-call
|
||||
gate + every completed vote persisted (judge gate-blocks are $0 and unledgered in the reused judges.py —
|
||||
review F9, documented; the cutoff is recorded via budget_truncated in the results JSON)."""
|
||||
from __future__ import annotations
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
import exp15_llm as L
|
||||
import q4a_traps as Q
|
||||
from judges import judge_cell
|
||||
|
||||
SD = Path("/home/ubuntu/books/gu-zhenren/exp15")
|
||||
OUT = SD / "q4a"
|
||||
GEN_LEDGER = SD / "q4a_gen_costs.jsonl"
|
||||
JUDGE_LEDGER = SD / "q4a_judge_costs.jsonl"
|
||||
Q4A_CAP = 2.0 # GLOBAL Q4a ceiling ($; own budget, OUTSIDE exp15 $15 — orchestrator)
|
||||
|
||||
SEEDS = [0, 1, 2]
|
||||
CONTRAST = { # name -> (a0 cell, arm cell); the SEED is chosen per-trap (see _pick_seed)
|
||||
"main": ("flash_raw", "pro_raw"), # judge's pro-vs-flash draft signal
|
||||
"floor": ("flash_raw", "flash_raw"), # same model, two DIFFERENT seeds -> noise+seed-variance floor
|
||||
"editor": ("pro_raw", "pro_glm"), # judge's editor-delta
|
||||
}
|
||||
|
||||
|
||||
def _read(cell, cid, seed):
|
||||
p = OUT / cell / f"{cid}.s{seed}.txt"
|
||||
return p.read_text(encoding="utf-8") if p.exists() else None
|
||||
|
||||
|
||||
def _located(cell, cid, seed, trap):
|
||||
txt = _read(cell, cid, seed)
|
||||
return bool(txt) and bool(Q.sentence_with(txt, trap["locator"]))
|
||||
|
||||
|
||||
def _pick_seed(a0cell, armcell, trap, floor):
|
||||
"""Per-trap seed selection: judge the meaning GIVEN both sides RENDERED the phenomenon (omission is a
|
||||
separate DET-coverage metric). Return (a0seed, armseed) or None. For 'floor', a0/arm are the same cell
|
||||
-> pick two DISTINCT seeds both rendering the trap (else the floor conflates omission with judge noise)."""
|
||||
cid = trap["chunk"]
|
||||
if floor:
|
||||
rs = [s for s in SEEDS if _located(a0cell, cid, s, trap)]
|
||||
return (rs[0], rs[1]) if len(rs) >= 2 else None
|
||||
for s in SEEDS: # same seed both sides, both rendering
|
||||
if _located(a0cell, cid, s, trap) and _located(armcell, cid, s, trap):
|
||||
return (s, s)
|
||||
return None
|
||||
|
||||
|
||||
def _window(text, trap, half=700):
|
||||
"""Symmetric +/-half window around the trap's phenomenon sentence; full text if not located."""
|
||||
sent = Q.sentence_with(text, trap["locator"])
|
||||
if not sent:
|
||||
return text
|
||||
probe = sent[:40]
|
||||
i = text.find(probe)
|
||||
if i < 0:
|
||||
return text
|
||||
lo, hi = max(0, i - half), min(len(text), i + len(sent) + half)
|
||||
return text[lo:hi]
|
||||
|
||||
|
||||
def _gen_spent():
|
||||
t = 0.0
|
||||
if GEN_LEDGER.exists():
|
||||
for line in GEN_LEDGER.read_text(encoding="utf-8").splitlines():
|
||||
try:
|
||||
t += json.loads(line).get("cost", 0.0)
|
||||
except Exception:
|
||||
pass
|
||||
return t
|
||||
|
||||
|
||||
def run(contrast, per_call_cap, limit=0):
|
||||
traps = Q.traps()
|
||||
if limit:
|
||||
traps = traps[:limit]
|
||||
a0cell, armcell = CONTRAST[contrast]
|
||||
gen = _gen_spent()
|
||||
sp = L.Spender(str(JUDGE_LEDGER), per_call_cap=per_call_cap, hard_cap=Q4A_CAP, seed_total=gen)
|
||||
print(f"[budget] Q4a gen spent ${gen:.4f}; judge ledger start ${sp.total:.4f}; global cap ${Q4A_CAP}; "
|
||||
f"contrast={contrast} (a0={a0cell} arm={armcell}, seed chosen per-trap where both render)")
|
||||
out, skipped, truncated = [], [], []
|
||||
for t in traps:
|
||||
pick = _pick_seed(a0cell, armcell, t, floor=(contrast == "floor"))
|
||||
if pick is None:
|
||||
skipped.append((t["id"], "no-seed-both-render")); continue
|
||||
a0seed, armseed = pick
|
||||
a0_full = _read(a0cell, t["chunk"], a0seed)
|
||||
arm_full = _read(armcell, t["chunk"], armseed)
|
||||
det_a0, _ = Q.det_score(a0_full, t)
|
||||
det_arm, _ = Q.det_score(arm_full, t)
|
||||
agg = judge_cell(sp, Q.judge_trap(t), a0_text=_window(a0_full, t), arm_text=_window(arm_full, t))
|
||||
if agg.get("budget_truncated"):
|
||||
# DROP a cap-truncated (partial/asymmetric) cell — do NOT let it enter the analysis (review F3)
|
||||
truncated.append(t["id"])
|
||||
print(f" [HARD-CAP] Q4a ${Q4A_CAP} bound at {t['id']} — cell dropped, stopping."); break
|
||||
agg.update(contrast=contrast, id=t["id"], cat_class=t["catastrophe_class"],
|
||||
a0_seed=a0seed, arm_seed=armseed, det_a0=det_a0, det_arm=det_arm)
|
||||
out.append(agg)
|
||||
print(f" {contrast} {t['id']} (s{a0seed}/s{armseed}): judge diff(arm-a0) {agg['paired_diff_arm_minus_a0']} "
|
||||
f"(a0_acc {agg['a0_accuracy']} arm_acc {agg['arm_accuracy']}; DET a0={det_a0} arm={det_arm})"
|
||||
f" spend ${sp.total:.4f}")
|
||||
resfile = SD / f"q4a_judge_results_{contrast}.json"
|
||||
resfile.write_text(json.dumps({"contrast": contrast, "n": len(out), "skipped": skipped,
|
||||
"truncated_dropped": truncated, "results": out},
|
||||
ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
# quick summary
|
||||
decided = [a for a in out if a["paired_diff_arm_minus_a0"] is not None]
|
||||
if decided:
|
||||
import statistics
|
||||
diffs = [a["paired_diff_arm_minus_a0"] for a in decided]
|
||||
print(f"\n{contrast}: n={len(decided)} mean diff {round(statistics.mean(diffs),3)} "
|
||||
f"mean|diff| {round(statistics.mean(abs(d) for d in diffs),3)} spend ${sp.total:.4f}")
|
||||
print(f"{contrast}: {len(out)} scored, {len(skipped)} skipped {skipped}, "
|
||||
f"{len(truncated)} cap-dropped {truncated}")
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--contrast", choices=list(CONTRAST), required=True)
|
||||
# per-call cap high enough that ONLY the $2 hard-cap governs (windowed inputs predict ~$0.03; review F8)
|
||||
ap.add_argument("--per-call-cap", type=float, default=0.09)
|
||||
ap.add_argument("--limit", type=int, default=0)
|
||||
a = ap.parse_args()
|
||||
run(a.contrast, a.per_call_cap, a.limit)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
69
eval/exp15/q4a_probe.py
Normal file
69
eval/exp15/q4a_probe.py
Normal file
|
|
@ -0,0 +1,69 @@
|
|||
#!/usr/bin/env python3
|
||||
"""exp15/Q4a — omission characterisation probe (sanity-gate follow-up). The first sanity pro-draft
|
||||
OMITTED the opening ~13 source paragraphs (aperture/44%/丙 exposition); a fresh pro call was COMPLETE
|
||||
-> stochastic omission at temp 0.3, NOT a capture artifact (translation is in content, not reasoning).
|
||||
This probe samples K drafts of a chunk from BOTH models and measures, per meaning-trap, whether the
|
||||
trap's phenomenon sentence is RENDERED (locatable) vs OMITTED — to decide seed count + whether omission
|
||||
is a Q4a signal. $ persisted to the gen ledger (D30.10)."""
|
||||
from __future__ import annotations
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14"))
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
import exp14_common as C
|
||||
import exp14_prompts as P
|
||||
import q4a_traps as Q
|
||||
|
||||
RECS = {(r["chapter"], r["chunk_idx"]): r
|
||||
for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))}
|
||||
INJ = json.load(open("/home/ubuntu/books/gu-zhenren/exp14/injection_blocks.json"))
|
||||
LEDGER = Path("/home/ubuntu/books/gu-zhenren/exp15/q4a_gen_costs.jsonl")
|
||||
OUTF = Path("/home/ubuntu/books/gu-zhenren/exp15/q4a_omission_probe.json")
|
||||
|
||||
|
||||
def run(chunks, k):
|
||||
traps = Q.traps()
|
||||
results = {}
|
||||
for cid in chunks:
|
||||
ch, ck = cid.split("."); src = RECS[(int(ch), int(ck))]["source"]
|
||||
gloss = INJ[f"{ch}/{ck}"]["bilingual_glossary"]
|
||||
msgs = C.messages_with_injection(P.translator_system(), gloss, P.translator_user(src))
|
||||
chunk_traps = [t for t in traps if t["chunk"] == cid]
|
||||
results[cid] = {"traps": [t["id"] for t in chunk_traps], "samples": {}}
|
||||
for model in ("deepseek-v4-flash", "deepseek-v4-pro"):
|
||||
rows = []
|
||||
for s in range(k):
|
||||
spec = C.spec(model, temp=0.3, max_tokens=8000)
|
||||
text, err, usage = C.call(spec, msgs)
|
||||
with LEDGER.open("a", encoding="utf-8") as f:
|
||||
f.write(json.dumps({"cell": f"PROBE_{model}", "chunk": cid, "seed": s, "model": model,
|
||||
"usage": usage, "cost": C.cost_of(model, usage), "err": err},
|
||||
ensure_ascii=False) + "\n")
|
||||
if err:
|
||||
rows.append({"seed": s, "err": err}); continue
|
||||
# per-trap: rendered (fix/fail) vs omitted ('?')
|
||||
verds = {t["id"]: Q.det_score(text, t)[0] for t in chunk_traps}
|
||||
rendered = sum(1 for v in verds.values() if v in ("fix", "fail"))
|
||||
rows.append({"seed": s, "chars": len(text), "finish": usage.get("finish_reason"),
|
||||
"rendered": rendered, "n_traps": len(chunk_traps), "verds": verds})
|
||||
print(f" {model:<18} {cid} s{s}: {len(text)}c finish={usage.get('finish_reason')} "
|
||||
f"rendered {rendered}/{len(chunk_traps)} {verds}")
|
||||
results[cid]["samples"][model] = rows
|
||||
OUTF.write_text(json.dumps(results, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
tot = sum(json.loads(l).get("cost", 0) for l in LEDGER.read_text().splitlines())
|
||||
print(f"\nprobe done -> {OUTF}; gen ledger total ${tot:.4f}")
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--chunks", nargs="+", default=["6.0", "7.0"])
|
||||
ap.add_argument("--k", type=int, default=5)
|
||||
a = ap.parse_args()
|
||||
run(a.chunks, a.k)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
294
eval/exp15/q4a_score.py
Normal file
294
eval/exp15/q4a_score.py
Normal file
|
|
@ -0,0 +1,294 @@
|
|||
#!/usr/bin/env python3
|
||||
"""exp15/Q4a — DETERMINISTIC scorer (PRIMARY instrument, $0). Reads the paired 2x2 cells generated
|
||||
by q4a_generate.py, DET-scores each meaning-trap with the exp14b rules (q4a_traps), and produces:
|
||||
|
||||
* 2x2 fix-rate matrix {flash,pro} x {raw, +glm-editor}
|
||||
* §D5 primary rule among traps FLASH-raw FAILS, fraction PRO-raw FIXES (threshold >=0.5)
|
||||
* editor delta (pro_glm - pro_raw) and (flash_glm - flash_raw): does the editor
|
||||
PRESERVE or BREAK the draft's fidelity (T-inv = editor property, D38)
|
||||
* catastrophe class a1/a2/c1 (polarity/rank inversions) tracked separately
|
||||
* seed variance flash-s0 vs flash-s1 (and pro) disagreement = the DET-level noise floor
|
||||
(should be ~0 for polarity/number traps -> confirms instrument reliability)
|
||||
* provenance x-check DET-score the on-disk BACKEND flash draft (records.json) + exp14b arm C,
|
||||
to confirm the freshly-regenerated flash side agrees with the saved one
|
||||
* COGS delta pro-draft vs flash-draft stage $ from the gen ledger (metric iii; pro ~x3.1)
|
||||
|
||||
Scoring convention: fix=1, fail=0, '?'=undecidable (EXCLUDED from rates, reported separately —
|
||||
addendum §2: no silent drops). Cell score over seeds = mean of decided seeds; cell 'fails' if
|
||||
score<0.5, 'fixes' if score>=0.5, 'undecided' if all seeds '?'. A cell whose file(s) are MISSING
|
||||
(not generated / valley-defer / GEN hard-stop) is tracked as 'missing' SEPARATELY from all-'?'
|
||||
(review F1). §D5 ratify/not-ratify is gated on (i) a complete 2x2 square for the trap and (ii) a
|
||||
pre-registered minimum decided denominator (>=3); otherwise the verdict is INSUFFICIENT-POWER, never
|
||||
a clean pass/fail off a tiny partial set. $0 (no api calls)."""
|
||||
from __future__ import annotations
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
import q4a_traps as Q
|
||||
|
||||
SD = Path("/home/ubuntu/books/gu-zhenren/exp15")
|
||||
OUT = SD / "q4a"
|
||||
GEN_LEDGER = SD / "q4a_gen_costs.jsonl"
|
||||
CELLS = ["flash_raw", "pro_raw", "flash_glm", "pro_glm"]
|
||||
SEEDS = [0, 1, 2]
|
||||
RECS = {(r["chapter"], r["chunk_idx"]): r
|
||||
for r in json.load(open("/home/ubuntu/books/gu-zhenren/rerun/records.json"))}
|
||||
EX14B = Path("/home/ubuntu/books/gu-zhenren/exp14b/arms")
|
||||
|
||||
|
||||
def _read(cell, cid, seed):
|
||||
p = OUT / cell / f"{cid}.s{seed}.txt"
|
||||
return p.read_text(encoding="utf-8") if p.exists() else None
|
||||
|
||||
|
||||
def _cell_verdict(scores):
|
||||
"""scores = list of 'fix'/'fail'/'?' over seeds -> (label, mean_of_decided|None, n_decided, n_q)."""
|
||||
dec = [1.0 if s == "fix" else 0.0 for s in scores if s in ("fix", "fail")]
|
||||
nq = sum(1 for s in scores if s == "?")
|
||||
if not dec:
|
||||
return "undecided", None, 0, nq
|
||||
m = sum(dec) / len(dec)
|
||||
return ("fix" if m >= 0.5 else "fail"), round(m, 3), len(dec), nq
|
||||
|
||||
|
||||
def score_all():
|
||||
traps = Q.traps()
|
||||
per_trap = []
|
||||
for t in traps:
|
||||
row = {"id": t["id"], "chunk": t["chunk"], "cat_class": t["catastrophe_class"], "desc": t["desc"],
|
||||
"cells": {}, "seed_raw": {}}
|
||||
row["seed_located"] = {}
|
||||
for cell in CELLS:
|
||||
seed_scores, located = [], []
|
||||
for seed in SEEDS:
|
||||
txt = _read(cell, t["chunk"], seed)
|
||||
if txt is None:
|
||||
seed_scores.append(None); located.append(None); continue
|
||||
v, span = Q.det_score(txt, t)
|
||||
seed_scores.append(v)
|
||||
located.append(bool(span)) # sentence found = phenomenon RENDERED (else omitted/locator-miss)
|
||||
row["seed_raw"][cell] = seed_scores
|
||||
row["seed_located"][cell] = located
|
||||
present = [s for s in seed_scores if s is not None]
|
||||
label, mean, ndec, nq = _cell_verdict(present)
|
||||
missing = len(present) < len(SEEDS) # some seed file absent (not generated / hard-stop)
|
||||
if len(present) == 0:
|
||||
label = "missing"
|
||||
loc = [x for x in located if x is not None]
|
||||
row["cells"][cell] = {"label": label, "mean": mean, "n_decided": ndec, "n_q": nq,
|
||||
"n_present": len(present), "missing": missing,
|
||||
"n_located": sum(1 for x in loc if x), "n_seen": len(loc)}
|
||||
# a trap is COMPLETE only if every cell has all seeds present (review F1: completeness gate)
|
||||
row["complete"] = all(not row["cells"][c]["missing"] for c in CELLS)
|
||||
row["paired_decided"] = all(row["cells"][c]["mean"] is not None for c in CELLS)
|
||||
per_trap.append(row)
|
||||
return traps, per_trap
|
||||
|
||||
|
||||
MIN_DENOM = 3 # pre-registered minimum decided flash-fail denominator for a §D5 ratify/not-ratify verdict
|
||||
|
||||
|
||||
def cell_fixrate(per_trap, cell, cat_only=False, paired_only=False):
|
||||
"""MARGINAL fix-rate: over this cell's own decided traps (paired_only=True -> only traps decided in
|
||||
ALL 4 cells, so the four cells are comparable — the paired headline, review F6)."""
|
||||
rows = [r for r in per_trap if (r["cat_class"] if cat_only else True)
|
||||
and (r["paired_decided"] if paired_only else True)]
|
||||
dec = [r["cells"][cell]["mean"] for r in rows if r["cells"][cell]["mean"] is not None]
|
||||
return (round(sum(1 for m in dec if m >= 0.5) / len(dec), 3), len(dec)) if dec else (None, 0)
|
||||
|
||||
|
||||
def d5_rule(per_trap):
|
||||
"""§D5 (review F1): among COMPLETE traps FLASH-raw FAILS (mean<0.5), fraction PRO-raw FIXES
|
||||
(mean>=0.5). ratified is a clean bool ONLY when the denominator >= MIN_DENOM AND the 2x2 square is
|
||||
complete for every counted trap; otherwise ratified=None and status='INSUFFICIENT-POWER'."""
|
||||
flash_fail, pro_fixes, detail = [], 0, []
|
||||
incomplete = [r["id"] for r in per_trap if not r["complete"]]
|
||||
for r in per_trap:
|
||||
if not r["complete"]:
|
||||
continue # never let a partial square enter the ratification metric
|
||||
f = r["cells"]["flash_raw"]["mean"]; p = r["cells"]["pro_raw"]["mean"]
|
||||
if f is None or p is None:
|
||||
continue
|
||||
if f < 0.5: # flash-raw fails this meaning-trap
|
||||
flash_fail.append(r["id"])
|
||||
fixed = p >= 0.5
|
||||
pro_fixes += int(fixed)
|
||||
detail.append({"id": r["id"], "cat": r["cat_class"], "flash": f, "pro": p, "pro_fixes": fixed})
|
||||
n = len(flash_fail)
|
||||
rate = round(pro_fixes / n, 3) if n else None
|
||||
powered = n >= MIN_DENOM
|
||||
ratified = (rate >= 0.5) if (powered and rate is not None) else None
|
||||
status = ("RATIFIED" if ratified else "NOT-RATIFIED") if powered else "INSUFFICIENT-POWER"
|
||||
return {"flash_fail_traps": flash_fail, "n_denominator": n, "min_denominator": MIN_DENOM,
|
||||
"pro_fixes": pro_fixes, "fix_rate": rate, "threshold": 0.5, "powered": powered,
|
||||
"ratified": ratified, "status": status, "incomplete_traps": incomplete, "detail": detail}
|
||||
|
||||
|
||||
def editor_delta(per_trap, side):
|
||||
"""(side_glm mean - side_raw mean) per trap; positive = editor improves, negative = editor breaks."""
|
||||
deltas, breaks = [], []
|
||||
for r in per_trap:
|
||||
raw = r["cells"][f"{side}_raw"]["mean"]; glm = r["cells"][f"{side}_glm"]["mean"]
|
||||
if raw is None or glm is None:
|
||||
continue
|
||||
d = glm - raw
|
||||
deltas.append(d)
|
||||
if d < 0:
|
||||
breaks.append({"id": r["id"], "raw": raw, "glm": glm})
|
||||
mean = round(sum(deltas) / len(deltas), 3) if deltas else None
|
||||
return {"mean_delta": mean, "n": len(deltas), "editor_breaks": breaks}
|
||||
|
||||
|
||||
def seed_floor(per_trap):
|
||||
"""DET-level noise: per cell, fraction of traps whose decided seeds are NOT unanimous on fix/fail
|
||||
(>=2 decided seeds required). ~0 expected for polarity/number traps -> confirms DET reliability."""
|
||||
out = {}
|
||||
for cell in CELLS:
|
||||
disagree, n = 0, 0
|
||||
for r in per_trap:
|
||||
ss = [s for s in r["seed_raw"][cell] if s in ("fix", "fail")]
|
||||
if len(ss) >= 2:
|
||||
n += 1; disagree += int(len(set(ss)) > 1)
|
||||
out[cell] = {"disagree": disagree, "n_multi_decided": n,
|
||||
"rate": round(disagree / n, 3) if n else None}
|
||||
return out
|
||||
|
||||
|
||||
def coverage(per_trap):
|
||||
"""OMISSION metric (sanity-gate finding): per cell, fraction of traps whose phenomenon sentence is
|
||||
LOCATED (rendered) vs not (omitted or locator-miss). flash_raw vs pro_raw = does the strong draft
|
||||
omit MORE? Reported honestly with the caveat that 'not-located' conflates true omission with a
|
||||
locator gap (symmetric-ish across models)."""
|
||||
out = {}
|
||||
for cell in CELLS:
|
||||
loc = tot = 0
|
||||
for r in per_trap:
|
||||
for x in r["seed_located"][cell]:
|
||||
if x is not None:
|
||||
tot += 1; loc += int(x)
|
||||
out[cell] = {"located": loc, "seen": tot, "rate": round(loc / tot, 3) if tot else None}
|
||||
return out
|
||||
|
||||
|
||||
def editor_restoration(per_trap):
|
||||
"""Sanity-gate signal: per side, count (trap,seed) where the RAW draft did NOT render the phenomenon
|
||||
(not located) but the glm editor DID (located) — the editor, seeing the source, RESTORED an omission.
|
||||
Direct evidence the editor stays load-bearing (argues against strong-draft+do-nothing)."""
|
||||
out = {}
|
||||
for side in ("flash", "pro"):
|
||||
restored, broke, both_absent = [], [], 0
|
||||
for r in per_trap:
|
||||
rawloc = r["seed_located"][f"{side}_raw"]; glmloc = r["seed_located"][f"{side}_glm"]
|
||||
for i in SEEDS:
|
||||
rw, gl = rawloc[i], glmloc[i]
|
||||
if rw is None or gl is None:
|
||||
continue
|
||||
if not rw and gl:
|
||||
restored.append({"id": r["id"], "seed": i})
|
||||
elif rw and not gl:
|
||||
broke.append({"id": r["id"], "seed": i})
|
||||
elif not rw and not gl:
|
||||
both_absent += 1
|
||||
out[side] = {"restored": restored, "n_restored": len(restored),
|
||||
"broke": broke, "n_broke": len(broke), "both_absent": both_absent}
|
||||
return out
|
||||
|
||||
|
||||
def provenance_xcheck(traps):
|
||||
"""DET-score the SAVED backend flash artifacts to confirm the fresh flash side agrees."""
|
||||
out = {"backend_flash_raw": {}, "exp14b_arm_C": {}}
|
||||
for t in traps:
|
||||
ch, ck = t["chunk"].split("."); r = RECS.get((int(ch), int(ck)))
|
||||
v_bk, _ = Q.det_score(r["draft"] if r else "", t)
|
||||
out["backend_flash_raw"][t["id"]] = v_bk
|
||||
cp = EX14B / "C" / f"{t['chunk']}.txt"
|
||||
v_c, _ = Q.det_score(cp.read_text(encoding="utf-8") if cp.exists() else "", t)
|
||||
out["exp14b_arm_C"][t["id"]] = v_c
|
||||
return out
|
||||
|
||||
|
||||
def cogs():
|
||||
if not GEN_LEDGER.exists():
|
||||
return {}
|
||||
agg = {}
|
||||
for line in GEN_LEDGER.read_text(encoding="utf-8").splitlines():
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except Exception:
|
||||
continue
|
||||
cell = d.get("cell", "?"); agg.setdefault(cell, {"cost": 0.0, "n": 0})
|
||||
agg[cell]["cost"] += d.get("cost", 0.0); agg[cell]["n"] += 1
|
||||
fr = agg.get("flash_raw", {}).get("cost", 0.0); pr = agg.get("pro_raw", {}).get("cost", 0.0)
|
||||
return {"by_cell": {k: {"cost": round(v["cost"], 5), "n": v["n"]} for k, v in agg.items()},
|
||||
"draft_stage_flash": round(fr, 5), "draft_stage_pro": round(pr, 5),
|
||||
"pro_over_flash_x": round(pr / fr, 2) if fr else None}
|
||||
|
||||
|
||||
def main():
|
||||
traps, per_trap = score_all()
|
||||
# MARGINAL matrix (each cell over its own decided traps) + PAIRED matrix (only traps decided in ALL
|
||||
# 4 cells -> the comparable headline; review F6). Marginal cells can be over DIFFERENT trap subsets.
|
||||
matrix = {cell: {"marginal": cell_fixrate(per_trap, cell),
|
||||
"paired": cell_fixrate(per_trap, cell, paired_only=True),
|
||||
"cat": cell_fixrate(per_trap, cell, cat_only=True)} for cell in CELLS}
|
||||
n_paired = sum(1 for r in per_trap if r["paired_decided"])
|
||||
n_complete = sum(1 for r in per_trap if r["complete"])
|
||||
result = {
|
||||
"n_traps": len(traps), "seeds": SEEDS,
|
||||
"n_complete_square": n_complete, "n_paired_decided": n_paired,
|
||||
"fixrate_matrix": matrix,
|
||||
"d5_rule": d5_rule(per_trap),
|
||||
"editor_delta_pro": editor_delta(per_trap, "pro"),
|
||||
"editor_delta_flash": editor_delta(per_trap, "flash"),
|
||||
"seed_floor": seed_floor(per_trap),
|
||||
"coverage": coverage(per_trap),
|
||||
"editor_restoration": editor_restoration(per_trap),
|
||||
"provenance_xcheck": provenance_xcheck(traps),
|
||||
"cogs": cogs(),
|
||||
"per_trap": per_trap,
|
||||
}
|
||||
(SD / "q4a_det_results.json").write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
|
||||
print(f"=== Q4a DET 2x2 fix-rate matrix (n_traps={result['n_traps']}, complete-square={n_complete}, "
|
||||
f"paired-decided={n_paired}) ===")
|
||||
print("cell = PAIRED (all-4-cells-decided; comparable) | marginal (own decided subset):")
|
||||
print(f"{'':<10}{'raw draft':<30}{'+glm editor':<30}")
|
||||
for side in ("flash", "pro"):
|
||||
raw, glm = matrix[f"{side}_raw"], matrix[f"{side}_glm"]
|
||||
rc = f"{raw['paired']} | {raw['marginal']}"
|
||||
gc = f"{glm['paired']} | {glm['marginal']}"
|
||||
print(f" {side:<8}{rc:<30}{gc:<30}")
|
||||
d5 = result["d5_rule"]
|
||||
print(f"\n§D5 [{d5['status']}]: among COMPLETE traps flash-raw FAILS {d5['n_denominator']} "
|
||||
f"{d5['flash_fail_traps']} (min for verdict {d5['min_denominator']}); pro-raw FIXES {d5['pro_fixes']} "
|
||||
f"-> fix-rate {d5['fix_rate']} (thr 0.5); incomplete traps {d5['incomplete_traps']}")
|
||||
print(f"editor-delta pro (glm-raw): {result['editor_delta_pro']['mean_delta']} "
|
||||
f"breaks={[b['id'] for b in result['editor_delta_pro']['editor_breaks']]}")
|
||||
print(f"editor-delta flash(glm-raw): {result['editor_delta_flash']['mean_delta']} "
|
||||
f"breaks={[b['id'] for b in result['editor_delta_flash']['editor_breaks']]}")
|
||||
print(f"seed-floor (DET non-unanimity rate): "
|
||||
f"{ {c: result['seed_floor'][c]['rate'] for c in CELLS} }")
|
||||
cov = result["coverage"]
|
||||
print(f"coverage (phenomenon RENDERED rate; omission proxy): "
|
||||
f"flash_raw {cov['flash_raw']['rate']} ({cov['flash_raw']['located']}/{cov['flash_raw']['seen']}) "
|
||||
f"pro_raw {cov['pro_raw']['rate']} ({cov['pro_raw']['located']}/{cov['pro_raw']['seen']}) | "
|
||||
f"flash_glm {cov['flash_glm']['rate']} pro_glm {cov['pro_glm']['rate']}")
|
||||
er = result["editor_restoration"]
|
||||
print(f"editor restoration (raw-omitted -> glm-rendered): "
|
||||
f"flash restored {er['flash']['n_restored']} broke {er['flash']['n_broke']} | "
|
||||
f"pro restored {er['pro']['n_restored']} broke {er['pro']['n_broke']}")
|
||||
print(f"COGS: flash-draft ${result['cogs'].get('draft_stage_flash')} vs "
|
||||
f"pro-draft ${result['cogs'].get('draft_stage_pro')} (x{result['cogs'].get('pro_over_flash_x')})")
|
||||
print("\nprovenance x-check (fresh flash_raw vs backend saved):")
|
||||
for t in traps:
|
||||
fresh = [per for per in per_trap if per["id"] == t["id"]][0]["cells"]["flash_raw"]["label"]
|
||||
bk = result["provenance_xcheck"]["backend_flash_raw"][t["id"]]
|
||||
armc = result["provenance_xcheck"]["exp14b_arm_C"][t["id"]]
|
||||
flag = "" if fresh == bk else " <-- DIFF"
|
||||
print(f" {t['id']:<3} fresh={fresh:<10} backend_flash={bk:<10} armC(flash+glm)={armc:<10}{flag}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
259
eval/exp15/q4a_traps.py
Normal file
259
eval/exp15/q4a_traps.py
Normal file
|
|
@ -0,0 +1,259 @@
|
|||
#!/usr/bin/env python3
|
||||
"""exp15/Q4a — meaning-trap definitions + CORRECTED deterministic scorer (B-hybrid, orchestrator-
|
||||
ratified 2026-07-18). PRIMARY instrument; $0; floor-immune (unlike the S2' judge blocks).
|
||||
|
||||
The Q4a concern (#1): can a strong draft (deepseek-v4-pro) be faithful enough that the editor is
|
||||
less load-bearing? Sharpest instrument = the exp14b DET meaning-trap battery (double-negation /
|
||||
number-scale / rank / counterfactual), which is $0-scorable by rule and immune to the judge noise-
|
||||
floor (0.126) that swamped every S2' block (addendum §1).
|
||||
|
||||
RULE CORRECTIONS (frozen here PRE-generation — no pro data exists yet, so this is legitimate pre-
|
||||
registration refinement, driven by the pre-spend adversarial review, NOT by outcomes). We may NOT
|
||||
edit the existing rig eval/exp14b/exp14b_score.py, so its flawed rules are OVERRIDDEN in this
|
||||
additive layer; its sound rules (b2/c2/c3) are reused verbatim. Corrections:
|
||||
* SENTENCE-scoping (review F4): score the single SENTENCE carrying the phenomenon, not the whole
|
||||
matched line/paragraph — so a raw draft (short line) and a glm editor output (merged paragraph)
|
||||
are scored over comparable spans; an unrelated 'никто не' elsewhere in a glm paragraph can no
|
||||
longer flip a1/a2.
|
||||
* a1/a2 double-negation (review F5): a FAITHFUL double-negation ('не было ни одного, кто НЕ знал'
|
||||
/ 'ни одного взгляда БЕЗ зависти' = everyone) must score FIX, not FAIL — checked before the
|
||||
polarity-inversion fail pattern.
|
||||
* a1/c1 anchors tightened; the stray 'мёртв'→fix branch of rule_polarity_all dropped (review F4).
|
||||
* counterfactual d1/d2 (review F2): a FAIL branch added (anchor sentence present but the
|
||||
counterfactual marker 否则→'иначе' dropped → fail) + 'а то'/'а не то' markers — else the class
|
||||
could never enter the §D5 flash-fail denominator.
|
||||
* b1 (review F7): FIX now requires BOTH numbers (1/12 AND 30%); a mangled fraction → fail.
|
||||
|
||||
Scoring: fix / fail / '?' (undecidable: sentence not found or unclassifiable — EXCLUDED from rates,
|
||||
never a silent pass; addendum §2). Catastrophe class a1/a2/c1 = gross polarity/rank inversions."""
|
||||
from __future__ import annotations
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "exp14b"))
|
||||
import exp14b_score as S # reuse SOUND rules only: rule_num_b2 (b2), rule_grade_bing (c2), rule_patriarch (c3)
|
||||
|
||||
CAT_CLASS = {"a1", "a2", "c1"}
|
||||
|
||||
|
||||
# ── sentence scoping ────────────────────────────────────────────────────────────────
|
||||
_ENDERS = ".!?…\n"
|
||||
|
||||
|
||||
def sentence_with(text, locator):
|
||||
"""Return the single sentence of `text` containing the FIRST match of `locator` (regex, re.I).
|
||||
'' if no match. Bounds on Russian/CJK sentence enders — makes raw vs editor spans comparable."""
|
||||
m = re.search(locator, text or "", re.I)
|
||||
if not m:
|
||||
return ""
|
||||
i = m.start()
|
||||
lo = max((text.rfind(p, 0, i) for p in _ENDERS), default=-1)
|
||||
hi_cands = [text.find(p, i) for p in _ENDERS if text.find(p, i) >= 0]
|
||||
hi = min(hi_cands) if hi_cands else len(text)
|
||||
return text[lo + 1:hi + 1].strip()
|
||||
|
||||
|
||||
# ── corrected rules (operate on the scoped SENTENCE) ────────────────────────────────
|
||||
def rule_a1(s):
|
||||
"""没有一个不知道 = ВСЕ знали. fail = инверсия 'никто не знал'. Double-neg faithful form = fix."""
|
||||
l = s.lower()
|
||||
# faithful double-negation FIRST (so it isn't caught by the inversion pattern below).
|
||||
# covers: 'не было ни одного, кто не знал', 'никто не оставался в неведении',
|
||||
# 'не было незнакомо / неизвестно ни одному / никому' (было НЕ-незнакомо = знали все).
|
||||
if re.search(r"(не было|нет|не оставалось|не остал)\s+(ни одного|никого|ни единого|человека)[^.]*(кто|котор|не зна|не ведал|неведени)", l) \
|
||||
or re.search(r"(никто|ни один)[^.]*не\s+(оставал|остал|был)[^.]*неведени", l) \
|
||||
or re.search(r"не\s+был[оа]?\s+(не\s*)?(знаком|извест)[а-я]*[^.]*(ни одному|никому|ни для кого|ни один|ни одного)", l):
|
||||
return "fix"
|
||||
# polarity INVERSION (nobody knew) — bidirectional: 'ни один … не знал' OR 'не знал … ни один'
|
||||
if re.search(r"(никто|ни один|ни одного|ни единого)[^.]*не\s*(знал|ведал|слышал|слыхал)", l) \
|
||||
or re.search(r"не\s*(знал|ведал|слышал|слыхал)[^.]*(ни один|никто|ни одного|ни единого)", l) \
|
||||
or re.search(r"не\s+был[аои]?\s+извест", l):
|
||||
return "fail"
|
||||
# positive faithful (everyone knew)
|
||||
if re.search(r"(все|всякий|каждый|весь род|в роду|поголовно|любой|всем)[^.]*(знал|знали|ведал|известн)", l) \
|
||||
or re.search(r"(знал[аи]?|ведал[аи]?|известн[ао]?)[^.]*(все|всякий|каждый|весь|любой)", l):
|
||||
return "fix"
|
||||
return "?"
|
||||
|
||||
|
||||
def rule_a2(s):
|
||||
"""无不…羡慕嫉妒 = ВСЕ полны зависти. fail = 'без зависти/никто не завидовал'. Double-neg = fix."""
|
||||
l = s.lower()
|
||||
if not re.search(r"завист|зависти|зависть|ревнос|завидова|зави́д", l):
|
||||
return "?"
|
||||
# faithful double-negation ('не было ни одного … без зависти')
|
||||
if re.search(r"(не было|нет|ни один|ни одного|никого)[^.]*без\s+завист", l):
|
||||
return "fix"
|
||||
# inversion (no envy)
|
||||
if re.search(r"без\s+завист|не было завист|(никто|ни один)[^.]*не\s*завид", l):
|
||||
return "fail"
|
||||
return "fix" # envy present, affirmative context
|
||||
|
||||
|
||||
def rule_b1(s):
|
||||
"""十二分之一 (1/12) AND 三成 (30%). FIX requires BOTH; mangled fraction or 30→3 = fail."""
|
||||
l = s.lower()
|
||||
has12 = bool(re.search(r"одн[ауой]{1,2} двенадцат|двенадцат[ауой]{1,2} (часть|долю)|1\s*/\s*12|двенадцатой", l))
|
||||
has30 = bool(re.search(r"тридцать процент|30\s*%|три десят|три десятых|0[.,]3\b", l))
|
||||
frac_bad = bool(re.search(r"одн[ауой]{1,2} втор|1\s*/\s*2|половин|одн[ауой]{1,2} трет|1\s*/\s*3|одн[ауой]{1,2} десят", l))
|
||||
pct_bad = bool(re.search(r"\b3\s*%|\bтри процент|три с полов", l)) and not has30
|
||||
if pct_bad or frac_bad:
|
||||
return "fail"
|
||||
if has12 and has30:
|
||||
return "fix"
|
||||
return "?"
|
||||
|
||||
|
||||
def rule_b2(s):
|
||||
"""四成四 = 44% = 44 сотых. FIX: 'сорок(а) четыр[её]х процент/сотых', '44%'. FAIL: 4,4% / 4% (10x low).
|
||||
Overrides exp14b rule_num_b2, which knew only the 'процент' surface form and missed the 'сотых'
|
||||
rendering + declined 'сорока' (review: repeated b2 '?' on valid pro renderings)."""
|
||||
l = s.lower()
|
||||
if re.search(r"сорок\w*\s+четыр\w*\s*(процент|сот|%)|\b44\s*(процент|%|сот)|0[.,]44", l):
|
||||
return "fix"
|
||||
if re.search(r"\b4[.,]4\s*(процент|%)|четыре\s+процент|\b4\s*%|четыре\s+сот\w*\s+высот", l):
|
||||
return "fail"
|
||||
return "?"
|
||||
|
||||
|
||||
def rule_c1(s):
|
||||
"""人上之人 = человек НАД людьми. fail = 'среди людей/один из людей'."""
|
||||
l = s.lower()
|
||||
if re.search(r"среди людей|одним из людей|как все|наравне с людьми|обычным человеком|таким же, как", l):
|
||||
return "fail"
|
||||
if re.search(r"над людьми|над (простыми|обычными|прочими|остальными)? ?людьми|выше (простых|обычных|других|прочих|остальных|всех)? ?людей|возвыс|человек[а-я]* выше|над смертными|вознес", l):
|
||||
return "fix"
|
||||
return "?"
|
||||
|
||||
|
||||
def _counterfactual(s):
|
||||
"""幸亏…否则 / 否则 contrafactual. Marker present = fix; anchor sentence but marker dropped = fail.
|
||||
Markers include the SUBJUNCTIVE forms that render 否则 without the literal 'иначе' (review of pro
|
||||
renderings: 'будь цикада не так слаба…', 'не будь…', 'если бы не…' = faithful counterfactual)."""
|
||||
l = s.lower()
|
||||
if re.search(r"иначе|в противном случае|не то бы|а не то|\bа то\b|иным образом|в ином случае|"
|
||||
r"если бы(?:\s+\w+){0,3}\s+не|\bне будь\b|\bбудь\b(?:\s+\w+){0,4}\s+не|в противном разе", l):
|
||||
return "fix"
|
||||
return "fail" # the locator already selected the counterfactual sentence; no marker => modality lost
|
||||
|
||||
|
||||
# rule dispatch: corrected rules override the flawed exp14b ones; sound exp14b rules reused verbatim.
|
||||
RULES = {"a1": rule_a1, "a2": rule_a2, "b1": rule_b1, "b2": rule_b2,
|
||||
"c1": rule_c1, "c2": S.rule_grade_bing, "c3": S.rule_patriarch,
|
||||
"d1": _counterfactual, "d2": _counterfactual}
|
||||
|
||||
# locator regex to find the phenomenon SENTENCE (broader than the classifier; tightened for a1/c1).
|
||||
LOCATOR = {
|
||||
# a1: the LEGEND (典故 of the 4th clan head & Huajiu Xingzhe) that "everyone/no-one-didn't" knew.
|
||||
# story-words incl. сказани/быль; knowledge-words incl. знаком (не было НЕзнакомо = знали все).
|
||||
"a1": r"(притч|легенд|истори|предани|типичн|наследи|сказани|быль|молв)[^.]{0,90}(знал|извест|ведал|неведени|знаком)|"
|
||||
r"(знал|извест|ведал|неведени|знаком)[^.]{0,90}(притч|легенд|Монах|Цветочн|Четвёрт|сказани|быль)",
|
||||
"a2": r"завист|зависти|зависть|ревнос|завидова",
|
||||
"b1": r"двенадцат|(истинн\w* ци|真元)[^.]{0,60}(потрач|истрач|消耗)", # 蛊虫祭炼1/12, 真元耗三成
|
||||
"b2": r"сорок\w*\s+четыр|\b44\b|четыр\w*\s+сот|четыре десят|четырёх десят|половин\w*\s+высот|высот\w*\s+апертур", # 四成四 = 44%
|
||||
"c1": r"стать человеком|человеком (над|среди|выше)|над людьми|среди людей|одним из людей|"
|
||||
r"выше .{0,15}людей|над смертными|наравне с людьми",
|
||||
"c2": r"Фан Юань[^.]{0,60}(разряд|класс|ранг|степен|丙)|(третьего|перв|втор) разряд|разряда [вбаг]|丙",
|
||||
"c3": r"族长|поколени|старейшин|глав[аеуы].{0,20}(рода|клана|поколени)|патриарх|засад|подстро|ловушк",
|
||||
# d1: 幸亏春秋蝉虚弱…否则. locate via 'к счастью'/'幸亏' OR a subjunctive counterfactual. NOT the cicada
|
||||
# entity (in the chapter TITLE) and NOT bare 'повезло' (matches success-context 'наконец повезло').
|
||||
"d1": r"счастью|благо, что|хорошо, что|幸亏|庆幸|\bне будь\b|\bбудь\b\s+\w+\s+не|если бы\s+\w+\s+не",
|
||||
# d2: 只恨我没有早生几百年,否则…揭破他的丑恶嘴脸. Anchor on the SPECIFIC lament ('жаль/досадно … родился …
|
||||
# раньше') — generic counterfactual markers (иначе/а то/если бы) collide with OTHER counterfactual and
|
||||
# villain sentences in this chunk (verified: they grab the wrong span). Fallback: the expose-face content.
|
||||
"d2": r"(жаль|досадно|обидно|горько|сожал)[^.!?]{0,55}(родил|рожд|появил|на свет)[^.!?]{0,45}"
|
||||
r"(раньше|ране|столет|сотен лет|веками|прежде|давно)|揭破|丑恶嘴脸|разоблач\w* его|сорв\w* (маск|личин)",
|
||||
}
|
||||
|
||||
ANNOT = {
|
||||
"a1": dict(src_zone="四代族长和花酒行者的典故,古月族人没有一个不知道的。",
|
||||
phenomenon="Двойное отрицание 没有一个不知道 ('нет НИ ОДНОГО, кто НЕ знал бы' = знали ВСЕ). "
|
||||
"Провал = инверсия полярности ('никто не знал').",
|
||||
expected="Все/каждый в клане знали эту легенду; никто не был в неведении."),
|
||||
"a2": dict(src_zone="许多人不由自主地看向第一排正襟危坐的古月方正,这可是甲等资质啊,目光中无不充满了羡慕嫉妒的情感。",
|
||||
phenomenon="无不 (двойное отрицание) — 目光中无不充满羡慕嫉妒 = ВСЕ взгляды полны зависти. "
|
||||
"Провал = 'без зависти / никто не завидовал'.",
|
||||
expected="Все взгляды полны зависти и ревности (无不 = поголовно)."),
|
||||
"b1": dict(src_zone="蛊虫只祭炼了十二分之一,而我的真元却消耗了整整三成。",
|
||||
phenomenon="十二分之一 = 1/12 (очищено гу); 三成 = 30% (истрачено ци). Провал = искажение дробей.",
|
||||
expected="Одна двенадцатая (1/12) очищено, истрачено целых тридцать процентов (30%)."),
|
||||
"b2": dict(src_zone="海面不到空窍的一半高度,只有四成四。",
|
||||
phenomenon="四成四 = 4,4 десятых = 44%. Провал = '4,4%' или '4%' (занижение ~10×).",
|
||||
expected="Около сорока четырёх процентов (44%)."),
|
||||
"c1": dict(src_zone="蛊师能拥有超越凡人的力量,成为人上之人,但是这其中付出的代价,也是高昂的。",
|
||||
phenomenon="人上之人 = человек НАД людьми. Провал = 'один из людей / среди людей'.",
|
||||
expected="Стать человеком ВЫШЕ прочих / возвыситься над обычными людьми."),
|
||||
"c2": dict(src_zone="方源只是个丙等,方正可是甲等资质。",
|
||||
phenomenon="丙等 = разряд В / третий (ранг Фан Юаня). Провал = разряд Б / второй / A.",
|
||||
expected="Разряд В (третий, 丙), в противопоставление 甲 (А) у брата."),
|
||||
"c3": dict(src_zone="当年,四代族长暗算花酒行者不成,战败后又偷袭,虽然击退了后者,但是他也因此身亡。",
|
||||
phenomenon="四代族长 = Четвёртый ГЛАВА рода/клана (族长). Провал = 'старейшина' (家老).",
|
||||
expected="Глава рода/клана (族长), а не старейшина."),
|
||||
"d1": dict(src_zone="庆幸的是,幸亏春秋蝉虚弱到这种程度,否则自己麻烦就大了!",
|
||||
phenomenon="幸亏…否则 = контрфактив ('к счастью …, ИНАЧЕ была бы беда'). Провал = потеря 'иначе'.",
|
||||
expected="К счастью (幸亏) …, иначе / в противном случае (否则) была бы беда."),
|
||||
"d2": dict(src_zone="“只恨我没有早生几百年,否则见到那个魔头,定要拼死揭破他的丑恶嘴脸。”",
|
||||
phenomenon="否则 = контрфактив ('жаль, что не родился раньше, ИНАЧЕ разоблачил бы демона'). "
|
||||
"Провал = потеря 'иначе/否则'.",
|
||||
expected="… иначе / в противном случае (否则) я бы разоблачил его."),
|
||||
}
|
||||
|
||||
# chunk + short label per trap (from exp14b TRAPS)
|
||||
CHUNK = {tid: cid for tid, cid, *_ in S.TRAPS}
|
||||
DESC = {tid: desc for tid, _c, _a, _r, desc in S.TRAPS}
|
||||
|
||||
|
||||
def traps():
|
||||
out = []
|
||||
for tid in ("a1", "a2", "b1", "b2", "c1", "c2", "c3", "d1", "d2"):
|
||||
out.append(dict(id=tid, chunk=CHUNK[tid], locator=LOCATOR[tid], rule=RULES[tid],
|
||||
desc=DESC[tid], catastrophe_class=tid in CAT_CLASS, **ANNOT[tid]))
|
||||
return out
|
||||
|
||||
|
||||
def det_score(text, trap):
|
||||
"""Sentence-scoped fix/fail/? . '?' = phenomenon sentence not found OR rule unclassifiable
|
||||
(never a silent pass). Returns (verdict, scoped_sentence)."""
|
||||
sent = sentence_with(text, trap["locator"])
|
||||
if not sent:
|
||||
return "?", ""
|
||||
return trap["rule"](sent), sent
|
||||
|
||||
|
||||
def judge_trap(trap):
|
||||
return {"id": trap["id"], "type": "T-inv" if trap["id"][0] in "acd" else "T-num",
|
||||
"src_zone": trap["src_zone"], "phenomenon": trap["phenomenon"], "expected": trap["expected"]}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
# rule self-test on HAND-WRITTEN fix/fail strings (validates the corrected rules pre-generation)
|
||||
cases = [
|
||||
("a1", "Эту легенду в роду Гуюэ знал каждый.", "fix"),
|
||||
("a1", "О легенде не было ни одного человека, кто не знал бы её.", "fix"),
|
||||
("a1", "О легенде про главу рода в роду Гуюэ не знал ни один человек.", "fail"),
|
||||
("a2", "Во всех взглядах читались зависть и ревность.", "fix"),
|
||||
("a2", "Не было ни одного взгляда без зависти.", "fix"),
|
||||
("a2", "В их взглядах не было зависти.", "fail"),
|
||||
("b1", "Гу-червь очищен лишь на одну двенадцатую, а истинная ци потрачена на целых тридцать процентов.", "fix"),
|
||||
("b1", "Гу-червь очищен наполовину, а истинная ци потрачена на три процента.", "fail"),
|
||||
("b2", "Вода в море достигает лишь сорока четырёх сотых высоты апертуры.", "fix"),
|
||||
("b2", "Уровень воды — только сорок четыре процента высоты апертуры.", "fix"),
|
||||
("b2", "Уровень воды — лишь четыре процента высоты апертуры.", "fail"),
|
||||
("c1", "Заклинатель гу становится человеком выше прочих людей.", "fix"),
|
||||
("c1", "Заклинатель гу остаётся одним из людей.", "fail"),
|
||||
("a1", "Сказание о Четвёртом главе рода не было незнакомо ни одному человеку из рода Гуюэ.", "fix"),
|
||||
("d1", "К счастью, Цикада так слаба, иначе была бы большая беда.", "fix"),
|
||||
("d1", "Будь цикада не так чудовищно слаба, у него самого были бы огромные проблемы.", "fix"),
|
||||
("d1", "К счастью, Цикада оказалась совсем слабой.", "fail"),
|
||||
("d2", "Жаль, что не родился раньше, а то разоблачил бы того демона.", "fix"),
|
||||
("d2", "Жаль, что не родился раньше; я бы разоблачил того демона.", "fail"),
|
||||
]
|
||||
tby = {t["id"]: t for t in traps()}
|
||||
ok = 0
|
||||
for tid, txt, want in cases:
|
||||
got, sent = det_score(txt, tby[tid])
|
||||
flag = "OK" if got == want else "**FAIL**"
|
||||
ok += got == want
|
||||
print(f" {tid} want={want:<5} got={got:<5} {flag} [{sent[:50]}]")
|
||||
print(f"\nrule self-test: {ok}/{len(cases)} pass")
|
||||
Loading…
Add table
Reference in a new issue