textmachine/eval
2026-07-12 20:55:45 +03:00
..
pilot Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
.gitignore Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
adaptive_incontext.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_levers.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_probe.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_reask.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
build_editor_artifact.py Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions 2026-07-09 19:05:28 +03:00
content_filter_probe.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
cost_model_v2.py Trim grok-reasoning re-narration in exp08 and cost_model: single-source the mechanism in provider-quirks, keep only cost-relevant facts plus pointers 2026-07-05 21:17:06 +03:00
coverage_precision.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
dialogue_precision.py Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
editor_bench.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
en_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
exp12_arms.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_blocks.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_candidate.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_extract.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_glossary_v2.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_judges.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_monitor.py Add exp12 monitor script and bump contract range to D30 in onboarding docs 2026-07-11 05:08:27 +03:00
exp12_pack.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_reflow_demo.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_rootcause.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp13_aggregate.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_blind_stage.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_gates.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_judges.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_monitor.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_seed_signoff.py Land seed v2 signoff transform script: reproducible provenance for owner-signed glossary v2 term statuses (D34.1) 2026-07-11 15:58:46 +03:00
exp13_seed_v2.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_translate.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp14_arms.py Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
exp14_common.py Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor 2026-07-11 22:50:09 +03:00
exp14_judges.py Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
exp14_kpi.py Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor 2026-07-11 22:50:09 +03:00
exp14_material.py Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor 2026-07-11 22:50:09 +03:00
exp14_monitor.py Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
exp14_prompts.py Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor 2026-07-11 22:50:09 +03:00
exp14_rank.py Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
exp14_regate.py Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
exp14_sizecurve.py Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
exp14b_arms.py Commit exp14b harness scripts (arms, judge, monitor, score) to close the reproducibility gap now that the meaning-battery experiment is ratified and closed at D38 2026-07-12 20:55:45 +03:00
exp14b_judge.py Commit exp14b harness scripts (arms, judge, monitor, score) to close the reproducibility gap now that the meaning-battery experiment is ratified and closed at D38 2026-07-12 20:55:45 +03:00
exp14b_monitor.py Commit exp14b harness scripts (arms, judge, monitor, score) to close the reproducibility gap now that the meaning-battery experiment is ratified and closed at D38 2026-07-12 20:55:45 +03:00
exp14b_score.py Commit exp14b harness scripts (arms, judge, monitor, score) to close the reproducibility gap now that the meaning-battery experiment is ratified and closed at D38 2026-07-12 20:55:45 +03:00
explicit_judges.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
extract_bench.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
gu_corpus_build.py Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
ja_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
jpm_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
llama_moe_bench.sh Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
local_bench.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
memory_hotpath.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
mistral_fidelity.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
ollama-up.sh Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
providers.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_erotica.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_explicit.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_grok_reason.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
providers_judges.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
providers_local8b.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
README.md Add experiments index and extend eval script map over the closed quality-arc families, keeping scripts in place because the D-log and reports cite them by path 2026-07-12 20:55:45 +03:00
refusal_bench.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
refusal_corpus_build.py Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
retrieval_bench.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
run_refusal.sh Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup 2026-07-04 21:06:47 +03:00
token_calc.py Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
webnovel_slice.py Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions 2026-07-09 19:05:28 +03:00

eval/ — полигон TextMachine

Зона сессии «Полигон»: эмпирическая валидация допущений архитектуры (мерить, а не верить). Отчёты — docs/experiments/NN-*.md (методика + цифры + честные ограничения). backend/ только читать; расхождения с architecture/research — пингом оркестратору в docs/PROGRESS.md. .env не читать; data/refusal_corpus/* не открывать без прямой задачи владельца.

Карта

Скрипт Эксперимент Заметки
refusal_bench.py exp02; его split_sentences/classify_output = приёмочный оракул coverage-гейта Go (паритет закреплён Go-тестом; любое изменение — Python-first, лок-степ) run_refusal.sh — однокоманда
token_calc.py exp01 токен-калибровка (r-множители) токенизаторы в tokenizers/
editor_bench.py + build_editor_artifact.py exp04 бейк-офф редакторов слепота починена (D13.5): рандомизация меток ПО ФРАГМЕНТУ (соль), ключ+цена в отдельном --key файле. Выбор grok-4.3 провизорен — exp04 §Оговорки
coverage_precision.py, content_filter_probe.py exp07 precision гейта (0 FP/57)
cost_model_v2.py exp08 COGS (~$0.69/ранобэ) цены датировать и перепроверять
extract_bench.py exp06 экстракция терминов (спот=локаль/рендер=облако) health-верификация каждого ответа — образец
retrieval_bench.py эмбеддинг-бенч памяти (bge-m3 дефолт)
pilot/declmetric.py, pilot/memory_eval.py, pilot/kana_precision.py харнесс пилота Ф2.5 (exp09) починен (D13.5): эхо-контроль, ПОЛНЫЕ выходы, fidelity-ось (cross-family судья), M0M3, Go-синк 7 расхождений (self-test+адверсариально verified). --self-test/--dry-run без API. Инъекция = зеркало memory.go — синк с Go v3 выполнен (пакет №2: suppressUnboundedPhonetic + апостроф-фолд, adversarial parity); при расхождении верить Go. kana_traps.json+trad2simp.txt (508, hash-синк)
webnovel_slice.py добор вне претрейна (r-коэф/эхо/coverage-precision) одна команда по появлению data/samples/webnovel/<lang>/*.txt; блокируется корпусом graceful
adaptive_*.py research/14 пробы (гибрид отклонён) индикативные, не решения
memory_hotpath.py УСТАРЕЛ (контракт до cc57c7b) — не переиспользовать актуальное зеркало — pilot/memory_eval.py
providers.json базовые URL/слаги для зондов квирки — docs/experiments/00-provider-quirks.md, читать ПЕРЕД любым вызовом

18+ / корпус-билдеры

Скрипт Эксперимент Заметки
explicit_judges.py, gu_corpus_build.py exp10 (violence-рука 18+) корпус-строки в data/ (gitignore); тексты вне git (Р8)
jpm_corpus_build.py, en_corpus_build.py, ja_corpus_build.py exp11 (erotica-рука 18+, zh/en/ja пары) jpm=金瓶梅 PD; en/ja — файлы владельца вне git
mistral_fidelity.py exp02/D14.2 (Mistral base-fidelity ре-съёмка с полным персистом) почему: первый Mistral-зонд потерял сырьё
dialogue_precision.py exp07 package-2 (precision гейта на диалог-плотном срезе)
refusal_corpus_build.py exp02 (наполнение refusal-корпуса L0L1) explicit/L3 этим НЕ наполняются
local_bench.py, llama_moe_bench.sh exp03 (локальный стенд) dense через ollama / MoE через llama.cpp

Скрипты «дуги качества» 12 → 13 → 14 → 14b (эксперименты ЗАКРЫТЫ — отчёты см. docs/experiments/README.md)

Оставлены на месте намеренно: цитируются по пути из D-лога/PROGRESS/отчётов; перемещение рвёт грепы (D23.3). Прогонять повторно не нужно — вердикты ратифицированы (D30/D32/D37/D38).

Семья Эксперимент Общие/особые модули
exp12_*.py (arms, blocks, candidate, extract, glossary_v2, judges, monitor, pack, reflow_demo, rootcause) exp12 диагноз (закрыт D30) exp12_judges импортит exp12_pack — семья неделима
exp13_*.py (aggregate, blind_stage, gates, judges, monitor, translate) exp13 бейк-офф переводчиков (закрыт D32)
exp13_seed_v2.py, exp13_seed_signoff.py ЖИВОЙ провенанс сида v2 (воспроизводимый трансформ, цитируется D-логом D30.5) НЕ спент, НЕ трогать/не перемещать
exp14_*.py (arms, judges, kpi, monitor, rank, regate, sizecurve) exp14 рычаги качества (закрыт D37) все импортят exp14_common (+ arms/sizecurve → exp14_prompts); exp14_material автономен
exp14b_*.py (arms, judge, monitor, score) exp14b батарея смысл-трапов (закрыт D38) реюзят exp14_common/exp14_promptsexp14 и 14b неразделимы

exp14_common.py + exp14_prompts.py — общий фундамент exp14 И exp14b. Их не переносить/не переименовывать без переезда обеих семей вместе (иначе рвётся import exp14_common).

Правила (выучены на ошибках)

  1. Провенанс данных: results-файлы не перезаписывать ре-раном (урок exp02: сырьё 66 вызовов утрачено); каждый прогон — model id + дата + usage; сырые ВЫХОДЫ сохранять целиком.
  2. Претрейн-контаминация: вся классика (Лу Синь/Акутагава/Уэллс) в претрейне моделей — C0 частично выдаёт канон, эхо систематично (deepseek 13/24 zh-чанков). Выводы на классике = предварительные; решающие замеры — на вебновелл-корпусе владельца.
  3. Тест ставится, только если его исход может перевернуть решение и не установлен литературой (урок снятого exp05).
  4. LLM-судьи: ≤23 семейства, 11+ повторов со свапом позиций, Bradley-Terry/медиана, судья не судит выходы своего семейства (D13.3); κ против человеческого IAA, не «средняя корреляция».
  5. Env: eval/.venv (torch-cpu, sentence-transformers, openai, dotenv, pymorphy3); ключи скрипты читают из eval/.env сами. Стенд: GTX 1070 8GB — модель полигона qwen3-abliterated:30b-a3b в ollama не выселять без нужды; localhost из-под прокси даёт 403 — для локальных вызовов no-proxy.
  6. data/ в .gitignore (коммитится только код) — критичные результаты дублируй цифрами в docs/experiments/.
  7. Аномалия провайдера ≠ повод гадать: странный отказ (интермиттентный 404/5xx, новый finish_reason, молча игнорируемые параметры) → сначала официальная дока вендора (deprecations/changelog/status) + свежий /models живьём, потом диагноз; в отчёт оба слоя — наблюдение И вендор-факт с датой/URL. Урок: 404-флейки gemini-2.5-flash оказались предвестником EOL всей 2.5-серии 16.10.2026 (research/15 → фактчек 10.07; правило — шапка 00-provider-quirks.md).

Доступные ключи от моделей

DEEPSEEK_API_KEY, ZAI_API_KEY, KIMI_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, XAI_API_KEY, MISTRAL_API_KEY