textmachine/eval
2026-07-18 22:08:20 +03:00
..
exp12 Group closed experiment scripts into eval/exp12|13|14|14b folders with import shims preserved, keep live seed and oracles in eval/, and update all path citations 2026-07-12 21:48:55 +03:00
exp13 Group closed experiment scripts into eval/exp12|13|14|14b folders with import shims preserved, keep live seed and oracles in eval/, and update all path citations 2026-07-12 21:48:55 +03:00
exp14 Group closed experiment scripts into eval/exp12|13|14|14b folders with import shims preserved, keep live seed and oracles in eval/, and update all path citations 2026-07-12 21:48:55 +03:00
exp14b Commit the reseed transform in eval/exp14b and note the seed-provenance lineage in the eval README, accepting the folder placement since it carries no contract citation 2026-07-12 23:34:16 +03:00
exp15 Land Q4a: fidelity is not bought at the translate stage, the pro arm is not ratified, the glm editor is the confirmed weak link, with corrected trap rules and the truncation-accepting harness bug documented 2026-07-18 22:08:20 +03:00
pilot Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
.gitignore Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
adaptive_incontext.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_levers.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_probe.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_reask.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
build_editor_artifact.py Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions 2026-07-09 19:05:28 +03:00
content_filter_probe.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
cost_model_v2.py Trim grok-reasoning re-narration in exp08 and cost_model: single-source the mechanism in provider-quirks, keep only cost-relevant facts plus pointers 2026-07-05 21:17:06 +03:00
coverage_precision.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
dialogue_precision.py Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
editor_bench.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
en_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
exp13_seed_signoff.py Land seed v2 signoff transform script: reproducible provenance for owner-signed glossary v2 term statuses (D34.1) 2026-07-11 15:58:46 +03:00
exp13_seed_v2.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
explicit_judges.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
extract_bench.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
gu_corpus_build.py Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
ja_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
jpm_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
llama_moe_bench.sh Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
local_bench.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
memory_hotpath.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
mistral_fidelity.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
ollama-up.sh Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
providers.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_erotica.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_explicit.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_grok_reason.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
providers_judges.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
providers_local8b.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
README.md Actualize the polygon README and the D15.2 spec status: tmctl export extraction rule, superseded editor caveat, judge EOL residual, deferred-implementation banner 2026-07-17 03:20:53 +03:00
refusal_bench.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
refusal_corpus_build.py Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
retrieval_bench.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
run_refusal.sh Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup 2026-07-04 21:06:47 +03:00
token_calc.py Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
webnovel_slice.py Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions 2026-07-09 19:05:28 +03:00

eval/ — полигон TextMachine

Зона сессии «Полигон»: эмпирическая валидация допущений архитектуры (мерить, а не верить). Отчёты — docs/experiments/NN-*.md (методика + цифры + честные ограничения). backend/ только читать; расхождения с architecture/research — пингом оркестратору в docs/PROGRESS.md. .env не читать; data/refusal_corpus/* не открывать без прямой задачи владельца.

Очередь полигона (пост-D39.6, 17.07): exp15 сегментация-эмпирика — промт docs/POLYGON_SEGMENTATION_EMPIRICS_SESSION_PROMPT.md ВЫДАН, ждёт запуска (фриз = первый коммит; скрипты — семьёй eval/exp15/ по конвенции D38.2) → полигон-стадия банк-майнинга (research/20 §D, отдельный промт ПОСЛЕ exp15). Residual-трекер пилота/18+/echo — docs/POLYGON_PACKAGE4_SESSION_PROMPT.md (отложен).

Карта

Скрипт Эксперимент Заметки
refusal_bench.py exp02; его split_sentences/classify_output = приёмочный оракул coverage-гейта Go (паритет закреплён Go-тестом; любое изменение — Python-first, лок-степ) run_refusal.sh — однокоманда
token_calc.py exp01 токен-калибровка (r-множители) токенизаторы в tokenizers/
editor_bench.py + build_editor_artifact.py exp04 бейк-офф редакторов слепота починена (D13.5): рандомизация меток ПО ФРАГМЕНТУ (соль), ключ+цена в отдельном --key файле. ⚠ Выбор grok-4.3 SUPERSEDED D30.1: боевой редактор = БИЛИНГВ glm-5, grok reasoning-off из редакторских ролей снят (no-op); exp04-ранг был на дефолт-reasoning
coverage_precision.py, content_filter_probe.py exp07 precision гейта (0 FP/57)
cost_model_v2.py exp08 COGS (~$0.69/ранобэ — ДО-флиповый; после флипа D1 билингв ≈$0.85, D30.4) цены датировать и перепроверять
extract_bench.py exp06 экстракция терминов (спот=локаль/рендер=облако) health-верификация каждого ответа — образец. ⚠ recall 0.981.0 мерен на PD-классике; вебновелл-ре-замер = research/20 §D-A6 (полигон-стадия банк-майнинга)
retrieval_bench.py эмбеддинг-бенч памяти (bge-m3 дефолт)
pilot/declmetric.py, pilot/memory_eval.py, pilot/kana_precision.py харнесс пилота Ф2.5 (exp09) починен (D13.5): эхо-контроль, ПОЛНЫЕ выходы, fidelity-ось (cross-family судья), M0M3, Go-синк 7 расхождений (self-test+адверсариально verified). --self-test/--dry-run без API. ⚠ судья-дефолт memory_eval --judge = gemini-2.5-flash (EOL 16.10.2026, research/15-фактчек) — мигрировать слаг при следующем прогоне (residual POLYGON_PACKAGE4). Инъекция = зеркало memory.go — синк с Go v3 выполнен (пакет №2: suppressUnboundedPhonetic + апостроф-фолд, adversarial parity); при расхождении верить Go. kana_traps.json+trad2simp.txt (508, hash-синк)
webnovel_slice.py добор вне претрейна (r-коэф/эхо/coverage-precision) одна команда по появлению data/samples/webnovel/<lang>/*.txt; блокируется корпусом graceful
adaptive_*.py research/14 пробы (гибрид отклонён) индикативные, не решения
memory_hotpath.py УСТАРЕЛ (контракт до cc57c7b) — не переиспользовать актуальное зеркало — pilot/memory_eval.py
providers.json базовые URL/слаги для зондов квирки — docs/experiments/00-provider-quirks.md, читать ПЕРЕД любым вызовом

18+ / корпус-билдеры

Скрипт Эксперимент Заметки
explicit_judges.py, gu_corpus_build.py exp10 (violence-рука 18+) корпус-строки в data/ (gitignore); тексты вне git (Р8)
jpm_corpus_build.py, en_corpus_build.py, ja_corpus_build.py exp11 (erotica-рука 18+, zh/en/ja пары) jpm=金瓶梅 PD; en/ja — файлы владельца вне git
mistral_fidelity.py exp02/D14.2 (Mistral base-fidelity ре-съёмка с полным персистом) почему: первый Mistral-зонд потерял сырьё
dialogue_precision.py exp07 package-2 (precision гейта на диалог-плотном срезе)
refusal_corpus_build.py exp02 (наполнение refusal-корпуса L0L1) explicit/L3 этим НЕ наполняются
local_bench.py, llama_moe_bench.sh exp03 (локальный стенд) dense через ollama / MoE через llama.cpp

Скрипты «дуги качества» 12 → 13 → 14 → 14b — сгруппированы по папкам eval/expNN/ (эксперименты ЗАКРЫТЫ; отчёты см. docs/experiments/README.md)

Спент-семьи структурированы по папкам eval/exp12/, eval/exp13/, eval/exp14/, eval/exp14b/ (решение владельца — D38.2): declutter top-level eval/, логическая группировка кода (это код-репо, не «архив»/не заморозка). Прогонять повторно не нужно — вердикты ратифицированы (D30/D32/D37/D38). Импорты сохранены (проверено py_compile+find_spec): скрипты, тянущие refusal_bench, дотягиваются до eval/ шимом .parent.parent; exp14b_* тянут общий exp14_common/exp14_prompts из ../exp14/ (exp14b построен на харнессе exp14). ЖИВОЙ провенанс сида v2 (exp13_seed_v2.py/exp13_seed_signoff.py) НЕ перемещён — остался в eval/ (цитируется D-логом D30.5, регенерирует боевой сид).

Папка Эксперимент Скрипты / особые модули
eval/exp12/ exp12 диагноз (закрыт D30) arms, blocks, candidate, extract, glossary_v2, judges, monitor, pack, reflow_demo, rootcause; exp12_judgesexp12_pack
eval/exp13/ exp13 бейк-офф переводчиков (закрыт D32) aggregate, blind_stage, gates, judges, monitor, translate
eval/exp14/ exp14 рычаги качества (закрыт D37) arms, common, judges, kpi, material, monitor, prompts, rank, regate, sizecurve; фундамент exp14_common/exp14_prompts
eval/exp14b/ exp14b батарея смысл-трапов (закрыт D38) arms, judge, monitor, score; реюзят exp14_common/exp14_prompts из ../exp14/; + exp14b_reseed_promote.py = ЖИВОЙ reseed-провенанс сида (промоут грейдов+四代族长, D38.3/4, идемпотентен)
eval/ (НЕ перемещён) ЖИВОЙ провенанс сида v2 exp13_seed_v2.py, exp13_seed_signoff.py (D30.5, регенерируют боевой сид); лестница reseed → exp14b/exp14b_reseed_promote.py (D38.3/4; в папке — нет контракт-пина, в отличие от exp13_seed_*)

exp14_common.py + exp14_prompts.py живут в eval/exp14/ и служат фундаментом exp14 И exp14b; exp14b_* тянут их через шим ../exp14, поэтому переименование/перенос exp14 без правки шимов exp14b порвёт import exp14_common.

Правила (выучены на ошибках)

  1. Провенанс данных: results-файлы не перезаписывать ре-раном (урок exp02: сырьё 66 вызовов утрачено); каждый прогон — model id + дата + usage; сырые ВЫХОДЫ сохранять целиком.
  2. Претрейн-контаминация: вся классика (Лу Синь/Акутагава/Уэллс) в претрейне моделей — C0 частично выдаёт канон, эхо систематично (deepseek 13/24 zh-чанков). Выводы на классике = предварительные; решающие замеры — на вебновелл-корпусе владельца.
  3. Тест ставится, только если его исход может перевернуть решение и не установлен литературой (урок снятого exp05).
  4. LLM-судьи: ≤23 семейства, 11+ повторов со свапом позиций, Bradley-Terry/медиана, судья не судит выходы своего семейства (D13.3); κ против человеческого IAA, не «средняя корреляция».
  5. Env: eval/.venv (torch-cpu, sentence-transformers, openai, dotenv, pymorphy3); ключи скрипты читают из eval/.env сами. Стенд: GTX 1070 8GB — модель полигона qwen3-abliterated:30b-a3b в ollama не выселять без нужды; localhost из-под прокси даёт 403 — для локальных вызовов no-proxy.
  6. data/ в .gitignore (коммитится только код) — критичные результаты дублируй цифрами в docs/experiments/.
  7. Экстракция прод-выходов — ТОЛЬКО через tmctl export (D39.5, инвариант №8 backend/README): сырое чтение чекпоинтов store (records.json-стиль) отдаёт НЕнормализованные байты (без export-contract слоя) и молча-неполную книгу (без manifest-join/drift-гарда). --plaintext для глаз, дефолтный JSON для скриптов. Прежний путь final_hash→response_text — только для форензики сырья, не для оценки качества.
  8. Аномалия провайдера ≠ повод гадать: странный отказ (интермиттентный 404/5xx, новый finish_reason, молча игнорируемые параметры) → сначала официальная дока вендора (deprecations/changelog/status) + свежий /models живьём, потом диагноз; в отчёт оба слоя — наблюдение И вендор-факт с датой/URL. Урок: 404-флейки gemini-2.5-flash оказались предвестником EOL всей 2.5-серии 16.10.2026 (research/15 → фактчек 10.07; правило — шапка 00-provider-quirks.md).

Доступные ключи от моделей

DEEPSEEK_API_KEY, ZAI_API_KEY, KIMI_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, XAI_API_KEY, MISTRAL_API_KEY