textmachine/eval
2026-07-11 08:50:55 +03:00
..
pilot Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
.gitignore Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
adaptive_incontext.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_levers.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_probe.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_reask.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
build_editor_artifact.py Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions 2026-07-09 19:05:28 +03:00
content_filter_probe.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
cost_model_v2.py Trim grok-reasoning re-narration in exp08 and cost_model: single-source the mechanism in provider-quirks, keep only cost-relevant facts plus pointers 2026-07-05 21:17:06 +03:00
coverage_precision.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
dialogue_precision.py Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
editor_bench.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
en_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
exp12_arms.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_blocks.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_candidate.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_extract.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_glossary_v2.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_judges.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_monitor.py Add exp12 monitor script and bump contract range to D30 in onboarding docs 2026-07-11 05:08:27 +03:00
exp12_pack.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_reflow_demo.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp12_rootcause.py Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
exp13_aggregate.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_blind_stage.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_gates.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_judges.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_monitor.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_seed_v2.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
exp13_translate.py Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
explicit_judges.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
extract_bench.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
gu_corpus_build.py Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
ja_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
jpm_corpus_build.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
llama_moe_bench.sh Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
local_bench.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
memory_hotpath.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
mistral_fidelity.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
ollama-up.sh Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
providers.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_erotica.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_explicit.json Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
providers_grok_reason.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
providers_judges.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
providers_local8b.json Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
README.md Add owner rule: provider anomalies require official vendor-doc fact-check before interpretation, recorded in CLAUDE.md, quirks header, and eval rules 2026-07-10 00:03:50 +03:00
refusal_bench.py Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
refusal_corpus_build.py Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
retrieval_bench.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
run_refusal.sh Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup 2026-07-04 21:06:47 +03:00
token_calc.py Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
webnovel_slice.py Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions 2026-07-09 19:05:28 +03:00

eval/ — полигон TextMachine

Зона сессии «Полигон»: эмпирическая валидация допущений архитектуры (мерить, а не верить). Отчёты — docs/experiments/NN-*.md (методика + цифры + честные ограничения). backend/ только читать; расхождения с architecture/research — пингом оркестратору в docs/PROGRESS.md. .env не читать; data/refusal_corpus/* не открывать без прямой задачи владельца.

Карта

Скрипт Эксперимент Заметки
refusal_bench.py exp02; его split_sentences/classify_output = приёмочный оракул coverage-гейта Go (паритет закреплён Go-тестом; любое изменение — Python-first, лок-степ) run_refusal.sh — однокоманда
token_calc.py exp01 токен-калибровка (r-множители) токенизаторы в tokenizers/
editor_bench.py + build_editor_artifact.py exp04 бейк-офф редакторов слепота починена (D13.5): рандомизация меток ПО ФРАГМЕНТУ (соль), ключ+цена в отдельном --key файле. Выбор grok-4.3 провизорен — exp04 §Оговорки
coverage_precision.py, content_filter_probe.py exp07 precision гейта (0 FP/57)
cost_model_v2.py exp08 COGS (~$0.69/ранобэ) цены датировать и перепроверять
extract_bench.py exp06 экстракция терминов (спот=локаль/рендер=облако) health-верификация каждого ответа — образец
retrieval_bench.py эмбеддинг-бенч памяти (bge-m3 дефолт)
pilot/declmetric.py, pilot/memory_eval.py, pilot/kana_precision.py харнесс пилота Ф2.5 (exp09) починен (D13.5): эхо-контроль, ПОЛНЫЕ выходы, fidelity-ось (cross-family судья), M0M3, Go-синк 7 расхождений (self-test+адверсариально verified). --self-test/--dry-run без API. Инъекция = зеркало memory.go — синк с Go v3 выполнен (пакет №2: suppressUnboundedPhonetic + апостроф-фолд, adversarial parity); при расхождении верить Go. kana_traps.json+trad2simp.txt (508, hash-синк)
webnovel_slice.py добор вне претрейна (r-коэф/эхо/coverage-precision) одна команда по появлению data/samples/webnovel/<lang>/*.txt; блокируется корпусом graceful
adaptive_*.py research/14 пробы (гибрид отклонён) индикативные, не решения
memory_hotpath.py УСТАРЕЛ (контракт до cc57c7b) — не переиспользовать актуальное зеркало — pilot/memory_eval.py
providers.json базовые URL/слаги для зондов квирки — docs/experiments/00-provider-quirks.md, читать ПЕРЕД любым вызовом

Правила (выучены на ошибках)

  1. Провенанс данных: results-файлы не перезаписывать ре-раном (урок exp02: сырьё 66 вызовов утрачено); каждый прогон — model id + дата + usage; сырые ВЫХОДЫ сохранять целиком.
  2. Претрейн-контаминация: вся классика (Лу Синь/Акутагава/Уэллс) в претрейне моделей — C0 частично выдаёт канон, эхо систематично (deepseek 13/24 zh-чанков). Выводы на классике = предварительные; решающие замеры — на вебновелл-корпусе владельца.
  3. Тест ставится, только если его исход может перевернуть решение и не установлен литературой (урок снятого exp05).
  4. LLM-судьи: ≤23 семейства, 11+ повторов со свапом позиций, Bradley-Terry/медиана, судья не судит выходы своего семейства (D13.3); κ против человеческого IAA, не «средняя корреляция».
  5. Env: eval/.venv (torch-cpu, sentence-transformers, openai, dotenv, pymorphy3); ключи скрипты читают из eval/.env сами. Стенд: GTX 1070 8GB — модель полигона qwen3-abliterated:30b-a3b в ollama не выселять без нужды; localhost из-под прокси даёт 403 — для локальных вызовов no-proxy.
  6. data/ в .gitignore (коммитится только код) — критичные результаты дублируй цифрами в docs/experiments/.
  7. Аномалия провайдера ≠ повод гадать: странный отказ (интермиттентный 404/5xx, новый finish_reason, молча игнорируемые параметры) → сначала официальная дока вендора (deprecations/changelog/status) + свежий /models живьём, потом диагноз; в отчёт оба слоя — наблюдение И вендор-факт с датой/URL. Урок: 404-флейки gemini-2.5-flash оказались предвестником EOL всей 2.5-серии 16.10.2026 (research/15 → фактчек 10.07; правило — шапка 00-provider-quirks.md).

Доступные ключи от моделей

DEEPSEEK_API_KEY, ZAI_API_KEY, KIMI_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, XAI_API_KEY, MISTRAL_API_KEY