textmachine/eval
2026-07-05 15:44:44 +03:00
..
pilot Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
.gitignore Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
adaptive_incontext.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_levers.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_probe.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
adaptive_reask.py Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
build_editor_artifact.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
content_filter_probe.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
cost_model_v2.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
coverage_precision.py Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
editor_bench.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
extract_bench.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
llama_moe_bench.sh Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
local_bench.py Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
memory_hotpath.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
ollama-up.sh Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
providers.json Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup 2026-07-04 21:06:47 +03:00
refusal_bench.py Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup 2026-07-04 21:06:47 +03:00
refusal_corpus_build.py Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
retrieval_bench.py Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
run_refusal.sh Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup 2026-07-04 21:06:47 +03:00
token_calc.py Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00