..
pilot
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
.gitignore
Initial commit: documentation, eval polygon, backend step 0 verdict
2026-07-04 06:50:59 +03:00
adaptive_incontext.py
Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve
2026-07-05 01:34:18 +03:00
adaptive_levers.py
Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve
2026-07-05 01:34:18 +03:00
adaptive_probe.py
Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve
2026-07-05 01:34:18 +03:00
adaptive_reask.py
Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve
2026-07-05 01:34:18 +03:00
build_editor_artifact.py
Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions
2026-07-09 19:05:28 +03:00
content_filter_probe.py
Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval
2026-07-05 15:44:44 +03:00
cost_model_v2.py
Trim grok-reasoning re-narration in exp08 and cost_model: single-source the mechanism in provider-quirks, keep only cost-relevant facts plus pointers
2026-07-05 21:17:06 +03:00
coverage_precision.py
Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval
2026-07-05 15:44:44 +03:00
dialogue_precision.py
Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review
2026-07-09 21:16:32 +03:00
editor_bench.py
Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config
2026-07-04 19:26:36 +03:00
en_corpus_build.py
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
exp12_arms.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp12_blocks.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp12_candidate.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp12_extract.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp12_glossary_v2.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp12_judges.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp12_monitor.py
Add exp12 monitor script and bump contract range to D30 in onboarding docs
2026-07-11 05:08:27 +03:00
exp12_pack.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp12_reflow_demo.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp12_rootcause.py
Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged
2026-07-11 05:05:40 +03:00
exp13_aggregate.py
Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32
2026-07-11 08:50:55 +03:00
exp13_blind_stage.py
Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32
2026-07-11 08:50:55 +03:00
exp13_gates.py
Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32
2026-07-11 08:50:55 +03:00
exp13_judges.py
Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32
2026-07-11 08:50:55 +03:00
exp13_monitor.py
Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32
2026-07-11 08:50:55 +03:00
exp13_seed_signoff.py
Land seed v2 signoff transform script: reproducible provenance for owner-signed glossary v2 term statuses (D34.1)
2026-07-11 15:58:46 +03:00
exp13_seed_v2.py
Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32
2026-07-11 08:50:55 +03:00
exp13_translate.py
Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32
2026-07-11 08:50:55 +03:00
exp14_arms.py
Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better
2026-07-12 03:43:47 +03:00
exp14_common.py
Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor
2026-07-11 22:50:09 +03:00
exp14_judges.py
Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better
2026-07-12 03:43:47 +03:00
exp14_kpi.py
Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor
2026-07-11 22:50:09 +03:00
exp14_material.py
Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor
2026-07-11 22:50:09 +03:00
exp14_monitor.py
Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better
2026-07-12 03:43:47 +03:00
exp14_prompts.py
Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor
2026-07-11 22:50:09 +03:00
exp14_rank.py
Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better
2026-07-12 03:43:47 +03:00
exp14_regate.py
Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better
2026-07-12 03:43:47 +03:00
exp14_sizecurve.py
Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better
2026-07-12 03:43:47 +03:00
exp14b_arms.py
Commit exp14b harness scripts (arms, judge, monitor, score) to close the reproducibility gap now that the meaning-battery experiment is ratified and closed at D38
2026-07-12 20:55:45 +03:00
exp14b_judge.py
Commit exp14b harness scripts (arms, judge, monitor, score) to close the reproducibility gap now that the meaning-battery experiment is ratified and closed at D38
2026-07-12 20:55:45 +03:00
exp14b_monitor.py
Commit exp14b harness scripts (arms, judge, monitor, score) to close the reproducibility gap now that the meaning-battery experiment is ratified and closed at D38
2026-07-12 20:55:45 +03:00
exp14b_score.py
Commit exp14b harness scripts (arms, judge, monitor, score) to close the reproducibility gap now that the meaning-battery experiment is ratified and closed at D38
2026-07-12 20:55:45 +03:00
explicit_judges.py
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
extract_bench.py
Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference
2026-07-05 01:34:18 +03:00
gu_corpus_build.py
Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review
2026-07-09 21:16:32 +03:00
ja_corpus_build.py
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
jpm_corpus_build.py
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
llama_moe_bench.sh
Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config
2026-07-04 19:26:36 +03:00
local_bench.py
Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config
2026-07-04 19:26:36 +03:00
memory_hotpath.py
Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference
2026-07-05 01:34:18 +03:00
mistral_fidelity.py
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
ollama-up.sh
Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config
2026-07-04 19:26:36 +03:00
providers.json
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
providers_erotica.json
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
providers_explicit.json
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
providers_grok_reason.json
Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review
2026-07-09 21:16:32 +03:00
providers_judges.json
Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review
2026-07-09 21:16:32 +03:00
providers_local8b.json
Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review
2026-07-09 21:16:32 +03:00
README.md
Add owner rule: provider anomalies require official vendor-doc fact-check before interpretation, recorded in CLAUDE.md, quirks header, and eval rules
2026-07-10 00:03:50 +03:00
refusal_bench.py
Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections
2026-07-10 04:02:10 +03:00
refusal_corpus_build.py
Initial commit: documentation, eval polygon, backend step 0 verdict
2026-07-04 06:50:59 +03:00
retrieval_bench.py
Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference
2026-07-05 01:34:18 +03:00
run_refusal.sh
Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup
2026-07-04 21:06:47 +03:00
token_calc.py
Initial commit: documentation, eval polygon, backend step 0 verdict
2026-07-04 06:50:59 +03:00
webnovel_slice.py
Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions
2026-07-09 19:05:28 +03:00