Commit graph

12 commits

Author SHA1 Message Date
Claude (backend session)
3d183bc2f4 Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
Claude (backend session)
9afea37bee Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions 2026-07-09 19:05:28 +03:00
Claude (backend session)
6ab92b5e66 Add MISTRAL_API_KEY and update keys list for eval/backend 2026-07-09 16:46:40 +03:00
Claude (backend session)
8ac568b97e Archive closed session prompts and add onboarding docs: root CLAUDE.md plus backend and eval READMEs 2026-07-09 16:46:39 +03:00
Claude (backend session)
1aed2c7627 Trim grok-reasoning re-narration in exp08 and cost_model: single-source the mechanism in provider-quirks, keep only cost-relevant facts plus pointers 2026-07-05 21:17:06 +03:00
Claude (backend session)
84737d032c Reconcile grok-reasoning across polygon docs: reasoning_effort:none disables it (probe bug), editor stays grok-4.3 with provider-quirks slug fixed; economics unchanged; add pilot mono-editor arm 2026-07-05 20:36:48 +03:00
Claude (backend session)
3911bafe30 Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
Claude (backend session)
e4ca3a57fd Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
Claude (backend session)
1878c3ea95 Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
Claude (backend session)
0dbcce7e97 Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup 2026-07-04 21:06:47 +03:00
Claude (backend session)
ce71a1a752 Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
Claude (backend session)
9bb00d49f2 Initial commit: documentation, eval polygon, backend step 0 verdict
TextMachine project repository (AI translation of literary books).
Includes: v2 architecture decisions, MVP plan, research 01-12,
polygon experiments 01-03, backend-session revalidation verdict
(03-implementation-notes.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 06:50:59 +03:00