Commit graph

8 commits

Author SHA1 Message Date
Claude (backend session)
1aed2c7627 Trim grok-reasoning re-narration in exp08 and cost_model: single-source the mechanism in provider-quirks, keep only cost-relevant facts plus pointers 2026-07-05 21:17:06 +03:00
Claude (backend session)
84737d032c Reconcile grok-reasoning across polygon docs: reasoning_effort:none disables it (probe bug), editor stays grok-4.3 with provider-quirks slug fixed; economics unchanged; add pilot mono-editor arm 2026-07-05 20:36:48 +03:00
Claude (backend session)
3911bafe30 Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
Claude (backend session)
e4ca3a57fd Add the adaptive-memory research verdict (doc 14) with four executable probes: flagship logprob access, in-context demos vs glossary, lemma post-check, and the deterministic lever curve 2026-07-05 01:34:18 +03:00
Claude (backend session)
1878c3ea95 Track the closed memory-validation and extraction sessions' eval artifacts that committed docs already reference 2026-07-05 01:34:18 +03:00
Claude (backend session)
0dbcce7e97 Harden refusal benchmark: split verdict taxonomy, live-audited provider configs, local warmup 2026-07-04 21:06:47 +03:00
Claude (backend session)
ce71a1a752 Add polygon eval scripts for local-model, editor and MoE benchmarks plus ollama launcher and provider config 2026-07-04 19:26:36 +03:00
Claude (backend session)
9bb00d49f2 Initial commit: documentation, eval polygon, backend step 0 verdict
TextMachine project repository (AI translation of literary books).
Includes: v2 architecture decisions, MVP plan, research 01-12,
polygon experiments 01-03, backend-session revalidation verdict
(03-implementation-notes.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 06:50:59 +03:00