Commit graph

41 commits

Author SHA1 Message Date
Claude (backend session)
09ca24b50a Land the ratified implementation plan as D39.12 with design11 verifications, record the mistral rate-limit quirk, and issue the backend single-pack session prompt 2026-07-19 04:03:39 +03:00
Claude (backend session)
667eec0dc5 Land exp16 bank-mining: the which-what split with the 0.97 recall code detector, banknote as the dst channel, the clean exp14b re-audit, and the stop-only truncation errata 2026-07-19 00:24:24 +03:00
Claude (backend session)
96345f589d Freeze exp16 bank-mining pre-registration: miner-v1 detectors V-A/V-B/V-C, alias/canon/banknote code, GT filter and thresholds tuned on ch1-15, before any paid call 2026-07-18 23:32:53 +03:00
Claude (backend session)
cfcf0f08e3 Land Q4a: fidelity is not bought at the translate stage, the pro arm is not ratified, the glm editor is the confirmed weak link, with corrected trap rules and the truncation-accepting harness bug documented 2026-07-18 22:08:20 +03:00
Claude (backend session)
034660ce1d Append the blind-read verdict addendum: no arm passes the owner bar, defects distributed across arms confirming the metric null, defect-class ledger and instrument fixes recorded 2026-07-18 17:24:30 +03:00
Claude (backend session)
fe34a16896 Land exp15 REV.2: floor-gated corrected results with per-vote judge persistence, the retracted boundary headline, and orchestrator landing fixes to the deviations ledger 2026-07-18 17:06:11 +03:00
Claude (backend session)
78a9a2dd9e Record exp15 Q2b instrument-ceiling result plus money and self-review sections, and append the paid Q1/Q3 handoff block for the next focused session 2026-07-17 23:09:41 +03:00
Claude (backend session)
b9272a19d0 Freeze exp15 pre-registration with verified slugs/prices, sub-$15 budget, grok+mistral judges, A0 injection method, synthetic S2-prime seam ledger, and gender-signed seed 2026-07-17 21:13:58 +03:00
Claude (backend session)
44bc4f8332 Repair the live ACCEPTANCE citation in research/17 to its archived path and drop the over-attributed D-numbers on the exp04 and exp12 index rows 2026-07-12 22:03:07 +03:00
Claude (backend session)
f468f27de6 Group closed experiment scripts into eval/exp12|13|14|14b folders with import shims preserved, keep live seed and oracles in eval/, and update all path citations 2026-07-12 21:48:55 +03:00
Claude (backend session)
9feb1151e8 Add experiments index and extend eval script map over the closed quality-arc families, keeping scripts in place because the D-log and reports cite them by path 2026-07-12 20:55:45 +03:00
Claude (backend session)
33c79d7a01 Ratify D38 accepting exp14b: refute only-gpt-5.4 (mistral and deepseek-pro fix meaning inversions), find term-drift root is seed draft-status, hold cheap-editor swap for a prose bake-off 2026-07-12 20:06:03 +03:00
Claude (backend session)
48e3f39eb8 Freeze exp14b pre-registration: meaning-trap battery (5 construction classes, deterministic + cross-family judges) to test whether the within-chunk comprehension fix is gpt-5.4-specific, plus P1a-vs-frontier ranking 2026-07-12 15:15:03 +03:00
Claude (backend session)
7ee668fba0 Ratify D37 accepting exp14 quality-empirics with fixes: claim-1 paragraphs is a cheap prompt lever, claim-2 gpt-5.4 escalation held under-powered, archive spent exp14 prompts 2026-07-12 05:15:55 +03:00
Claude (backend session)
b59ca66b83 Land external blind-eval corroboration: two zero sessions confirm F-disc is best on both prose and fidelity, flag floating grade glossary, and catch an A-2pass fidelity catastrophe that redirects near-term to single-pass discourse 2026-07-12 04:29:32 +03:00
Claude (backend session)
73a45848db Complete Gemini test with proper max_tokens: gemini-3.1-pro also inverts the double-negation trap, so only gpt-5.4 of three frontier families fixes within-chunk comprehension 2026-07-12 04:06:25 +03:00
Claude (backend session)
a19107479d Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
Claude (backend session)
174192b1eb Fold exp14 batch-1 trap-judge corroboration: cross-family panel unanimously confirms the T1 comprehension catastrophe on all cheap arms and clears both frontier arms 2026-07-11 23:29:52 +03:00
Claude (backend session)
dc5cf3bb9d Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor 2026-07-11 22:50:09 +03:00
Claude (backend session)
353bb52676 Freeze exp14 pre-registration: quality-lever empirics on slicing×prompt 2×2 plus ablations, frontier capability-vs-activation, size curve, annotated trap set, pre-registered metrics/gates and per-key budget caps 2026-07-11 22:06:31 +03:00
Claude (backend session)
7487b782f8 Land exp13 translator bake-off with review banner: seed v2 accepted pending owner sign-off, draft flash retained (pro rejected — manufactured convergence, core-term catastrophe, no fidelity gain), ratify D32 2026-07-11 08:50:55 +03:00
Claude (backend session)
3faecdee03 Ratify D31 documentation revision: actuality map, banners and supersession marks across architecture/experiments/research, fixed archive relative links, bumped contract range 2026-07-11 05:50:23 +03:00
Claude (backend session)
5bc75d0e9a Pre-register exp13 translator bake-off: six arms, cross-family span-citing judges, output gates, frozen before paid calls 2026-07-11 05:30:19 +03:00
Claude (backend session)
9e3762eaa8 Split PROGRESS journal: move 330KB of closed 04-10.07 chronology to archive with verbatim headings, keep live sections and anchors, repoint incoming references 2026-07-11 05:28:52 +03:00
Claude (backend session)
75db9e144f Land exp12 quality diagnosis with orchestrator review banner: mono-editor passivity confirmed, unanimity and fidelity-gate claims corrected, glossary v2 status trap flagged 2026-07-11 05:05:40 +03:00
Claude (backend session)
a96248f14b Ratify D22 accepting both packages with register caveat, 18+ judge revision, L3 pre-screen task, and D15.2 v3.1 amendments; sync status docs, archive prompts 2026-07-10 04:07:19 +03:00
Claude (backend session)
3551997292 Land polygon erotica package: exp11 three-register benchmark closing D14.4, seed finalized gender-hidden, fix-list D19.5, journal entries for both sessions with orchestrator review corrections 2026-07-10 04:02:10 +03:00
Claude (backend session)
8643fdfe96 Add owner rule: provider anomalies require official vendor-doc fact-check before interpretation, recorded in CLAUDE.md, quirks header, and eval rules 2026-07-10 00:03:50 +03:00
Claude (backend session)
7828647319 Fact-check Gemini live: 2.5-series shutdown 2026-10-16, undocumented 404 flakes, PROHIBITED_CONTENT is non-configurable; update quirks calendar and erotica prompt 2026-07-09 23:57:47 +03:00
Claude (backend session)
be7388128b Sync status docs with D18-D21, narrow grok quirks line per D19.1, archive voice-state prompt, issue backend package-3 handoff prompt 2026-07-09 23:39:22 +03:00
Claude (backend session)
3d183bc2f4 Land polygon package 2: mirror sync with Go v3, explicit benchmark exp10, Mistral fidelity rerun, terse dialogue precision, with orchestrator corrections from external review 2026-07-09 21:16:32 +03:00
Claude (backend session)
9afea37bee Land pilot-harness package: memory_eval echo control and fidelity axis, M0-M3 rename, artifact blinding, exp09 v2 with D13, kana precision, Mistral probe, webnovel slice; journal entries for both sessions 2026-07-09 19:05:28 +03:00
Claude (backend session)
1aed2c7627 Trim grok-reasoning re-narration in exp08 and cost_model: single-source the mechanism in provider-quirks, keep only cost-relevant facts plus pointers 2026-07-05 21:17:06 +03:00
Claude (backend session)
9792285e76 Correct exp04's 'grok without reasoning' characterization — grok-4.3 defaults to low reasoning (~13s), production editor sets reasoning_effort:none for 0 tokens; ranking unchanged 2026-07-05 21:07:39 +03:00
Claude (backend session)
84737d032c Reconcile grok-reasoning across polygon docs: reasoning_effort:none disables it (probe bug), editor stays grok-4.3 with provider-quirks slug fixed; economics unchanged; add pilot mono-editor arm 2026-07-05 20:36:48 +03:00
Claude (backend session)
3911bafe30 Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
Claude (backend session)
d15cdb3b71 Add D12 failure-resilience model: three-class failure taxonomy, F3 idempotency fix bounded to at-most-once, progress-surfacing spec grounded in durable-execution prior art 2026-07-04 23:10:32 +03:00
Claude (backend session)
87b92ba696 Add decisions log and provider-quirks, set grok-4.3 editor default and OpenAI-retained model stack per owner and editor bake-off 2026-07-04 18:47:45 +03:00
Claude (backend session)
9d78385c37 Apply external review fixes: snapshot local overrides, gate-enabled guard, TM-key doc note, db-lock gitignore, techdebt disposition 2026-07-04 16:29:25 +03:00
Claude (backend session)
7954325b7e Add Phase 0 backend skeleton: LLM core with native Anthropic adapter, money ledger, SQLite store with atomic chunk checkpoints, deterministic C1 mini-runner, tmctl CLI, configs and prompts 2026-07-04 09:04:06 +03:00
Claude (backend session)
9bb00d49f2 Initial commit: documentation, eval polygon, backend step 0 verdict
TextMachine project repository (AI translation of literary books).
Includes: v2 architecture decisions, MVP plan, research 01-12,
polygon experiments 01-03, backend-session revalidation verdict
(03-implementation-notes.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 06:50:59 +03:00