Commit graph

153 commits

Author SHA1 Message Date
Claude (backend session)
cabdfb35b4 Issue the Q4a mini-session prompt: deferred strong-translator arm with its own two-dollar budget outside the exp15 freeze, valley-only DeepSeek, fixed judging rig 2026-07-18 14:04:29 +03:00
Claude (backend session)
fdbee84378 Issue the exp16 bank-mining polygon prompt executing research/20 section D with the owner's budget, canon, no-genre-packs, and ja-ru replica decisions baked in, held until exp15 closes 2026-07-18 02:56:31 +03:00
Claude (backend session)
78a9a2dd9e Record exp15 Q2b instrument-ceiling result plus money and self-review sections, and append the paid Q1/Q3 handoff block for the next focused session 2026-07-17 23:09:41 +03:00
Claude (backend session)
b9272a19d0 Freeze exp15 pre-registration with verified slugs/prices, sub-$15 budget, grok+mistral judges, A0 injection method, synthetic S2-prime seam ledger, and gender-signed seed 2026-07-17 21:13:58 +03:00
Claude (backend session)
5d752e93e0 Doc hygiene per owner: archive the closed 10-13.07 chronicle out of PROGRESS, archive spent prompts, actualize CLAUDE, README and the orchestrator role banner to D39.6 2026-07-17 03:10:18 +03:00
Claude (backend session)
994987d8b7 Land research/20 bank-mining with orchestrator review header: section-A hash reproduced, repo claims spot-verified, injection confound corroborated, hold-gate deviation accepted 2026-07-17 03:05:40 +03:00
Claude (backend session)
bb4142f503 Ratify D39.5 landing pack-1.5 with the F1-F9 fix-list after execution verification, closing the polygon export surface question 2026-07-17 02:16:48 +03:00
Claude (backend session)
9dae0ed13d Ratify D39.4: adversarial debt repaid across both packs, hold pack-1.5 landing behind fix-list F1-F9 folded into the prompt's reserve slot 2026-07-17 01:33:22 +03:00
Claude (backend session)
2cfb7cdb3d Amend the research/20 prompt per researcher-19 review: require detector variants with a trade-off table compared as arms, and alias clustering as an explicit subtask with a code-vs-blind-zone boundary 2026-07-16 22:52:42 +03:00
Claude (backend session)
084c4c68c6 Ratify D39.3 bank-mining package dispositioning H15 and V4-p4, issue the research/20 prompt held until the exp15 freeze 2026-07-16 22:47:14 +03:00
Claude (backend session)
7e1e6780d1 Add the consolidated prompt-architecture note answering the owner's four recurring prompting questions with statuses and routing, linked from the layer-2 section 2026-07-16 22:24:28 +03:00
Claude (backend session)
11589d3935 Issue the small track-A pack-1.5 backend prompt: read-only tmctl export surface, owner-decided CJK gloss whitelist, runnable example, adversarial-debt reserve slot 2026-07-16 22:05:32 +03:00
Claude (backend session)
8ba0f720e8 Issue the exp15 segmentation-empirics polygon prompt executing research/19 section D with the owner's budget, gender-field, and gloss decisions baked in 2026-07-16 22:01:43 +03:00
Claude (backend session)
6c0010ee35 Ratify D39.2 accepting track-A pack-1 after direct execution verification and an independent golden re-capture reproduction, registering the 529-blocked adversarial pass as debt 2026-07-16 21:56:00 +03:00
Claude (backend session)
3d1dc8d8cb Ratify D39.1: land research/19 as the track-B design of record with the wave architecture and pre-registered empirics, review waived by the owner 2026-07-16 21:52:17 +03:00
Claude (backend session)
c5171b47c5 Land research/19 chunking-cohesion report with sources and archive its executed prompt: emission-window-cohesion split, Go chunker spec, wave architecture, pre-registered Q0-Q5 design 2026-07-16 21:48:02 +03:00
Claude (backend session)
2d9f70bb84 Correct a narrowing in the target-architecture and track-B prompt: add the generality invariant so the chunker stays general across any book, language, and structure 2026-07-13 01:53:07 +03:00
Claude (backend session)
d42d01e75a Issue the track-A backend build-now pack and the track-B chunking-cohesion researcher prompt for the architectural reset 2026-07-13 01:45:42 +03:00
Claude (backend session)
225abc0930 Ratify D39: land the broken-telephone sync-audit ledger and the target seven-layer architecture with the phased build-now and research tracks, folding in the polygon sync 2026-07-13 01:45:42 +03:00
Claude (backend session)
79477173f8 Archive three executed session prompts with outcome banners: the arch-reset (now D39), the discourse editor, and the reseed transform 2026-07-13 01:45:29 +03:00
Claude (backend session)
c9c9df7364 Pause the build round for an architectural reset: issue the orchestrator handoff prompt capturing the owner's five concerns and the polygon-orchestrator-backend sync plan 2026-07-13 00:32:09 +03:00
Claude (backend session)
183c8ad24f Ratify D38.5 landing the resnapshot components: the reseed transform and the discourse editor with chengyu and the few_shot toggle, both execution-verified 2026-07-12 23:48:49 +03:00
Claude (backend session)
e781f7b08f Update reseed handoff in PROGRESS: correct transform path to exp14b folder and record the stale grade-note tail was cleaned before the resnapshot per orchestrator 2026-07-12 23:27:56 +03:00
Claude (backend session)
38c5fb461c Ratify D38.4 completeness audit of exp14/14b: add the forgotten chengyu atom to the discourse prompt, record the unbuilt inversion guard raising editor-swap priority, flag the chunking arm and guard toggle 2026-07-12 23:10:15 +03:00
Claude (backend session)
7e6f9a5a4f Issue the discourse-prompt (backend) and reseed (polygon) handoff prompts for the single resnapshot, and archive the landed infra-pack prompt 2026-07-12 22:49:51 +03:00
Claude (backend session)
ffaf076fd9 Ratify D38.3 accepting the landed readability infra-pack after execution-verifying build, vet, race tests and the three span-defect pins 2026-07-12 22:41:39 +03:00
Claude (backend session)
b14ba6f20c Apply verification fixes: correct the archived ACCEPTANCE line anchor, the D34.3 re-gate attribution in both diagrams, the Mistral 18/18 provenance, and the actuality-map revision stamp 2026-07-12 22:21:02 +03:00
Claude (backend session)
feb1848b9c Ratify D38.2 in the decision log and record the consolidation in PROGRESS: eval regrouped, docs and diagrams actualized to D38.1, links verified, backend untouched 2026-07-12 22:07:24 +03:00
Claude (backend session)
fc94d6725a Archive the executed POLYGON_CLEANUP prompt with an outcome banner and record it in the README closed-prompts list 2026-07-12 22:04:40 +03:00
Claude (backend session)
44bc4f8332 Repair the live ACCEPTANCE citation in research/17 to its archived path and drop the over-attributed D-numbers on the exp04 and exp12 index rows 2026-07-12 22:03:07 +03:00
Claude (backend session)
5d24ba401c Sync architecture diagrams to the D1-D38.1 head: bilingual editor flip, output-sanitizer and quality-diagnostics nodes, channel-B v2, seed v2, and provider role restatements 2026-07-12 21:59:56 +03:00
Claude (backend session)
6f12fc0244 Actualize onboarding and status docs to the D38.1 head and bump decision-log range references, extending the actuality map with D38.1 and D38.2 2026-07-12 21:52:21 +03:00
Claude (backend session)
f468f27de6 Group closed experiment scripts into eval/exp12|13|14|14b folders with import shims preserved, keep live seed and oracles in eval/, and update all path citations 2026-07-12 21:48:55 +03:00
Claude (backend session)
af469140dc Report polygon cleanup in PROGRESS: fact-check refuted the naive archive because closed scripts are cited by path, so consolidate via index and script map instead 2026-07-12 20:57:19 +03:00
Claude (backend session)
9feb1151e8 Add experiments index and extend eval script map over the closed quality-arc families, keeping scripts in place because the D-log and reports cite them by path 2026-07-12 20:55:45 +03:00
Claude (backend session)
2d760fdb4c Issue polygon cleanup prompt: fact-check consolidation direction, build experiments index, carefully archive spent eval scripts and commit exp14b, flag books and puml hygiene 2026-07-12 20:47:10 +03:00
Claude (backend session)
8b8a9ec18d Land blind h2h and pivot to building: ChatGPT/Claude split on frontier-vs-parity but agree the gap is glossary+sanitizer infra, so issue the cheap infra-pack backend prompt (D38.1) 2026-07-12 20:30:41 +03:00
Claude (backend session)
33c79d7a01 Ratify D38 accepting exp14b: refute only-gpt-5.4 (mistral and deepseek-pro fix meaning inversions), find term-drift root is seed draft-status, hold cheap-editor swap for a prose bake-off 2026-07-12 20:06:03 +03:00
Claude (backend session)
c5c277615a Add owner-directed guardrail: session prompts must mandate self-review of code, model requests, and results because polygon errs and skews experiments 2026-07-12 17:44:48 +03:00
Claude (backend session)
48e3f39eb8 Freeze exp14b pre-registration: meaning-trap battery (5 construction classes, deterministic + cross-family judges) to test whether the within-chunk comprehension fix is gpt-5.4-specific, plus P1a-vs-frontier ranking 2026-07-12 15:15:03 +03:00
Claude (backend session)
7a30ab4fb6 Issue exp14b polygon prompt: meaning-trap battery to settle whether gpt-5.4 escalation is needed plus P1a-vs-frontier ranking, per owner direction 2026-07-12 05:21:56 +03:00
Claude (backend session)
7ee668fba0 Ratify D37 accepting exp14 quality-empirics with fixes: claim-1 paragraphs is a cheap prompt lever, claim-2 gpt-5.4 escalation held under-powered, archive spent exp14 prompts 2026-07-12 05:15:55 +03:00
Claude (backend session)
b59ca66b83 Land external blind-eval corroboration: two zero sessions confirm F-disc is best on both prose and fidelity, flag floating grade glossary, and catch an A-2pass fidelity catastrophe that redirects near-term to single-pass discourse 2026-07-12 04:29:32 +03:00
Claude (backend session)
73a45848db Complete Gemini test with proper max_tokens: gemini-3.1-pro also inverts the double-negation trap, so only gpt-5.4 of three frontier families fixes within-chunk comprehension 2026-07-12 04:06:25 +03:00
Claude (backend session)
a19107479d Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
Claude (backend session)
174192b1eb Fold exp14 batch-1 trap-judge corroboration: cross-family panel unanimously confirms the T1 comprehension catastrophe on all cheap arms and clears both frontier arms 2026-07-11 23:29:52 +03:00
Claude (backend session)
dc5cf3bb9d Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor 2026-07-11 22:50:09 +03:00
Claude (backend session)
353bb52676 Freeze exp14 pre-registration: quality-lever empirics on slicing×prompt 2×2 plus ablations, frontier capability-vs-activation, size curve, annotated trap set, pre-registered metrics/gates and per-key budget caps 2026-07-11 22:06:31 +03:00
Claude (backend session)
ee259bd4d0 Archive five spent session prompts, record confirmed key budgets, and keep partially-run POLYGON_PACKAGE4 active as pilot-track residual (D36.1) 2026-07-11 21:33:15 +03:00
Claude (backend session)
bf24d64e53 Ratify D36 accepting research/18 (Claude with fixes, ChatGPT clean) and issue polygon quality-empirics session prompt grafting both recommendation tables 2026-07-11 21:14:00 +03:00