Commit graph

6 commits

Author SHA1 Message Date
Claude (backend session)
b59ca66b83 Land external blind-eval corroboration: two zero sessions confirm F-disc is best on both prose and fidelity, flag floating grade glossary, and catch an A-2pass fidelity catastrophe that redirects near-term to single-pass discourse 2026-07-12 04:29:32 +03:00
Claude (backend session)
73a45848db Complete Gemini test with proper max_tokens: gemini-3.1-pro also inverts the double-negation trap, so only gpt-5.4 of three frontier families fixes within-chunk comprehension 2026-07-12 04:06:25 +03:00
Claude (backend session)
a19107479d Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better 2026-07-12 03:43:47 +03:00
Claude (backend session)
174192b1eb Fold exp14 batch-1 trap-judge corroboration: cross-family panel unanimously confirms the T1 comprehension catastrophe on all cheap arms and clears both frontier arms 2026-07-11 23:29:52 +03:00
Claude (backend session)
dc5cf3bb9d Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor 2026-07-11 22:50:09 +03:00
Claude (backend session)
353bb52676 Freeze exp14 pre-registration: quality-lever empirics on slicing×prompt 2×2 plus ablations, frontier capability-vs-activation, size curve, annotated trap set, pre-registered metrics/gates and per-key budget caps 2026-07-11 22:06:31 +03:00