Claude (backend session)
|
b59ca66b83
|
Land external blind-eval corroboration: two zero sessions confirm F-disc is best on both prose and fidelity, flag floating grade glossary, and catch an A-2pass fidelity catastrophe that redirects near-term to single-pass discourse
|
2026-07-12 04:29:32 +03:00 |
|
Claude (backend session)
|
73a45848db
|
Complete Gemini test with proper max_tokens: gemini-3.1-pro also inverts the double-negation trap, so only gpt-5.4 of three frontier families fixes within-chunk comprehension
|
2026-07-12 04:06:25 +03:00 |
|
Claude (backend session)
|
a19107479d
|
Land exp14 batch-2: 3-family frontier shows the comprehension fix is gpt-5.4-specific (grok inverts, gemini inconclusive), re-gate flags ~10 fidelity chunks, bigger chunks are cheaper and better
|
2026-07-12 03:43:47 +03:00 |
|
Claude (backend session)
|
174192b1eb
|
Fold exp14 batch-1 trap-judge corroboration: cross-family panel unanimously confirms the T1 comprehension catastrophe on all cheap arms and clears both frontier arms
|
2026-07-11 23:29:52 +03:00 |
|
Claude (backend session)
|
dc5cf3bb9d
|
Land exp14 batch-1: empirics show claim-1 paragraphs is a cheap prompt/pipeline lever while claim-2 within-chunk comprehension is a frontier-model-capability floor
|
2026-07-11 22:50:09 +03:00 |
|
Claude (backend session)
|
353bb52676
|
Freeze exp14 pre-registration: quality-lever empirics on slicing×prompt 2×2 plus ablations, frontier capability-vs-activation, size curve, annotated trap set, pre-registered metrics/gates and per-key budget caps
|
2026-07-11 22:06:31 +03:00 |
|