textmachine/docs/experiments
2026-07-05 21:07:39 +03:00
..
00-provider-quirks.md Reconcile grok-reasoning across polygon docs: reasoning_effort:none disables it (probe bug), editor stays grok-4.3 with provider-quirks slug fixed; economics unchanged; add pilot mono-editor arm 2026-07-05 20:36:48 +03:00
01-token-calibration.md Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
02-refusal-benchmark.md Initial commit: documentation, eval polygon, backend step 0 verdict 2026-07-04 06:50:59 +03:00
03-local-stand.md Add decisions log and provider-quirks, set grok-4.3 editor default and OpenAI-retained model stack per owner and editor bake-off 2026-07-04 18:47:45 +03:00
04-editor-quality.md Correct exp04's 'grok without reasoning' characterization — grok-4.3 defaults to low reasoning (~13s), production editor sets reasoning_effort:none for 0 tokens; ranking unchanged 2026-07-05 21:07:39 +03:00
06-local-extraction.md Add D12 failure-resilience model: three-class failure taxonomy, F3 idempotency fix bounded to at-most-once, progress-surfacing spec grounded in durable-execution prior art 2026-07-04 23:10:32 +03:00
07-coverage-precision.md Land polygon session 3: coverage-gate precision (0 FP/57, flip safe), cost-model v2 on grok-editor stack, and Phase-2.5 pilot protocol with a Go-mirror memory-eval 2026-07-05 15:44:44 +03:00
08-cost-model-v2.md Reconcile grok-reasoning across polygon docs: reasoning_effort:none disables it (probe bug), editor stays grok-4.3 with provider-quirks slug fixed; economics unchanged; add pilot mono-editor arm 2026-07-05 20:36:48 +03:00
09-pilot-protocol.md Reconcile grok-reasoning across polygon docs: reasoning_effort:none disables it (probe bug), editor stays grok-4.3 with provider-quirks slug fixed; economics unchanged; add pilot mono-editor arm 2026-07-05 20:36:48 +03:00