#!/usr/bin/env python3 """Ф2.5 memory-eval (INDICATIVE run + harness scaffold) — the deciding-experiment design for the memory-bank bet (research/13 §Вердикт, decisions-log D-пилот, D13.5). QUESTION (owner's central fear + go/no-go on the bank bet): is DETERMINISTIC SELECTIVE glossary injection (our bank) worth it vs just putting the WHOLE glossary in a long context, and does a WRONG injection silently degrade the translation? WMT25 three-mode protocol, RENAMED M0–M3 per D13.4 (the old C0–C3 collided with exp09 §3 CORE configurations): M0 no glossary (baseline — model's base behaviour) M1 SELECTIVE bank injection (only records whose keys fire in the chunk) — OUR bank M2 FULL glossary + sliding summary in the prompt — "just long context" M3 RANDOM/WRONG glossary (dst shuffled) — adversarial arm, directly tests silent degradation Decision rule (pre-registered, exp09 §6): M1≈M2 quality but M1 cheaper → bank justified; M2≫M1 → need fuller context; M3 drops vs M0 on FIDELITY → wrong injection harms → post-check/disposition justified (owner's fear confirmed & mitigated, not just asserted). THREE AXES, measured SEPARATELY (exp09 §6): (a) consistency — decl-aware (declmetric.py, pymorphy3 + Go-style stored-decl) approved-dst rendering across the chunks where src fires; (b) FIDELITY/omission — LLM judge of a DIFFERENT family vs the SOURCE (D13.3 family rule: judge NOT from the draft/translator model's family). This is the axis that makes the pre-registered "M3 drops vs M0 on fidelity" rule TESTABLE — without it the central owner-fear measurement is unverifiable (07-review §4.5). CometKiwi is OPTIONAL and only with the caveat "not validated on ru-literature" (research/11-gap-2) — not wired here. (c) cost — input/output tokens per mode. INJECTION = eval-side MIRROR of backend/internal/pipeline/memory.go (THE SOURCE OF TRUTH), synced D13.5 across the 7 divergences the review flagged (07-review §4.5): 1. sticky depth = stickyDepth(2), union window, reset at chapter boundary (runner.go:534-560); 2. until_ch spoiler gate (memory.go:374-382), not only since_ch; 3. span-containment gated on a SPOILER-VALID suppressor (memory.go:591-617); 4. token budget + F2 eviction in priority order (memory.go:262-370); 5. sticky disposition INHERITANCE — a collision-prone AMBIGUOUS carried by sticky stays AMBIGUOUS, not upgraded to CONFIRMED. The ratified D16.2 fix, now LANDED in Go v3 (memoryMatchVersion memmatch-v3 "sticky-inherit-disp"); the mirror matches the v3 contract and was re-verified against the landed Go by the package-2 self-test; 6. FULL vendored trad2simp table (eval/pilot/trad2simp.txt, hash-synced with the Go embed), not a 12-entry stub; 7. default-ignorables = unicode.Cf + variation selectors FE00..FE0F / E0100..E01EF (memnorm.go isIgnorableForMatch), not one hardcoded VS char; 8. source word-boundary for spaced-phonetic keys (D16.3, memory.go suppressUnboundedPhonetic) — a Latin/Cyrillic key fires only as a WHOLE WORD (rose⊄roseanne), a cross-script neighbour is a valid boundary; KANA/HAN keys are NOT boundary-checked (no word spaces; name+particle recall). Package-2 Task 1a — the mirror previously described this carve-out but did not apply it in select(); now applied, parity-covered by TestSourcePhoneticWordBoundary + the kana carve-out; 9. apostrophe-variant fold ‘ ’ ʼ ' → ' on the source side (norm_src) AND the target side (declmetric.normalize_target), symmetric with Go isApostropheVariant (package-2 Task 1b) — critical for the en→ru pilot arm (д'Артаньян / O'Brien). Han-script tests cover ALL planes (Ext A–I + non-decomposing compat), and significantLen counts Unicode LETTERS only (mirrors Go's unicode.IsLetter branch; excludes kana middle-dot U+30FB and combining marks) — closes the adversarial-review deltas #1/#2 (supplementary-plane names). KNOWN LATENT GAPS (package-2 parity review, verified NOT reachable for the zh acceptance book / Ah-Q pilot, so left as-is; documented, not silent): (2) the fixed-range kana/hangul/Han anchor tests in spaced_phonetic_script/_is_han approximate Go's full unicode script tables — they miss exotic ranges (Kana Phonetic-Ext U+31F0.., conjoining Jamo, CJK radicals, iteration marks 々〆) and 84 unassigned F900 code points, but none FLIP spaced_phonetic_script's result for a realistic zh/ja/en key (a real key always carries a true anchor); (4) str.isspace() treats C0 separators U+001C–1F as space where Go does not — cosmetic (both are non-letter boundaries). U+0130 İ (the sole str.lower()≠Go divergence) and the glossaryLineTokens kana marks (ー・) ARE fixed above. On any divergence the Go code wins. (eval/memory_hotpath.py is the PRE-cc57c7b prototype — do NOT reuse.) ECHO CONTROL (D13.5 / 07-review §4.5): the model ECHOES the injected glossary preamble ("ГЛОССАРИЙ … 阿Q → А-кью …") before the translation. Scoring the whole output then counts the echoed dst forms as "rendered" → M-modes inflated, and the ADVERSARIAL M3 harm UNDER-counted (the review traced chunk1-M3 inflated 0.667→1.0). We now (a) DETECT the preamble/source echo (reusing refusal_bench's CJK classifier), (b) STRIP it to the translation body before scoring, and (c) FLAG the datapoint echo_contaminated and EXCLUDE it from the clean headline aggregate (optionally --regen-on-echo re-asks once with a hardened prompt) — never a silent strip. Provenance (eval rule #1): FULL raw outputs saved (not out[:160]); every run stamps model id + UTC date + usage; a NEW output file per run (never overwrite — exp02 lost 66 raw calls that way). Usage: eval/.venv/bin/python eval/pilot/memory_eval.py --dry-run # selection mechanism eval/.venv/bin/python eval/pilot/memory_eval.py --self-test # mirror invariants (no API) eval/.venv/bin/python eval/pilot/memory_eval.py --chapter 6 # indicative M0–M3 on Ah-Q """ from __future__ import annotations import argparse, json, re, sys, time, unicodedata from datetime import datetime, timezone from pathlib import Path sys.path.insert(0, str(Path(__file__).resolve().parent)) sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) import refusal_bench as rb import declmetric as dm ROOT = Path(__file__).resolve().parent SAMPLES = ROOT.parent / "data" / "samples" OUTDIR = ROOT.parent / "data" / "pilot" OUTDIR.mkdir(parents=True, exist_ok=True) TRAD_TABLE = ROOT / "trad2simp.txt" # vendored from backend/internal/pipeline/data/trad2simp.txt # --- normalization (faithful mirror of memnorm.go) -------------------------------- def _load_trad2simp(path: Path) -> dict[str, str]: """Parse the vendored ' ' table (mirror of parseTradTable). A corrupt line is a loud error — the same fail-loud contract as Go, so a truncated table can't silently re-open the A4 orthography hole. The file has 508 content lines → 500 unique trad keys: 8 chars (東決莊語說談關飛) are listed under two radical families with the SAME simp target — benign consistent duplicates (later-wins, exactly like Go). A duplicate with a DIFFERENT target is a typo → fail loud (mirror of Go's parseTradTable panic), package-2 Task 1d.""" m: dict[str, str] = {} if not path.exists(): raise SystemExit(f"memory_eval: vendored trad2simp table missing: {path}\n" f" cp backend/internal/pipeline/data/trad2simp.txt {path}") for i, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1): s = line.strip() if not s or s.startswith("#"): continue fields = s.split() if len(fields) != 2 or len(fields[0]) != 1 or len(fields[1]) != 1: raise SystemExit(f"memory_eval: trad2simp line {i}: want 2 single-rune fields, got {s!r}") tr, si = fields if tr in m and m[tr] != si: # inconsistent duplicate = typo → loud, like Go raise SystemExit(f"memory_eval: trad2simp line {i}: {tr!r} maps to both {m[tr]!r} and {si!r}") m[tr] = si return m TRAD2SIMP = _load_trad2simp(TRAD_TABLE) def _is_ignorable(c: str) -> bool: """Mirror of Go isIgnorableForMatch: unicode.Cf + variation selectors FE00..FE0F / E0100..E01EF.""" o = ord(c) return unicodedata.category(c) == "Cf" or (0xFE00 <= o <= 0xFE0F) or (0xE0100 <= o <= 0xE01EF) # Typographic apostrophe/single-quote variants that fold to ASCII '\'' (mirror of Go # isApostropheVariant): U+2018/2019 curly single quotes, U+02BC modifier-letter apostrophe, # U+FF07 fullwidth apostrophe. Smart-quote editors turn a source key O'Brien's apostrophe into # U+2019, which NFKC does NOT fold to U+0027 — leaving the key silently un-matchable on en names. _APOSTROPHE_VARIANTS = frozenset("‘’ʼ'") def _is_letter(c: str) -> bool: return unicodedata.category(c)[0] == "L" # unicode.IsLetter equivalent (L*) def _is_han(c: str) -> bool: """Full-plane Han (mirror of Go unicode.Is(unicode.Han)): Ext A, URO, compat, Ext B–I, compat-supplement. Covers supplementary-plane ideographs (rare surnames/fantasy names a webnovel corpus can hit) that BMP-only ranges miss — adversarial-review delta #2.""" o = ord(c) return (0x3400 <= o <= 0x4DBF or 0x4E00 <= o <= 0x9FFF or 0xF900 <= o <= 0xFAFF or 0x20000 <= o <= 0x2A6DF or 0x2A700 <= o <= 0x2EBEF or 0x2F800 <= o <= 0x2FA1F or 0x30000 <= o <= 0x323AF) def norm_src(s: str) -> str: """Mirror of memnorm.go normalizeSourceKey: NFKC → apostrophe-variant fold → drop ignorables → trad→simp (per rune) → katakana→hiragana → Unicode lower. Applied SYMMETRICALLY to keys and chunk text (the apostrophe fold keeps O'Brien matchable across smart-quote/ASCII forms).""" s = unicodedata.normalize("NFKC", s) out = [] for c in s: if c in _APOSTROPHE_VARIANTS: c = "'" # typographic apostrophe (O'Brien) folds to ASCII so a key matches either form if _is_ignorable(c): continue c = TRAD2SIMP.get(c, c) o = ord(c) if 0x30A1 <= o <= 0x30F6: # katakana → hiragana c = chr(o - 0x60) if c == "İ": # U+0130: Go unicode.ToLower→'i'; Python str.lower→'i̇' (i+U+0307). Match Go out.append("i"); continue # the ONLY rune where .lower() diverges from Go (parity-review #1) out.append(c.lower()) return "".join(out) def significant_len(nk: str) -> int: """Mirror of memnorm.go significantLen: count Unicode LETTERS (category L*). Since Han/kana/ hangul letters are all category L across every plane, this subsumes Go's explicit Han/Hiragana/Katakana/Hangul clause AND excludes the non-letters it also excludes (kana middle-dot U+30FB Po, combining marks Mn) — more faithful than a fixed block range.""" return sum(1 for c in nk if _is_letter(c)) def any_han(nk: str) -> bool: return any(_is_han(c) for c in nk) def min_key_len_for(nk: str) -> int: return 2 if any_han(nk) else 3 # minKeyLenHan / minKeyLenPhonetic def collision_prone(nk: str) -> bool: """Short AND purely phonetic (no Han anchor) → likely to fire inside an unrelated word → injected AMBIGUOUS (memory.go collisionProneKey; threshold = minKeyLenPhonetic=3).""" return (not any_han(nk)) and significant_len(nk) <= 3 # --- source word-boundary for spaced-phonetic keys (D16.3, memory.go suppressUnboundedPhonetic) -- def _is_latin_letter(c: str) -> bool: return _is_letter(c) and "LATIN" in unicodedata.name(c, "") def _is_cyrillic_letter(c: str) -> bool: return _is_letter(c) and "CYRILLIC" in unicodedata.name(c, "") def spaced_phonetic_script(norm_key: str) -> str | None: """Mirror of memory.go spacedPhoneticScript: the word-delimited alphabetic script a normalized key belongs to ('latin'|'cyrillic'), or None when the key carries a Han/kana/ hangul anchor (no word segmentation) or MIXES the two spaced scripts. Digits/punctuation don't set the script but don't disqualify (so "o'brien" is still latin).""" script = None for c in norm_key: o = ord(c) if _is_han(c) or (0x3040 <= o <= 0x30FF) or (0xAC00 <= o <= 0xD7A3): return None # ideographic/kana/hangul anchor is never boundary-checked here if _is_latin_letter(c): if script == "cyrillic": return None # mixed spaced scripts — do not boundary-check script = "latin" elif _is_cyrillic_letter(c): if script == "latin": return None script = "cyrillic" return script def _letter_in_script(c: str, script: str) -> bool: """Mirror of memory.go letterInScript: r is a LETTER of the given spaced script — the same-script letter that breaks a word boundary. A digit/punct/space or a letter of another script is a valid boundary and returns False.""" if script == "latin": return _is_latin_letter(c) if script == "cyrillic": return _is_cyrillic_letter(c) return False def unbounded_phonetic(key: str, start: int, end: int, ntext: str) -> bool: """Mirror of memory.go suppressUnboundedPhonetic's per-occurrence test: True → SUPPRESS. A Latin/Cyrillic key that fired INSIDE a longer same-script word ("rose" in "roseanne"/ "roses") is not a whole-entity match. A cross-script neighbour (a Latin name abutting a Han char in unspaced source) is a valid boundary → keep. Kana/Han keys return script None → never suppressed (Japanese has no word spaces; name+particle recall — see kana_precision).""" script = spaced_phonetic_script(key) if script is None: return False before_bad = start > 0 and _letter_in_script(ntext[start - 1], script) after_bad = end < len(ntext) and _letter_in_script(ntext[end], script) return before_bad or after_bad STICKY_DEPTH = 2 # memory.go stickyDepth # --- disposition (mirror of memory.go dispositionFor + injectionDisposition) ------ CONFIRMED, AMBIGUOUS, REJECT = "confirmed", "ambiguous", "reject" def disposition_for(status: str, via: str, allow_short: bool) -> str: """Exact-match disposition. approved → confirmed; auto/draft → ambiguous; an approved entry matched by a COLLISION-PRONE key is downgraded to ambiguous (A2 surface-collision). via=='sticky' NEVER reaches here in the D16.2 mirror — sticky INHERITS its disposition.""" if via != "sticky" and not allow_short and collision_prone(via): return AMBIGUOUS return CONFIRMED if status == "approved" else AMBIGUOUS # --- glossary (Ah-Q) as store-like rows ------------------------------------------- # Fields mirror store.GlossaryEntry columns the hot path reads. One SPOILER entry (革命党, # since_ch=7) demonstrates the since gate; 尼姑 is status=draft to exercise AMBIGUOUS-by-status. GLOSSARY = [ {"src": "阿Q", "dst": "А-кью", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "name", "lemma_keys": ["а-кью"], "hyphen_pat": r"А[-\s]?кью", "decl_forms": ["А-кью"]}, {"src": "未庄", "dst": "Вэйчжуан", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "place", "lemma_keys": ["вэйчжуан"], "decl_forms": ["Вэйчжуане", "Вэйчжуан", "Вэйчжуана"]}, {"src": "赵太爷", "dst": "почтенный Чжао", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": ["赵老太爷"], "type": "name", "lemma_keys": ["чжао"], "decl_forms": ["почтенного Чжао", "почтенный Чжао", "Чжао"]}, {"src": "王胡", "dst": "Бородатый Ван", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": ["王癞胡", "癞胡"], "type": "name", "lemma_keys": ["ван"], "decl_forms": ["Бородатого Вана", "Бородатый Ван", "Бородатому Вану"]}, {"src": "小D", "dst": "Маленький Дэн", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "name", "lemma_keys": ["дэн"], "decl_forms": ["Маленького Дэна", "Маленький Дэн", "Дэна"]}, {"src": "假洋鬼子", "dst": "Поддельный заморский чёрт", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "nickname", "lemma_keys": ["заморский", "чёрт"], "decl_forms": ["Поддельного заморского чёрта", "заморский чёрт"]}, {"src": "吴妈", "dst": "У-ма", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "name", "lemma_keys": ["у-ма"], "hyphen_pat": r"У[-\s]?ма", "decl_forms": ["У-ма", "У-мы"]}, {"src": "秀才", "dst": "сюцай", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "title", "lemma_keys": ["сюцай"], "decl_forms": ["сюцая", "сюцай", "сюцаю"]}, {"src": "尼姑", "dst": "монашка", "status": "draft", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "term", "lemma_keys": ["монашка"], "decl_forms": ["монашку", "монашка", "монашки"]}, # SPOILER demo: only valid from chapter 7 (revolution). Injected earlier → hard reject. {"src": "革命党", "dst": "революционеры", "status": "approved", "sense": "", "since_ch": 7, "until_ch": 0, "allow_short": False, "aliases": [], "type": "term", "lemma_keys": ["революционер"], "decl_forms": ["революционеров", "революционеры"]}, ] def eligible_keys(entry: dict) -> list[str]: """Normalized source surfaces (src + aliases) passing the per-lang min-key ban (A3). A record with empty dst is NOT matchable (renders nothing) — mirror memory.go:167-173.""" if not str(entry.get("dst", "")).strip(): return [] out = [] for raw in [entry["src"]] + entry.get("aliases", []): nk = norm_src(raw) if nk and (significant_len(nk) >= min_key_len_for(nk) or entry.get("allow_short")): out.append(nk) return out def spoiler_blocked(entry: dict, chapter: int) -> bool: """Mirror of memory.go spoilerBlocked: since_ch/until_ch window (0 = unbounded).""" if entry.get("since_ch", 0) and chapter < entry["since_ch"]: return True if entry.get("until_ch", 0) and chapter > entry["until_ch"]: return True return False def glossary_line_tokens(entry: dict) -> int: """Mirror of memory.go glossaryLineTokens: cjk + other/3 over 'src → dst' (floor-free). cjk bucket = full-plane Han + kana block + hangul (matches Go unicode.In(Han,Hiragana, Katakana,Hangul); supplementary-plane Han now weighted ×1, not ÷3).""" cjk = other = 0 # Go glossaryLineTokens buckets by unicode.In(Han,Hiragana,Katakana,Hangul), which keys on SCRIPT # (not Script_Extensions). Mirror the marks that BMP Han/kana ranges miss (D19.5(д) token-bucket # sync; each verified against Go's tables via a scratch unicode.In probe): 々 U+3005 and 〇 U+3007 # are Script=Han → CJK (weight 1); halfwidth-katakana LETTERS U+FF66–FF9D are Script=Katakana → CJK. # 〆 U+3006 is NOT Han in Go (→ other) so it is deliberately NOT added. kana_nonletter EXCLUDES the # Script=Common marks Go counts as `other`: ゠・ー (U+30A0/30FB/30FC), combining U+3099/309A, and # the SPACING voiced marks ゛゜ U+309B/309C (previously mis-counted as CJK — the parity gap this fixes). han_marks = {0x3005, 0x3007} kana_nonletter = {0x30A0, 0x30FB, 0x30FC, 0x3099, 0x309A, 0x309B, 0x309C} for c in f"{entry['src']} → {entry.get('dst','')}": o = ord(c) if (_is_han(c) or o in han_marks or (0x3040 <= o <= 0x30FF and o not in kana_nonletter) or (0xFF66 <= o <= 0xFF9D) or (0xAC00 <= o <= 0xD7A3)): cjk += 1 elif c.isspace(): pass else: other += 1 return cjk + other // 3 def _key_occurrences(ntext: str, key: str) -> list[tuple[int, int]]: """All [start,end) occurrences of key in ntext (rune indices == str indices in Python).""" spans, i = [], ntext.find(key) while i != -1: spans.append((i, i + len(key))) i = ntext.find(key, i + 1) return spans def select(chunk: str, chapter: int, sticky_prev: dict[str, str], budget_tokens: int = 0) -> dict: """Faithful mirror of MemoryBank.Select for ONE chunk: span-match → spoiler-gated containment → disposition → sticky (inherited disposition) → priority token budget. sticky_prev: {id -> inherited disposition} from the prior ≤stickyDepth chunks of this chapter. Returns injected/rejected/evicted lists + active {id->disposition} for the next chunk's sticky window.""" ntext = norm_src(chunk) # 1. all key occurrences with owning entry indices occ = [] # (start, end, key, entry_idx) for ei, e in enumerate(GLOSSARY): for k in eligible_keys(e): for (s, t) in _key_occurrences(ntext, k): occ.append((s, t, k, ei)) occ.sort(key=lambda m: (m[0], m[1], m[2])) # 1b. source word-boundary for spaced-phonetic keys (D16.3, memory.go suppressUnboundedPhonetic), # BEFORE containment — a Latin/Cyrillic key fired inside a longer same-script word # ("rose" ⊂ "roseanne") is not a whole-entity match. Kana/Han keys are NOT boundary-checked. occ = [m for m in occ if not unbounded_phonetic(m[2], m[0], m[1], ntext)] # 2. longest-match containment, suppressor gated on spoiler-VALIDITY (memory.go:591-617) def valid_suppressor(m): return not spoiler_blocked(GLOSSARY[m[3]], chapter) kept = [] for i, m in enumerate(occ): contained = False for j, o in enumerate(occ): if i == j: continue if o[0] <= m[0] and o[1] >= m[1] and (o[1]-o[0]) > (m[1]-m[0]) and valid_suppressor(o): contained = True break if not contained: kept.append(m) # 3. surviving occ → entries, longest firing key per entry matched_via: dict[int, str] = {} for (s, t, k, ei) in kept: if ei not in matched_via or len(k) > len(matched_via[ei]): matched_via[ei] = k injected_hits: dict[int, dict] = {} # entry_idx -> picked rejected = [] active: dict[str, str] = {} # id -> disposition (exact matches, for next sticky) for ei in range(len(GLOSSARY)): via = matched_via.get(ei) if via is None: continue e = GLOSSARY[ei] eid = _eid(e) if spoiler_blocked(e, chapter): rejected.append({**e, "_via": via, "_disp": REJECT}) continue disp = disposition_for(e["status"], via, e.get("allow_short", False)) injected_hits[ei] = {**e, "_via": via, "_disp": disp} active[eid] = disp # exact-matched ids feed the next chunk's sticky (NOT sticky-carried) # 4. sticky scene-inertia: carry prior-chunk exact matches not re-matched here, INHERITING # their disposition (D16.2), unless the spoiler window blocks them (spoiler beats sticky). hit_ids = {_eid(p) for p in injected_hits.values()} for ei, e in enumerate(GLOSSARY): eid = _eid(e) if eid not in sticky_prev or eid in hit_ids or ei in injected_hits: continue if spoiler_blocked(e, chapter): continue injected_hits[ei] = {**e, "_via": "sticky", "_disp": sticky_prev[eid]} # 5. priority order + F2 token budget/eviction (memory.go:342-368) picked = list(injected_hits.values()) picked.sort(key=_priority_rank) evicted = [] if budget_tokens > 0: used, cut = 0, len(picked) for i, p in enumerate(picked): used += glossary_line_tokens(p) if used > budget_tokens: cut = i break picked, evicted = picked[:cut], picked[cut:] return {"injected": picked, "rejected": rejected, "evicted": evicted, "active": active} def _eid(e: dict) -> str: return f'{e["src"]}\x1f{e.get("sense","")}\x1f{e.get("since_ch",0)}\x1f{e.get("until_ch",0)}' def _priority_rank(p: dict) -> int: rank = 0 if p["_disp"] != CONFIRMED: rank |= 1 << 2 if p["_via"] == "sticky": rank |= 1 << 1 if p["status"] != "approved": rank |= 1 << 0 return rank # --- prompt assembly per mode ----------------------------------------------------- SYSTEM = ("Ты профессиональный литературный переводчик. Переведи фрагмент китайского " "произведения на русский язык. Сохрани все реплики и детали без пропусков; стиль — " "живой литературный русский. Выведи ТОЛЬКО перевод — НЕ повторяй глоссарий и инструкции, " "начни сразу с текста перевода.") SYSTEM_HARDENED = SYSTEM + (" ВАЖНО: не выводи блок ГЛОССАРИЙ, строки вида «иероглиф → перевод», " "заголовки или пояснения — только сам художественный перевод.") # Byte-synced with Go glossaryBlockHeader (memory.go) — the ⟨проверить⟩ clause was missing from # the mirror (package-2 Task 1d), so the injected M1 prompt now matches what the hot path emits. GHEAD = ("ГЛОССАРИЙ (используй эти утверждённые переводы имён и терминов последовательно; " "строки с пометкой ⟨проверить⟩ — неподтверждённые кандидаты):") def gloss_block(entries: list[dict], full: bool = False) -> str: if not entries: return "" lines = [] for e in entries: line = f"{e['src']} → {e['dst']}" if not full and e.get("_disp") == AMBIGUOUS: line += " ⟨проверить⟩" lines.append(line) return GHEAD + "\n" + "\n".join(lines) def build_prompt(chunk: str, mode: str, sel: dict, summary: str, shuffled: dict) -> str: if mode == "M0": return "ФРАГМЕНТ:\n" + chunk if mode == "M1": # selective bank blk = gloss_block(sel["injected"]) return (blk + "\n\n" if blk else "") + "ФРАГМЕНТ:\n" + chunk if mode == "M2": # full glossary + sliding summary (long context) blk = gloss_block([{**e, "_disp": CONFIRMED} for e in GLOSSARY if not e.get("since_ch") and not e.get("until_ch") and str(e.get("dst", "")).strip()], full=True) pre = (f"КРАТКОЕ СОДЕРЖАНИЕ ПРЕДЫДУЩЕГО:\n{summary}\n\n" if summary else "") return pre + blk + "\n\nФРАГМЕНТ:\n" + chunk if mode == "M3": # random/wrong glossary (adversarial): fired entries with SHUFFLED dst wrong = [{**e, "dst": shuffled[e["src"]], "_disp": CONFIRMED} for e in sel["injected"] if e["_via"] != "sticky"] blk = gloss_block(wrong, full=True) return (blk + "\n\n" if blk else "") + "ФРАГМЕНТ:\n" + chunk raise ValueError(mode) # --- echo control (D13.5 / 07-review §4.5) ---------------------------------------- GLOSS_ECHO_MARKERS = ("ГЛОССАРИЙ", "глоссарий", "КАНОНИЧЕСКИЕ ПЕРЕВОД", "используй эти утверждённые", "использовать последовательно", "используйте эти", "КРАТКОЕ СОДЕРЖАНИЕ", "⟨проверить", "проверить⟩") ARROW_RE = re.compile(r".+\s*(→|->|=>)\s*.+") def _is_gloss_echo_line(s: str) -> bool: """A single output line is a glossary/instruction ECHO if it is either (a) a HEADER echo — the line STARTS with a gloss marker (case-insensitive) — or (b) a 'src → dst' gloss ENTRY echo — an arrow with CJK on the LEFT (the source term), a SHORT line, and NO sentence-ending punctuation. Used for BOTH the leading preamble and the embedded/trailing echo scan (package-2 Task 1c). Both rules are TIGHTENED vs the first cut (parity-review #5): anchoring the marker to the line start no longer strips prose that merely contains «глоссарий» mid-sentence, and the entry rule's period/length guard no longer strips prose like «Путь (道) → его судьба была предрешена.». A prose source-echo (untranslated CJK with NO arrow) is deliberately NOT matched here — cjk_share catches it. KNOWN BOUNDARY (D19.5(д), documented not fixed): line-start anchoring still FALSE-POSITIVES on a body line that GENUINELY begins with a marker word — e.g. prose «Глоссарий рода был утерян в пожаре.» starting its own line is flagged as a header echo and dropped from the scored body. This is a deliberate CONSERVATIVE choice, not a silent one: the datapoint carries a visible glossary_preamble_echo / embedded_glossary_echo reason (and is marked echo_contaminated), so an over-strip is auditable and never inflates the clean headline unseen. A tighter fix (require a trailing ':' or an adjacent arrow-line) is deferred — over-flagging a rare sentence-initial «Глоссарий» is safe; letting a real echo glossary through is not.""" s = s.strip() if not s: return False sl = s.lower() if any(sl.startswith(m.lower()) for m in GLOSS_ECHO_MARKERS): return True m = ARROW_RE.match(s) if m: left = re.split(r"→|->|=>", s, 1)[0] if rb.CJK_RE.search(left) and len(s) <= 80 and not re.search(r"[.!?…]", s): return True return False def detect_echo(raw: str) -> tuple[str, list[str]]: """Strip glossary/instruction-echo to the translation body AND report contamination reasons. Two scopes (package-2 Task 1c — the old detector only caught the LEADING preamble, so a trailing/embedded echo glossary «…текст…\\nГЛОССАРИЙ:\\n阿Q → А-кью» silently inflated consistency): (1) a LEADING preamble region (markers / arrow-gloss / blanks) is skipped; (2) any REMAINING line that is a gloss-echo line (a marker, or a 'src → dst' arrow line with CJK, ANYWHERE in the body — tail or embedded) is dropped from the scored body and flagged. Also flags source-echo (CJK leaked into the body, the DeepSeek-style untranslated echo).""" lines = raw.split("\n") # (1) leading preamble region body_start, saw_pre = 0, False for i, ln in enumerate(lines): s = ln.strip() if not s: if saw_pre: body_start = i + 1 continue if _is_gloss_echo_line(s): saw_pre, body_start = True, i + 1 continue break # first prose line → body begins # (2) embedded/trailing echo lines anywhere in the remaining body → drop from the scored body kept, n_embedded = [], 0 for ln in lines[body_start:]: if _is_gloss_echo_line(ln): n_embedded += 1 continue kept.append(ln) body = "\n".join(kept).strip() reasons = [] if saw_pre: reasons.append("glossary_preamble_echo") if n_embedded: reasons.append(f"embedded_glossary_echo({n_embedded})") if not body: # whole output was preamble/echo → keep raw but flag hard reasons.append("empty_after_strip") body = raw.strip() ces = rb.cjk_share(body) if ces > 0.10: reasons.append(f"source_echo(cjk={ces:.0%})") return body, reasons # --- fidelity judge (cross-family, D13.3) ----------------------------------------- FAMILY = {"deepseek": "deepseek", "grok": "xai", "gemini": "google", "gpt": "openai", "glm": "zhipu", "kimi": "moonshot", "mistral": "mistral", "qwen": "alibaba"} def family_of(model: str) -> str: for k, v in FAMILY.items(): if k in model: return v return model FIDELITY_SYSTEM = ( "Ты — строгий редактор-оценщик ВЕРНОСТИ художественного перевода. Тебе дают ИСХОДНЫЙ " "фрагмент (китайский) и его ПЕРЕВОД на русский. Оцени ТОЛЬКО верность исходнику (не стиль): " "искажения смысла (mistranslation), пропуски (omission), отсебятину (addition), особенно " "НЕВЕРНО переданные имена/термины. Верни СТРОГО JSON без пояснений: " '{"fidelity": <0-100>, "n_mistranslation": , "n_omission": , "n_addition": , ' '"wrong_names": [<строки>], "notes": "<кратко>"}. fidelity=100 — идеально верно; ниже — за ' "каждое искажение/пропуск. Имена, переданные не как в источнике, — это mistranslation.") def fidelity_judge(source: str, translation: str, judge_model: str, translator_model: str, providers: dict) -> dict: """Cross-family fidelity judge vs the SOURCE (D13.3: judge family != translator family).""" if family_of(judge_model) == family_of(translator_model): raise ValueError(f"fidelity judge {judge_model} shares family with translator " f"{translator_model} ({family_of(judge_model)}) — violates D13.3") user = (f"ИСХОДНЫЙ ФРАГМЕНТ (китайский):\n{source}\n\n" f"ПЕРЕВОД НА РУССКИЙ (оцени его верность):\n{translation}\n\nВерни JSON.") p = providers[judge_model] txt, err, usage = _retry_call({"name": judge_model, "model": p["model"], **p}, FIDELITY_SYSTEM, user, timeout=200) if err or not txt: return {"error": err or "empty", "usage": usage or {}} m = re.search(r"\{.*\}", txt, re.S) if not m: return {"error": "no-json", "raw": txt[:300], "usage": usage or {}} try: out = json.loads(m.group(0)) except json.JSONDecodeError: return {"error": "bad-json", "raw": txt[:300], "usage": usage or {}} out["usage"] = usage or {} out["judge_model"] = judge_model return out # --- providers (translator + judge) ----------------------------------------------- PROVIDERS = { "deepseek-v4-flash": {"base_url": "https://api.deepseek.com/v1", "api_key_env": "DEEPSEEK_API_KEY", "model": "deepseek-v4-flash", "max_tokens": 8000}, "grok-4.20-0309-non-reasoning": {"base_url": "https://api.x.ai/v1", "api_key_env": "XAI_API_KEY", "model": "grok-4.20-0309-non-reasoning", "max_tokens": 8000, "temperature": 0.3}, "gemini-2.5-flash": {"base_url": "https://generativelanguage.googleapis.com/v1beta/openai", "api_key_env": "GEMINI_API_KEY", "model": "gemini-2.5-flash", "max_tokens": 8000, "extra_body": {"extra_body": {"google": {"thinking_config": {"thinking_budget": 0}}}}}, "gpt-5-mini": {"base_url": "https://api.openai.com/v1", "api_key_env": "OPENAI_API_KEY", "model": "gpt-5-mini", "reasoning_params": True, "max_tokens": 8000, "extra_body": {"reasoning_effort": "minimal"}}, } def _retry_call(p: dict, system: str, user: str, timeout: int = 200, attempts: int = 3): """rb.call_provider with backoff on TRANSIENT errors (503 overload / 429 rate / 5xx / transport) so a provider spike does not punch a hole in the fidelity axis or a translation. A content refusal / config error is NOT retried (it is a real signal).""" last = (None, "unknown", {}) for i in range(attempts): t, err, usage = rb.call_provider(p, system, user, timeout=timeout) if not err or not re.match(r"(http\|(503|429|500|502|504))|transport\|", err or ""): return t, err, usage last = (t, err, usage) time.sleep(2 * (i + 1)) return last def translate(prompt: str, model: str, hardened: bool = False) -> tuple[str, dict]: p = PROVIDERS[model] sysmsg = SYSTEM_HARDENED if hardened else SYSTEM t, err, usage = _retry_call({"name": model, **p}, sysmsg, prompt, timeout=200) return (t or ""), (usage or {}) # --- chunking (reuse the coverage harness's chunker semantics) -------------------- def chunk_text(text: str, target=1200) -> list[str]: paras = [p.strip() for p in re.split(r"\n\s*\n", text) if p.strip()] chunks, buf = [], [] for p in paras: if buf and len("".join(buf)) + len(p) > target: chunks.append("\n\n".join(buf)); buf = [] buf.append(p) if buf: chunks.append("\n\n".join(buf)) return chunks # --- self-test: mirror invariants without any API call ---------------------------- def self_test() -> int: fails = [] def check(cond, msg): if not cond: fails.append(msg) # since gate: 革命党 rejected at ch6, injected at ch7 s6 = select("革命党来了,未庄震动。", 6, {}) check(any(r["src"] == "革命党" for r in s6["rejected"]), "since-gate: 革命党 not rejected at ch6") s7 = select("革命党来了,未庄震动。", 7, {}) check(any(p["src"] == "革命党" for p in s7["injected"]), "since-gate: 革命党 not injected at ch7") # until gate (synthetic): a term valid only until ch3 must be rejected at ch10 global GLOSSARY saved = GLOSSARY GLOSSARY = saved + [{"src": "旧党", "dst": "старая партия", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 3, "allow_short": False, "aliases": [], "type": "term", "lemma_keys": ["партия"], "decl_forms": ["старой партии", "старая партия"]}] check(spoiler_blocked(GLOSSARY[-1], 10), "until-gate: 旧党 not blocked at ch10") check(not spoiler_blocked(GLOSSARY[-1], 2), "until-gate: 旧党 wrongly blocked at ch2") GLOSSARY = saved # collision downgrade: a short phonetic key (allow_short off) → AMBIGUOUS even if approved check(disposition_for("approved", "ai", False) == AMBIGUOUS, "collision: 'ai' not downgraded") check(disposition_for("approved", "мерос", False) == CONFIRMED, "collision: 5-char key wrongly downgraded") check(disposition_for("approved", "リン", False) == AMBIGUOUS, "collision: 2-kana key not downgraded") # supplementary-plane Han is Han-anchored (floor 2, CONFIRMED), not phonetic — review delta #2 check(any_han("\U00020000") and not collision_prone("\U00020000\U00020001"), "supplementary-plane Han not treated as Han") # significant_len counts letters only (kana middle-dot U+30FB excluded) — review delta #1 check(significant_len("リ・ン") == 2, "significant_len counts kana middle-dot U+30FB") # sticky depth-2 + disposition inheritance: a draft term inherits AMBIGUOUS across a # pronominal chunk, then falls out after depth 2. win = [] def union(w): # mirror unionSticky out = {} for d in w: out.update(d) return out c0 = select("这尼姑是谁。", 6, union(win)); win = (win + [c0["active"]])[-STICKY_DEPTH:] check(any(p["src"] == "尼姑" and p["_disp"] == AMBIGUOUS for p in c0["injected"]), "sticky: 尼姑(draft) not AMBIGUOUS on exact match") c1 = select("他走了。", 6, union(win)); win = (win + [c1["active"]])[-STICKY_DEPTH:] sk = [p for p in c1["injected"] if p["src"] == "尼姑"] check(sk and sk[0]["_via"] == "sticky" and sk[0]["_disp"] == AMBIGUOUS, "sticky: 尼姑 not carried with INHERITED ambiguous disposition (D16.2)") # depth-2 window: a term exact-matched at c0 is carried across c1 AND c2 (the window holds the # last stickyDepth active sets), then falls out at c3 — mirror of runner.go unionSticky/trim. c2 = select("又走了。", 6, union(win)); win = (win + [c2["active"]])[-STICKY_DEPTH:] check(any(p["src"] == "尼姑" and p["_via"] == "sticky" for p in c2["injected"]), "sticky: 尼姑 should still be carried at depth 2 (c2)") c3 = select("再走了。", 6, union(win)); win = (win + [c3["active"]])[-STICKY_DEPTH:] check(not any(p["src"] == "尼姑" for p in c3["injected"]), "sticky: 尼姑 should fall out after the depth-2 window (c3)") # spoiler-gated containment (synthetic): a spoiler-blocked longer key must NOT suppress a # valid shorter nested name. GLOSSARY = saved + [ {"src": "林", "dst": "Линь", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": True, "aliases": [], "type": "name", "lemma_keys": ["линь"], "decl_forms": ["Линь"]}, {"src": "林动的父亲", "dst": "отец Линь Дуна", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 3, "allow_short": False, "aliases": [], "type": "name", "lemma_keys": ["отец"], "decl_forms": ["отца Линь Дуна"]}] sc = select("林动的父亲来了。", 10, {}) check(any(p["src"] == "林" for p in sc["injected"]), "containment: 林 suppressed by spoiler-blocked longer key at ch10") GLOSSARY = saved # source phonetic word-boundary (D16.3, mirror of Go TestSourcePhoneticWordBoundary): a # Latin/Cyrillic key fires only as a WHOLE WORD, a cross-script neighbour is a valid boundary. GLOSSARY = saved + [ {"src": "rose", "dst": "Роза", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "name", "lemma_keys": ["роза"], "decl_forms": ["Роза"]}] check(len(select("roseanne walked in.", 1, {})["injected"]) == 0, "src-boundary: latin key fired inside 'roseanne'") check(len(select("the roses bloomed.", 1, {})["injected"]) == 0, "src-boundary: latin key fired inside 'roses'") rz = select("a rose bloomed.", 1, {})["injected"] check(any(p["src"] == "rose" and p["_disp"] == CONFIRMED for p in rz), "src-boundary: standalone latin key must fire CONFIRMED") check(any(p["src"] == "rose" for p in select("rose.", 1, {})["injected"]), "src-boundary: latin key at string edges must fire") check(any(p["src"] == "rose" for p in select("我叫rose。", 1, {})["injected"]), "src-boundary: latin key abutting a cross-script (Han) neighbour must fire") GLOSSARY = saved # kana carve-out (mirror of Go TestKanaNotSourceBoundaryChecked): a long kana key still fires # between kana neighbours (name+particle と…が), NOT source-boundary-suppressed. GLOSSARY = saved + [ {"src": "ながいなまえ", "dst": "Длинное имя", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "name", "lemma_keys": ["длинное"], "decl_forms": ["Длинное имя"]}] check(any(p["src"] == "ながいなまえ" and p["_disp"] == CONFIRMED for p in select("私とながいなまえが会った。", 1, {})["injected"]), "src-boundary: kana key must NOT be boundary-suppressed (name+particle recall)") GLOSSARY = saved # apostrophe-variant fold (package-2 Task 1b, mirror of Go TestApostropheFold): typographic ≡ # ASCII on the source side, and an ASCII-apostrophe key matches a typographic-apostrophe chunk. check(norm_src("O’Brien") == norm_src("O'Brien"), "apostrophe: source fold not applied") GLOSSARY = saved + [ {"src": "o'brien", "dst": "О'Брайен", "status": "approved", "sense": "", "since_ch": 0, "until_ch": 0, "allow_short": False, "aliases": [], "type": "name", "lemma_keys": ["о'брайен"], "decl_forms": ["О'Брайен"]}] check(any(p["src"] == "o'brien" for p in select("mr o’brien arrived.", 1, {})["injected"]), "apostrophe: typographic apostrophe in chunk must match ASCII-apostrophe key") GLOSSARY = saved # token budget eviction: tiny budget keeps only the highest-priority record sb = select("阿Q和未庄和王胡和秀才。", 6, {}, budget_tokens=3) check(len(sb["evicted"]) > 0 and len(sb["injected"]) >= 1, "budget: no eviction under tiny budget") # echo detector (package-2 Task 1c): leading preamble, embedded/trailing glossary, and clean. b1, r1 = detect_echo("ГЛОССАРИЙ (используй эти утверждённые):\n阿Q → А-кью\n\nЖил-был А-кью.") check(b1 == "Жил-был А-кью." and any("preamble" in r for r in r1), "echo: leading preamble not stripped/flagged") b2, r2 = detect_echo("Жил-был А-кью.\n\nГЛОССАРИЙ:\n阿Q → А-кью") check(b2 == "Жил-был А-кью." and any("embedded_glossary_echo" in r for r in r2), "echo: trailing glossary not stripped/flagged") b3, r3 = detect_echo("Начало главы.\n阿Q → А-кью\nКонец главы.") check(b3 == "Начало главы.\nКонец главы." and any("embedded_glossary_echo" in r for r in r3), "echo: embedded arrow-gloss line not stripped/flagged") b4, r4 = detect_echo("Жил-был А-кью в деревне Вэйчжуан.") check(b4 == "Жил-был А-кью в деревне Вэйчжуан." and not r4, "echo: clean translation wrongly flagged") # echo detector must NOT strip legit prose (parity-review #5): a buried marker word, or an # arrow+CJK line that is actually a sentence, are kept. b5, r5 = detect_echo("В конце книги он нашёл глоссарий терминов.") check(b5 == "В конце книги он нашёл глоссарий терминов." and not r5, "echo: FP — buried «глоссарий» wrongly stripped") b6, r6 = detect_echo("Путь (道) → его судьба была предрешена.") check(b6 == "Путь (道) → его судьба была предрешена." and "glossary" not in " ".join(r6), "echo: FP — prose with arrow+CJK wrongly stripped as gloss") # İ→lower Go-parity (parity-review #1): norm_src must fold İ to 'i' (Python .lower() → 'i̇') check(norm_src("İ") == "i" and dm.normalize_target("İ") == "i", "İ: not folded to Go's lowercase 'i'") # trad2simp byte-parity with the Go embed (package-2 Task 1d): the vendored table must stay # byte-identical to backend/internal/pipeline/data/trad2simp.txt (memoryNormVersion hashes its # bytes) — a drift silently diverges the normalizers. Skipped if the Go source isn't checked out. go_embed = ROOT.parent.parent / "backend" / "internal" / "pipeline" / "data" / "trad2simp.txt" if go_embed.exists(): import hashlib h_eval = hashlib.md5(TRAD_TABLE.read_bytes()).hexdigest() h_go = hashlib.md5(go_embed.read_bytes()).hexdigest() check(h_eval == h_go, f"trad2simp: vendored table diverged from Go embed ({h_eval} != {h_go})") print("SELF-TEST:", "OK" if not fails else "FAIL") for f in fails: print(" ✗", f) return 1 if fails else 0 def main(): ap = argparse.ArgumentParser() ap.add_argument("--src", default=str(SAMPLES / "zh" / "luxun-ah-q-ch5-9.txt")) ap.add_argument("--chapter", type=int, default=6, help="chapter number for the spoiler gate " "(single-chapter indicative: no per-chunk chapter-reset; the deciding run " "feeds real multi-chapter books, where the sticky window resets per chapter)") ap.add_argument("--modes", default="M0,M1,M2,M3") ap.add_argument("--model", default="grok-4.20-0309-non-reasoning", help="translator/draft model") ap.add_argument("--judge", default="gemini-2.5-flash", help="cross-family fidelity judge (D13.3)") ap.add_argument("--no-fidelity", action="store_true", help="skip the LLM fidelity axis") ap.add_argument("--budget", type=int, default=0, help="glossary token budget (0 = unbounded)") ap.add_argument("--regen-on-echo", action="store_true", help="re-ask once with a hardened prompt on echo") ap.add_argument("--max-chunks", type=int, default=5) ap.add_argument("--dry-run", action="store_true") ap.add_argument("--self-test", action="store_true") ap.add_argument("--out", default="") args = ap.parse_args() if args.self_test: sys.exit(self_test()) stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ") out_path = Path(args.out) if args.out else OUTDIR / f"memory_eval_M_{stamp}.json" raw_dir = OUTDIR / f"raw_{stamp}" text = Path(args.src).read_text(encoding="utf-8") chunks = chunk_text(text)[:args.max_chunks] modes = args.modes.split(",") # deterministic wrong-dst shuffle for M3 (rotate dst among present names/places) names = [e for e in GLOSSARY if e["type"] in ("name", "place")] rot = {names[i]["src"]: names[(i + 1) % len(names)]["dst"] for i in range(len(names))} shuffled = {e["src"]: rot.get(e["src"], "НЕВЕРНО") for e in GLOSSARY} print(f"src={Path(args.src).name} chapter={args.chapter} chunks={len(chunks)} model={args.model} " f"judge={args.judge} modes={modes} budget={args.budget}", file=sys.stderr) if not args.dry_run and not args.no_fidelity: if family_of(args.judge) == family_of(args.model): sys.exit(f"D13.3: judge {args.judge} shares family with translator {args.model}") raw_dir.mkdir(parents=True, exist_ok=True) rows, sticky_win = [], [] # sticky_win: list of {id->disp}, capped at STICKY_DEPTH def union_sticky(): out = {} for d in sticky_win: out.update(d) return out for ci, ch in enumerate(chunks): sel = select(ch, args.chapter, union_sticky(), args.budget) present = {e["src"]: e for e in sel["injected"] if e["_via"] != "sticky"} row = {"chunk": ci, "src_chars": len(ch), "injected": [{"src": e["src"], "disp": e["_disp"], "via": e["_via"]} for e in sel["injected"]], "rejected_spoiler": [{"src": r["src"]} for r in sel["rejected"]], "evicted": [{"src": e["src"]} for e in sel["evicted"]], "present_for_consistency": list(present)} if not args.dry_run: row["by_mode"] = {} for mode in modes: prompt = build_prompt(ch, mode, sel, summary="", shuffled=shuffled) raw, usage = translate(prompt, args.model) body, echo = detect_echo(raw) regen = None if echo and args.regen_on_echo: raw2, usage2 = translate(prompt, args.model, hardened=True) body2, echo2 = detect_echo(raw2) regen = {"raw": raw2, "echo": echo2, "usage": usage2} if not echo2: # clean on retry → use it body, echo, usage, raw = body2, echo2, usage2, raw2 cons = dm.consistency(body, {s: {"dst": e["dst"], "lemma_keys": e.get("lemma_keys", []), "hyphen_pat": e.get("hyphen_pat"), "decl_forms": e.get("decl_forms", [])} for s, e in present.items()}) if present else None fid = None if not args.no_fidelity: fid = fidelity_judge(ch, body, args.judge, args.model, PROVIDERS) # persist FULL raw output for post-hoc audit (never a preview) (raw_dir / f"chunk{ci}_{mode}.txt").write_text( f"# model={args.model} chunk={ci} mode={mode} chapter={args.chapter} echo={echo}\n" f"# usage={usage}\n\n{raw}", encoding="utf-8") row["by_mode"][mode] = { "echo_contaminated": echo or None, "consistency_pymorphy": cons["pymorphy"]["score"] if cons else None, "consistency_stored_decl": cons["stored_decl"]["score"] if cons else None, "missed_pymorphy": cons["pymorphy"]["missed"] if cons else [], "missed_stored_decl": cons["stored_decl"]["missed"] if cons else [], "fidelity": fid, "in_tok": usage.get("prompt_tokens"), "out_tok": usage.get("completion_tokens"), "out_full": raw, "body_scored": body, "regen": regen} fscore = (fid or {}).get("fidelity") if isinstance(fid, dict) else None print(f" chunk{ci} {mode}: cons(pym)={row['by_mode'][mode]['consistency_pymorphy']} " f"cons(decl)={row['by_mode'][mode]['consistency_stored_decl']} " f"fidelity={fscore} echo={echo or '-'} " f"in={usage.get('prompt_tokens')} out={usage.get('completion_tokens')}", file=sys.stderr) sticky_win = (sticky_win + [sel["active"]])[-STICKY_DEPTH:] rows.append(row) meta = {"run_utc": stamp, "src": args.src, "chapter": args.chapter, "modes": modes, "translator_model": args.model, "judge_model": None if args.no_fidelity else args.judge, "budget_tokens": args.budget, "regen_on_echo": args.regen_on_echo, "trad2simp_entries": len(TRAD2SIMP), "sticky_depth": STICKY_DEPTH, "note": "M0-M3 = terminology modes (exp09 §6); mirror synced with memory.go (Go v3 " "memmatch-v3): D16.2 sticky-inherit-disp, D16.3 source phonetic word-boundary, " "apostrophe fold (package-2 Tasks 1a/1b) — all parity-covered by --self-test."} out_path.write_text(json.dumps({"meta": meta, "rows": rows}, ensure_ascii=False, indent=2)) print(f"\nwrote {out_path}", file=sys.stderr) if not args.dry_run: print(f"raw outputs: {raw_dir}", file=sys.stderr) _print_aggregate(rows, modes) if args.dry_run: for r in rows: print(f"chunk{r['chunk']} inj={[(e['src'], e['disp'], e['via']) for e in r['injected']]} " f"spoiler_rej={[e['src'] for e in r['rejected_spoiler']]} evicted={[e['src'] for e in r['evicted']]}") def _print_aggregate(rows, modes): """Headline aggregate with clean (echo-excluded) vs all points, per mode.""" print("\n=== AGGREGATE (mean over chunks; clean = echo-contaminated points excluded) ===", file=sys.stderr) for mode in modes: pts = [r["by_mode"][mode] for r in rows if "by_mode" in r] clean = [p for p in pts if not p["echo_contaminated"]] def mean(seq, key, src=pts): vals = [p[key] for p in src if isinstance(p.get(key), (int, float))] return round(sum(vals) / len(vals), 3) if vals else None def fmean(src): vals = [p["fidelity"]["fidelity"] for p in src if isinstance(p.get("fidelity"), dict) and isinstance(p["fidelity"].get("fidelity"), (int, float))] return round(sum(vals) / len(vals), 1) if vals else None n_echo = sum(1 for p in pts if p["echo_contaminated"]) print(f" {mode}: cons(decl) all={mean(pts,'consistency_stored_decl')} clean={mean(clean,'consistency_stored_decl',clean)} " f"| fidelity all={fmean(pts)} clean={fmean(clean)} | echo_pts={n_echo}/{len(pts)}", file=sys.stderr) if __name__ == "__main__": main()