Branch `insights/goal-d2` from the tip of `insights/goal-d1` (PR 100, unmerged). Candidate revision c4f76cf069abfc98f10a9b1fe8846555f2117990: main plus D1 plus this branch, generator claude-opus-5 through the in-app Claude CLI adapter, verifier gpt-6-astra through the in-app Codex CLI adapter, the D1 calendar guard, and the promoted definitions. Baseline: the same code with the glm generator and no verifier. Method: the goal-3 runtime, scorer and clock (2026-09-06 12:00 UTC, Europe/London), two workers, arms sequential. Rendered 2026-09-08T21:23:55.643761+00:00; every number is in `docs/evidence/insights-goal-d2-2026-09-08/results.json`.
| ruling | knowledge entry | doc | verifier default | synonyms | leg | definition record (physical predicate) |
|---|---|---|---|---|---|---|
| routing | `source-routing-ct-default` | `docs/bi/knowledge/source-routing-default.md` | yes | 9 | both | knowledge entry only |
| 1 | `staff-house-exclusion-ct` | `docs/bi/knowledge/staff-house-exclusion.md` | yes | 12 | ct | `ct_staff_house_excluded`: excludes:public.accounts.email contains @currencytransfer.com, excludes:public.accounts.id in [9871, 1090], excludes:public.accounts.skip_commission_calculations equals True; `pd_non_employee_person`: excludes:pipedrive.persons_enriched.pd_is_ct_employee equals True |
| 1 | `staff-house-exclusion-pd` | `docs/bi/knowledge/staff-house-exclusion.md` | yes | 9 | gold | `ct_staff_house_excluded`: excludes:public.accounts.email contains @currencytransfer.com, excludes:public.accounts.id in [9871, 1090], excludes:public.accounts.skip_commission_calculations equals True; `pd_non_employee_person`: excludes:pipedrive.persons_enriched.pd_is_ct_employee equals True |
| 2 | `person-identity-ct-user-id` | `docs/bi/knowledge/person-identity.md` | yes | 9 | gold | knowledge entry only |
| 3 | `quoted-not-booked-window` | `docs/bi/knowledge/quoted-not-booked.md` | no | 10 | ct | `client_requested_quote`: all_of:public.quotes.type equals by_client |
| 4 | `acquisition-source-comparison` | `docs/bi/knowledge/acquisition-source-comparison.md` | no | 9 | both | knowledge entry only |
| 5 | `signup-to-first-trade-cohort` | `docs/bi/knowledge/signup-to-first-trade.md` | no | 7 | ct | knowledge entry only |
| 6 | `sales-stream-teams` | `docs/bi/knowledge/sales-stream.md` | no | 14 | ct | knowledge entry only |
| 7 | `refer-and-earn-new-referrer` | `docs/bi/knowledge/refer-and-earn-referrers-and-payouts.md` | no | 5 | ct | knowledge entry only |
| 8 | `refer-and-earn-payouts-two-columns` | `docs/bi/knowledge/refer-and-earn-referrers-and-payouts.md` | no | 7 | ct | `refer_and_earn_paid_payout`: all_of:public.user_commission_statements.paid_at is_not_null None |
| 9 | `revenue-by-source-window` | `docs/bi/knowledge/revenue-by-source-window.md` | no | 7 | ct | knowledge entry only |
| 10 | `sales-stream-teams` | `docs/bi/knowledge/sales-stream.md` | no | 14 | ct | knowledge entry only |
| 11 | `monthly-pace-full-prior-months` | `docs/bi/knowledge/monthly-pace.md` | no | 7 | ct | knowledge entry only |
| 12 | `period-series-zero-fill` | `docs/bi/knowledge/period-series-zero-fill.md` | no | 12 | both | knowledge entry only |
| 13 | `margin-basis-points` | `docs/bi/knowledge/ct-revenue-measures.md` | no | 6 | ct | knowledge entry only |
| 14 | `period-comparison-to-date` | `docs/bi/knowledge/period-over-period.md` | no | 6 | both | knowledge entry only |
| 15 | `book-resolution` | `docs/bi/knowledge/book-resolution.md` | no | 9 | both | knowledge entry only |
| 16 | `high-value-metal-band` | `docs/bi/knowledge/high-value-clients.md` | no | 15 | ct | `high_value_client`: all_of:public.accounts.group in ['gold', 'platinum'] |
| 17 | `declared-vs-traded-differs` | `docs/bi/knowledge/declared-currency-intent.md` | no | 7 | ct | knowledge entry only |
| 18 | `calendar-window-defaults` | `docs/bi/knowledge/calendar-window-defaults.md` | yes | 15 | both | knowledge entry only |
| 19 | `two-class-breakdown-wide` | `docs/bi/knowledge/breakdown-shape.md` | no | 11 | both | knowledge entry only |
Partial or knowledge-only families, with the reason:
Every definition record is created certified and global by `scripts/seed_ratified_definitions.py`, compiles through the application's own definition compiler to exactly the predicate above, and is found by the definition resolver by its name and by a synonym over the deterministic lexical embedder and a local synonym index (`tests/test_insights_rulings_promotion.py`). The house-account ids in the predicates are the constants the repository already documents (`spec/pd-querying-design.md`, `bi/fabricated_cohort_guard.py`). The knowledge entries land in the generation prompt only in leftover budget (priority 50 and above); the founder-pinned prompt selections in `tests/test_generation_curriculum.py` still pass unchanged, which is the evidence that they stayed intact. The verifier receives an entry whenever the question names it or it is a house default.
Holdout v11-r2 (250 attempts): every case corrected to rulings 1 and 2 by wrapping its sealed cohort query (gold: employees dropped through `persons_enriched.pd_is_ct_employee`, identifier `COALESCE(persons_enriched.ct_id, person id)`; CT: the house terms on `public.accounts`). References changed for 114 cases: {'ct/count/same': 42, 'ct/count/changed': 48, 'gold/count/same': 60, 'ct/listing/same': 28, 'ct/listing/changed': 32, 'gold/listing/changed': 34, 'gold/listing/same': 6}. The synthetic fixture flags no Pipedrive person as an employee (so the Pipedrive half of ruling 1 changes nothing there), most persons carry a CT user id, and a handful of synthetic staff-mailbox and statistics-flagged accounts carry the CT changes; 15 CT references change only because of the statistics flag, which stands in for the missing test-account marker and may exclude legitimate accounts or miss unflagged test accounts.
| retained run | label changes | correct completed | silent wrong | precision | correct first turns | first-turn errors | full conversations |
|---|---|---|---|---|---|---|---|
| E5 original e479e81 (goal 3) | 25 | 36 → 23 | 31 → 48 | 53.7% → 32.4% | 14 → 9 | 60 → 60 | 1 → 1 |
| V18 final candidate | 38 | 57 → 46 | 54 → 74 | 51.4% → 38.3% | 18 → 19 | 35 → 35 | 4 → 3 |
| Goal A run 2 | 21 | 37 → 32 | 32 → 39 | 53.6% → 45.1% | 16 → 12 | 54 → 54 | 4 → 2 |
| Goal A run 3 | 19 | 33 → 29 | 38 → 43 | 46.5% → 40.3% | 14 → 10 | 58 → 58 | 1 → 1 |
| Goal B control glm-5.2 r1 | 24 | 39 → 33 | 32 → 42 | 54.9% → 44.0% | 16 → 13 | 56 → 56 | 2 → 2 |
| Goal B control glm-5.2 r2 | 23 | 34 → 29 | 40 → 45 | 45.9% → 39.2% | 15 → 10 | 57 → 57 | 2 → 2 |
| Goal B control glm-5.2 r3 | 22 | 33 → 28 | 31 → 41 | 51.6% → 40.6% | 14 → 10 | 54 → 54 | 3 → 2 |
| Goal B glm-5.3 r1 | 21 | 34 → 26 | 29 → 42 | 54.0% → 38.2% | 13 → 10 | 62 → 62 | 2 → 1 |
| Goal B glm-5.3 r2 | 24 | 28 → 21 | 29 → 39 | 49.1% → 35.0% | 14 → 11 | 61 → 61 | 2 → 2 |
| Goal B glm-5.3 r3 | 21 | 42 → 29 | 34 → 49 | 55.3% → 37.2% | 16 → 11 | 56 → 56 | 4 → 2 |
| Goal B2 claude-fable-5-1-cli r1 | 32 | 49 → 35 | 47 → 61 | 51.0% → 36.5% | 19 → 12 | 49 → 49 | 6 → 2 |
| Goal B2 claude-fable-5-1-cli r2 | 31 | 52 → 35 | 45 → 64 | 53.6% → 35.4% | 20 → 14 | 46 → 46 | 5 → 2 |
| Goal B2 claude-fable-5-1-cli r3 | 33 | 52 → 33 | 43 → 64 | 54.7% → 34.0% | 20 → 12 | 43 → 43 | 5 → 2 |
| Goal B2 claude-opus-5-cli r1 | 37 | 61 → 38 | 48 → 71 | 56.0% → 34.9% | 27 → 16 | 38 → 38 | 5 → 2 |
| Goal B2 claude-opus-5-cli r2 | 35 | 60 → 39 | 51 → 72 | 54.1% → 35.1% | 25 → 15 | 37 → 37 | 6 → 2 |
| Goal B2 claude-opus-5-cli r3 | 37 | 63 → 40 | 50 → 73 | 55.8% → 35.4% | 26 → 16 | 41 → 41 | 5 → 2 |
| Goal B2 claude-sonnet-5-cli r1 | 21 | 34 → 23 | 29 → 42 | 54.0% → 35.4% | 15 → 10 | 62 → 62 | 3 → 2 |
| Goal B2 claude-sonnet-5-cli r2 | 18 | 33 → 23 | 35 → 47 | 48.5% → 32.9% | 15 → 11 | 61 → 61 | 4 → 2 |
| Goal B2 claude-sonnet-5-cli r3 | 18 | 31 → 25 | 34 → 42 | 47.7% → 37.3% | 17 → 14 | 58 → 58 | 2 → 2 |
| Goal B2 gpt-5.6-sol-codex r1 | 17 | 41 → 30 | 51 → 66 | 44.6% → 31.2% | 16 → 12 | 43 → 43 | 3 → 2 |
| Goal B2 gpt-5.6-sol-codex r2 | 20 | 40 → 29 | 40 → 56 | 50.0% → 34.1% | 16 → 13 | 46 → 46 | 3 → 2 |
| Goal B2 gpt-5.6-sol-codex r3 | 20 | 40 → 32 | 47 → 61 | 46.0% → 34.4% | 15 → 13 | 49 → 49 | 2 → 3 |
| Goal B2 gpt-6-astra-codex r1 | 20 | 26 → 20 | 10 → 20 | 72.2% → 50.0% | 13 → 10 | 73 → 73 | 2 → 2 |
| Goal B2 gpt-6-astra-codex r2 | 21 | 26 → 19 | 9 → 20 | 74.3% → 48.7% | 11 → 8 | 75 → 75 | 2 → 2 |
| Goal B2 gpt-6-astra-codex r3 | 20 | 28 → 22 | 7 → 17 | 80.0% → 56.4% | 14 → 11 | 71 → 71 | 2 → 2 |
| D1 step 0 main glm | 18 | 28 → 26 | 35 → 39 | 44.4% → 40.0% | 14 → 12 | 59 → 59 | 2 → 2 |
| D1 step 1 guard glm r1 (429s) | 23 | 34 → 26 | 23 → 32 | 59.6% → 44.8% | 15 → 11 | 64 → 64 | 2 → 1 |
| D1 step 1 guard glm r2 | 28 | 40 → 30 | 30 → 44 | 57.1% → 40.5% | 17 → 12 | 57 → 57 | 3 → 2 |
| D1 step 2 opus-5 r1 | 37 | 59 → 40 | 57 → 76 | 50.9% → 34.5% | 25 → 16 | 39 → 39 | 5 → 3 |
| D1 step 2 opus-5 r2 | 34 | 55 → 36 | 55 → 75 | 50.0% → 32.4% | 25 → 15 | 40 → 40 | 5 → 2 |
| D1 step 2 opus-5 r3 | 33 | 59 → 40 | 55 → 74 | 51.8% → 35.1% | 26 → 17 | 40 → 40 | 4 → 2 |
Transition classes over all retained runs (old → new):
Fresh sealed slice (revision fresh-v1 → fresh-v1-r2, sealed before opening): 19 amendments, 15 references changed, 2 contracts changed.
| case | family | rulings | reference | contract |
|---|---|---|---|---|
| fresh:F04:realized-monthly-2025 | F04 | [12] | changed | same |
| fresh:F04:booked-monthly-last-6 | F04 | [12] | changed | same |
| fresh:F04:billable-monthly-2026 | F04 | [12] | changed | same |
| fresh:F04:booked-monthly-corporate-2025 | F04 | [12] | changed | same |
| fresh:F16:per-month-h1-2026 | F16 | [12] | changed | same |
| fresh:F11:quotes-per-month-2025 | F11 | [12] | changed | same |
| fresh:F14:rate-by-month-6m | F14 | [12] | changed | same |
| fresh:F26:referrals-by-month-12m | F26 | [12] | changed | same |
| fresh:F27:notified-by-month-2025 | F27 | [12] | changed | same |
| fresh:F33:per-month-2025 | F33 | [12] | changed | same |
| fresh:F38:leads-per-month-2025 | F38 | [12] | changed | same |
| fresh:P03:won-by-month-2025 | P03 | [12] | changed | same |
| fresh:F24:signups-q3-vs-q2 | F24 | [14] | same | same |
| fresh:F12:tier-counts | F12 | [19] | changed | changed |
| fresh:P05:done-vs-open-2026 | P05 | [19] | changed | changed |
| fresh:F16:per-quarter-2026 | F16 | [12] | changed | same |
| fresh:X02:won-2025-never-traded | X02 | [1] | same | same |
| fresh:X02:open-already-traded-count | X02 | [1] | same | same |
| fresh:X02:won-traded-30d | X02 | [1] | same | same |
Real-usage sample (real-usage-v1 → real-usage-v1-r2): 27 amendments, 22 references changed, 12 contracts and 4 answer modes changed.
Retained real-usage runs rescored under the corrected oracle (Goal C and D1; the old outcome is reproduced first, then the corrected references, contracts and modes are applied):
| retained run | old outcomes reproduced | label changes | correct completed | silent wrong | precision | correct first turns | first-turn errors |
|---|---|---|---|---|---|---|---|
| Goal C original r1 | 101/101 | 8 | 4 → 4 | 27 → 33 | 12.9% → 10.8% | 4 → 4 | 16 → 16 |
| Goal C original r2 | 101/101 | 9 | 5 → 5 | 29 → 34 | 14.7% → 12.8% | 5 → 5 | 14 → 14 |
| Goal C landed r1 | 101/101 | 9 | 3 → 3 | 24 → 29 | 11.1% → 9.4% | 2 → 2 | 9 → 9 |
| Goal C landed r2 | 101/101 | 7 | 5 → 5 | 29 → 34 | 14.7% → 12.8% | 4 → 4 | 10 → 10 |
| Goal C gpt-6-astra-codex r1 | 101/101 | 6 | 2 → 2 | 22 → 24 | 8.3% → 7.7% | 1 → 1 | 23 → 23 |
| Goal C claude-opus-5-cli r1 | 101/101 | 10 | 0 → 0 | 31 → 35 | 0.0% → 0.0% | 0 → 0 | 3 → 3 |
| D1 step 0 main glm | 101/101 | 9 | 2 → 2 | 17 → 22 | 10.5% → 8.3% | 2 → 2 | 11 → 11 |
| D1 step 1 guard glm r1 (409 cascade) | 101/101 | 7 | 4 → 4 | 25 → 30 | 13.8% → 11.8% | 4 → 4 | 17 → 17 |
| D1 step 1 guard glm r2 (429s) | 101/101 | 7 | 3 → 3 | 29 → 32 | 9.4% → 8.6% | 2 → 2 | 21 → 21 |
| D1 step 1 guard glm r3 | 101/101 | 9 | 3 → 3 | 32 → 37 | 8.6% → 7.5% | 2 → 2 | 12 → 12 |
| D1 step 2 opus-5 r1 (409 cascade) | 101/101 | 12 | 0 → 0 | 33 → 37 | 0.0% → 0.0% | 0 → 0 | 9 → 9 |
| D1 step 2 opus-5 r2 | 101/101 | 9 | 0 → 0 | 35 → 40 | 0.0% → 0.0% | 0 → 0 | 8 → 8 |
| case | family | rulings | reference | mode | contract |
|---|---|---|---|---|---|
| real:F04:21d2be6834ad | F04 | [12] | changed | rows | same |
| real:F11:925779a425e5 | F11 | [12] | changed | rows | same |
| real:F14:93e30e7000cb | F14 | [12] | changed | rows | same |
| real:F27:dbbb7ca706b8 | F27 | [12] | changed | rows | same |
| real:F14:1606045b2a17 | F14 | [12] | changed | rows | same |
| real:F10:878f1692df8e | F10 | [3] | same | rows | same |
| real:F10:8690cb315864 | F10 | [3] | same | rows | same |
| real:F10:60e720a241cb | F10 | [3] | same | scalar | same |
| real:F04:76163d80e715 | F04 | [19, 12] | changed | rows | changed |
| real:F23:05a0efd0ea26 | F23 | [12] | changed | rows | same |
| real:F28:e470f9a6e6ec | F28 | [12] | changed | rows | same |
| real:F12:5f49ca1be797 | F12 | [19] | changed | rows | changed |
| real:P05:500554bebc0b | P05 | [19] | changed | rows | changed |
| real:F01:f37f3677a187 | F01 | [19] | changed | rows | changed |
| real:F34:c8afa5f52bfc | F34 | [19] | changed | rows | changed |
| real:F32:4430cd7533fc | F32 | [19] | changed | rows | changed |
| real:F24:42b91cc9ecf2 | F24 | [14] | same | rows | same |
| real:F26:54b37dd80eb0 | F26 | [7] | same | rows | same |
| real:F26:b89c5524e753 | F26 | [8] | changed | rows | changed |
| real:F30:90434d56a0bb | F30 | [17] | changed | rows | changed |
| real:X05:c59daf5445c0 | X05 | [16] | changed | ordered | same |
| real:X01:337697c83d2f | X01 | [15] | changed | ordered → clarification | changed |
| real:F05:6ea2da828831 | F05 | [10, 6] | changed | rows | same |
| real:F09:10472e45f2a1 | F09 | [6] | changed | clarification → rows | changed |
| real:F09:a0f8c779a0a1 | F09 | [6, 5] | changed | clarification → rows | changed |
| real:F15:370d3a12ee50 | F15 | [6] | changed | clarification → rows | changed |
| real:F18:2e221d8aded2 | F18 | [6] | changed | ordered | same |
Per repeat (denominators: 116 supported first turns, 122 attempts, 3 conversations; precision = correct completed over correct completed plus silent wrong; verifier verdict counts are over initial and repeat checks, repairs applied are counted separately):
| arm / repeat | First-turn execution errors | Correct first turns | Precision (correct / answered) | Silent wrong | Correct completed | Refusals | Comparison unverified | Full conversations | refusals by guard | refusals by verifier | verifier verdicts accept / reject / clarify; repairs applied | p50 / p95 per question (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| candidate r1 | 22/116 | 0/116 | 0.0% (0/11) | 11/122 | 0/122 | 74/122 | 14/122 | 0/3 | 2 | 72 | 25 / 122 / 0 / 51 | 30.2 / 49.7 |
| candidate r2 | 24/116 | 0/116 | 0.0% (0/9) | 9/122 | 0/122 | 78/122 | 11/122 | 0/3 | 4 | 74 | 20 / 122 / 0 / 52 | 29.8 / 51.1 |
| candidate r3 | 25/116 | 0/116 | 0.0% (0/14) | 14/122 | 0/122 | 68/122 | 15/122 | 0/3 | 2 | 66 | 29 / 119 / 0 / 54 | 30.8 / 52.0 |
| baseline r1 | 30/116 | 0/116 | 0.0% (0/47) | 47/122 | 0/122 | 10/122 | 34/122 | 0/3 | 10 | 0 | 0 / 0 / 0 / 0 | 11.4 / 22.8 |
| baseline r2 | 32/116 | 3/116 | 7.1% (3/42) | 39/122 | 3/122 | 10/122 | 37/122 | 0/3 | 10 | 0 | 0 / 0 / 0 / 0 | 9.8 / 20.2 |
| baseline r3 | 32/116 | 1/116 | 2.2% (1/45) | 44/122 | 1/122 | 7/122 | 36/122 | 0/3 | 7 | 0 | 0 / 0 / 0 / 0 | 8.9 / 20.3 |
Arm means with min and max over repeats, the baseline range (max minus min over its three repeats, the observed spread used as the noise floor; it is not a bound on the candidate's variability) and the paired intent-family cluster bootstrap (candidate minus baseline on the metric's own scale, 10,000 draws; a difference of two all-zero arms gives a degenerate [0, 0] interval):
| metric | candidate mean [min to max] | baseline mean [min to max] | baseline range | paired difference [95% CI] |
|---|---|---|---|---|
| First-turn execution errors | 23.7 [22.0 to 25.0] | 31.3 [30.0 to 32.0] | 2.0 | -7.7 [-18.4 to 1.4] |
| Correct first turns | 0.0 [0.0 to 0.0] | 1.3 [0.0 to 3.0] | 3.0 | -1.3 [-3.0 to 0.0] |
| Precision (correct / answered) | 0.0% [0.0% to 0.0%] | 3.1% [0.0% to 7.1%] | 7.1% | -3.0% [-7.0% to 0.0%] |
| Silent wrong | 11.3 [9.0 to 14.0] | 43.3 [39.0 to 47.0] | 8.0 | -32.0 [-40.7 to -23.2] |
| Correct completed | 0.0 [0.0 to 0.0] | 1.3 [0.0 to 3.0] | 3.0 | not a paired metric |
| Refusals | 73.3 [68.0 to 78.0] | 9.0 [7.0 to 10.0] | 3.0 | not a paired metric |
| Comparison unverified | 13.3 [11.0 to 15.0] | 35.7 [34.0 to 37.0] | 3.0 | not a paired metric |
| Full conversations | 0.0 [0.0 to 0.0] | 0.0 [0.0 to 0.0] | 0.0 | 0.0 [0.0 to 0.0] |
Acceptance. The rules were committed in `scripts/insights_d2_report.py` at ead329bd (18:05 UTC on 2026-09-08), after the smoke runs' outcomes had been seen and while the later-aborted first attempt was running, and before any measurement repeat completed (the first completed at 18:49 UTC). Values are shown unrounded where they are thresholds or rates:
| target | rule | result | pass |
|---|---|---|---|
| precision_at_least_80 | candidate mean precision >= 0.80; gain over baseline > noise spread; paired 95% lower bound > 0 | {"candidate": "0", "baseline": "0.03122", "threshold": "0.8", "point": false, "gain": "-0.03122", "beyond_noise": false, "paired_bound": false} | FAIL |
| correct_completed_not_below_baseline | candidate mean correct completed >= baseline mean; paired 95% interval not entirely below 0 | {"candidate": 0, "baseline": "1.333", "point": false, "gain": "-1.333", "paired_metric": "first_turn_correct_rate (the paired scorer's correct rate)", "paired_not_entirely_below_zero": true} | FAIL |
| silent_wrong_down_by_half | candidate mean silent wrong <= 0.5 x baseline mean; reduction > noise spread; paired 95% upper bound < 0 | {"candidate": "11.33", "baseline": "43.33", "threshold": "21.67", "point": true, "reduction": "32", "beyond_noise": true, "paired_bound": true} | PASS |
| first_turn_errors_at_most_25_per_100 | candidate mean first-turn error rate <= 0.25; reduction over baseline > noise spread; paired 95% upper bound < 0 | {"candidate_rate": "0.204", "baseline_rate": "0.2701", "candidate": "23.67", "baseline": "31.33", "threshold_rate": "0.25", "point": true, "reduction": "7.667", "beyond_noise": true, "paired_bound": false} | FAIL |
| full_conversations_at_least_20_per_50 | candidate mean full-conversation rate >= 0.40; gain over baseline > noise spread; paired 95% lower bound > 0 (3 conversations only) | {"candidate_rate": "0", "baseline_rate": "0", "candidate": 0, "baseline": 0, "conversations": 3, "threshold_rate": "0.4", "point": false, "gain": 0, "beyond_noise": false, "paired_bound": false} | FAIL |
| zero_curated_invariant_violations | the curated critical test files of the goal-3 V18 report (those present on this branch) pass on the candidate revision with zero failures; a node that failed while the suite ran beside the live measurement counts only if it also fails when rerun in isolation | {"failed_as_run": 2, "failed_after_isolated_rerun": 0, "js_failed": 0} | PASS |
| p95_latency_at_most_30s | candidate p95 per-question latency (submission to terminal result) <= 30 s on every repeat | {"threshold_ms": 30000} | FAIL |
Stable-case view (baseline repeats as control): stable-correct in baseline 0; lost 0 (correct to wrong 0, correct to error 0); gained 0. With no stable-correct baseline case, 'lost 0' is no evidence of preservation.
What the verifier rejected (signal tags on non-operational reject verdicts, the three fresh repeats and the real-usage run pooled; a verdict may carry several tags; the tags are the verifier's own diagnoses, not independently confirmed defects, and an applied repair is not a confirmed correction):
| signal tag | reject verdicts |
|---|---|
| missing-staff-house-exclusion | 149 |
| wrong-population | 99 |
| missing-staff-house-exclusions | 67 |
| wrong-window | 45 |
| missing-staff-exclusion | 29 |
| missing-window-upper-bound | 25 |
| wrong-grain | 19 |
| missing-zero-fill | 15 |
| missing-deleted-filter | 15 |
| missing-house-exclusion | 13 |
| missing-deleted-trade-filter | 12 |
| wrong-prior-turn-scope | 10 |
| staff-house-exclusion | 10 |
| missing-house-exclusions | 8 |
| forced-empty-leg | 6 |
Repairs skipped, by reason (the three fresh repeats and the real-usage run): invocation budget: 81, regeneration returned no new SQL: 17, regeneration failed: CalendarSemanticsError: 12, regeneration failed: FanOutSumError: 2, repair failed the leg gate: CtLegValidationError: 2.
Verifier call accounting: completed verifier calls exceed persisted verdicts by 4, 4, 2 and 2 in the four candidate runs (12 of 576 completed calls); those calls belong to turns that ended in an execution error or a job timeout after the call, so no turn record carries their verdict; the calls are in the verifier trace files. Failed calls (deadline or provider) are the operational verdicts.
Per-call statistics. Runs were sequential; the curated Python suite ran niced with one worker beside the first minutes of candidate fresh repeat 1 (18:16 to 18:19 UTC) and nothing else shared the host. Wall time includes the CLI child's start-up. Model ids: Claude runs report the billed model from Claude Code's modelUsage (a served identity); Codex runs report the CLI banner echo (not a server-side identity); the zai baseline records the requested id only, because the sealed runtime does not record response.model, and Goal B observed that endpoint serving glm-5.3 for glm-5.2 requests. Tokens: Claude reports input and output per call; Codex reports one total only (input and output unknown); a blank cell means not reported, never zero.
| run | generator calls | failed | p50 ms | p95 ms | model id (source above) | generator tokens in / out / total | verifier calls | failed | p50 ms | p95 ms | model id | verifier tokens total | sandbox read-only every call | tool sections |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| cand-fresh-r1 | 270 | 0 | 6231.8 | 10194.3 | {'claude-opus-5': 270} | 4723102 / 138756 / | 152 | 1 | 11047.2 | 20790.0 | {'gpt-6-astra': 151} | 1069102 | True | 0 |
| cand-fresh-r2 | 270 | 0 | 6145.0 | 10274.7 | {'claude-opus-5': 270} | 4745365 / 138855 / | 149 | 3 | 10854.0 | 21965.6 | {'gpt-6-astra': 146} | 901095 | True | 0 |
| cand-fresh-r3 | 269 | 0 | 6285.8 | 10207.3 | {'claude-opus-5': 269} | 4695022 / 137879 / | 150 | 0 | 10803.2 | 21110.6 | {'gpt-6-astra': 150} | 887772 | True | 0 |
| cand-real-r1 | 211 | 2 | 8099.5 | 14074.7 | {'claude-opus-5': 209} | 4514029 / 149796 / | 134 | 5 | 12076.7 | 21553.8 | {'gpt-6-astra': 129} | 843134 | True | 0 |
| base-fresh-r1 | 229 | 1 | 5279.3 | 9058.3 | {'glm-5.2': 228} | 2046958 / 70160 / | 0 | 0 | n/a | n/a | 0 | |||
| base-fresh-r2 | 228 | 0 | 4890.6 | 8061.6 | {'glm-5.2': 228} | 1993194 / 72056 / | 0 | 0 | n/a | n/a | 0 | |||
| base-fresh-r3 | 229 | 0 | 4283.7 | 7786.2 | {'glm-5.2': 229} | 2047197 / 73306 / | 0 | 0 | n/a | n/a | 0 | |||
| base-real-r1 | 158 | 0 | 4666.0 | 9024.2 | {'glm-5.2': 158} | 1799701 / 59736 / | 0 | 0 | n/a | n/a | 0 |
One repeat per arm, so no repeat variability is estimated here; the candidate's precision rests on three scored answers, the baseline's on 38.
| arm | First-turn execution errors | Correct first turns | Precision (correct / answered) | Silent wrong | Correct completed | Refusals | Comparison unverified | Full conversations | refusals by guard | refusals by verifier | verifier verdicts accept / reject / clarify; repairs applied | p50 / p95 per question (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| candidate r1 | 8/86 | 0/86 | 0.0% (0/3) | 3/101 | 0/101 | 72/101 | 16/101 | 0/5 | 2 | 70 | 19 / 108 / 0 / 44 | 37.4 / 59.8 |
| baseline r1 | 10/86 | 3/86 | 7.9% (3/38) | 35/101 | 3/101 | 10/101 | 39/101 | 0/5 | 10 | 0 | 0 / 0 / 0 / 0 | 9.2 / 22.3 |
The Goal C and D1 real-usage runs were rescored under the corrected oracle in section 2 (all 12 reproduce their old outcomes on 101/101 attempts first); they ran without the D2 definitions or the verifier, so they orient, they do not pair with, the two runs above.