Insights Goal D2: rulings promoted, evaluation reconciled, generate-then-verify, combined candidate on the fresh sealed slice (2026-09-08)

Branch `insights/goal-d2` from the tip of `insights/goal-d1` (PR 100, unmerged). Candidate revision c4f76cf069abfc98f10a9b1fe8846555f2117990: main plus D1 plus this branch, generator claude-opus-5 through the in-app Claude CLI adapter, verifier gpt-6-astra through the in-app Codex CLI adapter, the D1 calendar guard, and the promoted definitions. Baseline: the same code with the glm generator and no verifier. Method: the goal-3 runtime, scorer and clock (2026-09-06 12:00 UTC, Europe/London), two workers, arms sequential. Rendered 2026-09-08T21:23:55.643761+00:00; every number is in `docs/evidence/insights-goal-d2-2026-09-08/results.json`.

Headline

1. Rulings promoted

rulingknowledge entrydocverifier defaultsynonymslegdefinition record (physical predicate)
routing`source-routing-ct-default``docs/bi/knowledge/source-routing-default.md`yes9bothknowledge entry only
1`staff-house-exclusion-ct``docs/bi/knowledge/staff-house-exclusion.md`yes12ct`ct_staff_house_excluded`: excludes:public.accounts.email contains @currencytransfer.com, excludes:public.accounts.id in [9871, 1090], excludes:public.accounts.skip_commission_calculations equals True; `pd_non_employee_person`: excludes:pipedrive.persons_enriched.pd_is_ct_employee equals True
1`staff-house-exclusion-pd``docs/bi/knowledge/staff-house-exclusion.md`yes9gold`ct_staff_house_excluded`: excludes:public.accounts.email contains @currencytransfer.com, excludes:public.accounts.id in [9871, 1090], excludes:public.accounts.skip_commission_calculations equals True; `pd_non_employee_person`: excludes:pipedrive.persons_enriched.pd_is_ct_employee equals True
2`person-identity-ct-user-id``docs/bi/knowledge/person-identity.md`yes9goldknowledge entry only
3`quoted-not-booked-window``docs/bi/knowledge/quoted-not-booked.md`no10ct`client_requested_quote`: all_of:public.quotes.type equals by_client
4`acquisition-source-comparison``docs/bi/knowledge/acquisition-source-comparison.md`no9bothknowledge entry only
5`signup-to-first-trade-cohort``docs/bi/knowledge/signup-to-first-trade.md`no7ctknowledge entry only
6`sales-stream-teams``docs/bi/knowledge/sales-stream.md`no14ctknowledge entry only
7`refer-and-earn-new-referrer``docs/bi/knowledge/refer-and-earn-referrers-and-payouts.md`no5ctknowledge entry only
8`refer-and-earn-payouts-two-columns``docs/bi/knowledge/refer-and-earn-referrers-and-payouts.md`no7ct`refer_and_earn_paid_payout`: all_of:public.user_commission_statements.paid_at is_not_null None
9`revenue-by-source-window``docs/bi/knowledge/revenue-by-source-window.md`no7ctknowledge entry only
10`sales-stream-teams``docs/bi/knowledge/sales-stream.md`no14ctknowledge entry only
11`monthly-pace-full-prior-months``docs/bi/knowledge/monthly-pace.md`no7ctknowledge entry only
12`period-series-zero-fill``docs/bi/knowledge/period-series-zero-fill.md`no12bothknowledge entry only
13`margin-basis-points``docs/bi/knowledge/ct-revenue-measures.md`no6ctknowledge entry only
14`period-comparison-to-date``docs/bi/knowledge/period-over-period.md`no6bothknowledge entry only
15`book-resolution``docs/bi/knowledge/book-resolution.md`no9bothknowledge entry only
16`high-value-metal-band``docs/bi/knowledge/high-value-clients.md`no15ct`high_value_client`: all_of:public.accounts.group in ['gold', 'platinum']
17`declared-vs-traded-differs``docs/bi/knowledge/declared-currency-intent.md`no7ctknowledge entry only
18`calendar-window-defaults``docs/bi/knowledge/calendar-window-defaults.md`yes15bothknowledge entry only
19`two-class-breakdown-wide``docs/bi/knowledge/breakdown-shape.md`no11bothknowledge entry only

Partial or knowledge-only families, with the reason:

Every definition record is created certified and global by `scripts/seed_ratified_definitions.py`, compiles through the application's own definition compiler to exactly the predicate above, and is found by the definition resolver by its name and by a synonym over the deterministic lexical embedder and a local synonym index (`tests/test_insights_rulings_promotion.py`). The house-account ids in the predicates are the constants the repository already documents (`spec/pd-querying-design.md`, `bi/fabricated_cohort_guard.py`). The knowledge entries land in the generation prompt only in leftover budget (priority 50 and above); the founder-pinned prompt selections in `tests/test_generation_curriculum.py` still pass unchanged, which is the evidence that they stayed intact. The verifier receives an entry whenever the question names it or it is a house default.

2. Reconciliation ledger

Holdout v11-r2 (250 attempts): every case corrected to rulings 1 and 2 by wrapping its sealed cohort query (gold: employees dropped through `persons_enriched.pd_is_ct_employee`, identifier `COALESCE(persons_enriched.ct_id, person id)`; CT: the house terms on `public.accounts`). References changed for 114 cases: {'ct/count/same': 42, 'ct/count/changed': 48, 'gold/count/same': 60, 'ct/listing/same': 28, 'ct/listing/changed': 32, 'gold/listing/changed': 34, 'gold/listing/same': 6}. The synthetic fixture flags no Pipedrive person as an employee (so the Pipedrive half of ruling 1 changes nothing there), most persons carry a CT user id, and a handful of synthetic staff-mailbox and statistics-flagged accounts carry the CT changes; 15 CT references change only because of the statistics flag, which stands in for the missing test-account marker and may exclude legitimate accounts or miss unflagged test accounts.

retained runlabel changescorrect completedsilent wrongprecisioncorrect first turnsfirst-turn errorsfull conversations
E5 original e479e81 (goal 3)2536 → 2331 → 4853.7% → 32.4%14 → 960 → 601 → 1
V18 final candidate3857 → 4654 → 7451.4% → 38.3%18 → 1935 → 354 → 3
Goal A run 22137 → 3232 → 3953.6% → 45.1%16 → 1254 → 544 → 2
Goal A run 31933 → 2938 → 4346.5% → 40.3%14 → 1058 → 581 → 1
Goal B control glm-5.2 r12439 → 3332 → 4254.9% → 44.0%16 → 1356 → 562 → 2
Goal B control glm-5.2 r22334 → 2940 → 4545.9% → 39.2%15 → 1057 → 572 → 2
Goal B control glm-5.2 r32233 → 2831 → 4151.6% → 40.6%14 → 1054 → 543 → 2
Goal B glm-5.3 r12134 → 2629 → 4254.0% → 38.2%13 → 1062 → 622 → 1
Goal B glm-5.3 r22428 → 2129 → 3949.1% → 35.0%14 → 1161 → 612 → 2
Goal B glm-5.3 r32142 → 2934 → 4955.3% → 37.2%16 → 1156 → 564 → 2
Goal B2 claude-fable-5-1-cli r13249 → 3547 → 6151.0% → 36.5%19 → 1249 → 496 → 2
Goal B2 claude-fable-5-1-cli r23152 → 3545 → 6453.6% → 35.4%20 → 1446 → 465 → 2
Goal B2 claude-fable-5-1-cli r33352 → 3343 → 6454.7% → 34.0%20 → 1243 → 435 → 2
Goal B2 claude-opus-5-cli r13761 → 3848 → 7156.0% → 34.9%27 → 1638 → 385 → 2
Goal B2 claude-opus-5-cli r23560 → 3951 → 7254.1% → 35.1%25 → 1537 → 376 → 2
Goal B2 claude-opus-5-cli r33763 → 4050 → 7355.8% → 35.4%26 → 1641 → 415 → 2
Goal B2 claude-sonnet-5-cli r12134 → 2329 → 4254.0% → 35.4%15 → 1062 → 623 → 2
Goal B2 claude-sonnet-5-cli r21833 → 2335 → 4748.5% → 32.9%15 → 1161 → 614 → 2
Goal B2 claude-sonnet-5-cli r31831 → 2534 → 4247.7% → 37.3%17 → 1458 → 582 → 2
Goal B2 gpt-5.6-sol-codex r11741 → 3051 → 6644.6% → 31.2%16 → 1243 → 433 → 2
Goal B2 gpt-5.6-sol-codex r22040 → 2940 → 5650.0% → 34.1%16 → 1346 → 463 → 2
Goal B2 gpt-5.6-sol-codex r32040 → 3247 → 6146.0% → 34.4%15 → 1349 → 492 → 3
Goal B2 gpt-6-astra-codex r12026 → 2010 → 2072.2% → 50.0%13 → 1073 → 732 → 2
Goal B2 gpt-6-astra-codex r22126 → 199 → 2074.3% → 48.7%11 → 875 → 752 → 2
Goal B2 gpt-6-astra-codex r32028 → 227 → 1780.0% → 56.4%14 → 1171 → 712 → 2
D1 step 0 main glm1828 → 2635 → 3944.4% → 40.0%14 → 1259 → 592 → 2
D1 step 1 guard glm r1 (429s)2334 → 2623 → 3259.6% → 44.8%15 → 1164 → 642 → 1
D1 step 1 guard glm r22840 → 3030 → 4457.1% → 40.5%17 → 1257 → 573 → 2
D1 step 2 opus-5 r13759 → 4057 → 7650.9% → 34.5%25 → 1639 → 395 → 3
D1 step 2 opus-5 r23455 → 3655 → 7550.0% → 32.4%25 → 1540 → 405 → 2
D1 step 2 opus-5 r33359 → 4055 → 7451.8% → 35.1%26 → 1740 → 404 → 2

Transition classes over all retained runs (old → new):

Fresh sealed slice (revision fresh-v1 → fresh-v1-r2, sealed before opening): 19 amendments, 15 references changed, 2 contracts changed.

casefamilyrulingsreferencecontract
fresh:F04:realized-monthly-2025F04[12]changedsame
fresh:F04:booked-monthly-last-6F04[12]changedsame
fresh:F04:billable-monthly-2026F04[12]changedsame
fresh:F04:booked-monthly-corporate-2025F04[12]changedsame
fresh:F16:per-month-h1-2026F16[12]changedsame
fresh:F11:quotes-per-month-2025F11[12]changedsame
fresh:F14:rate-by-month-6mF14[12]changedsame
fresh:F26:referrals-by-month-12mF26[12]changedsame
fresh:F27:notified-by-month-2025F27[12]changedsame
fresh:F33:per-month-2025F33[12]changedsame
fresh:F38:leads-per-month-2025F38[12]changedsame
fresh:P03:won-by-month-2025P03[12]changedsame
fresh:F24:signups-q3-vs-q2F24[14]samesame
fresh:F12:tier-countsF12[19]changedchanged
fresh:P05:done-vs-open-2026P05[19]changedchanged
fresh:F16:per-quarter-2026F16[12]changedsame
fresh:X02:won-2025-never-tradedX02[1]samesame
fresh:X02:open-already-traded-countX02[1]samesame
fresh:X02:won-traded-30dX02[1]samesame

Real-usage sample (real-usage-v1 → real-usage-v1-r2): 27 amendments, 22 references changed, 12 contracts and 4 answer modes changed.

Retained real-usage runs rescored under the corrected oracle (Goal C and D1; the old outcome is reproduced first, then the corrected references, contracts and modes are applied):

retained runold outcomes reproducedlabel changescorrect completedsilent wrongprecisioncorrect first turnsfirst-turn errors
Goal C original r1101/10184 → 427 → 3312.9% → 10.8%4 → 416 → 16
Goal C original r2101/10195 → 529 → 3414.7% → 12.8%5 → 514 → 14
Goal C landed r1101/10193 → 324 → 2911.1% → 9.4%2 → 29 → 9
Goal C landed r2101/10175 → 529 → 3414.7% → 12.8%4 → 410 → 10
Goal C gpt-6-astra-codex r1101/10162 → 222 → 248.3% → 7.7%1 → 123 → 23
Goal C claude-opus-5-cli r1101/101100 → 031 → 350.0% → 0.0%0 → 03 → 3
D1 step 0 main glm101/10192 → 217 → 2210.5% → 8.3%2 → 211 → 11
D1 step 1 guard glm r1 (409 cascade)101/10174 → 425 → 3013.8% → 11.8%4 → 417 → 17
D1 step 1 guard glm r2 (429s)101/10173 → 329 → 329.4% → 8.6%2 → 221 → 21
D1 step 1 guard glm r3101/10193 → 332 → 378.6% → 7.5%2 → 212 → 12
D1 step 2 opus-5 r1 (409 cascade)101/101120 → 033 → 370.0% → 0.0%0 → 09 → 9
D1 step 2 opus-5 r2101/10190 → 035 → 400.0% → 0.0%0 → 08 → 8
casefamilyrulingsreferencemodecontract
real:F04:21d2be6834adF04[12]changedrowssame
real:F11:925779a425e5F11[12]changedrowssame
real:F14:93e30e7000cbF14[12]changedrowssame
real:F27:dbbb7ca706b8F27[12]changedrowssame
real:F14:1606045b2a17F14[12]changedrowssame
real:F10:878f1692df8eF10[3]samerowssame
real:F10:8690cb315864F10[3]samerowssame
real:F10:60e720a241cbF10[3]samescalarsame
real:F04:76163d80e715F04[19, 12]changedrowschanged
real:F23:05a0efd0ea26F23[12]changedrowssame
real:F28:e470f9a6e6ecF28[12]changedrowssame
real:F12:5f49ca1be797F12[19]changedrowschanged
real:P05:500554bebc0bP05[19]changedrowschanged
real:F01:f37f3677a187F01[19]changedrowschanged
real:F34:c8afa5f52bfcF34[19]changedrowschanged
real:F32:4430cd7533fcF32[19]changedrowschanged
real:F24:42b91cc9ecf2F24[14]samerowssame
real:F26:54b37dd80eb0F26[7]samerowssame
real:F26:b89c5524e753F26[8]changedrowschanged
real:F30:90434d56a0bbF30[17]changedrowschanged
real:X05:c59daf5445c0X05[16]changedorderedsame
real:X01:337697c83d2fX01[15]changedordered → clarificationchanged
real:F05:6ea2da828831F05[10, 6]changedrowssame
real:F09:10472e45f2a1F09[6]changedclarification → rowschanged
real:F09:a0f8c779a0a1F09[6, 5]changedclarification → rowschanged
real:F15:370d3a12ee50F15[6]changedclarification → rowschanged
real:F18:2e221d8aded2F18[6]changedorderedsame

3. Verifier

4. Fresh sealed slice

Per repeat (denominators: 116 supported first turns, 122 attempts, 3 conversations; precision = correct completed over correct completed plus silent wrong; verifier verdict counts are over initial and repeat checks, repairs applied are counted separately):

arm / repeatFirst-turn execution errorsCorrect first turnsPrecision (correct / answered)Silent wrongCorrect completedRefusalsComparison unverifiedFull conversationsrefusals by guardrefusals by verifierverifier verdicts accept / reject / clarify; repairs appliedp50 / p95 per question (s)
candidate r122/1160/1160.0% (0/11)11/1220/12274/12214/1220/327225 / 122 / 0 / 5130.2 / 49.7
candidate r224/1160/1160.0% (0/9)9/1220/12278/12211/1220/347420 / 122 / 0 / 5229.8 / 51.1
candidate r325/1160/1160.0% (0/14)14/1220/12268/12215/1220/326629 / 119 / 0 / 5430.8 / 52.0
baseline r130/1160/1160.0% (0/47)47/1220/12210/12234/1220/31000 / 0 / 0 / 011.4 / 22.8
baseline r232/1163/1167.1% (3/42)39/1223/12210/12237/1220/31000 / 0 / 0 / 09.8 / 20.2
baseline r332/1161/1162.2% (1/45)44/1221/1227/12236/1220/3700 / 0 / 0 / 08.9 / 20.3

Arm means with min and max over repeats, the baseline range (max minus min over its three repeats, the observed spread used as the noise floor; it is not a bound on the candidate's variability) and the paired intent-family cluster bootstrap (candidate minus baseline on the metric's own scale, 10,000 draws; a difference of two all-zero arms gives a degenerate [0, 0] interval):

metriccandidate mean [min to max]baseline mean [min to max]baseline rangepaired difference [95% CI]
First-turn execution errors23.7 [22.0 to 25.0]31.3 [30.0 to 32.0]2.0-7.7 [-18.4 to 1.4]
Correct first turns0.0 [0.0 to 0.0]1.3 [0.0 to 3.0]3.0-1.3 [-3.0 to 0.0]
Precision (correct / answered)0.0% [0.0% to 0.0%]3.1% [0.0% to 7.1%]7.1%-3.0% [-7.0% to 0.0%]
Silent wrong11.3 [9.0 to 14.0]43.3 [39.0 to 47.0]8.0-32.0 [-40.7 to -23.2]
Correct completed0.0 [0.0 to 0.0]1.3 [0.0 to 3.0]3.0not a paired metric
Refusals73.3 [68.0 to 78.0]9.0 [7.0 to 10.0]3.0not a paired metric
Comparison unverified13.3 [11.0 to 15.0]35.7 [34.0 to 37.0]3.0not a paired metric
Full conversations0.0 [0.0 to 0.0]0.0 [0.0 to 0.0]0.00.0 [0.0 to 0.0]

Acceptance. The rules were committed in `scripts/insights_d2_report.py` at ead329bd (18:05 UTC on 2026-09-08), after the smoke runs' outcomes had been seen and while the later-aborted first attempt was running, and before any measurement repeat completed (the first completed at 18:49 UTC). Values are shown unrounded where they are thresholds or rates:

targetruleresultpass
precision_at_least_80candidate mean precision >= 0.80; gain over baseline > noise spread; paired 95% lower bound > 0{"candidate": "0", "baseline": "0.03122", "threshold": "0.8", "point": false, "gain": "-0.03122", "beyond_noise": false, "paired_bound": false}FAIL
correct_completed_not_below_baselinecandidate mean correct completed >= baseline mean; paired 95% interval not entirely below 0{"candidate": 0, "baseline": "1.333", "point": false, "gain": "-1.333", "paired_metric": "first_turn_correct_rate (the paired scorer's correct rate)", "paired_not_entirely_below_zero": true}FAIL
silent_wrong_down_by_halfcandidate mean silent wrong <= 0.5 x baseline mean; reduction > noise spread; paired 95% upper bound < 0{"candidate": "11.33", "baseline": "43.33", "threshold": "21.67", "point": true, "reduction": "32", "beyond_noise": true, "paired_bound": true}PASS
first_turn_errors_at_most_25_per_100candidate mean first-turn error rate <= 0.25; reduction over baseline > noise spread; paired 95% upper bound < 0{"candidate_rate": "0.204", "baseline_rate": "0.2701", "candidate": "23.67", "baseline": "31.33", "threshold_rate": "0.25", "point": true, "reduction": "7.667", "beyond_noise": true, "paired_bound": false}FAIL
full_conversations_at_least_20_per_50candidate mean full-conversation rate >= 0.40; gain over baseline > noise spread; paired 95% lower bound > 0 (3 conversations only){"candidate_rate": "0", "baseline_rate": "0", "candidate": 0, "baseline": 0, "conversations": 3, "threshold_rate": "0.4", "point": false, "gain": 0, "beyond_noise": false, "paired_bound": false}FAIL
zero_curated_invariant_violationsthe curated critical test files of the goal-3 V18 report (those present on this branch) pass on the candidate revision with zero failures; a node that failed while the suite ran beside the live measurement counts only if it also fails when rerun in isolation{"failed_as_run": 2, "failed_after_isolated_rerun": 0, "js_failed": 0}PASS
p95_latency_at_most_30scandidate p95 per-question latency (submission to terminal result) <= 30 s on every repeat{"threshold_ms": 30000}FAIL

Stable-case view (baseline repeats as control): stable-correct in baseline 0; lost 0 (correct to wrong 0, correct to error 0); gained 0. With no stable-correct baseline case, 'lost 0' is no evidence of preservation.

What the verifier rejected (signal tags on non-operational reject verdicts, the three fresh repeats and the real-usage run pooled; a verdict may carry several tags; the tags are the verifier's own diagnoses, not independently confirmed defects, and an applied repair is not a confirmed correction):

signal tagreject verdicts
missing-staff-house-exclusion149
wrong-population99
missing-staff-house-exclusions67
wrong-window45
missing-staff-exclusion29
missing-window-upper-bound25
wrong-grain19
missing-zero-fill15
missing-deleted-filter15
missing-house-exclusion13
missing-deleted-trade-filter12
wrong-prior-turn-scope10
staff-house-exclusion10
missing-house-exclusions8
forced-empty-leg6

Repairs skipped, by reason (the three fresh repeats and the real-usage run): invocation budget: 81, regeneration returned no new SQL: 17, regeneration failed: CalendarSemanticsError: 12, regeneration failed: FanOutSumError: 2, repair failed the leg gate: CtLegValidationError: 2.

Verifier call accounting: completed verifier calls exceed persisted verdicts by 4, 4, 2 and 2 in the four candidate runs (12 of 576 completed calls); those calls belong to turns that ended in an execution error or a job timeout after the call, so no turn record carries their verdict; the calls are in the verifier trace files. Failed calls (deadline or provider) are the operational verdicts.

Per-call statistics. Runs were sequential; the curated Python suite ran niced with one worker beside the first minutes of candidate fresh repeat 1 (18:16 to 18:19 UTC) and nothing else shared the host. Wall time includes the CLI child's start-up. Model ids: Claude runs report the billed model from Claude Code's modelUsage (a served identity); Codex runs report the CLI banner echo (not a server-side identity); the zai baseline records the requested id only, because the sealed runtime does not record response.model, and Goal B observed that endpoint serving glm-5.3 for glm-5.2 requests. Tokens: Claude reports input and output per call; Codex reports one total only (input and output unknown); a blank cell means not reported, never zero.

rungenerator callsfailedp50 msp95 msmodel id (source above)generator tokens in / out / totalverifier callsfailedp50 msp95 msmodel idverifier tokens totalsandbox read-only every calltool sections
cand-fresh-r127006231.810194.3{'claude-opus-5': 270}4723102 / 138756 /152111047.220790.0{'gpt-6-astra': 151}1069102True0
cand-fresh-r227006145.010274.7{'claude-opus-5': 270}4745365 / 138855 /149310854.021965.6{'gpt-6-astra': 146}901095True0
cand-fresh-r326906285.810207.3{'claude-opus-5': 269}4695022 / 137879 /150010803.221110.6{'gpt-6-astra': 150}887772True0
cand-real-r121128099.514074.7{'claude-opus-5': 209}4514029 / 149796 /134512076.721553.8{'gpt-6-astra': 129}843134True0
base-fresh-r122915279.39058.3{'glm-5.2': 228}2046958 / 70160 /00n/an/a0
base-fresh-r222804890.68061.6{'glm-5.2': 228}1993194 / 72056 /00n/an/a0
base-fresh-r322904283.77786.2{'glm-5.2': 229}2047197 / 73306 /00n/an/a0
base-real-r115804666.09024.2{'glm-5.2': 158}1799701 / 59736 /00n/an/a0

5. Real-usage sample (101 attempts, 86 supported first turns, 5 conversations; expected answers corrected to the rulings)

One repeat per arm, so no repeat variability is estimated here; the candidate's precision rests on three scored answers, the baseline's on 38.

armFirst-turn execution errorsCorrect first turnsPrecision (correct / answered)Silent wrongCorrect completedRefusalsComparison unverifiedFull conversationsrefusals by guardrefusals by verifierverifier verdicts accept / reject / clarify; repairs appliedp50 / p95 per question (s)
candidate r18/860/860.0% (0/3)3/1010/10172/10116/1010/527019 / 108 / 0 / 4437.4 / 59.8
baseline r110/863/867.9% (3/38)35/1013/10110/10139/1010/51000 / 0 / 0 / 09.2 / 22.3

The Goal C and D1 real-usage runs were rescored under the corrected oracle in section 2 (all 12 reproduce their old outcomes on 101/101 attempts first); they ran without the D2 definitions or the verifier, so they orient, they do not pair with, the two runs above.

6. Method

7. Unverified

8. Evidence