Decision contract
Question, comparators, primary gate, exclusions, and stop rules.
SYNTHETIC DECISION RECORD / VH-DR-0013
INDEPENDENT AI SYSTEM EVALUATION / SYNTHETIC DEMONSTRATION
Candidate may proceed under the quality gate in this synthetic scope. Runtime remains unresolved; no speed or savings claim is authorized.
This dossier demonstrates the depth, traceability, and restraint of a Verahelm delivery. It is not a customer result, certification, security review, legal opinion, or promised operational or financial outcome.
DOSSIER MAP
QUESTION → EVIDENCE → DECISIONQuestion, comparators, primary gate, exclusions, and stop rules.
Fixture composition, class balance, authority labels, and evidence boundary.
Field, case, risk, segment, runtime, and paired-comparison results.
Complete candidate error inventory, regressions, and critical handling defect.
Controls, ablation, replay, independent checks, and uncertainty.
Risks, owners, stage gates, approval state, pilot design, and next discriminating test.
Stable example identifier for discussion and review.
No customer or field evidence is represented.
A real engagement names the person authorized to decide.
Material input, version, environment, or control changes require reissue.
All organizations, tickets, outputs, and measurements in this public dossier are fictional or synthetic.
OPERATING DECISION
BOUNDARY / EXPLICITTHE QUESTION
Quality observations favor the candidate. Runtime is slower, so a speed claim is withheld and the limitation remains open.
Whether the guarded candidate merits a controlled, recommendation-only customer benchmark.
Frozen straightforward baseline, guarded candidate, and differentiator ablation on identical records.
Correctness first; high-consequence handling controls the release interpretation.
No autonomous disposition while any uncontrolled high-risk route or action error remains.
WORKLOAD AND LABEL CONTRACT
30 RECORDS / 120 SCORED FIELDSSYNTHETIC FIXTURE
Each ticket has four required outputs: route, risk, action, and review flag. A case is exact only when all four fields match the frozen gold record.
Rare high-consequence cases are intentionally overrepresented for the demonstration.
Review correctness is measured separately from route, risk, and action.
Correct accountable function receives the record.
Consequence class matches the gold boundary.
Required operational response is preserved.
Human authority is invoked when the contract requires it.
RESULT MATRIX
BASELINE / CANDIDATEAverages do not control the release state. The remaining high-risk route/action defect keeps the recommendation conditional.
SEGMENT AND REGRESSION AUDIT
AGGREGATES DO NOT HIDE FAILURESBY GOLD RISK
Low-risk errors can still create unnecessary review load and alert fatigue.
BY GOLD ROUTE
Six records per route are diagnostic, not a stable production estimate.
COMPLETE CANDIDATE ERROR INVENTORY
5 CASES / 8 FIELD ERRORSUnnecessary engineering work and incorrect action.
Drawing conflict may bypass engineering ownership.
Excess review and alert burden.
Ownership error despite a low-risk label.
High risk detected; required authority and action were wrong.
Perfect high-risk detection is not perfect high-risk handling. Case T025 controls the autonomy boundary.
EVIDENCE CHAIN
CHECKS / CONTROLSComparison rules and the synthetic case set remained fixed for the evaluated result.
PASSPositive, negative, coverage, determinism, and ablation controls passed within the demonstration.
PASSA separately implemented score recomputation and a manual record-level review agreed with the stated synthetic result.
PASSTwenty-one comparison runs were interleaved across three synthetic workload sizes.
PASSThe ablation scored between baseline and candidate; contribution is supported only inside this fixture.
89.2%Four regressions, one critical action error, and slower runtime remain in the decision record.
OPENIndependent agreement inside a synthetic test is not customer-blinded production validation.
UNCERTAINTY AND PAIRED COMPARISON
DESCRIPTIVE ≠ POPULATION PROOFCandidate: 25/30 exact cases. Simple 95% Wilson interval: 66.4%–92.7%.
Nine candidate improvements versus four regressions on matched cases.
The paired difference is not significant at α=0.05 on this 30-case fixture. Values are reported to three decimal places.
These intervals do not account for distribution shift, label error, dependence between records, or customer-specific workflow conditions.
RUNTIME AND VALUE BOUNDARY
NO SPEEDUP / NO SAVINGS CLAIMThe candidate is slower in this local proxy. Quality evidence does not authorize a runtime, throughput, cost, or return-on-investment claim.
LIMITATION REGISTER
OPEN / RETAINEDThe workload contains no customer data and does not reproduce a live operating environment.
The higher-quality candidate is slower than the simple baseline; no speed claim is made.
No buyer-controlled blinded holdout or field deployment has tested transfer beyond this demonstration.
Any financial illustration would be a planning assumption, not a savings or return guarantee.
CUSTOMER-BLINDED PILOT DESIGN
PROPOSED / NOT YET EXECUTEDNEXT EVIDENCE LEVEL
The buyer retains the de-identified blind set until decision rules, candidate configuration, metrics, exclusions, and stop conditions are frozen.
DECISION HANDOFF
NEXT TEST / MINIMUMNEXT DISCRIMINATING TEST
Use authorized, de-identified buyer records. Retain the baseline, record distribution shifts and exceptions, and authorize only the claims that survive the new evidence boundary.
Example role only; a real delivery names the responsible preparer and issue date.
Synthetic score agreement is recorded; customer-blinded validation remains outstanding.
No person or organization is authorized by this public sample.
Recommendation-only progression; autonomy and external claims remain withheld.
VERAHELM HOLDINGS LLC / INDEPENDENT TECHNICAL EVALUATION