Xunzhuo's picture
Publish clean Decision model repository
db8a7ed
|
Raw History Blame Contribute Delete
2.81 kB

Official typed-decisions TEST

The fixed final Lex checkpoint and the official laya-typed-decisions specialist were each evaluated once on the original 2,000 TEST decisions, from 400 state components. Both models support all 2,000 complete inputs within 1,024 tokens, their native windows and their untruncated official-default support. No TRAIN state component or full input overlaps this TEST support. The test contains four English workflows and 20 question schemas. This is not evidence of multilingual or unseen-schema generalization.

Metric Lex Official specialist (shipped temperature)
Original-label hard accuracy 78.15% (1,563/2,000) 76.60% (1,532/2,000)
Soft NLL 0.961902 0.884390
Mean soft Brier 0.030858 0.018452
Score RPS 0.022994 0.015602
Score expected-value MAE 0.251408 0.242408

Lex's accuracy gain is 1.55 percentage points (31 decisions). The paired 95% interval is [-0.20, +3.15125] percentage points, which includes zero. The interval uses 2,000 state-component bootstrap samples, seed 20260921, stratified by each component's workflow set; complete components are retained. This supports a higher observed hard-accuracy point, not statistically significant or uniformly better performance. Probability metrics are worse for Lex in this comparison. No calibration was fitted.

Type Decisions Lex accuracy Official accuracy
Choice 600 74.00% 73.33%
Noul 600 84.67% 85.67%
Score 800 76.375% 72.25%
Workflow Decisions Lex accuracy Official accuracy
Agent trace observability 500 72.4% 73.0%
Customer service 500 79.6% 76.4%
Invoice processing 500 84.4% 80.4%
Security incidents 500 76.2% 76.6%

Hard accuracy uses the source's original semantic hard label. Choice and Score use the first maximum in supplied order; Noul uses p_yes >= 0.5. Soft NLL uses the explicitly normalized original soft label distribution and natural logarithms (probability floor 1e-12). Brier averages squared probability error over the valid candidates; ordinal Score RPS averages cumulative-distribution squared error over K−1 boundaries. Score MAE compares the predicted expected value against the original source score. The official raw-temperature accuracy is also 76.60%; raw soft NLL is 0.869598.

The frozen TEST parquet SHA256 is 4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c. The actual saved-summary SHA256 is 1a21955b24be16cefa2963779e71edf4d6b5adcd14679fc438e6a87e3e604b88. No TEST text, labels or predictions are redistributed. The annotations are synthetic/teacher-derived; accuracy does not certify operational safety or calibrated business utility.