Download evaluation/EVALUATION.md from vllm-sr/Decision-1.0-Lex-0.6B: direct link, hf CLI and curl.
- Browser
- Download file 2.81 kB
-
https://huggingface.co/vllm-sr/Decision-1.0-Lex-0.6B/resolve/main/evaluation/EVALUATION.md
- Command line
-
hf download hf://vllm-sr/Decision-1.0-Lex-0.6B/evaluation/EVALUATION.md
-
curl -L -o EVALUATION.md https://huggingface.co/vllm-sr/Decision-1.0-Lex-0.6B/resolve/main/evaluation/EVALUATION.md
Official typed-decisions TEST
The fixed final Lex checkpoint and the official laya-typed-decisions specialist were each evaluated once on the original 2,000 TEST decisions, from 400 state components. Both models support all 2,000 complete inputs within 1,024 tokens, their native windows and their untruncated official-default support. No TRAIN state component or full input overlaps this TEST support. The test contains four English workflows and 20 question schemas. This is not evidence of multilingual or unseen-schema generalization.
| Metric | Lex | Official specialist (shipped temperature) |
|---|---|---|
| Original-label hard accuracy | 78.15% (1,563/2,000) | 76.60% (1,532/2,000) |
| Soft NLL | 0.961902 | 0.884390 |
| Mean soft Brier | 0.030858 | 0.018452 |
| Score RPS | 0.022994 | 0.015602 |
| Score expected-value MAE | 0.251408 | 0.242408 |
Lex's accuracy gain is 1.55 percentage points (31 decisions). The paired 95% interval is [-0.20, +3.15125] percentage points, which includes zero. The interval uses 2,000 state-component bootstrap samples, seed 20260921, stratified by each component's workflow set; complete components are retained. This supports a higher observed hard-accuracy point, not statistically significant or uniformly better performance. Probability metrics are worse for Lex in this comparison. No calibration was fitted.
| Type | Decisions | Lex accuracy | Official accuracy |
|---|---|---|---|
| Choice | 600 | 74.00% | 73.33% |
| Noul | 600 | 84.67% | 85.67% |
| Score | 800 | 76.375% | 72.25% |
| Workflow | Decisions | Lex accuracy | Official accuracy |
|---|---|---|---|
| Agent trace observability | 500 | 72.4% | 73.0% |
| Customer service | 500 | 79.6% | 76.4% |
| Invoice processing | 500 | 84.4% | 80.4% |
| Security incidents | 500 | 76.2% | 76.6% |
Hard accuracy uses the source's original semantic hard label. Choice and Score use the first maximum in supplied order; Noul uses p_yes >= 0.5. Soft NLL uses the explicitly normalized original soft label distribution and natural logarithms (probability floor 1e-12). Brier averages squared probability error over the valid candidates; ordinal Score RPS averages cumulative-distribution squared error over K−1 boundaries. Score MAE compares the predicted expected value against the original source score. The official raw-temperature accuracy is also 76.60%; raw soft NLL is 0.869598.
The frozen TEST parquet SHA256 is 4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c. The actual saved-summary SHA256 is 1a21955b24be16cefa2963779e71edf4d6b5adcd14679fc438e6a87e3e604b88. No TEST text, labels or predictions are redistributed. The annotations are synthetic/teacher-derived; accuracy does not certify operational safety or calibrated business utility.