Research
A fine-tuned 27B open model beat every frontier model we tested on workers' comp claims work
September 30, 2026 · Yash Gupta and Piyush Varanjani
We fine-tuned an open-weight 27B model on workers' compensation decisions. On 496 tasks it had never seen, it got 84.5% fully right. The best of the 14 frontier configurations we tested, GPT-6 Astra, got 66.9%. Claude Opus 5 got 50.4%.
It also costs a ninth as much to run: $5.71 per 1,000 tasks, against $50.90 for GPT-6 Astra. A second model we tuned, a 35B mixture-of-experts with 3B active parameters, scored 82.7% at $1.62 per 1,000 tasks, about 30 times cheaper.
Claims work runs on rules
Every claims decision is full of house rules: which document governs, which date starts the clock, what counts as a finding, how to cite it. General models know a great deal about the world and very little about these rules, so they paraphrase where the rule asks for the exact words, cite the wrong passage, or answer “not stated” when the finding is printed on the page. A model trained on the rules gets them right.
What the model does
Every task hands the model a real decision and a strict output contract. It has to find the facts the rules ask for, return them as structured data, and back each one with a quote copied exactly from the source. The test set spans eight task types from three workers' comp systems.
Two examples from the training data show what “the rules” means in practice.
Federal (ECAB): build the claim's procedural timeline. The rules register four event types, each in a fixed sentence form: claim filed, development letter, reconsideration request and denial decision. Anything else is noise, even when it carries a date.
On November 29, 2023 appellant, then a 36-year-old city cemetery caretaker, filed a
traumatic injury claim (Form CA-1) alleging that on October 27, 2023 he sustained a right
shoulder injury when backfilling a gravesite while in the performance of duty. [...]
Appellant stopped work on November 28, 2023. On January 12, 2024
OWCP accepted the claim for right shoulder sprain and right rotator cuff injury.
By decision dated January 12, 2024, OWCP denied appellant’s claim for COP, finding
that he had not reported his injury on an OWCP-approved form within 30 days [...]{"answerability": "answerable", "elapsed_days": null, "events": [
{"event_type": "claim_filed", "event_date": "2023-11-29", "actor": "appellant",
"date_basis": "explicit_filing_date",
"support_quote": "On November 29, 2023 appellant, then a 36-year-old city cemetery caretaker, filed a\ntraumatic injury claim (Form CA-1)"},
{"event_type": "claim_denial_decision", "event_date": "2024-01-12", "actor": "OWCP",
"date_basis": "decision_document_date",
"support_quote": "By decision dated January 12, 2024, OWCP denied appellant’s claim"}]}“OWCP accepted the claim” carries a date but is not a registered event, so it stays out of the timeline. The quotes keep the source's line break, because the rule asks for the literal clause.
Illinois (IWCC): read the arbitrator's findings, issue by issue. The findings block prints one sentence per disputed issue, and a single “not” flips the finding. The model returns every registered issue in order, with the exact sentence as evidence.
FINDINGS
On 04/23/2017, Respondent was operating under and subject to the provisions of the Act.
On this date, an employee-employer relationship did exist between Petitioner and Respondent.
On this date, Petitioner did sustain an accident that arose out of and in the course of employment.
Timely notice of this accident was given to Respondent.
Petitioner's current condition of ill-being is causally related to the accident.
In the year preceding the injury, Petitioner earned $43,956.00; the average weekly wage was $845.30.
On the date of accident, Petitioner was 61 years of age, single with 0 dependent children.
Petitioner has received all reasonable and necessary medical services.
Respondent has not paid all appropriate charges for all reasonable and necessary medical services.
Respondent is entitled to a credit of $7,626.48 under Section 8(j) of the Act.{"issue_findings": [
{"issue": "act_jurisdiction", "finding": "affirmative", "support_quote": "On 04/23/2017, Respondent was operating under and subject to the provisions of the Act."},
{"issue": "employment_relationship", "finding": "affirmative", "support_quote": "On this date, an employee-employer relationship did exist between Petitioner and Respondent."},
{"issue": "accident", "finding": "affirmative", "support_quote": "On this date, Petitioner did sustain an accident that arose out of and in the course of employment."},
{"issue": "notice", "finding": "affirmative", "support_quote": "Timely notice of this accident was given to Respondent."},
{"issue": "causal_connection", "finding": "affirmative", "support_quote": "Petitioner's current condition of ill-being is causally related to the accident."},
{"issue": "medical_services_received", "finding": "affirmative", "support_quote": "Petitioner has received all reasonable and necessary medical services."},
{"issue": "medical_charges_paid", "finding": "negative", "support_quote": "Respondent has not paid all appropriate charges for all reasonable and necessary medical services."}]}The wage, age and credit lines are findings too, but not registered issues, so they are skipped. The unpaid medical charges come back negative because of one word.
Results
The tuned models sit alone in the top-left corner: more accurate than every frontier configuration, and cheaper than most of them.
Tuning doubled the 27B, from 41.7% to 84.5%, and took the 35B from 16.1% to 82.7%. After only 64 training examples, the 27B already scored 57.5%, level with GPT-6 Sol at extra-high reasoning effort (56.9%). Two passes over the full training set did the rest.
How we built it
Data
- Sources. Public appeals decisions from the federal Employees' Compensation Appeals Board (149 test tasks), the Illinois Workers' Compensation Commission (153) and the Texas appeals panel (194).
- Labels by rule. Each task type registers the sentence forms that count. References are extracted deterministically from the decision text, and each one stores the character span and hash of its evidence. No model wrote a label.
- Splits. 1,474 training examples, 108 held back for calibration, 496 test tasks. The split is by matter: none of the 244 matters in the test set appears in training or calibration. The test set was frozen before training, and nothing learned from it went back into training.
Training
- Models. Qwen3.8-27B (dense) and Qwen3.6-35B-A3B (mixture of experts, 3B active parameters).
- Recipe. Rank-8 LoRA, learning rate 1e-4, batch size 16, two epochs (186 optimizer steps), 32k context, thinking disabled, loss on assistant tokens only. The data order rotates task families and avoids repeating a matter within a batch.
- Compute. Under $50 for the 27B.
Evaluation
- One shot. One attempt per task, no tools, no retries. Our models ran at temperature 0; frontier models at provider defaults, with native JSON-schema output where the provider offers it and an 8,192-token output cap.
- All or nothing. A task counts only when everything is right: the JSON parses and matches the schema, every field value is exact, and every quote is a literal span of the source that supports the answer. A partly right answer scores zero.
- Cost. Reported input and output tokens × list price; our models at hosted sampling rates.
| Model | Fully correct | Tasks | Cost per 1,000 tasks |
|---|---|---|---|
| Qwen3.8-27B, fine-tuned by Hypermodel | 84.5% | 419 / 496 | $5.71 |
| Qwen3.6-35B-A3B, fine-tuned by Hypermodel | 82.7% | 410 / 496 | $1.62 |
| GPT-6 Astra · medium | 66.9% | 332 / 496 | $50.90 |
| Gemini 3.8 Flash · low | 60.7% | 301 / 496 | $2.89 |
| GPT-5.6 Sol · medium | 60.3% | 299 / 496 | $21.57 |
| GPT-5.6 Terra · medium | 60.1% | 298 / 496 | $11.13 |
| Gemini 3.8 Flash · medium | 57.3% | 284 / 496 | — |
| GPT-6 Sol · extra-high | 56.9% | 282 / 496 | $12.10 |
| GPT-6 Sol · medium | 52.6% | 261 / 496 | $7.42 |
| Claude Opus 5 · medium | 50.4% | 250 / 496 | $36.18 |
| GPT-6 Luna · extra-high | 49.8% | 247 / 496 | $0.73 |
| GPT-6 Luna · medium | 47.0% | 233 / 496 | $0.44 |
| Claude Sonnet 5 | 44.6% | 221 / 496 | $14.23 |
| Qwen3.8-27B, before tuning | 41.7% | 207 / 496 | $6.17 |
| GPT-5.6 Luna · medium | 37.1% | 184 / 496 | — |
| GPT-5.6 Luna · no reasoning | 32.5% | 161 / 496 | $0.92 |
| Claude Haiku 4.5 | 19.4% | 96 / 496 | $18.20 |
| Qwen3.6-35B-A3B, before tuning | 16.1% | 80 / 496 | $1.76 |
What it means for claims teams
Most of the work in a claims operation is high-volume and rule-bound: reading files, pulling the facts, applying the organization's own standards the same way every time. That work does not need the largest model on the market. It needs a model that knows the rules, costs little enough to run on every file, and can run inside the organization's own environment.
That is what we build: models trained on a team's own rules, with frontier models reserved for the open-ended cases where they earn their cost.