hypermodel
Workers' CompCommercial AutoPropertySecurityCompanyBlogBook a free AI audit

Research

A fine-tuned 27B open model beat every frontier model we tested on workers' comp claims work

September 30, 2026 · Yash Gupta and Piyush Varanjani

We fine-tuned an open-weight 27B model on workers' compensation decisions. On 496 tasks it had never seen, it got 84.5% fully right. The best of the 14 frontier configurations we tested, GPT-6 Astra, got 66.9%. Claude Opus 5 got 50.4%.

It also costs a ninth as much to run: $5.71 per 1,000 tasks, against $50.90 for GPT-6 Astra. A second model we tuned, a 35B mixture-of-experts with 3B active parameters, scored 82.7% at $1.62 per 1,000 tasks, about 30 times cheaper.

Bar chart of fully correct answers on 496 held-out claims tasks. Fine-tuned Qwen3.8-27B 84.5%, fine-tuned Qwen3.6-35B-A3B 82.7%, GPT-6 Astra 66.9%, down to Qwen3.6-35B-A3B before tuning at 16.1%.
Fine-tuned open models against 14 frontier configurations from OpenAI, Google and Anthropic, same 496 tasks, one attempt each.

Claims work runs on rules

Every claims decision is full of house rules: which document governs, which date starts the clock, what counts as a finding, how to cite it. General models know a great deal about the world and very little about these rules, so they paraphrase where the rule asks for the exact words, cite the wrong passage, or answer “not stated” when the finding is printed on the page. A model trained on the rules gets them right.

What the model does

Every task hands the model a real decision and a strict output contract. It has to find the facts the rules ask for, return them as structured data, and back each one with a quote copied exactly from the source. The test set spans eight task types from three workers' comp systems.

Test set of 496 tasks across eight task types: Texas disputed issue codes 84, panel action per issue 55, panel holding ledger 55; Illinois arbitrator findings 79, commission review outcome 74; Federal ECAB latest recorded event 52, procedural timeline 49, days between events 48.
Every task type comes from public appeals decisions. None of the 244 matters in the test set appears in training.

Two examples from the training data show what “the rules” means in practice.

Federal (ECAB): build the claim's procedural timeline. The rules register four event types, each in a fixed sentence form: claim filed, development letter, reconsideration request and denial decision. Anything else is noise, even when it carries a date.

On November 29, 2023 appellant, then a 36-year-old city cemetery caretaker, filed a
traumatic injury claim (Form CA-1) alleging that on October 27, 2023 he sustained a right
shoulder injury when backfilling a gravesite while in the performance of duty. [...]
Appellant stopped work on November 28, 2023. On January 12, 2024
OWCP accepted the claim for right shoulder sprain and right rotator cuff injury.

By decision dated January 12, 2024, OWCP denied appellant’s claim for COP, finding
that he had not reported his injury on an OWCP-approved form within 30 days [...]
{"answerability": "answerable", "elapsed_days": null, "events": [
  {"event_type": "claim_filed", "event_date": "2023-11-29", "actor": "appellant",
   "date_basis": "explicit_filing_date",
   "support_quote": "On November 29, 2023 appellant, then a 36-year-old city cemetery caretaker, filed a\ntraumatic injury claim (Form CA-1)"},
  {"event_type": "claim_denial_decision", "event_date": "2024-01-12", "actor": "OWCP",
   "date_basis": "decision_document_date",
   "support_quote": "By decision dated January 12, 2024, OWCP denied appellant’s claim"}]}

“OWCP accepted the claim” carries a date but is not a registered event, so it stays out of the timeline. The quotes keep the source's line break, because the rule asks for the literal clause.

Illinois (IWCC): read the arbitrator's findings, issue by issue. The findings block prints one sentence per disputed issue, and a single “not” flips the finding. The model returns every registered issue in order, with the exact sentence as evidence.

FINDINGS
On 04/23/2017, Respondent was operating under and subject to the provisions of the Act.
On this date, an employee-employer relationship did exist between Petitioner and Respondent.
On this date, Petitioner did sustain an accident that arose out of and in the course of employment.
Timely notice of this accident was given to Respondent.
Petitioner's current condition of ill-being is causally related to the accident.
In the year preceding the injury, Petitioner earned $43,956.00; the average weekly wage was $845.30.
On the date of accident, Petitioner was 61 years of age, single with 0 dependent children.
Petitioner has received all reasonable and necessary medical services.
Respondent has not paid all appropriate charges for all reasonable and necessary medical services.
Respondent is entitled to a credit of $7,626.48 under Section 8(j) of the Act.
{"issue_findings": [
  {"issue": "act_jurisdiction", "finding": "affirmative", "support_quote": "On 04/23/2017, Respondent was operating under and subject to the provisions of the Act."},
  {"issue": "employment_relationship", "finding": "affirmative", "support_quote": "On this date, an employee-employer relationship did exist between Petitioner and Respondent."},
  {"issue": "accident", "finding": "affirmative", "support_quote": "On this date, Petitioner did sustain an accident that arose out of and in the course of employment."},
  {"issue": "notice", "finding": "affirmative", "support_quote": "Timely notice of this accident was given to Respondent."},
  {"issue": "causal_connection", "finding": "affirmative", "support_quote": "Petitioner's current condition of ill-being is causally related to the accident."},
  {"issue": "medical_services_received", "finding": "affirmative", "support_quote": "Petitioner has received all reasonable and necessary medical services."},
  {"issue": "medical_charges_paid", "finding": "negative", "support_quote": "Respondent has not paid all appropriate charges for all reasonable and necessary medical services."}]}

The wage, age and credit lines are findings too, but not registered issues, so they are skipped. The unpaid medical charges come back negative because of one word.

Results

Scatter plot of accuracy against cost per 1,000 tasks on a log scale. The two fine-tuned models sit alone at top left: 84.5% at $5.71 and 82.7% at $1.62. GPT-6 Astra is 66.9% at $50.90.
Each point is one model configuration. Up and to the left is better. Two frontier configurations without complete cost records are not plotted.

The tuned models sit alone in the top-left corner: more accurate than every frontier configuration, and cheaper than most of them.

Line chart of fully correct answers during training. Qwen3.8-27B goes from 41.7% before tuning to 57.5% after 64 examples to 84.5% after two epochs. Qwen3.6-35B-A3B goes from 16.1% to 51.0% to 82.7%. The best frontier model, GPT-6 Astra, is 66.9%.
The same models, same prompts, measured before tuning, after the first 64 training examples, and after two epochs over 1,474 examples.

Tuning doubled the 27B, from 41.7% to 84.5%, and took the 35B from 16.1% to 82.7%. After only 64 training examples, the 27B already scored 57.5%, level with GPT-6 Sol at extra-high reasoning effort (56.9%). Two passes over the full training set did the rest.

How we built it

Data

  • Sources. Public appeals decisions from the federal Employees' Compensation Appeals Board (149 test tasks), the Illinois Workers' Compensation Commission (153) and the Texas appeals panel (194).
  • Labels by rule. Each task type registers the sentence forms that count. References are extracted deterministically from the decision text, and each one stores the character span and hash of its evidence. No model wrote a label.
  • Splits. 1,474 training examples, 108 held back for calibration, 496 test tasks. The split is by matter: none of the 244 matters in the test set appears in training or calibration. The test set was frozen before training, and nothing learned from it went back into training.

Training

  • Models. Qwen3.8-27B (dense) and Qwen3.6-35B-A3B (mixture of experts, 3B active parameters).
  • Recipe. Rank-8 LoRA, learning rate 1e-4, batch size 16, two epochs (186 optimizer steps), 32k context, thinking disabled, loss on assistant tokens only. The data order rotates task families and avoids repeating a matter within a batch.
  • Compute. Under $50 for the 27B.

Evaluation

  • One shot. One attempt per task, no tools, no retries. Our models ran at temperature 0; frontier models at provider defaults, with native JSON-schema output where the provider offers it and an 8,192-token output cap.
  • All or nothing. A task counts only when everything is right: the JSON parses and matches the schema, every field value is exact, and every quote is a literal span of the source that supports the answer. A partly right answer scores zero.
  • Cost. Reported input and output tokens × list price; our models at hosted sampling rates.
ModelFully correctTasksCost per 1,000 tasks
Qwen3.8-27B, fine-tuned by Hypermodel84.5%419 / 496$5.71
Qwen3.6-35B-A3B, fine-tuned by Hypermodel82.7%410 / 496$1.62
GPT-6 Astra · medium66.9%332 / 496$50.90
Gemini 3.8 Flash · low60.7%301 / 496$2.89
GPT-5.6 Sol · medium60.3%299 / 496$21.57
GPT-5.6 Terra · medium60.1%298 / 496$11.13
Gemini 3.8 Flash · medium57.3%284 / 496—
GPT-6 Sol · extra-high56.9%282 / 496$12.10
GPT-6 Sol · medium52.6%261 / 496$7.42
Claude Opus 5 · medium50.4%250 / 496$36.18
GPT-6 Luna · extra-high49.8%247 / 496$0.73
GPT-6 Luna · medium47.0%233 / 496$0.44
Claude Sonnet 544.6%221 / 496$14.23
Qwen3.8-27B, before tuning41.7%207 / 496$6.17
GPT-5.6 Luna · medium37.1%184 / 496—
GPT-5.6 Luna · no reasoning32.5%161 / 496$0.92
Claude Haiku 4.519.4%96 / 496$18.20
Qwen3.6-35B-A3B, before tuning16.1%80 / 496$1.76

What it means for claims teams

Most of the work in a claims operation is high-volume and rule-bound: reading files, pulling the facts, applying the organization's own standards the same way every time. That work does not need the largest model on the market. It needs a model that knows the rules, costs little enough to run on every file, and can run inside the organization's own environment.

That is what we build: models trained on a team's own rules, with frontier models reserved for the open-ended cases where they earn their cost.

hypermodel

The AI transformation partner for carriers, MGAs and TPAs.

Based in California · VC-backed

Product

Workers' CompCommercial AutoPropertyThe Claims Audit

Company

CompanySecurityBlog