Peon, the agentic claim assessor.
Peon reads a claim end to end and proposes the payable amount per benefit — evaluated the way published model benchmarks are run: frozen datasets, versioned scorers, every number tied to a snapshot.
0%
Exact-match (2026-08-29-gemini37flash-wave16-overlay1)
253
Claims scored / 274 in op-vn-corpus-v1
3
Recorded runs
Charts
Accuracy, cost, and the trade-off between them.
Intelligence Index v1 (= exact-match %)
Exact-match rate per recorded run.
$/claim
Average cost per assessed claim, per run.
- omitted: openai/gpt-5.6-sol (2026-08-27-sol-wave16): no spend recorded
Speed (claims/hour)
Throughput per run.
- omitted: openai/gpt-5.6-sol (2026-08-27-sol-wave16): latency isn't published yet
Intelligence Index v1 (= exact-match %) vs $/claim
Accuracy/cost trade-off across runs.
- omitted: openai/gpt-5.6-sol (2026-08-27-sol-wave16): no cost recorded
Leaderboard
Every recorded run, named and dated.
| Run | Date | Model | Exact-match | NP | Coverage | $/claim | Spend |
|---|---|---|---|---|---|---|---|
| 2026-08-27-sol-wave16 | 2026-08-27 | openai/gpt-5.6-sol | 218/274 (79.6%) | 12 | 274/274 (100.0%) | — | — |
| 2026-08-29-gemini37flash-wave16 | 2026-08-29 | google-vertex/gemini-3.7-flash | 190/253 (75.1%) | 15 | 253/274 (92.3%) | $0.264 | $66.67 |
| 2026-08-29-gemini37flash-wave16-overlay1 | 2026-08-29 | google-vertex/gemini-3.7-flash | 192/253 (75.9%) | 15 | 253/274 (92.3%) | $0.271 | $68.47 |
Methodology
Run like a published model benchmark.
Frozen snapshots
Every number here comes from a committed benchmark snapshot — model, code version, transport, and scorer pinned together — never a number remembered from a doc or compared against shifting production data.
Exact-match verdicts
A claim scores OK only when paid, non-paid, and shortfall amounts match the assessor's final exactly, per benefit. Anything else is a miss.
No-proposal counts as a miss
A run that escalates or defers a claim (NP) is not excluded from the denominator — it counts against exact-match, the same as a wrong answer.
Coverage, not a shrunk denominator
A partial run reports scored/dataset-size coverage rather than silently scoring only the claims it happened to finish.
Datasets
op-vn-corpus-v1
OutPatient VN assessed-claims corpus
op-vn-flash20-v1
OutPatient VN quick-20 subset
Scorer: dry-compare fold+BAD_FINAL 2026-08-27. Generated 2026-08-30.
See Peon run on your own claims.
Same evaluation discipline, applied to your claim mix.