Benchmarks

Peon, the agentic claim assessor.

Peon reads a claim end to end and proposes the payable amount per benefit — evaluated the way published model benchmarks are run: frozen datasets, versioned scorers, every number tied to a snapshot.

0%

Exact-match (2026-08-29-gemini37flash-wave16-overlay1)

253

Claims scored / 274 in op-vn-corpus-v1

3

Recorded runs

Charts

Accuracy, cost, and the trade-off between them.

Intelligence Index v1 (= exact-match %)

Exact-match rate per recorded run.

openai/gpt-5.6-sol (2026-08-27-sol-wave16)79.6% (218/274)
google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16)75.1% (190/253)
google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16-overlay1)75.9% (192/253)

$/claim

Average cost per assessed claim, per run.

google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16)$0.264
google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16-overlay1)$0.271
  • omitted: openai/gpt-5.6-sol (2026-08-27-sol-wave16): no spend recorded

Speed (claims/hour)

Throughput per run.

google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16)1.05
google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16-overlay1)1.05
  • omitted: openai/gpt-5.6-sol (2026-08-27-sol-wave16): latency isn't published yet

Intelligence Index v1 (= exact-match %) vs $/claim

Accuracy/cost trade-off across runs.

$0.263$0.272$/claim75.0%76.0%Intelligence Index v1 (= exact-match %)google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16)google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16): $0.264, 75.1%google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16-overlay1)google-vertex/gemini-3.7-flash (2026-08-29-gemini37flash-wave16-overlay1): $0.271, 75.9%
  • omitted: openai/gpt-5.6-sol (2026-08-27-sol-wave16): no cost recorded

Leaderboard

Every recorded run, named and dated.

RunDateModelExact-matchNPCoverage$/claimSpend
2026-08-27-sol-wave162026-08-27openai/gpt-5.6-sol218/274 (79.6%)12274/274 (100.0%)
2026-08-29-gemini37flash-wave162026-08-29google-vertex/gemini-3.7-flash190/253 (75.1%)15253/274 (92.3%)$0.264$66.67
2026-08-29-gemini37flash-wave16-overlay12026-08-29google-vertex/gemini-3.7-flash192/253 (75.9%)15253/274 (92.3%)$0.271$68.47

Methodology

Run like a published model benchmark.

Frozen snapshots

Every number here comes from a committed benchmark snapshot — model, code version, transport, and scorer pinned together — never a number remembered from a doc or compared against shifting production data.

Exact-match verdicts

A claim scores OK only when paid, non-paid, and shortfall amounts match the assessor's final exactly, per benefit. Anything else is a miss.

No-proposal counts as a miss

A run that escalates or defers a claim (NP) is not excluded from the denominator — it counts against exact-match, the same as a wrong answer.

Coverage, not a shrunk denominator

A partial run reports scored/dataset-size coverage rather than silently scoring only the claims it happened to finish.

Datasets

op-vn-corpus-v1

OutPatient VN assessed-claims corpus

274 claims · dev

op-vn-flash20-v1

OutPatient VN quick-20 subset

20 claims · dev

Scorer: dry-compare fold+BAD_FINAL 2026-08-27. Generated 2026-08-30.

See Peon run on your own claims.

Same evaluation discipline, applied to your claim mix.