September 8, 2026

ClinReg benchmark leaderboard

Twenty-four models scored across three agentic documentation tasks — IND module 3 drafting, TLF for CSRs, and literature screening — with cost measured per completed run.

Category24 models · 10 open-weight

Overall score

mean of the three task scores · 0–100 · higher is better
0 20 40 60 80 100 GLM-5.3 · 89.389.3 GLM-5.3 Gemini 3.8 Flash · 88.688.6 Gemini 3.8 Flash GPT-5.6 Sol · 88.488.4 GPT-5.6 Sol GLM-5.3 Flash · 87.687.6 GLM-5.3 Flash GLM-5.2 · 87.487.4 GLM-5.2 Kimi K3 · 86.886.8 Kimi K3 Opus 5 · 86.386.3 Opus 5 Gemini 3.7 Flash · 85.485.4 Gemini 3.7 Flash GPT-5.5 · 84.984.9 GPT-5.5 Muse Spark 1.3 · 84.984.9 Muse Spark 1.3 GPT-5.6 Terra · 84.884.8 GPT-5.6 Terra Grok 4.5 · 84.384.3 Grok 4.5 Opus 4.8 · 82.682.6 Opus 4.8 Opus 4.7 · 81.681.6 Opus 4.7 Gemini 3.1 Pro · 80.880.8 Gemini 3.1 Pro Opus 4.6 · 79.279.2 Opus 4.6 GPT-5.6 Luna · 78.578.5 GPT-5.6 Luna Sonnet 5 · 76.376.3 Sonnet 5 Sonnet 4.6 · 76.376.3 Sonnet 4.6 DeepSeek V4 · 75.975.9 DeepSeek V4 Kimi K2.6 · 75.675.6 Kimi K2.6 MiniMax M3 · 73.673.6 MiniMax M3 Gemma 4 31B · 73.573.5 Gemma 4 31B Nemotron 3 Ultra · 70.370.3 Nemotron 3 Ultra
Table24 models · sortable by any column
ModelOverallINDTLFLit screenCost per run
avg of tasks
INDTLFLit
GLM-5.3opennew89.386.088.293.6$3.60$3.31$6.77$0.71
Gemini 3.8 Flashnew88.688.084.193.6$3.48$2.56$7.48$0.40
GPT-5.6 Sol88.486.385.293.6$6.48$3.69$12.82$2.92
GLM-5.3 Flashopennew87.685.085.392.5$0.30$0.15$0.62$0.12
GLM-5.2open87.482.786.093.6$2.19$2.86$3.27$0.44
Kimi K3open86.886.782.391.5$3.86$3.08$6.81$1.70
Opus 586.386.778.593.6$22.29$26.03$36.65$4.20
Gemini 3.7 Flashnew85.484.078.793.6$1.33$1.37$2.15$0.47
GPT-5.584.982.381.091.5$2.82$3.64$2.19$2.64
Muse Spark 1.3opennew84.990.073.391.5$1.83$1.13$3.81$0.55
GPT-5.6 Terra84.891.269.593.6$1.94$3.37$1.06$1.39
Grok 4.584.382.379.291.5$4.95$10.82$2.90$1.12
Opus 4.882.676.777.593.6$19.98$30.27$25.91$3.75
Opus 4.781.680.073.491.5$19.32$22.21$31.93$3.81
Gemini 3.1 Pro80.873.577.591.5$3.15$4.33$3.51$1.60
Opus 4.679.276.771.489.4$11.81$22.05$10.23$3.15
GPT-5.6 Luna78.587.054.993.6$1.58$3.52$0.61$0.61
Sonnet 576.376.767.285.1$10.91$6.77$23.70$2.27
Sonnet 4.676.365.072.591.5$7.51$13.43$7.35$1.75
DeepSeek V4open75.975.063.489.4$1.07$0.94$1.23$1.03
Kimi K2.6open75.671.763.591.5$1.30$1.71$1.52$0.67
MiniMax M3open73.670.063.587.2$0.98$1.82$0.91$0.21
Gemma 4 31Bopen73.572.758.589.4$0.97$0.43$2.24$0.25
Nemotron 3 Ultraopen70.354.365.091.5$2.25$4.70$1.55$0.52
openopen-weightnewadded in this releasescore shadinglow → high within the columnbold = best in the columnoverall = mean of the three task scores · cost = list price per completed run

Overall vs cost

average cost per taskopen-weightclosed
65 70 75 80 85 90 95 $0.5 $1 $2 $5 $10 $20 average cost per task (USD, log scale) Overall score GLM-5.3 · 89.3 · $3.60 per run · open-weight Gemini 3.8 Flash · 88.6 · $3.48 per run · closed GPT-5.6 Sol · 88.4 · $6.48 per run · closed GLM-5.3 Flash · 87.6 · $0.30 per run · open-weight GLM-5.2 · 87.4 · $2.19 per run · open-weight Kimi K3 · 86.8 · $3.86 per run · open-weight Opus 5 · 86.3 · $22.29 per run · closed Gemini 3.7 Flash · 85.4 · $1.33 per run · closed GPT-5.5 · 84.9 · $2.82 per run · closed Muse Spark 1.3 · 84.9 · $1.83 per run · open-weight GPT-5.6 Terra · 84.8 · $1.94 per run · closed Grok 4.5 · 84.3 · $4.95 per run · closed Opus 4.8 · 82.6 · $19.98 per run · closed Opus 4.7 · 81.6 · $19.32 per run · closed Gemini 3.1 Pro · 80.8 · $3.15 per run · closed Opus 4.6 · 79.2 · $11.81 per run · closed GPT-5.6 Luna · 78.5 · $1.58 per run · closed Sonnet 5 · 76.3 · $10.91 per run · closed Sonnet 4.6 · 76.3 · $7.51 per run · closed DeepSeek V4 · 75.9 · $1.07 per run · open-weight Kimi K2.6 · 75.6 · $1.30 per run · open-weight MiniMax M3 · 73.6 · $0.98 per run · open-weight Gemma 4 31B · 73.5 · $0.97 per run · open-weight Nemotron 3 Ultra · 70.3 · $2.25 per run · open-weight GLM-5.3 🏆 Gemini 3.8 Flash 🏆 GPT-5.6 Sol 🏆 GLM-5.3 Flash GLM-5.2 Kimi K3 Opus 5 Gemini 3.7 Flash GPT-5.5 Muse Spark 1.3 GPT-5.6 Terra Grok 4.5 Opus 4.8 Opus 4.7 Gemini 3.1 Pro Opus 4.6 GPT-5.6 Luna Sonnet 5 Sonnet 4.6 DeepSeek V4 Kimi K2.6 MiniMax M3 Gemma 4 31B Nemotron 3 Ultra
dotted lines join successive releases within one model line (GPT-5.6, Opus, Sonnet) · 🏆 top 3 on this score
02

What the board says.

LeadersGLM-5.3 leads at 89.3; Gemini 3.8 Flash (88.6) and GPT-5.6 Sol (88.4) follow. The top ten sit within 4.3 points.
Open weightsOpen-weight models hold 5 of the top 10. Best open-weight: GLM-5.3 at 89.3 ($3.60 per run). Best proprietary: Gemini 3.8 Flash at 88.6 ($3.48).
CostInside the top ten, average cost per run goes from $0.30 (GLM-5.3 Flash, 87.6) to $22.29 (Opus 5, 86.3): 75× the price for 1.3 points less.
Task dependenceGPT-5.6 Luna scores 87.0 on IND drafting but 54.9 on TLF programming, the widest spread on the board. Read the task columns, not only the overall.
03

How it is scored.

IND module 3 drafting

Draft the chemistry, manufacturing and controls sections of an IND from a public regulator's assessment report. A judge panel rates the extracted facts for faithfulness and coverage, 0–10; score = that rating × 10.

TLF for CSRs

From raw CRF data and specs, write and run the SDTM → ADaM → TLF pipeline for the CDISC pilot study in one session. Score = (share of output cells matching the published reference × 100 + reviewer panel × 10) / 2.

Literature screening

Full-text eligibility decisions on 47 papers from a Cochrane review, scored against the human reviewers. Score = percent of decisions correct.

Judges

Judged scores use models from three labs, never the model under test alone. IND fabrication and omission calls count only when three of four finders agree; TLF cell matches are exact comparisons to the reference; screening decisions are compared to the Cochrane reviewers' own.

Harness

Claude-family models ran under Claude Code, GPT-family models under Codex, Gemini and open-weight models under OpenCode. A score belongs to a model together with its harness and task skill.

Cost

List price at run time: input, output and reasoning tokens, cache reads at the provider's cached rate, averaged over the three tasks. Prices are verified on the provider's pricing page before a row is posted.
Read the full methods and results on the blogtask construction, judge panels, harness settings and per-model notes