Measured, not claimed

Against the others, on the same tasks

Recorded task outcomes from 2026-09-17. This comparison is exploratory: matching models, resource budgets and independent task coverage have not been verified. Numerical leaders describe these recorded results; they do not establish a competitor win percentage.

The improvement race

See the lead. See the next challenge.

Explore the latest published measurements. Select a category, then focus or tap a bar for its evidence. Highlighted segments show a measured advantage.

Loading the published comparison…

Read the recorded benchmark tables.

Published results refresh here every minute. A new score requires a completed benchmark and publication; animation does not indicate a new measurement.

Comparison on corpus 1.0.0 (2026-09-17)

Exploratory recorded results. Metric leaders describe the available numbers, not confirmed product superiority. Matching model, resource budgets, independent tasks and uncertainty have not been verified. Missing metrics are not wins.

metricbetterpriestaiclaude-codeollama-aloneopenclaw-localopenclaw-cloudleader
correctnesshigher0.8671.00.00.5330.6Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader.
completionRatehigher1.01.01.01.01.0Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader.
secondsPerVerifiedTasklower161.6114.553no verified work130.06683.932Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader.
endToEndSecondsMedianlower115.84314.437no verified work62.81247.187Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader.
timeToFirstResultSecondslower90.656not measuredno verified worknot measurednot measuredIndicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader.
strayFilesPerTasklower0.01.6no verified work0.00.0Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader.
invalidToolCallRatelower0.226not measuredno verified worknot measurednot measuredIndicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader.
regressionRatelower0.1330.0no verified work0.0670.0Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader.

Per task (passed/trials)

taskpriestaiclaude-codeollama-aloneopenclaw-localopenclaw-cloud
fix-typo3/31/10/31/33/3
write-test2/31/10/31/33/3
read-only-answer3/31/10/30/30/3
rename-across-files2/31/10/33/30/3
make-the-test-pass3/31/10/33/33/3

Strengths and weaknesses

Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No strength or weakness is named.

Disclosure

The harness that produced this page ships inside every download: python priest_cli.py eval run runs the corpus on this machine, --agent-command runs any other agent on it, and eval compare writes a page like this one.