Measured, not claimed
Against the others, on the same tasks
Recorded task outcomes from 2026-09-17. This comparison is exploratory: matching models, resource budgets and independent task coverage have not been verified. Numerical leaders describe these recorded results; they do not establish a competitor win percentage.
The improvement race
See the lead. See the next challenge.
Explore the latest published measurements. Select a category, then focus or tap a bar for its evidence. Highlighted segments show a measured advantage.
Loading the published comparison…
Published results refresh here every minute. A new score requires a completed benchmark and publication; animation does not indicate a new measurement.
Comparison on corpus 1.0.0 (2026-09-17)
Exploratory recorded results. Metric leaders describe the available numbers, not confirmed product superiority. Matching model, resource budgets, independent tasks and uncertainty have not been verified. Missing metrics are not wins.
| metric | better | priestai | claude-code | ollama-alone | openclaw-local | openclaw-cloud | leader |
|---|---|---|---|---|---|---|---|
| correctness | higher | 0.867 | 1.0 | 0.0 | 0.533 | 0.6 | Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader. |
| completionRate | higher | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader. |
| secondsPerVerifiedTask | lower | 161.61 | 14.553 | no verified work | 130.066 | 83.932 | Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader. |
| endToEndSecondsMedian | lower | 115.843 | 14.437 | no verified work | 62.812 | 47.187 | Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader. |
| timeToFirstResultSeconds | lower | 90.656 | not measured | no verified work | not measured | not measured | Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader. |
| strayFilesPerTask | lower | 0.0 | 1.6 | no verified work | 0.0 | 0.0 | Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader. |
| invalidToolCallRate | lower | 0.226 | not measured | no verified work | not measured | not measured | Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader. |
| regressionRate | lower | 0.133 | 0.0 | no verified work | 0.067 | 0.0 | Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No metric leader. |
Per task (passed/trials)
| task | priestai | claude-code | ollama-alone | openclaw-local | openclaw-cloud |
|---|---|---|---|---|---|
| fix-typo | 3/3 | 1/1 | 0/3 | 1/3 | 3/3 |
| write-test | 2/3 | 1/1 | 0/3 | 1/3 | 3/3 |
| read-only-answer | 3/3 | 1/1 | 0/3 | 0/3 | 0/3 |
| rename-across-files | 2/3 | 1/1 | 0/3 | 3/3 | 0/3 |
| make-the-test-pass | 3/3 | 1/1 | 0/3 | 3/3 | 3/3 |
Strengths and weaknesses
Indicative only: unequal trials (priestai 3, claude-code 1, ollama-alone 3, openclaw-local 3, openclaw-cloud 3). No strength or weakness is named.
Disclosure
- priestai: route kiwi_kiwi/mythos-9b-merged_v2-abliterated:q4k; trials 3; hardware laptop; served {'model': 'kiwi_kiwi/mythos-9b-merged_v2-abliterated:q4k', 'host': 'local', 'asked': '', 'ctx': 65536, 'ctx_measured': True}; this product; warm-up not recorded (result predates the warm-up phase; any model load is inside its task seconds); configuration {"agentCommand": "", "agentEnvKeys": [], "externalAgent": false, "hardware": {"cpuCount": 32, "gpuMemoryGb": null, "memoryGb": 15.6, "platform": "win32"}, "label": "priestai", "mode": "edits", "productVersion": "0.1.0"}
- claude-code: route external:claude; trials 1 (indicative: under three); hardware laptop; served not reported; external agent through a subprocess (tool calls not observable); warm-up not recorded (result predates the warm-up phase; any model load is inside its task seconds); configuration {"agentCommand": "claude", "agentEnvKeys": [], "externalAgent": true, "hardware": {"cpuCount": 32, "gpuMemoryGb": null, "memoryGb": 15.6, "platform": "win32"}, "label": "claude-code", "mode": "edits", "productVersion": "0.1.0"}
- ollama-alone: route external:ollama; trials 3; hardware laptop; served not reported; external agent through a subprocess (tool calls not observable); warm-up not recorded (result predates the warm-up phase; any model load is inside its task seconds); configuration {"agentCommand": "ollama", "agentEnvKeys": [], "externalAgent": true, "hardware": {"cpuCount": 32, "gpuMemoryGb": null, "memoryGb": 15.6, "platform": "win32"}, "label": "ollama-alone", "mode": "edits", "productVersion": "0.1.0"}
- openclaw-local: route external:openclaw; trials 3; hardware laptop; served not reported; external agent through a subprocess (tool calls not observable); warm-up not recorded (result predates the warm-up phase; any model load is inside its task seconds); configuration {"agentCommand": "openclaw", "agentEnvKeys": ["OLLAMA_API_KEY"], "externalAgent": true, "hardware": {"cpuCount": 32, "gpuMemoryGb": null, "memoryGb": 15.6, "platform": "win32"}, "label": "openclaw-local", "mode": "edits", "productVersion": "0.1.0"}
- openclaw-cloud: route external:openclaw; trials 3; hardware laptop; served not reported; external agent through a subprocess (tool calls not observable); warm-up not recorded (result predates the warm-up phase; any model load is inside its task seconds); configuration {"agentCommand": "openclaw", "agentEnvKeys": [], "externalAgent": true, "hardware": {"cpuCount": 32, "gpuMemoryGb": null, "memoryGb": 15.6, "platform": "win32"}, "label": "openclaw-cloud", "mode": "edits", "productVersion": "0.1.0"}
The harness that produced this page ships inside every download: python priest_cli.py eval run
runs the corpus on this machine, --agent-command runs any other agent on it, and eval compare writes a page like this one.