ClickNext Bench · Internal Benchmark
Pick your AI model with real test data: accuracy, speed and cost per run through ClickNext Agent. Every number on this page comes from real runs, is auto-graded, and updates itself weekly.
Last updated 28 September 2026Every model gets the identical suite, run two ways — Raw = the bare model with no system prompt, Via App = the exact same system layer employees use in ClickNext Agent
| Model | Via App | Raw | Δ | Agent (4 tasks) | Avg time/item | Cost/run |
|---|---|---|---|---|---|---|
| gpt-5.6-lunabest value | 100% (45/45) | 97.8% | +2.2 | 4/4 | 1.3s | $0.0345 (~฿1.24) |
| deepseek-reasoner | 100% (45/45) | 100% | 0 | 4/4 | 1.1s | $0.0720 (~฿2.59) |
| deepseek-v4-flash | 100% (45/45) | 100% | 0 | 4/4 | 1.1s | $0.0721 (~฿2.60) |
| grok-build-0.1 | 100% (45/45) | 100% | 0 | 4/4 | 9.2s | $0.1688 (~฿6.08) |
| gpt-5.6-sol | 100% (45/45) | 100% | 0 | 4/4 | 1.4s | $0.3254 (~฿11.71) |
| gpt-5.6-terra | 100% (45/45) | 100% | 0 | 4/4 | 1.0s | $0.3312 (~฿11.92) |
| grok-4.5 | 100% (45/45) | 100% | 0 | 4/4 | 3.8s | $0.3398 (~฿12.23) |
| grok-4.6 | 100% (45/45) | 100% | 0 | 4/4 | 11.1s | $0.3531 (~฿12.71) |
| kimi-k3 | 100% (45/45) | 100% | 0 | 4/4 | 7.1s | $0.7960 (~฿28.66) |
| gpt-6-astra | 100% (45/45) | 100% | 0 | — | 1.8s | $1.6409 (~฿59.07) |
| autoapp router | 97.8% (44/45) | 91.1% | +6.7 | — | 1.1s | $0.0346 (~฿1.25) |
| deepseek-v4-pro | 97.8% (44/45) | 100% | −2.2 | 4/4 | 2.2s | $0.2244 (~฿8.08) |
| gpt-4o | 88.9% (40/45) | 84.4% | +4.4 | 4/4 | 0.9s | $0.4032 (~฿14.51) |
| gpt-4o-mini | 84.4% (38/45) | 80% | +4.4 | 3/4 | 0.7s | $0.0242 (~฿0.87) |
| Inferact/Qwen3.8-Flash-Next-NVFP4 | 82.2% (37/45) | 77.8% | +4.4 | 4/4 | 0.7s | — |
| deepseek-chat | 77.8% (35/45) | 88.9% | −11.1 | 4/4 | 0.6s | $0.0358 (~฿1.29) |
| openai/gpt-oss-120b | Incomplete data (rate limit) | — | — | incomplete | — | — |
| openai/gpt-oss-20b | Incomplete data (rate limit) | — | — | incomplete | — | — |
| qwen/qwen3.6-27b | Incomplete data (rate limit) | — | — | incomplete | — | — |
Δ = Via-App score minus Raw score (how many points our layer adds or costs) · Agent = 4 real file-editing tasks in a sandbox · Cost/run = real API cost of one full 45-item Via-App run (from actual token usage; baht estimated at 36 THB/USD) · auto = the app's router that picks a model per question automatically
Besides the main suite, an agent benchmark measures real file-editing work — the ability to act through the system, not just reply:
config.json and change a value as instructed without breaking the structureEach task runs in an isolated sandbox and is graded from the resulting files, never from the model's own claims.
Results show the Via-App score against Raw, plus average time per item and cost per run where available — models with near-identical scores can differ in cost by an order of magnitude, which is exactly what you need to see before picking one.
In auto mode, ClickNext Agent's router picks a model per question instead of locking everything to the most expensive one, delivering near-frontier quality at a far lower cost per run — see the real numbers in the table above.
ClickNext Bench is ClickNext's internal benchmark measuring how well each AI model performs when used through ClickNext Agent: a 45-item auto-graded suite covering math, logic, Thai language, data extraction, instruction following, coding, technical knowledge and hard problems, plus 4 real file-editing agent tasks.
Raw asks the model directly with no system prompt. Via App asks through ClickNext Agent's system layer — the same system prompt, working discipline and coaching employees use in production. The Δ column shows how many points our layer adds or costs; a positive Δ means the model performs better through the app than bare.
The results table shows each model's real cost per run, which can differ by several times at near-identical scores. The app's auto mode uses a router that picks a suitable model per question automatically, delivering quality close to frontier models at a much lower cost per run — users never have to pay top-model prices for every question.
It is not a global ranking — these are results from ClickNext's internal suite, for the model versions and conditions at test time. Its strength is transparency: fixed-answer auto-grading, identical criteria for every model, and published dates, methodology and real costs. Results may differ on other workloads.
The benchmark runs automatically every week and this page updates its numbers from the latest run automatically, with internal alerts when any model's score drops significantly from the previous round.
Download the app at clicknexttest.biz/agent/download (Windows, macOS, Linux), install with one click and sign in with your company email. Full usage documentation is on the docs and guide pages.
Free for employees, one-click install on Windows, macOS and Linux