aiplans.dev
Data & Pricing Methodology
Our comparisons separate source facts from our own calculations, keep every score attached to the benchmark it came from, and name the public sources below.
Pricing data sources
API token prices and subscription plan details come primarily from each provider’s own public pricing pages, API documentation, or provider-supplied machine-readable pricing endpoints. Cloud platforms (Azure OpenAI, AWS Bedrock, Vertex AI), aggregators (OpenRouter, SiliconFlow, Together AI and similar), and resellers such as XycAi are tracked as distinct channels and never presented as the model producer. When a channel exposes multiple public prices for the same model, such as XiuRouter service groups, XycAi account tiers, or OpenRouter upstream-provider endpoints, the comparison table uses a documented headline price and keeps the other variants attached to that channel.
Benchmark and leaderboard sources
We do not run the evaluations ourselves. Scores and ranks are synced nightly from the following public sources, and every displayed value retains its original benchmark name, version context, and snapshot date.
- Artificial Analysis — task benchmarks on the model leaderboard (GPQA Diamond, Humanity’s Last Exam, SciCode, Terminal-Bench Hard, IFBench, MMMU-Pro), the Coding Agent Index (an equal-weight combination of DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA pass@1, with measured cost and runtime per task), and the official per-million-token list prices used for price verification. Each coding-agent row is one agent × host-model configuration (e.g. Claude Code or Codex running a specific model), not a model-only score.
- LMArena — crowdsourced human-preference Arena ELO scores for chat models.
- VBench and VBench++ — standardized text-to-video and image-to-video benchmark scores, plus Arena AI’s public video leaderboard ranks for video models.
Normalization
- API prices are shown per one million tokens when the source supports token billing.
- Native currency remains visible; USD-normalized values are used for cross-currency ordering and savings calculations, using cached published reference exchange rates.
- Subscription prices keep their billing period and annual discount assumptions visible.
- “Official baseline” means the lowest tracked official or producer channel, not the cheapest channel overall.
Updates and verification
An automated pipeline runs every night and syncs prices, plans, benchmark scores, and coding-agent results. A read-only audit flags missing, stale, zero, inverted, or statistically unusual prices, and cross-checks our official-channel prices against the independently published Artificial Analysis list prices. When an external sync fails, the previous snapshot remains published until the next successful run. Every leaderboard table shows the snapshot date.
How to read benchmark scores
Benchmark scores are comparison signals measured on fixed task suites, not guarantees of quality for your workload. Leaderboards are re-run as agents and models ship versions, so values change over time; results from different benchmark versions are not directly comparable. Automation can still be wrong, so every purchase decision should be checked against the linked provider source.
Corrections
Corrections can be submitted through GitHub with a source URL. Material price changes are kept in the project price-history data where available.