Full methodology: where our data comes from, how we verify it, when we update it, and why transparency is the foundation of every ranking we publish.
Every number on our LLM Models and Cost Comparison pages comes from a traceable source. We publish this methodology page for the same reason we require citation from the models we evaluate: accountability.
This page covers where we get our data, how we weight our rankings, how often we update, what we won't publish, and how you can verify our claims.
We use a four-tier sourcing hierarchy. Tier 1 is the most authoritative; Tier 4 is supplemental context only.
Model developers publish their own benchmark results, pricing, and technical specifications. These are our primary sources for:
We treat provider-claimed benchmarks as the starting point, not the final word. When independent evaluations (Tier 2) disagree with provider claims by more than 5%, we flag the discrepancy and default to the independent score.
Third-party evaluators provide cross-model comparisons under standardized conditions. Our Tier 2 sources:
We track academic papers that include model evaluations, particularly for:
When we cite a research paper, we link to it directly. If the paper is behind a paywall, we link to the preprint on arXiv or the author's institutional repository.
For metrics not covered by the above — particularly hallucination rates in long-context scenarios and real-world latency measurements — we run our own standardized test suites directly against each model's API. Our test methodology:
Our overall model rankings use a weighted composite of six factors:
| Factor | Weight | Source |
|---|---|---|
| GPT Benchmark (composite) | 30% | Tier 1 (provider-claimed, verified) |
| MMLU Accuracy | 20% | Tier 1 + Tier 2 (HELM cross-check) |
| LMSys ELO Rating | 20% | Tier 2 (LMSys Chatbot Arena) |
| Context Window Size | 10% | Tier 1 (official specs) |
| Hallucination Rate | 10% | Tier 4 (our own testing) + TruthfulQA |
| Cost Efficiency (score/$) | 10% | Tier 1 (official pricing) |
Each factor is normalized to a 0–100 scale. The weighted average produces the overall score. We publish the raw factor scores on our models page so readers can re-weight according to their own priorities — if cost matters more to you than context window, the data is there to recalculate.
| Data Type | Update Frequency | Trigger for Off-Cycle Update |
|---|---|---|
| Pricing | Monthly | Any provider price change >10% |
| Benchmark scores | Quarterly | New major benchmark release |
| Model roster | Continuous | New flagship model from any provider |
| Hallucination rates | Quarterly | Major model update or new safety findings |
| News articles | As published | N/A — each article has its own date |
The "last updated" date is visible on every data page. If a page says "Updated June 2026," every number on that page was verified against its source in June 2026.
Transparency also means being clear about what's not on our pages:
We want readers to fact-check us. Here's how:
If you find an error, tell us. We publish corrections prominently and update the "last modified" date. Accuracy matters more than being right the first time.
Google's EEAT guidelines (Experience, Expertise, Authoritativeness, Trustworthiness) reward sites that demonstrate transparency about their information sources. More importantly, trust is earned through verification, not claims. Every number on our ranking pages should be independently verifiable by a reader with an internet connection and a few minutes.
If you can't verify it, we haven't done our job.
Updated June 18, 2026. Methodology last reviewed June 2026. See the current rankings → or compare pricing →