Capability–Accessibility Gap Tracker
How much AI capability growth is actually accessible?
Overview
Discussions of model significance typically focus on the technical capabilities of models alone. That makes sense for predicting the trajectory of model capabilities but doesn’t do much to give an intuitive sense of the real world significance of existing models. A model that is 95% as capable as the cutting edge for 10% of the cost is more likely to be impactful because it is a more competitive candidate for widespread adoption. This project is designed to give an intuitive sense for how competitive models are when weighed against their accessibility to give a clearer sense of which models are worth tracking closely.
The diffusibility score below combines a capability score, derived from Epoch AI's Epoch Capabilities Index (ECI), and an accessibility score that takes the better of two paths: how cheap the model is to use over an API, or — only for open-weight models — how modest its VRAM requirement is at 4-bit quantization.
Explore the data
Each point below is one model. The x-axis is capability; the y-axis is the better of its cost-based or VRAM-based accessibility score. Point size is the Diffusibility score, and the model with the highest Diffusibility is marked with a star and outlined in the chart below. Adjust the parameters underneath to see how sensitive the ranking is to each assumption.
Capability vs. Accessibility
Diffusibility = √(Capability × Accessibility)
Sources and data
Capability: Epoch AI — ECI score, fit from eci_benchmarks.csv via the eci-public package (not a raw CSV column).
Pricing: llm-prices.com (github.com/simonw/llm-prices), with OpenRouter's public API as a fallback for models not listed there.
Parameter counts (VRAM basis): Epoch AI's all_ai_models.csv "Parameters" field, open-weight models only.
Independent research tool. Not affiliated with or endorsed by Our World in Data; page styling is inspired by their Grapher tool.
Tunable parameters
Methodology
Capability score normalizes Epoch's raw ECI (an unbounded, Elo-like scale) to a 0–1 range using two tunable anchors, eci_floor and eci_ceiling. Raw ECI is never used directly anywhere downstream.
Cost score applies a logistic curve to a blended per-token price (a 3:1 input:output mix by default, matching a typical chat workload), centered on a reference price meant to approximate a $20/month consumer subscription converted to an effective $/M-token rate.
VRAM score applies the same logistic shape to an estimated VRAM requirement at 4-bit quantization (≈0.5–0.6 GB per billion parameters, plus fixed runtime overhead), centered on a single high-end consumer GPU's memory ceiling. It is only computed for open-weight models — closed models have no self-hosting path, so it is left null and has no bearing on their score.
Accessibility, final takes the maximum of the two: a model only needs one viable path to count as accessible.
Diffusibility is the geometric mean of capability and accessibility, scaled by regulatory ease (currently a constant 1.0 for every model). It's meant to capture how likely a model is to actually diffuse through real-world use, not just how capable it is in the abstract.
Assumptions and Design Choices
Diffusibility is described here as the geometric mean of capability and accessibility. This both reduces the total potential variance from error in determining accessibility and models the assumption that a model’s diffusibility will ultimately be bottlenecked by either of these elements if one is sufficiently far behind the other. Even an incredibly capable model will not be widespread if it’s difficult to access and even a highly accessible model will be ignored if there are models that offer considerably greater utility. Accessibility is currently described only in terms of how expensive it is to run or, in the case of open weight models, the cost and how easy it is to run on hardware. Accessibility is modeled on a log scale on a sigmoid relative to a given reference point. This assumes that there are diminishing returns on cost and hardware accessibility the more or less accessible you go. The reference point is set by default to center on the rough cost/token of a consumer on a paid subscription plan ($20/month) and how much VRAM is accessible to a high end consumer. You can adjust the parameters according to your own assumptions. Capability is normalized according to where they sit between a baseline that all the models exceed and a ceiling that surpasses all available models (by default set to 100 and 170 respectively.) It’s worth noting that these assumptions only hold for higher income countries: 24GB of VRAM and $8/M is not a reasonable midpoint for a low or middle income country.
Interesting Findings and Future Work
This is a first version, meaning that the calculation for accessibility is incredibly basic for the time being. Memory is not the only requirement to run an LLM locally but it is a decent proxy for the broader set of compute requirements which will require (or otherwise be bottlenecked by) memory. Other hurdles such as regulatory restrictions and geographic constraints (such as broadband or electricity access) are also significant parts of accessibility that should be included in future models.
One interesting finding is that, based on the default assumptions of this model, cost/token is the easier way to access even low compute models across all open source models. This would make sense given the cost advantages available to hyperscalers from returns to scale.
Calculating for pareto optimality shows that there are a wide variety of models on the pareto curve but also several that are strictly dominated by their competitors.
Sources
Epoch AI (epoch.ai) for capability scores and model metadata; llm-prices.com and OpenRouter for pricing. See the expandable "Sources and data" strip under the chart above for the exact fields used, and the data-quality summary printed by transform.py for any models excluded due to missing data.