Current Run
Estimated spend
—
Comparison unavailable
| Task | Input chars | Words |
|---|
(updated …)
Our claim is that we are in the subsidiary phase of Artificial Intelligence providers. Providers are scrambling to build as much infrastructure as they can afford to accommodate increased demand down the road as AI becomes more and more ubiquitous. As we saw in ride-sharing and delivery services companies in the early and mid-2010s, prices for these services were artificially low in an effort to gain market share.
These pricing metrics are all publicly-available as a ‘cost per 1m tokens’ on each provider’s pricing pages and documentation. However, this is only half the metric to determine which models are indeed ‘low cost’. Total cost factors in how efficient they are in providing responses, as the price per token multiplied by tokens consumed is the true measure of consumer spend.
We’ve standardized six different tasks that test different dimensions of language models. Tasks of rationalization, logic, call-and-response are issued to both a flagship and workhorse model for each provider every day. We do not judge the arbitrary measure of quality – how well the respective models complete their tasks – but instead track the model’s efficiency in completing these tasks over time.
It is this, a comparative token expenditure, that is less transparent to the public than a standardized ‘cost per 1m tokens’ metric. Our hypothesis is as follows – we will observe differentiation (are DeepSeek and Qwen really low cost?) and ‘drift’ between the relative level of output from each provider as both their flagship and workhorse models evolve.
Each provider typically offers multiple, general-purpose models. We test their frontier grade (flagship) model, and their daily-driver (workhorse) in order to map the discrepancy in token expenditures between the two. The upgrading of the different flagship vs. workhorse models versions over time are reflected by the broken segments of the lines associated with each provider. What is of interest is that, assuming the number of parameters for each model increases with every version release, does the number of expended tokens increase as well?
We are still in the early stages of our analysis, but there exists a significant difference in the token expenditure across these half-dozen providers. A 10x differential in token expenditure for Full Suite of Tasks between the models from Grok/xAI and Google suggests that, despite a relative pricing alignment on Flagship vs Workhorse (~15% variance) there exists a significant difference in cost-per-usage.
DeepSeek is an absolute ‘token hog’ when it comes to Task C, the coding task. Its consumption of ~33k tokens is nearly 10x of the majority of other providers. Again, we make no assumption of quality through these exercises, but is this discrepancy an indication of the proficiency for different models against specific task types?
| Date | Provider | Model | Task | Prompt tokens | / 1K chars |
|---|
Estimated spend
—
Comparison unavailable
Estimated spend
—
Comparison unavailable
Estimated spend
—
Comparison unavailable
| Date | Provider | Model | Task | Input | Output | Supporting | Total |
|---|
Estimated spend = observed usage × the same-day published rate for the served model. It may differ from provider invoices.