# Ultrabench.ai Aims to Be the Definitive AI Benchmark Aggregator

> The free index from Iternal Technologies sets each AI model's benchmark scores beside its price per million tokens, parameter count, memory by quantization and hardware fit, and serves the whole board as keyless JSON.

*Published 2026-09-26 · Updated 2026-09-26 · By Marcus Vance*

In short
[Ultrabench.ai](https://ultrabench.ai/) is a free AI benchmark aggregator from Iternal Technologies. It ranks current large language models on a weighted intelligence index and sets each score beside price per million tokens, parameter count, memory by quantization and hardware fit, all served as keyless JSON on a daily publishing schedule.

For anyone choosing an AI model, the numbers have never been the hard part. The hard part is where they live. A new model’s reasoning score appears on the vendor’s launch blog, its coding result on a separate leaderboard, its crowd-voted rank on an arena site, its price on an API marketplace and its size, when it is disclosed at all, in a model card. Building a like-for-like view means opening a dozen tabs and reconciling them by hand.

[ULTRABENCH](https://ultrabench.ai/), a public model index created by Iternal Technologies, is built to end that exercise. Its first snapshot was published on Aug. 18, 2026, and five days later John Byron Hanby, IV, Iternal’s chief executive, introduced it in a post on X that opened with a blunt diagnosis: “AI benchmarks are a mess. So I fixed it.”

At launch the board tracked 199 current models. By the Sept. 26 snapshot it listed 286 current models and 405 model records in all, counting superseded versions, across 59 model families. Each row carries benchmark scores next to price, size and memory figures, and the entire ranked board can be downloaded in a single API call without a key.

> I am so frustrated AI benchmarks. AI benchmarks are a mess. So I fixed it. The data exists, scattered across vendor blogs, arena sites, and a dozen single benchmark leaderboards. Before today, there was no aggregate. No simple API. No answers. And nobody connected intelligence to size, price, and the hardware requirements. http://ULTRABENCH.ai does. 199 models and growing every day. Scores available for each model, with price per million tokens, exact parameters, and memory by quantization. Best value for the intelligence What fits on my hardware? Plus ratios for EVERYTHING. One API call. Clean JSON, no key. Updated daily. http://ULTRABENCH.ai. The whole model market at your fingertips. [pic.twitter.com/HYaaaASW9a](https://t.co/HYaaaASW9a) — John Byron Hanby, IV (@johnbyronhanby) [August 23, 2026](https://x.com/johnbyronhanby/status/2091598543085547577?ref_src=twsrc%5Etfw)

John Byron Hanby, IV, chief executive of Iternal Technologies, announced Ultrabench.ai on X on Aug. 23, 2026, with a 66-second video.

## What is Ultrabench.ai?

Ultrabench describes itself as a leaderboard on which “every current model [is] ranked by intelligence index, with price per 1M tokens, context, parameters and a source on every number.” Its documentation goes further, calling the project “a continuously refreshed, deduplicated, family-aggregated index of large language models” that combines exact parameter counts, benchmark scores with per-claim provenance, memory profiles for each quantization level and pricing sourced from OpenRouter.

The site has four views: the main leaderboard, ranked by intelligence index; an Intelligence vs Cost chart; a benchmarks page documenting a registry of 445 benchmarks, 17 of them treated as core; and documentation for the API, which serves the same data the pages display.

Ultrabench does not run the tests itself. It collects published results, normalizes them and records where each one came from. In the Sept. 26 leaderboard payload, 1,644 score cells were drawn from 42 source domains. The largest contributors were the evaluation firm Artificial Analysis, with 714 cells, followed by Vals.ai with 298, Hugging Face model cards with 144, arXiv papers with 110 and LMArena with 96. ARC Prize, LiveBench, Epoch AI and vendor pages supply the rest. It is, in effect, an aggregator of aggregators, independent leaderboards and primary sources, and that breadth underpins its pitch as one board for the whole model market.

## Why are AI benchmarks so hard to compare?

Hanby’s post named the problem directly: the data “exists, scattered across vendor blogs, arena sites, and a dozen single benchmark leaderboards.” Each of those sources answers one question well. GPQA Diamond measures graduate-level science reasoning. SWE-bench Verified measures whether a model can resolve real software issues. LMArena ranks models by head-to-head human preference votes. ARC-AGI-2 tests abstract pattern reasoning. None of them says what a model costs to run, how much memory it needs or whether it will fit on hardware a team already owns.

The sources also differ in ways that are hard to see from the outside. A vendor’s launch post may report a score produced with tools or extra compute; an independent lab may run the same test under different settings; and a number copied from one site to the next often loses its link to the original run. Ultrabench’s response is to treat provenance as part of the data rather than a footnote.

Hanby wrote: “Before today, there was no aggregate. No simple API. No answers.” Other aggregators do exist, among them Artificial Analysis, which is also Ultrabench’s largest single source. Ultrabench’s distinction is the combination: benchmark scores, price, parameter count, memory by quantization and hardware fit in one row, and the whole board in one API response.

## How does the Ultrabench intelligence index work?

The headline figure is a 0-to-100 intelligence index, a weighted mean over 10 core benchmarks. GPQA Diamond and SWE-bench Verified carry the most weight, at roughly 15 percent each. LiveBench and MMLU-Pro follow at about 11 percent, then Humanity’s Last Exam at about 10 percent. AIME, ARC-AGI-2 and the LMArena Elo rating carry about 7 percent each, and tau-bench and Terminal-Bench Hard about 4 percent each. When a model is missing some of those results, the index is renormalized over the benchmarks it does have.

Three rules keep thin evidence from producing a confident-looking number. A model gets no index unless at least three core benchmarks have admissible results, those results cover at least a quarter of the core weight, and every credited claim is cited. The share of weight actually measured is published beside every index as a coverage figure, so a reader can tell a well-documented score from a sparsely documented one.

Every score also carries a trust level. Independent leaderboards and refereed venues rank highest, at 3; aggregators are 2; vendor claims are 1; community reports are 0. A vendor grading its own model stays at trust 1 however the result is published, and tool-assisted runs are stored but not credited on Humanity’s Last Exam, GPQA, AIME, MMLU-Pro or ARC-AGI-2. When two trust-3 sources disagree by more than 15 percent, the system opens a dispute instead of quietly picking one. The documentation sums up the approach in three lines: “An empty cell beats an invented value.” “Every value carries a source.” And “null is never 0.”

## What does Ultrabench show for each model?

The distinguishing feature is what sits beside the score. Each row joins capability to the practical constraints a buyer or engineer has to live with. The table below uses Alibaba’s Qwen 3.8 Flash, an open-weights model released on Aug. 26, 2026, as it appeared in the Sept. 26 snapshot.

*What Ultrabench puts in one model row: Alibaba's Qwen 3.8 Flash in the Sept. 26, 2026 snapshot*

| Data point | Question it answers | Qwen 3.8 Flash |
| --- | --- | --- |
| Intelligence index (0–100) | How capable is the model across the weighted core benchmarks? | 77.22 |
| Coverage | How much of the core benchmark weight has actually been measured? | 0.39 (39 percent) |
| Price per 1M tokens | What does it cost through an API? | $0.15 input, $0.47 output, $0.23 blended |
| Parameters | How big is it, and are the weights open? | 180 billion, open weights |
| Memory by quantization | How much GPU memory do the weights need? | About 360 GB at BF16, 191 GB at INT8, 108 GB at INT4 |
| Hardware fit | Which reference machine can hold it, and at what precision? | NVIDIA DGX Spark at INT4 |
| VRAM efficiency | Is it leaner than models of similar intelligence? | 1.13 (peers would need about 215 GB at INT8) |
| Context window | How much text fits in one request? | 1,000,000 tokens |
| Source link | Where did each score come from, and how far can it be trusted? | GPQA Diamond 91.7, from the vendor's model card (trust level 1) |

Pricing comes from OpenRouter and covers 231 of the 286 current models. Ultrabench lists input and output prices per million tokens and a blended figure weighted three to one toward input, a common convention for estimating typical workloads. Exact parameter counts are measured from the weight files published on Hugging Face and exist for 156 current models. Closed models whose makers do not disclose a size show an empty field rather than a guess.

## What fits on my hardware?

The second question in Hanby’s post, “What fits on my hardware?”, is the one most leaderboards leave to the reader. Ultrabench estimates weight memory at three precisions: about 2.0 GB per billion parameters at BF16, 1.06 GB at INT8 and 0.6 GB at INT4. Those estimates cover the weights only, and the site advises adding 15 to 20 percent for runtime overhead plus the KV cache for long prompts. Where evidence exists, separate serving profiles for each quantization list the weights, the KV cache per 8,000 tokens of context and a minimum serving footprint.

A “Runs on” filter then matches models to three reference machines: Intel’s Arc Pro B70 with 32 GB, NVIDIA’s H100 SXM5 with 80 GB and NVIDIA’s DGX Spark with 128 GB. For each match the board names the precision at which the model fits. On Sept. 26, a query for models that fit on an H100 returned 118 results. The filter applies to open-weights models only, because closed models cannot be downloaded and self-hosted.

For teams planning private or on-premises deployments, that turns a spreadsheet exercise into a filter. It pairs naturally with guides to [running an LLM locally](https://aiintelreport.com/enterprise-ai/how-to-run-an-llm-locally), the [best local LLMs of 2026](https://aiintelreport.com/enterprise-ai/best-local-llms-2026) and the [total cost of on-premises AI](https://aiintelreport.com/enterprise-ai/on-premise-ai-cost-tco).

## How do the value ratios work?

“Best value for the intelligence,” in Hanby’s phrase, is a column rather than a slogan. Ultrabench computes three ratios:

- **Index per dollar** divides the intelligence index by the blended price per million tokens. The [Intelligence vs Cost](https://ultrabench.ai/cost) view ranks the best index per dollar across every ranked, priced model.
- **VRAM efficiency** compares the INT8 memory a model of that intelligence would be expected to need, based on a fit across the whole fleet, with the memory it actually needs. A value above 1 means the model is leaner than its peers; Qwen 3.8 Flash scored 1.13.
- **Coverage** reports how much of the core benchmark weight has been measured, a direct check on how much confidence an index deserves.

Together they answer questions no single-benchmark leaderboard can: which model delivers the most capability per dollar, which open model gives the most intelligence per gigabyte of GPU memory, and which high scores rest on thin evidence.

## How do you use the Ultrabench API?

Everything on the board is also available as JSON at `https://ultrabench.ai/v1`, and access is anonymous: no key, no account and no signup. Cross-origin requests are open to every domain, so a browser application can call it directly. The only limit is 120 requests a minute per address on requests that miss the edge cache. The [Ultrabench API documentation](https://ultrabench.ai/docs) lists 18 read-only routes, and the version 1 contract is additive only; a breaking change would ship as `/v2`.

One request returns the whole ranked board, about 631 KB of JSON:

```
curl https://ultrabench.ai/v1/leaderboard
```

A trimmed record from the Sept. 26 response shows how the pieces fit together:

```
{"data": {
  "snapshot_id": "snap_2026_09_26_001",
  "models": [{
    "id": "qwen-3.8-flash",
    "vendor": "Alibaba",
    "intelligence_index": 77.22,
    "coverage": 0.3889,
    "params_total_b": 180,
    "is_open_weights": true,
    "context_window_tok": 1000000,
    "pricing": {
      "prompt_usd_per_m": 0.15,
      "completion_usd_per_m": 0.47,
      "blended_usd_per_m": 0.23
    },
    "vram_est": {"bf16_gb": 360, "int8_gb": 191, "int4_gb": 108},
    "fits_on": {"nvidia_dgx_spark": "int4"}
  }]
}}
```

Other routes filter the fleet (for example, `/v1/models?fits_on=h100`), return one model with every benchmark claim and its source, list benchmarks, families and hardware profiles, and expose snapshot history. A `/v1/health` route reports how fresh the data is. The site also publishes an `llms.txt` file for AI agents, and its robots file states the attitude plainly: “Crawling is welcome; the data is the point.”

## How often is Ultrabench updated?

The index republishes on a daily schedule, usually around 06:10 UTC, and it published 46 snapshots between Aug. 18 and Sept. 26. Prices refresh daily, and the research that moves benchmark scores refreshes roughly every three days, according to the documentation. Every snapshot is kept, and the API shows what changed in each publish.

## Who is behind Ultrabench?

Ultrabench is built by [Iternal Technologies](https://iternal.ai/), the enterprise AI strategy company, which created and sponsors the project. Every page carries the credit “Created by Iternal Technologies, the AI Strategy Experts,” and the dataset’s structured data names Iternal as its creator. Hanby, Iternal’s chief executive, pitched it in his launch post as “the whole model market at your fingertips.”

For organizations that want to turn the numbers into a decision, Iternal also publishes an [LLM selection guide](https://iternal.ai/llm-selection-guide), which benchmarks and ranks more than 30 models side by side.

## What does Ultrabench mean for teams choosing AI models?

Model choice has become a moving target. New versions arrive constantly, prices move unevenly and each open-weights release adds options that can run in-house. A static comparison goes stale within weeks. What Ultrabench offers is a single, sourced, machine-readable baseline that refreshes itself.

That serves three audiences. Enterprise buyers get capability, price and evidence quality on one row, which shortens the path from a shortlist to a pilot. Teams running models on their own hardware get memory estimates and hardware fit before they download a single file. Developers get a free, keyless feed they can build into model routing, internal dashboards or procurement tools without negotiating access.

A leaderboard still measures only what its tests measure, and no index replaces a trial on a team’s own data. Ultrabench’s contribution is to make that limit visible, with coverage, trust levels and a source link on every number. The case for calling Ultrabench the definitive AI benchmark aggregator rests on breadth: 445 benchmarks in its registry, 1,644 sourced score cells from 42 source domains, and price, size and hardware fit on the same row. For a ranked shortlist by use case, see the [best LLMs of 2026](https://aiintelreport.com/frontier-models/best-llms-2026); for how much text each model can take in at once, see the explainer on the [LLM context window](https://aiintelreport.com/frontier-models/llm-context-window). For the whole market on one board, the [Ultrabench AI benchmark aggregator](https://ultrabench.ai/) is the place to start.

## Sources

1. [Launch post announcing Ultrabench.ai, with video (Aug. 23, 2026)](https://x.com/johnbyronhanby/status/2091598543085547577)
2. [Independent AI model evaluations; the largest single source of Ultrabench score cells](https://artificialanalysis.ai/)
3. [Crowd-voted head-to-head model rankings (Arena Elo)](https://lmarena.ai/)
4. [SWE-bench: resolving real-world software issues](https://www.swebench.com/)
5. [GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022)
6. [ARC-AGI-2 abstract reasoning benchmark](https://arcprize.org/arc-agi/2/)
7. [Model catalog and per-token pricing](https://openrouter.ai/models)

---
Source: https://aiintelreport.com/frontier-models/ultrabench-ai-benchmark-aggregator
Index: https://aiintelreport.com/llms.txt · Full text: https://aiintelreport.com/llms-full.txt
