Thursday, July 23, 2026

Today’s Edition

AI Intel Report

MARKETS

Frontier Models

Celeris-1 Matches Near-GPT-5 Intelligence at 15x Faster Response Times

Celeris Labs, founded by former Marqo engineers, releases a diffusion-based model that reaches 75.9 percent MMLU-Pro accuracy with 158 ms p50 latency and immediate OpenAI-compatible API access.

4 MIN READ
In a sleek modern technology research laboratory filled with rows of high-performance server racks and computational hardware arrays, several anonymous engineers wearing casual professional attire work collaboratively at multiple workstations equipped with large flat-panel displays and dense cabling systems connected to specialized GPU clusters and networking equipment, one engineer seated at a central desk examines detailed performance readouts on a monitor while another stands nearby adjusting connections on an open server chassis revealing internal components including cooling fans and circuit boards, the environment features clean white walls with subtle blue accent lighting, scattered technical notebooks and diagnostic tools on countertops, a large window in the background overlooking an urban cityscape at dusk with soft natural light mixing with interior illumination, the scene captures the essence of rapid AI model development and deployment through visible hardware infrastructure representing advanced diffusion-based systems achieving high accuracy benchmarks comparable to leading frontier models yet optimized for substantially reduced inference latency, elements include multiple interconnected computers running parallel processes with indicator lights blinking steadily to signify active computation, anonymized figures focused intently on their tasks without any visible personal identifiers, the overall composition emphasizes the integration of legacy engineering expertise from prior organizations into new ventures focused on delivering immediate API-compatible access for developers seeking efficient high-intelligence solutions, detailed textures on metallic server surfaces, rubberized cable insulation, matte-finished monitor bezels, ergonomic office chairs, and polished flooring reflecting overhead fixtures create a realistic photojournalistic atmosphere grounded in the daily operations of a startup laboratory advancing machine learning capabilities beyond previous standards in speed and precision while maintaining strong benchmark results on complex knowledge evaluation tasks.
Illustration: AI Intel Report

Celeris-1 is a general-purpose language model delivering near-GPT-5 level intelligence with 15x faster response times via a new diffusion-based inference architecture.

Celeris Labs has made Celeris-1 available starting today through an OpenAI-compatible API under the model name celeris-1. The release follows several months of development by the research team at the lab focused on frontier intelligence with reduced inference times.

What background led to the founding of Celeris Labs?

Celeris Labs was established by Tom Hamer and Jesse Clark after their prior roles at Marqo. The organization operates from celeris.ai with a stated mission to create the world's fastest language models. A team of engineers and researchers contributed to the project in the months leading up to the public announcement.

The lab positions its work around a future in which frontier intelligence responds in microseconds. This focus stems from the founders' experience in AI infrastructure and model development at their previous company.

What performance results has Celeris-1 demonstrated on benchmarks?

Celeris-1 was evaluated on the full MMLU-Pro test set using 5-shot chain-of-thought prompting and strict scoring protocols. The model recorded 75.9 percent accuracy on this benchmark. GPT-5 variants achieved scores in the 78 to 81 percent range under comparable conditions according to the published comparisons.

Additional metrics include a p50 response time of 158 milliseconds and throughput of 1,664 tokens per second. These figures were reported alongside tests against GPT-5 and GPT-5 mini variants to place speed and quality on equal footing.

Benchmark comparison of Celeris-1 against GPT-5 variants on the MMLU-Pro test set.
ModelMMLU-Pro Accuracyp50 LatencyTokens per Second
Celeris-175.9%158 ms1,664
GPT-5 variants78-81%Not reportedNot reported

How does the diffusion architecture support the reported speed improvements?

Celeris-1 relies on a diffusion architecture that performs non-autoregressive token generation. This method enables parallel processing during inference rather than sequential token prediction used in many conventional language models. The design contributes directly to the observed reductions in response latency.

Documentation from the lab describes the approach as a new architecture for language model inference built on diffusion principles. This structure supports the dual goals of maintaining high accuracy while delivering the measured throughput and latency values.

Who are the individuals responsible for Celeris Labs and Celeris-1?

Tom Hamer serves as co-founder and has publicly detailed the launch of the lab. Jesse Clark is the other co-founder, with both individuals drawing on prior experience at Marqo to advance the diffusion-based models. Their combined background in AI research underpins the technical decisions behind the architecture.

What market implications follow from the Celeris-1 release?

The OpenAI-compatible API format allows straightforward integration for existing applications that already support similar endpoints. Developers gain access to a model that balances benchmark performance with substantially lower response times compared to autoregressive alternatives.

Enterprise users and real-time application builders may evaluate the model for workloads where latency directly affects user experience. The availability through sign-up provides an entry point for testing these characteristics against current production systems.

What statements have been made by the founders regarding the launch?

Big announcement - today we're launching Celeris Labs, an AI research lab building the world's fastest LLM.Tom Hamer, Co-founder

The announcement noted that a team of engineers and researchers worked on the project for the preceding months. It framed the lab as dedicated to building the world's fastest LLM through the new inference methods.

What steps are required to begin using Celeris-1?

  1. Complete the sign-up process on the celeris.ai platform to obtain access credentials.
  2. Configure client code to call the OpenAI-compatible endpoint while specifying the model name celeris-1.
  3. Review the full benchmark methodology and results published by Celeris Labs for validation.
  4. Monitor the lab's updates for any subsequent model iterations or expanded feature availability.

The model became accessible on the day of the announcement for qualified users who complete the registration.

Frequently asked

How can developers integrate Celeris-1 into existing applications?

Developers access Celeris-1 through the OpenAI-compatible API by using the model identifier celeris-1 after completing sign-up at celeris.ai. The endpoint supports standard request formats already familiar to many integration teams.

Sources

  1. Celeris Labs — Celeris-1 is designed to be fast and accurate at the same time... We ran it against GPT-5, GPT-5 mini... on MMLU-Pro...
  2. LinkedIn — Big announcement - today we're launching Celeris Labs, an AI research lab building the world's fastest LLM.
  3. Celeris Labs — A new architecture for language model inference... built on diffusion... model="celeris-1"
  4. Celeris Labs — Celeris is an artificial intelligence research lab building the world's fastest language models. We are creating a future where frontier intelligence responds in microseconds.