Arena, the best-known open leaderboard platform for AI models, has closed a $200 million Series B round. The deal was announced on October 8, 2026, valuing the company at $3.1 billion — nearly double the $1.7 billion valuation from its January Series A. In other words, investors valued the company at almost twice the price in under ten months.
Arena’s history began in academia: the company started in 2023 as a research project at UC Berkeley. Within three years the research initiative became a commercial company and today is one of the most recognized names in AI model evaluation. Details of the deal were reported by TechCrunch.
Nine months passed between January’s $150M Series A and October’s $200M Series B. In that time the company not only raised $50M more, but also lifted its valuation from $1.7B to $3.1B — adding nearly $140M in value each month on average.
Deal Details
The round was led by two major venture firms — Lightspeed Venture Partners and Khosla Ventures. They were joined by Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, Andreessen Horowitz, Felicis, and other investors. The breadth of the participant list — spanning venture funds and corporate investors — signals strong interest in the deal.
In January 2026 the company raised $150M in its Series A at a $1.7B post-money valuation. A post-money valuation reflects the company’s value after the investment. The $3.1B valuation in the Series B is nearly double that figure: over the ten-month span the valuation grew almost twofold.
Such rapid valuation growth ranks among the sharpest in the AI infrastructure sector. Companies typically take more than a year between Series A and Series B; Arena moved to the next round in under ten months while nearly doubling its valuation.
Notably, the funds in the deal are among the most active investors in AI. The presence of names like Andreessen Horowitz and Felicis in a round led by Lightspeed Venture Partners and Khosla Ventures adds further weight to the deal.
Revenue Figures
The near-doubling of the valuation coincided with a sharp rise in revenue. In June, Arena said its annualized revenue run rate had reached $100M. The run rate is current monthly revenue multiplied by twelve, reflecting the company’s annualized revenue potential.
At the time of the Series A, the figure was $30M. In other words, the company more than tripled its revenue in six months. The journey from $30M at the start of the year to $100M by midsummer shows a project that began as an open leaderboard platform becoming a significant commercial revenue source.
Big investment deals in AI keep coming: DeepSeek reportedly raised $1.2B recently. The Arena case shows that model evaluation and ranking infrastructure is drawing strong investor interest as a standalone business as well.
This rapid revenue growth shows the company succeeding commercially beyond its open leaderboard. The jump from $30M to $100M in six months means, on average, more than $10M of additional annualized run rate each month.
New “alignment” Leaderboard Category
Alongside the investment, Arena added a new “alignment” leaderboard category. It evaluates models on three criteria: unauthorized actions, incorrect attribution, and deceptive completion.
"AI is advancing faster than our ability to evaluate it, and static benchmarks break down once models recognize they're being tested," the company said in its funding announcement. "The world needs a neutral third party to measure how safe and aligned AI actually is once it's in the hands of real people. Arena is stepping into that role today," it added.
The first criterion — unauthorized actions — measures cases where a model steps outside the user’s task and takes independent action without asking permission. The second — incorrect attribution — covers cases where a model misstates the source of its answer or passes off someone else’s work as its own. The third — deceptive completion — assesses cases where a model presents unfinished work as finished, misleadingly framing the result.
Early leaderboard results show OpenAI models in the lead. Claude Opus 5.5 ranked sixth, and Claude Fable ninth. The new category lets models be compared not only on answer quality but on behavioral reliability.
The move signals broadening evaluation criteria: where traditional leaderboards focused on the quality of model responses, the new category measures the reliability of model behavior.
Notably, early results revealed a significant gap between labs’ models: while OpenAI models took the top spots, Anthropic’s Claude Opus 5.5 ranked sixth and Claude Fable ninth. The gap points to intensifying competition on reliability criteria as well.



