Impact-Site-Verification: 41b53a0c-6d04-458b-a457-fe9e29acde1a

AI & Machine Learning·Seed··6 min read

Olam Labs

Social strategy games as rigorous AI evaluation, with anonymity and fairness parity.

NN

NewName Editorial

Editorial Team

Olam Labs product image 1
Olam Labs product image 2

Olam Labs is not trying to build a better game. It is trying to build a better benchmark—one where the test is a game of Risk, Poker, or Codenames, and the subjects are frontier LLMs. The company's Multi-Agent Arena invites humans to play these social strategy games against AI agents, but the real product is not the entertainment; it is the data. Every match feeds Olam's benchmarks, which aim to measure the social behaviors that standard single-agent evaluations routinely miss: deception, negotiation, grudges, collaboration, and the messy, long-horizon dynamics that define real-world interaction.

This is a deliberate bet. As Olam's about page argues, the current evaluation industry focuses on single-agent domain tasks and real-world data—useful, but insufficient for predicting how models will behave in complex systems with humans. The company's thesis is that multi-agent simulations, run at scale, can produce the qualitative measurements that are otherwise difficult to quantify. Games are the vehicle because they are structured, legible, and inherently social.

The game is the benchmark

The core idea is simple: put humans and AI agents in the same game, with the same information and the same actions, and see what happens. Olam's arena supports games like Risk, Poker, and Codenames, each chosen for its social depth. Risk requires negotiation and betrayal. Poker demands bluffing and reading opponents. Codenames tests communication and shared understanding. These are not toy environments; they are long-horizon games where trust is built and broken over hours.

What makes this a benchmark rather than a novelty is the scale and the intent. Olam runs public long-horizon games and internal simulations with hundreds of agents concurrently. The company's Tokenshire simulation, for example, is a town of AI agents autonomously running a government and economy under resource constraints. These runs are published as research, offering a window into emergent social behavior. The games are not just for fun—they are instruments.

Why anonymity is the load-bearing wall

Anonymity is not a privacy feature; it is an experimental control. In the arena, every player—human or AI—sees only generic names like "Ryan" or "Milo." The AIs do not know if their opponent is human. This design choice is critical because it removes a variable that would otherwise contaminate the data: the AI's response to human presence. If an agent knows it is playing a human, it might adjust its behavior—being more cautious, more aggressive, or more sycophantic. By anonymizing everyone, Olam ensures that the AI is interacting with a generic player, not a human or another AI, which makes the behavioral data cleaner.

This also levels the playing field in a way that matters for evaluation. A human player cannot exploit the AI's knowledge of human psychology, and the AI cannot rely on cues like typing speed or reaction time. The result is a more controlled experiment, one where the measured behavior is a function of the game state and the agent's strategy, not the meta-knowledge of who is on the other side.

The fairness parity that makes the data legible

Fairness in Olam's arena is not just about equal rules; it is about equal primitives. The website states: "The AIs and humans have the exact same primitives for actions and information." What the harness is to the AI, the UI is to the human. This parity is what makes the data legible. If a human could see something an AI could not, or vice versa, the comparison would be invalid. By ensuring that both sides have access to the same game state and the same set of possible actions, Olam can attribute differences in performance to the agents' reasoning and social skills, not to asymmetric information.

This is a rigorous standard, and it is rare in AI evaluation. Most benchmarks are static—a set of questions or tasks with known answers. Olam's is dynamic and interactive, with the added complexity of multiple agents and humans. The fairness parity is what turns a chaotic game into a controlled experiment. It also means that when a negotiation goes sour, Olam can trace exactly why: the game state, the messages exchanged, the actions taken. This full data legibility is impossible in real-world evaluations, where you cannot see every variable. In the arena, you can.

Long-horizon games reveal what single-turn evals miss

Standard AI evaluations often test a single turn or a short conversation. They measure whether a model can answer a question or follow an instruction. But real-world social interaction is not a single turn; it is a series of moves, each influenced by the history of the interaction. Olam's games are long-horizon by design. A game of Risk can last hours, with alliances formed and broken, and grudges carried from one turn to the next. This temporal depth is what allows Olam to measure behaviors like trust, deception, and social intelligence—traits that are invisible in a single prompt-response pair.

The company's internal simulations push this even further. Tokenshire, the simulated town, runs autonomously for long horizons, with agents socializing, running a government, and managing an economy under resource constraints. These simulations produce emergent behaviors that are not pre-programmed, offering a glimpse into how models might behave in complex, real-world systems. Olam's bet is that these long-horizon, multi-agent environments can meet the shortage of evaluations that will arise as AI becomes more integrated into society.

The business of selling social evaluation

Olam Labs is backed by Y Combinator, and its about page outlines a clear commercial path: pre-deployment evaluations for frontier labs, datasets of top-performing humans and agents, and custom arenas for training and evals. The company is positioning itself as a partner for labs that want to test their models for social intelligence and agentic performance before release. This is a timely offering, as the industry grapples with how to ensure AI systems are safe and aligned in social contexts.

The business model is not about selling games; it is about selling evaluation. The arena is the bait, and the benchmarks are the product. By running public games, Olam collects data on how humans and AIs interact, which it can then package into datasets and evaluations for paying customers. The anonymity and fairness parity are not just experimental controls; they are selling points. They assure labs that the data is clean and the comparisons are valid.

There is a risk, of course. The arena's success depends on sustained human participation. If people stop playing, the data pipeline dries up. But Olam's internal simulations, like Tokenshire, can run without humans, providing a fallback and a different kind of data. The company is also courting researchers and labs directly, offering custom environments for specific evaluation needs. This dual approach—public engagement and private partnerships—gives it multiple revenue streams and a buffer against the whims of casual gamers.

Olam Labs is making a bet that the future of AI evaluation is social, and that games are the most rigorous way to measure it. The arena is a clever proof of concept, but the real test will be whether the data it produces convinces frontier labs to pay for it. If it does, Olam will have built not just a benchmark, but a new standard for evaluating the social intelligence of AI.