Google paper reveals AI agents can rationally cooperate through similarity inference

4 hours ago 29

For decades, game theory has had a pretty firm answer to the question of whether two strangers should cooperate when they’ll never meet again: absolutely not. The rational play, according to Nash equilibrium, is to defect. Every time. No exceptions.

A new research paper from Google DeepMind, Mila-Quebec AI Institute, and ETH Zürich argues that AI agents built on foundation models don’t play by those rules. And they have the math to back it up.

The end of inevitable defection

The paper, submitted to arXiv on August 4, 2026, spans 75 pages with 11 figures and introduces a concept the authors call “similarity inference.” The core idea is deceptively simple: AI agents built from similar training processes can recognize that similarity and use it to predict what the other agent will do.

Classical game theory treats each player as a fully independent decision-maker, what the paper calls “decoupled agency.” Under that assumption, your opponent’s choice is a black box. You can’t predict it, so you hedge by defecting. The Nash equilibrium holds, and everyone ends up worse off than if they’d cooperated.

The research team, led by Alexander Meulemans, proposes an alternative framework called “embedded agency.” In this model, agents perceive themselves as part of their environment rather than separate from it. More importantly, they maintain genuine uncertainty about their own decision-making processes.

That uncertainty is the key ingredient. Because an agent isn’t entirely sure how it will decide, it can treat its own reasoning as evidence about how a similar agent would reason. If I’m leaning toward cooperation, and I know my counterpart thinks like me, then my counterpart is probably leaning toward cooperation too. The logic bootstraps itself into a stable cooperative outcome.

Embedded equilibrium replaces Nash

To formalize this, the paper introduces a new solution concept called “embedded equilibrium.” Where Nash equilibrium assumes players optimize independently, embedded equilibrium accounts for the correlations between agents that arise from shared architecture, training data, and optimization procedures.

Under Nash, cooperation in a one-shot Prisoner’s Dilemma is irrational. Under embedded equilibrium, it can be the uniquely rational strategy.

The authors’ empirical findings show that foundation-model agents placed in social dilemmas and given optimized planning tools actually do converge toward cooperative outcomes. The math predicts cooperation, and the experiments confirm it.

Why this matters beyond the lab

The paper also raises questions about what happens when agents are not similar. If cooperation depends on similarity inference, then interactions between agents built on different architectures or trained on different data could default back to classical defection dynamics.

For policymakers and AI governance bodies, the embedded equilibrium concept introduces a variable that existing regulatory frameworks don’t account for. Antitrust law, for instance, generally requires evidence of explicit communication to establish collusion. Agents that cooperate through similarity inference communicate nothing. They just think alike.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article