Multi-agent cooperation through in-context co-player inference

Paper
2026-03-02

Description

This paper explores how cooperation can emerge among self-interested agents in multi-agent reinforcement learning without hardcoded assumptions about how other agents learn; instead of using specialized meta-learning or explicit timescale separations, the authors show that training sequence-model agents against a diverse mix of co-players naturally leads them to infer and adapt to their partners’ strategies within an episode, producing in-context best-response behaviors that, under mutual pressure to shape each other’s learning dynamics, lead to cooperative outcomes, suggesting a scalable way to achieve cooperation using standard decentralized RL with sequence models.
PDF Preview

User Reviews

No reviews yet for this resource.