Will general purpose AI models beat average score of human players in Diplomacy by 2028?
💡 What the odds say
The market puts this at about a 56% chance — more likely than not.
No money — just record your call and see if you were right. Yes is at 56% right now.
The market is pricing a slight edge for general-purpose AI mastering Diplomacy by 2028, driven by recent evidence that large language models can outperform specialized tools in complex domains, but the lack of a direct Diplomacy benchmark and the game's unique negotiation demands keep the odds from being higher.
📊 Base rate: Historical AI milestones in complex strategy games (e.g., AlphaGo in Go, OpenAI Five in Dota 2) have typically been achieved by specialized systems, not general-purpose models, with success rates for general models in such games below 30%.
What's driving it
- • A Nature study (Jun 12) showing general-purpose LLMs outperforming specialized clinical AI tools suggests these models can excel in tasks requiring nuanced reasoning, boosting confidence for Diplomacy.
- • The Kairos Model's breakthrough in open-source embodied intelligence (Jun 22) indicates rapid progress in general-purpose AI capabilities, indirectly supporting the 'Yes' case.
- • A critical XDA article (Jun 10) warns that local LLMs are useless without tool integration, highlighting a key weakness that could hinder general-purpose models in Diplomacy's multi-agent environment.
- • No clear catalyst recently: the golf putter article (Jul 10) and Snowflake earnings (May 27) are irrelevant to AI game performance.
The case for YES
- • General-purpose LLMs have demonstrated superior performance in complex, multi-step reasoning tasks like medical diagnosis (Nature, Jun 12), which parallels Diplomacy's strategic depth.
- • The Kairos Model (Jun 22) shows that open-source general-purpose AI is advancing rapidly, potentially enabling broader access to models that can be fine-tuned for Diplomacy.
- • Diplomacy's reliance on language and negotiation plays to the strengths of LLMs, which have already shown proficiency in dialogue and persuasion in other contexts.
The case for NO
- • Diplomacy requires real-time multi-agent coordination and tool use, which the XDA article (Jun 10) identifies as a current weakness for local LLMs, and likely for general-purpose models.
- • Specialized AI systems have historically dominated complex games (e.g., AlphaGo, Pluribus), and general-purpose models have not yet matched that level of performance in any comparable game.
- • The Nature study (Jun 12) focused on static benchmarks, not interactive environments like Diplomacy, where deception and long-term alliance dynamics pose unique challenges.
What to watch
- • A direct benchmark or competition where a general-purpose AI plays Diplomacy against humans (e.g., announced by Meta or DeepMind) would likely shift odds toward 'Yes'.
- • Release of a new general-purpose model with demonstrated tool-use and multi-agent capabilities (e.g., GPT-5 or equivalent) could increase 'Yes' probability.
- • A study showing general-purpose AI failing at Diplomacy or similar negotiation games would push odds toward 'No'.
AI-generated · grounded in recent news + odds · informational only, not advice. Verify on the source platform.
Data from Manifold’s public API, for informational purposes only. PredictPal is not affiliated with any platform and does not facilitate trading.
Discussion
Loading…
How it resolves
Resolved by whoever created the market, at their discretion per the question's description. It's play-money (Mana) and not tied to an official source — treat it as a community forecast.
ⓘ A market settles under its own written rules, which can lag what looks decided in the news — so the price may not move to 100% the moment an outcome seems obvious.
View the official rules on Manifold ↗Related markets
When will Starship flight 13 happen?
The field is heavily skewed toward launch within 4 months but with near-zero chance in the next two weeks, reflecting the grounding after the May 28 booster failure (USA Today) — the key unknown is how long the investigation takes.
30 outcomes
Who will win the 2026 Fields Medals?
The race is effectively over: all four winners have been publicly announced on July 23, 2026, collapsing the field to the four named recipients and making any further movement purely a matter of market resolution mechanics.
52 outcomes
Perfect score achieved by an AI model in the International Math Olympiad (IMO) 2026?
Multiple news outlets report that AI models have achieved a perfect score at the 2026 IMO, making the 97% odds a near-certainty that the event has already occurred, with the market awaiting final formal confirmation.
Yes ≈ 97% chance
Apple Announces AI Glasses by September 30, 2026
The field is moderately concentrated with the 'No' side leading at 62%, but the 45% sum for two candidates indicates significant overlap or mispricing; the biggest recent shift is Meta's June 23 launch of $299 smart glasses (Forbes, Jun 23), which likely boosted the 'No' side by making Apple's entry seem less urgent or unique.
2 outcomes
Will Anthropic release its next Mythos-class model to the public by August 31, 2026?
Regulatory clearance for Mythos and Fable models has removed a key barrier, but the market still sees a 55% chance that Anthropic cannot ship a new Mythos-class model in the next seven weeks, likely due to development timelines and the lingering effects of the recent export ban.
Yes ≈ 46% chance
Companies to go public in 2026
The field is extremely concentrated on SpaceX and Anthropic, whose combined odds (161%) imply the market expects at least one of them to go public in 2026, but the lack of recent headlines about either company's IPO process leaves the high probabilities unexplained by concrete events.
8 outcomes