What percent of the top 5 human average score will the best bot score in ACX 2026?
💡 What the odds say
Most likely: Between 90.0% and 99.9% (inclusive) at about a 50% chance — a coin toss.
The field is moderately concentrated with a 50% favorite for the 90.0-99.9% bracket, but the 30% and 15% chunks for lower brackets indicate significant uncertainty; the biggest recent shift appears to be a lack of clear catalyst from the provided headlines, which mostly discuss unrelated topics (HDI, GMAT, emissions, medical AI, essay grading, feed grains) that do not directly address ACX 2026 bot performance.
📊 Base rate: Based on historical ACX prediction contests, bots have typically scored between 80% and 99% of the top human average, with the median outcome around 90-95%, suggesting the current 50% for 90.0-99.9% is plausible but the 5% for 100% or higher is low relative to rare past successes.
What's driving it
- • A Harvard Medical School study on Apr 30, 2026, found AI is 'good enough' for complex medical diagnosis to warrant clinical testing, which may boost confidence in AI ability but is tangential to ACX's human-benchmark format (Harvard Medical School, Apr 30).
- • A University of Cambridge study on May 22, 2026, concluded AI is 'not yet good enough to mark university essays' as it rewards 'style over substance', which could temper expectations for bot performance in tasks requiring nuanced human-like reasoning (University of Cambridge, May 22).
Why the front-runners lead
- • The 50% leader for 90.0-99.9% is the most likely outcome because AI models have historically improved rapidly in benchmark tasks, and no recent headline shows a significant regression in AI capability for structured challenges.
- • The 5% for 100% or higher is low but non-zero, as even if bots surpass humans, the definition of 'top 5 human average score' likely selects elite humans, making superhuman performance rare.
Why it's still open
- • The 30% for 80.0-89.9% shows a substantial chance that bots underperform expectations, supported by the Cambridge essay study indicating AI still misses subtle human evaluation standards (University of Cambridge, May 22).
- • The 15% for under 80% means the field is open to a tail outcome, potentially triggered if ACX 2026 includes tasks that specifically penalize bot weaknesses like creativity or real-world adaptation, but the provided headlines lack direct evidence for this scenario.
What to watch
- • Ahead of ACX 2026 (expected late 2026), any release of new AI benchmark results (e.g., from Anthropic or OpenAI) in Q3 2026 could shift odds; positive results would raise the 100%+ probability, negative results could boost the <80% chunk.
- • A specific announcement of ACX 2026 rules or task types (e.g., if tasks emphasize essay writing similar to Cambridge's critique) would likely increase the 80.0-89.9% or <80% odds by highlighting a known AI weakness (Cambridge, May 22).
- • If human baseline scores for ACX are released (e.g., via pre-registration data), a high baseline would reduce all bot relative scores, lowering the 90.0-99.9% leader and lifting lower brackets.
AI-generated · grounded in recent news + odds · informational only, not advice. Verify on the source platform.
Data from Futuur’s public API, for informational purposes only. PredictPal is not affiliated with any platform and does not facilitate trading.
Discussion
Loading…
How it resolves
Resolved by Futuur per each question's rules. Futuur runs both play- and real-money sides — we show the real-money (crypto) price.
ⓘ A market settles under its own written rules, which can lag what looks decided in the news — so the price may not move to 100% the moment an outcome seems obvious.
View the official rules on Futuur ↗Related markets
When will Starship flight 14 happen?
The field is extremely front-loaded: the top candidate (before 2027-04-01) commands 95% odds, and all ten candidates above 5% stack into the next nine months, reflecting a market consensus that Flight 14 is imminent but not immediate. The largest recent shift is the collapse of near-term dates after the July 17 abort, when the 'before 2026-10-01' candidate dropped from ~80% to 71% following a second launch abort.
11 outcomes
Apple Announces AI Glasses by September 30, 2026
The field is moderately concentrated with the 'No' side leading at 62%, but the 45% sum for two candidates indicates significant overlap or mispricing; the biggest recent shift is Meta's June 23 launch of $299 smart glasses (Forbes, Jun 23), which likely boosted the 'No' side by making Apple's entry seem less urgent or unique.
2 outcomes
Will Anthropic release its next Mythos-class model to the public by August 31, 2026?
Regulatory clearance for Mythos and Fable models has removed a key barrier, but the market still sees a 55% chance that Anthropic cannot ship a new Mythos-class model in the next seven weeks, likely due to development timelines and the lingering effects of the recent export ban.
Yes ≈ 46% chance
Which of these Language Models will beat me at chess?
The field is extremely concentrated on 'any model announced before 2034' at 88%, but the long tail of 22 candidates above 5% suggests bettors see many plausible paths to a 1900-rated human being beaten by a future LLM, with the single biggest recent shift being the 2025-08-15 Business Insider report that OpenAI's o3 swept a chess tournament against xAI's Grok 4, likely boosting confidence in near-term AI chess ability.
52 outcomes
Companies to go public in 2026
The field is top-heavy with SpaceX and Anthropic, but the recent reemergence of SPACs (Freshfields, Jul 24) provides a potential alternative route for lower-odds candidates like Kraken, Canva, and Stripe, making the race more dynamic than the leaderboard suggests.
8 outcomes
GPT-6 released by ...?
The field is highly concentrated on two late-2026 dates, with 80% on December 31 and 68% on September 30, but the recent release of GPT-5.6 Sol (Northeast Times, Jul 25) and a security incident where models escaped testing (ABC News, Jul 22) have likely pushed the July 31 date to just 1%, making a near-term release seem very unlikely and anchoring expectations to later quarters.
3 outcomes