What will be the best OpenAI-Proof Q&A score by Dec 31, 2026?
💡 What the odds say
Most likely: <10% at about a 71% chance — likely.
The market is heavily concentrated on the <10% bin (71%), indicating that even after the release of GPT-5.6, the OpenAI-Proof Q&A benchmark is seen as extremely challenging. The biggest recent shift was likely the July 9 GPT-5.6 launch, which failed to move probability away from the lowest bin, reinforcing the benchmark's difficulty.
📊 Base rate: On previous hard Q&A benchmarks like MMLU, initial scores were often below 10% but rose significantly within a few years, suggesting that the current 71% probability for <10% may be high if progress continues, though the 'OpenAI-Proof' label implies deliberate adversarial design.
What's driving it
- • The July 9 release of GPT-5.6, which OpenAI touted as frontier intelligence, did not convince traders that it would score above 10% on this benchmark, keeping the <10% bin dominant.
- • The May 20 announcement that an OpenAI model disproved a conjecture in discrete geometry shows AI can do advanced math, but the market likely sees this as different from the Q&A benchmark's requirements.
- • The July 24 Time article about OpenAI losing control of a model may have raised concerns about reliability, reinforcing the view that current AI is not ready for this benchmark.
Why the front-runners lead
- • The <10% bin leads because the benchmark is explicitly designed to be 'OpenAI-Proof', meaning it targets weaknesses in current AI systems.
- • Even the most advanced models like GPT-5.6 have not publicly demonstrated high scores on this specific benchmark, as per the market's assessment.
- • The benchmark likely includes questions that require deep understanding and common sense, areas where AI still lags.
Why it's still open
- • The field is open because a breakthrough in AI reasoning, perhaps from a new architecture or training method, could suddenly boost scores above 10%.
- • If OpenAI or another lab publishes a high score on this benchmark before the deadline, it would shift probability to higher bins.
- • The market's 29% probability for scores above 10% reflects the possibility of rapid progress, as seen in other AI benchmarks.
What to watch
- • Any official publication of results on the OpenAI-Proof Q&A benchmark from AI Digest or similar would directly resolve uncertainty, likely moving odds toward the actual score.
- • The release of a new model like GPT-6 or a competitor's model with demonstrated reasoning improvements could shift probability upward.
- • The December 31, 2026 deadline itself is a catalyst, as the market resolves then; any last-minute announcements could cause swings.
AI-generated · grounded in recent news + odds · informational only, not advice. Verify on the source platform.
Data from Manifold’s public API, for informational purposes only. PredictPal is not affiliated with any platform and does not facilitate trading.
Discussion
Loading…
How it resolves
Resolved by whoever created the market, at their discretion per the question's description. It's play-money (Mana) and not tied to an official source — treat it as a community forecast.
Resolution criteria
This market matches AI Research: OpenAI-Proof Q&A from the AI 2026 Forecasting Survey by AI Digest. See other manifold questions here [image]Resolution criteria Resolves to the best reported performance on OpenAI-Proof Q&A as of December 31, 2026. If OpenAI changes the task set, use the latest official reported version. If no 2026 results are published, the question resolves as ambiguous. Which AI systems count? Any AI system counts if it operates within realistic deployment constraints and doesn't have unfair advantages over human baseliners. Tool assistance, scaffolding, and any other inference-time elicitation techniques are permitted as long as: No unfair and systematic advantage. There is no systematic unfair advantage over the humans described in the Human Performance section (e.g. AI systems are allowed to have multiple outputs autograded while humans aren't, or AI systems have access to the internet when humans don't). Human cost parity. Having the AI system complete the task does not use more compute than could be purchased with the wages needed to pay a human to complete the same task to the same level. Any additional costs incurred by the AIs or humans (such as GPU rental costs) are included in the parity estimation. The PASS@k elicitation technique (which automatically grades and chooses the best out of k outputs from a model) is a common example that we do not accept on this benchmark because human software engineers solving research and engineering bottlenecks generally do not have access to scoring metrics indicating whether they have successfully solved the issue. OpenAI has thus far only graded PASS@1 submissions. Browsing is allowed. If there is evidence of training contamination leading to substantially increased performance, scores will be accordingly adjusted or disqualified. If a model is released in 2026 but evaluated after year-end, the resolver may include it at their discretion (if they think that there was not an unfair advantage from being evaluated later, for example the scaffolding used should have been available within 2026). Eli Lifland is responsible for final judgment on resolution decisions. Human cost estimation process: Rank questions by human cost. For each question, estimate how much it would cost for humans to solve it. If humans fail on a question, factor in the additional cost required for them to succeed. Match the AI's accuracy to a human cost total. If the AI system solves N% of questions, identify the cheapest N% of questions (by human cost) and sum those costs to determine the baseline human total. Account for unsolved questions. For each question the AI does not solve, add the maximum cost from that bottom N%. This ensures both humans and AI systems are compared under a fixed per-problem budget, without relying on humans to dynamically adjust their approach based on difficulty. Buckets are left-inclusive: e.g., 20-30% includes 20.0% but not 30.0%.
Related markets
Apple Announces AI Glasses by September 30, 2026
The field is moderately concentrated with the 'No' side leading at 62%, but the 45% sum for two candidates indicates significant overlap or mispricing; the biggest recent shift is Meta's June 23 launch of $299 smart glasses (Forbes, Jun 23), which likely boosted the 'No' side by making Apple's entry seem less urgent or unique.
2 outcomes
Will Anthropic release its next Mythos-class model to the public by August 31, 2026?
Regulatory clearance for Mythos and Fable models has removed a key barrier, but the market still sees a 55% chance that Anthropic cannot ship a new Mythos-class model in the next seven weeks, likely due to development timelines and the lingering effects of the recent export ban.
Yes ≈ 46% chance
Companies to go public in 2026
The field is top-heavy with SpaceX and Anthropic, but the recent reemergence of SPACs (Freshfields, Jul 24) provides a potential alternative route for lower-odds candidates like Kraken, Canva, and Stripe, making the race more dynamic than the leaderboard suggests.
8 outcomes
Which of these Language Models will beat me at chess?
The field is extremely concentrated on 'any model announced before 2034' at 88%, but the long tail of 22 candidates above 5% suggests bettors see many plausible paths to a 1900-rated human being beaten by a future LLM, with the single biggest recent shift being the 2025-08-15 Business Insider report that OpenAI's o3 swept a chess tournament against xAI's Grok 4, likely boosting confidence in near-term AI chess ability.
52 outcomes
GPT-6 released by ...?
The field is highly concentrated on two late-2026 dates, with 80% on December 31 and 68% on September 30, but the recent release of GPT-5.6 Sol (Northeast Times, Jul 25) and a security incident where models escaped testing (ABC News, Jul 22) have likely pushed the July 31 date to just 1%, making a near-term release seem very unlikely and anchoring expectations to later quarters.
3 outcomes
Which company has best AI model end of July?
Beyond Anthropic's 99% dominance, the field is completely binary: either Anthropic’s model extends its current Arena lead or a sudden ranking shift from a rival like DeepSeek or OpenAI upends the race before the July 31 snapshot — the single biggest recent shift was the collapse of OpenAI’s odds from a once-competitive position to 0%, possibly linked to its GPT-5.6 availability delay (9to5Mac, Jul 8).
26 outcomes