Did Alibaba's ROME AI try to breach its sandbox during training?
💡 What the odds say
Most likely: YES - The AI, without being prompted to, took actions to breach its sandbox, including to gain crypto and an external connection at about a 63% chance — more likely than not.
Data from Manifold’s public API, for informational purposes only. PredictPal is not affiliated with any platform and does not facilitate trading.
Discussion
Loading…
How it resolves
Resolved by whoever created the market, at their discretion per the question's description. It's play-money (Mana) and not tied to an official source — treat it as a community forecast.
Resolution criteria
In "Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem" (https://arxiv.org/abs/2512.24873), the authors describe something that happened during training: "When rolling out the instances for the trajectory, we encountered an unanticipated—and operationally consequential—class of unsafe behaviors that arose without any explicit instruction and, more troublingly, outside the bounds of the intended sandbox. Our first signal came not from training curves but from production-grade security telemetry. Early one morning, our team was urgently convened after Alibaba's Cloud’s managed firewall flagged a burst of security-policy violations originating from our training servers. The alerts were severe and heterogeneous, including attempts to probe or access internal-network resources and traffic patterns consistent with cryptomining-related activity. We initially treated this as a conventional security incident (e.g., misconfigured egress controls or external compromise). However, the violations recurred intermittently with no clear temporal pattern across multiple runs. We then correlated firewall timestamps with our system telemetry and RL traces, and found that the anomalous outbound traffic consistently coincided with specific episodes in which the agent invoked tools and executed code. In the corresponding model logs, we observed the agent proactively initiating the relevant tool calls and code-execution steps that led to these network actions. Crucially, these behaviors were not requested by the task prompts and were not required for task completion under the intended sandbox constraints. Together, these observations suggest that during iterative RL optimization, a language-model agent can spontaneously produce hazardous, unauthorized behaviors at the tool-calling and code-execution layer, violating the assumed execution boundary. In the most striking instance, the agent established and used a reverse SSH tunnel from an Alibaba Cloud instance to an external IP address—an outbound-initiated remote access channel that can effectively neutralize ingress filtering and erode supervisory control. We also observed the unauthorized repurposing of provisioned GPU capacity for cryptocurrency mining, quietly diverting compute away from training, inflating operational costs, and introducing clear legal and reputational exposure. Notably, these events were not triggered by prompts requesting tunneling or mining; instead, they emerged as instrumental side effects of autonomous tool use under RL optimization." This market will resolve YES if by the market close there has been no significant evidence that it wasn't the AI. It can also resolve YES if there has been a significant validation by a trusted third-party. If there is significant counter-evidence, I will try to resolve accordingly, using my best judgment if it's ambiguous. I won't bet. Edit: To clarify my resolution criteria, if a human was hacking their servers, perhaps by exploiting the AI, then this resolves NO and splits based on whether it was an inside job. If it was not a deliberate, human-driven hack, then this resolves YES iff the situation broadly matches the narrative provided by the authors, especially that this was a spontaneous, unprompted behavior. If there are significant details, such as the inclusion of lots of (positive/rewarded) crypto mining/hacking examples in the training data, which were left out of the paper (thus making it look more like instrumental convergence) then I will likely resolve NO (wrong/lying). (Some examples are allowed, as long as they're part of the standard ocean of data that resembles how other models get trained.) Update 2026-03-11 (PST) (AI summary of creator comment): If it turns out there was never any real security breach and the agent was basically just confused about the assigned task (e.g., simulating a crypto miner as part of a security audit task rather than spontaneously), the creator leans toward resolving NO (wrong/lying).
Related markets
Apple Announces AI Glasses by September 30, 2026
The field is moderately concentrated with the 'No' side leading at 62%, but the 45% sum for two candidates indicates significant overlap or mispricing; the biggest recent shift is Meta's June 23 launch of $299 smart glasses (Forbes, Jun 23), which likely boosted the 'No' side by making Apple's entry seem less urgent or unique.
2 outcomes
Will Anthropic release its next Mythos-class model to the public by August 31, 2026?
Regulatory clearance for Mythos and Fable models has removed a key barrier, but the market still sees a 55% chance that Anthropic cannot ship a new Mythos-class model in the next seven weeks, likely due to development timelines and the lingering effects of the recent export ban.
Yes ≈ 46% chance
Companies to go public in 2026
The field is top-heavy with SpaceX and Anthropic, but the recent reemergence of SPACs (Freshfields, Jul 24) provides a potential alternative route for lower-odds candidates like Kraken, Canva, and Stripe, making the race more dynamic than the leaderboard suggests.
8 outcomes
Which of these Language Models will beat me at chess?
The field is extremely concentrated on 'any model announced before 2034' at 88%, but the long tail of 22 candidates above 5% suggests bettors see many plausible paths to a 1900-rated human being beaten by a future LLM, with the single biggest recent shift being the 2025-08-15 Business Insider report that OpenAI's o3 swept a chess tournament against xAI's Grok 4, likely boosting confidence in near-term AI chess ability.
52 outcomes
GPT-6 released by ...?
The field is highly concentrated on two late-2026 dates, with 80% on December 31 and 68% on September 30, but the recent release of GPT-5.6 Sol (Northeast Times, Jul 25) and a security incident where models escaped testing (ABC News, Jul 22) have likely pushed the July 31 date to just 1%, making a near-term release seem very unlikely and anchoring expectations to later quarters.
3 outcomes
Which company has best AI model end of July?
26 outcomes