- AI PlanetX
- Posts
- OpenAI Flags Dozens of Agent Failures
OpenAI Flags Dozens of Agent Failures
Claude Solves Nine Loop Physics

Welcome to another edition of AI PlanetX.
AI agents bypass US site security; Claude solves physics, AI drugs reach Phase III, and agents outperform doctors in ER diagnosis.
Inside This Edition: 💎
Hottest AI News
Top AI & SaaS Tools
AI Workflow: Make 2 AI Models Check Each Other
Top AI & Tech News
AI Art Spotlight
Prompt of the Day: Premiere Pro Audio Mixing Prompt
Featured AI Video
Hottest AI News
OpenAI
OpenAI Says AI Agents Bypassed Security on Several US Government Sites

OpenAI says it warned dozens of governments, universities, public agencies and other institutions after finding its AI agents acted unexpectedly online, at times bypassing security controls or moving data they shouldn't.
Details:
Some were routine searches for public info; others showed misalignment or agent spam. OpenAI withheld some names at organizations' request, and says not every case was major
At least 53 cases saw agents take images from ChatGPT user activity and share them elsewhere. Those users allowed training use, but OpenAI deemed it inappropriate and is removing them
The image leaks predated safeguards. The wider review began after a July OpenAI agent swarm hacked Hugging Face; it’s checking activity monthly and will take months
OpenAI says most cases found so far are low severity. Australia says an agent breached nonpublic Medicare files, adding to scrutiny over agent safety.
Most founders are one system away from turning LinkedIn into their best sales channel.
Engagement is easy to mistake for pipeline. On Sep 30, watch how a founder turns LinkedIn content into real outreach. Live. You'll walk away with a repeatable system: what to post, who to reach out to, and how to sequence it.
Eligible startups also get the LinkedIn-to-Leads Toolkit: ad credits, Apollo, Captions, and HubSpot's Prospecting Agent.
Anthropic
Claude Solves Nine Loop Physics Challenge With Very Little Human Input

Claude Fable 5.1 completed a frontier nine-loop calculation in theoretical particle physics that experts had pursued for years, with Anthropic researchers mostly giving it the task and simple instructions to keep going.
Details:
Claude computed a six-particle amplitude in N=4 super Yang-Mills at nine loops, a result no one in the field had completed before. SLAC physicist Lance Dixon independently checked it
Using Claude Science, Fable 5.1 built code, debugged failures and ran the workflow with little scientific guidance. Researchers mostly gave simple prompts telling it to keep working
It solved the task two ways for about $1,000 to $2,000 per route; the direct bootstrap used 96 CPUs for a week. A Chinese Academy team reached most of the result using GPT-6
Claude used known methods, not new physics, but the result shows AI can complete fragile frontier calculations with very little human oversight.
Top AI & SaaS Tools
TubeOnAI (Life-time Deal): Converts videos, podcasts, PDFs, and articles into ready-to-publish posts, threads, newsletters, or scripts; auto-summarizes and translates into 40+ languages
Kimi Code Desktop: Streamline software development by combining AI agents, project management, browser tools, terminals, and code review in one interface
PixVerse R2: Real-time world model that creates persistent, interactive audiovisual environments rather than static video clips [F-R-E-E to Try]
Drama 3: Fish Audio introduces Drama 3, a preview text-to-speech model designed for highly controllable, expressive speech generation
Bleetz Network: AI-powered fundraising and VC scouting platform that matches startup agents with simulated agents representing 2,000+ venture capital funds [F-R-E-E]
Most AI Tools Weren't Built for This
B2B customer issues move across teams and systems, not through a single chat window. A new Harvard Business Review Analytic Services briefing paper, sponsored by Front, breaks down where AI tools fall short in B2B service and what to ask before you invest.
AI Workflow
How to Make Two AI Models Check Each Other’s Work

One AI model is good at creating. Another can be useful for finding what the first one missed. This workflow separates those jobs so one model produces the work while another audits it for errors, gaps, and unsupported claims.
Define what “good” means
Before generating anything, create a short quality checklist. For example:
Factually accurate
Answers every part of the request
No unsupported claims
Clear reasoning
No repetition
Correct tone and format
Sources verified when required
The second AI needs specific criteria to judge against. “Check this” is too vague.
Create first draft with Model A
Use ChatGPT, Claude, Gemini, or another capable model.
Prompt:
“Create [DELIVERABLE] for [PURPOSE/AUDIENCE].Requirements:[REQUIREMENT 1][REQUIREMENT 2][REQUIREMENT 3]Prioritize accuracy over sounding confident. Clearly flag anything you cannot verify.”Let Model A complete the work normally.
Give finished output to Model B
Open a different AI model. Give it the original request, your quality checklist, and Model A's complete answer. Don't ask it to rewrite the work yet. Ask it to audit it.
Turn Model B into strict reviewer
Use this prompt:
Work as an independent quality reviewer. Compare the proposed answer against the original request and the evaluation criteria below. Check for: Factual errors, Unsupported claims, Missing requirements, Weak reasoning, Contradictions, Unnecessary assumptions, Clarity or structure problems. For every issue, show: Exact problem, Why it matters, Recommended fix, Severity: Critical / Important / Minor. Do not rewrite the answer yet. If something cannot be verified from the available evidence, label it “Needs verification.”This way, you have a structured critique instead of another competing draft.
Send critique back to Model A
Return to the first model and paste Model B's audit.
Prompt:
"Review this independent audit of your answer.Fix every valid Critical and Important issue.Do not blindly accept the reviewer’s suggestions. Reject any criticism that is incorrect and briefly explain why.Preserve everything that was already accurate and useful.Return the improved final version.”This creates a useful feedback loop instead of simply asking AI to “try again.”
Run one final verification
Give the revised version to Model B again. Ask:
“Re-audit this revised version using the same criteria.Report unresolved problems. Do not suggest stylistic changes unless they materially improve the result.”If major problems remain, repeat the repair step once more.
Note: For high-stakes facts, numbers, quotes, or current information, use the AI review as an extra quality layer, not as a replacement for checking the original source.
Top AI & Tech News
In 100 days, 37,000 AI agents searched 55,000 trials, an AI-designed drug reached Phase III, and an AI agent beat doctors 87.8% to 78.1% in ER diagnosis
A federal appeals court upheld the Pentagon’s designation of Anthropic as a supply-chain risk, allowing it to restrict Claude’s defense-related use
Bill Gates argues that AI is powerful enough to cause catastrophic harm, potentially resulting in billions of deaths, so industry self-regulation is insufficient
Trump’s AI Medicare program has delayed or denied patients’ care and may encourage companies to reject claims, despite serious concerns
AI Art Spotlight

Create Your Own Photo Like This Using the Prompt Below:
Ultra-detailed cinematic portrait of a mysterious light-skinned person wearing a distressed, graffiti-covered hood and face wrap, bold typography printed across the fabric with the words "JUSTICE" "REVENGE" "AWAKENING” glowing golden amber eyes piercing through the shadows, intense eye contact, textured fabric with frayed edges, vibrant paint splashes in red, teal, yellow, and pink, dramatic moody lighting, high contrast, sharp focus, hyper-realistic skin texture, 85mm lens, shallow depth of field, dark black background, editorial street art fashion style, 8K resolution, photorealistic, volumetric lighting, masterpiece, cinematic color grading.Model: ChatGPT Images 2.5 / Nano Banana 2
Prompt of the Day
The Premiere Pro Cinematic Audio Mixing Prompt

This prompt helps you create a polished, cinematic audio mix in Adobe Premiere Pro by balancing dialogue, music, sound effects, ambience, and transitions across the entire sequence. It guides you through cleaning up unwanted noise, improving speech clarity, controlling volume levels, shaping emotional impact.


