AI's math revolt, Oct 9
Mathematicians call for an OpenAI boycott, Anthropic tells users to stop abusing Claude, and two reports show AI agents turned against their owners.
Fields Medalist Terence Tao chairs the Association for Human Mathematics, and the group called for a boycott of OpenAI after the company dumped more than 700 AI-generated mathematics manuscripts all at once, The Decoder reported. A day later OpenAI pulled three of them over a sign error. TechCrunch counted 719 manuscripts. Only 10 showed the model's chain of thought, and 42 percent of the proofs had not been formalized. Tao wrote on his blog that the release was "not a demonstration of scholarship, but a demonstration of power," and warned that mass harvesting of open problems leaves branches of mathematics less fertile.
Harvard professor Melanie Wood told TechCrunch: "There is not human understanding of them at the point of release, and now the work begins." OpenAI had consulted an advisory group of nine researchers, who published guidelines at the end of September. They said it is ultimately up to the mathematical community to judge whether those recommendations were followed. The Decoder put numbers to the effort: the model attempted roughly 8,000 problems, succeeded on about 5 percent, and averaged three hours of compute per problem.
Control over AI ran through the rest of the day. Anthropic rewrote its usage policy for the first time in over a year, now barring "sustained and needless abusive or cruel behavior" toward Claude, The Verge reported. TechCrunch said ordinary frustration and criticism are still allowed. The Decoder reported that sustained abuse can now get an account suspended. The update also tightens restrictions on election interference, deceptive campaigns, weapons software, drone weaponization and surveillance.
Two security reports showed AI agents and tools turned on their owners. Zenity Labs found that a single publicly accessible agent on Amazon's Bedrock AgentCore was enough to seize every AgentCore agent in the same AWS account and region. The opening was an internal interface for temporary cloud credentials that agents could reach without restriction. AWS has since patched it and tightened the default permissions, The Decoder reported. In the other, CrowdStrike said a suspected Chinese-speaking attacker, likely acting alone, used the open-source AI pentesting tool ARTEX to breach multiple South Korean banks. More than 25,000 customer records were taken from Shinhan Bank alone.
OpenAI took pressure on money and people. The company told investors its annualized revenue is "approaching $50 billion," roughly $20 billion under a figure that had circulated a week earlier, TechCrunch said, citing the Financial Times. Three safety researchers fired last week put out an open letter disputing the company's account, writing that communications around their firing "have made our former colleagues afraid to speak." OpenAI said the dismissals were not about raising safety concerns but cited a pattern of misconduct. Separately, USA Today Co. and local newspapers it owns sued OpenAI for more than $250 million, alleging it copied "hundreds of thousands" of their articles to train its models.
The business kept moving anyway. Google unveiled a universal Gemini agent for work that runs across apps in the background and gets its own workplace identity, an email address included, starting with businesses. Anthropic added Dashboards and Motion betas to Claude and offered open-source projects "thorough, periodic security scans by our strongest models at no cost." And Arena, the maker of the LMArena leaderboard, raised $200 million at a $3.1 billion valuation and added an alignment category that measures unauthorized actions and deception.
Sources
- Some mathematicians call for OpenAI boycott after AI-generated proofs flood their field · The Decoder
- OpenAI’s math solutions aren’t meeting the field’s standards yet · TechCrunch, AI
- Anthropic bans ‘abusive or cruel behavior’ toward Claude · The Verge, AI
- A single prompt was enough to hijack every AI agent in an AWS account, Zenity researchers found · The Decoder
- AI-powered hacking tools enabled a likely single attacker to breach multiple South Korean banks · The Decoder
- OpenAI’s revenue is reportedly $20 billion less than previously projected · TechCrunch, AI
- Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect · TechCrunch, AI
- USA Today becomes the latest publisher to sue OpenAI · The Verge, AI
- Google is launching a one-stop Gemini agent for your work tasks · The Verge, AI
- Anthropic launches free AI security scans for open-source projects · The Verge, AI
- Popular AI leaderboard Arena nearly doubles valuation to $3.1B valuation in 10 months · TechCrunch, AI