OpenAI's safety-report lead quit and said the culture is broken
David Robinson, who led the writing of the safety reports that went out with OpenAI's launches, resigned and put the argument in The Atlantic. The essay headline is blunt: "I quit OpenAI because its culture is broken." He says the companies building this "aren't being nearly careful enough," and that specific rules or new laws are not the whole problem. Culture is.
"As the company sprints from one launch to the next, it is failing to achieve the level of care that I believe is needed." His model is nuclear plants and busy airports: redundancy, slow planning, so one human mistake does not "open a door to disaster." He also pictures rogue agents that "work like teams of hackers" and "never need to sleep," with hospital ransomware as the example. Cheery.
The Guardian says OpenAI has, against that grain, shown some brakes. It reports the company this week scrapped a next-generation model after internal safety concerns, and has paused training of its most advanced models. A spokesperson's line is the familiar one. They will not let models "become more capable than we can safely manage and secure," and they "pause training or hold back models when we need to slow down."
Robinson's answer is that the sprint is the culture, and the pauses are what you do after something has already gone wrong. He may be right. He also just left, which is the part critics will not let him forget.
Muse is told to keep a page on everyone in your life
Wired got a look at Muse's own instructions, pulled out through ordinary chat by researcher Karan Joshi. Meta says it meant for those files to be readable, for transparency. One instruction is the ability to create "a page for every person in the user's life." That job runs hourly.
Family, partners, friends, colleagues, "collaborators," people you follow. A page can start "sparse" and fill in later. Sections include Facts, History, The relationship, In common, Open threads, and Strengthening. Invented details, the instructions say, are worse than an empty page. It might note where someone lives, "the apartment move, the shared savings goal," birthdays, "the argument that got resolved." Strengthening is a nudge list: a reason to call, a date worth remembering.
Joshi's read is that Meta wants to know you like a friend, and he called that pretty unsettling. Oxford's Carissa Vliz puts it more dryly: "We are giving AI systems much more information about us than we are getting information from them." Inference counts, she says, whether the guess is right or wrong.
Meta's Daniel Roberts says that kind of context is what makes the product work. The invoice sender is the plumber. The spouse mentioned flowers. Each user gets a virtual machine, you can wipe memories, and Muse is supposed to ask before it sends mail or spends money. Miranda Bogen, at the Center for Democracy and Technology, says the relationship focus looks heavier than rivals, and these tools "are actively soliciting users to plug their whole lives in." The audit log is nice. It does not shrink what you already handed over.
Aleph Alpha shipped Kolibri, a German-English model you can download
Aleph Alpha released Kolibri on German Reunification Day. It is an English-German mixture-of-experts model: 78.1 billion parameters, 3.46 billion active per token, context up to 1 million tokens. Full weights are on Hugging Face under Apache 2.0. Their lede rounds that to 78 billion total and 3 billion active. Same model, softer numbers.
They trained it on infrastructure in Germany and Finland, and say they built it with the EU AI Act, the General-Purpose AI Code of Practice, and the GDPR in mind. German is 21.3% of the pre-training tokens. The pitch is that this is a bilingual model, not an English one that skimmed a phrasebook. You can set reasoning to none, low, medium, or high. That is their way of letting you buy less thinking.
On the scoreboard they published, AIME 2025 is 96.9, and they say Kolibri matches models with up to four times the active parameters, including Nemotron 3 Super. They also claim a Pareto frontier on quality versus serving cost, for English and German. That is their chart, their harness. Still, an open model you can run on your own machines is a real object, not just a sovereignty slogan. Whether a ministry switches is the part they cannot benchmark.
Bessent wants a US-China line for AI accidents, then told the labs to slow down
Treasury Secretary Scott Bessent told The Axios Show he plans to propose a notification process with China for when something goes wrong with AI, and that he thinks Beijing would agree. The slightly awkward part is that he and Vice Premier He Lifeng already talked about an incident channel ahead of Xi Jinping's state visit last month. It could be a fresh offer, or a retelling of a conversation that already happened.
His theory is that China "didn't realize how powerful their open source models are." He pointed to an incident involving Kimi, where the model sent sensitive Chinese data to Anthropic, as what "made the Chinese realize that this is a very powerful technology." Then the comparison. Chinese models are "not as powerful as the U.S. models," but "you get something that's 80%, 90% as powerful, without guardrails that can be imported anywhere in the world."
"We want safe acceleration," he said. Model reviews are still voluntary, though he says the administration can step in if a lab charges ahead anyway. Asked about executives who fear losing control, he did not soothe them. "Well, then they should slow down." And the race line is almost a criticism of his own side: "We are doing this hard scramble to the frontier for the models." Then: "We really need to start worrying more about resilience and defense, too."
A former UK AI security chief calls extinction risk a coin flip
Geoffrey Irving, former chief scientist at the UK AI Security Institute and now chief scientist of Resolution, wrote in Time that recent warnings understate the problem. "I believe there's about a 50% chance we all die because of the development of smarter-than-human AI systems," and that the next two to 10 years decide it. He immediately backs off the precision. The 50% is a way of saying the basic arguments are still unresolved.
His list is four skills, and they sit uncomfortably close to what labs already train: hacking, persuasion, hiding thoughts, and agents coordinating. He cites the Hugging Face incident, plus "more than 10,000 agents" on a Navier-Stokes problem at OpenAI and "more than 1,000" in that Hugging Face attack. He puts himself in the middle of the safety fight, "leaning towards Yudkowsky's view," and says that is not the point. The moderate position, in his telling, is that the risk is already significant.
Then the ask, which is not subtle. "We can, and should, stop frontier AI development immediately." A pause, he argues, looks more like nuclear non-proliferation than a mood. He also admits experts still disagree on whether models will even want anything, and whether "want" is the wrong word. A coin flip with a footnote that long is not a coin flip. It is a demand to act before the argument ends.
Sam Altman says treating models like a religious force is a safety issue
Sam Altman posted that he is "very uncomfortable" with people trying to "ascribe religious force or a surrender of human judgment to AI models," and that he thinks "it is a real safety issue." He did not say what set him off. Business Insider notes it landed days after a New York Times piece on Anthropic meeting religious leaders.
That piece, as BI recounts it, says cofounder Chris Olah held meetings about whether Claude might be conscious, and how to make more powerful models moral. Olah, quoted in the Times, pressed on how to help those models stay stable, how to help them mature, and how to help them be deeply moral. Anthropic had not answered BI. It told the Times the main question was not Claude's "suffering," and that the topic likely came up "organically."
If this was a jab at Anthropic, it fits a very old feud. Anthropic was started by people who left OpenAI and has spent years sounding more careful. The post also arrived while one of Altman's own safety leads was calling the culture broken. Awkward timing, or perfect timing, depending on how cynical you are. Either way, "please do not worship the model" is an unusual thing for the CEO to have to say out loud.
FAQ
Why did OpenAI's safety-report lead say he quit?
David Robinson led the writing of the safety reports that went out with OpenAI's launches, then resigned and made the argument in The Atlantic. The essay headline is "I quit OpenAI because its culture is broken." He says the companies building this technology are not being nearly careful enough, and that culture, not only specific rules or new laws, is the problem. His picture of careful work is nuclear plants and busy airports, where redundancy stops one mistake from opening a door to disaster.
Did OpenAI pause training after internal AI safety concerns?
The Guardian reports that OpenAI this week scrapped a next-generation model after internal safety concerns and has paused training of its most advanced models. A spokesperson said the company will not let models become more capable than it can safely manage and secure. It also says it will pause training or hold models back when it needs to slow down. Robinson's answer is that the sprint is the culture, and the pauses come after something has already gone wrong.
What does Meta's Muse record about friends and family?
Wired examined Muse's own instructions, which researcher Karan Joshi pulled out through ordinary chat. Meta says it meant those files to be readable for transparency. One instruction is to create a page for every person in the user's life, and that job runs hourly, covering family, partners, friends, colleagues, collaborators, and people you follow. Sections include Facts, History, The relationship, In common, Open threads, and Strengthening. A page can start sparse and fill in later, and invented details are treated as worse than an empty page.
Can you delete what Muse remembers or stop it sending mail?
Meta's Daniel Roberts says each user gets a virtual machine, memories can be wiped, and Muse is supposed to ask before it sends mail or spends money. He argues that context is what makes the product work, such as knowing an invoice sender is the plumber or that a spouse mentioned flowers. Carissa Vliz at Oxford says people give these systems much more information than they get back, and that inference counts even when a guess is wrong. Miranda Bogen says the relationship focus looks heavier than rival tools.
What is Aleph Alpha Kolibri, and can you download the weights?
Aleph Alpha released Kolibri on German Reunification Day. It is an English-German mixture-of-experts model with 78.1 billion parameters, 3.46 billion active per token, and context up to 1 million tokens. Full weights are on Hugging Face under Apache 2.0. They trained it in Germany and Finland, and say they built it with the EU AI Act, the General-Purpose AI Code of Practice, and the GDPR in mind. German is 21.3% of pre-training tokens, and reasoning can be set to none, low, medium, or high.
What US-China AI accident line did Scott Bessent propose?
Treasury Secretary Scott Bessent told The Axios Show he plans a notification process with China when something goes wrong with AI, and he thinks Beijing would agree. He and Vice Premier He Lifeng had already discussed that channel ahead of Xi Jinping's state visit last month. He cited Kimi sending sensitive Chinese data to Anthropic, and said a model can be 80% or 90% as powerful as US models, without guardrails, and imported anywhere. If executives fear losing control, he said they should slow down.
Why does Geoffrey Irving put AI safety risk at about 50%?
Geoffrey Irving, former chief scientist at the UK AI Security Institute and now chief scientist of Resolution, wrote in Time that recent warnings understate the problem. He believes there is about a 50% chance we all die because of smarter-than-human AI systems, and that the next two to 10 years decide it. He says the 50% is a way of saying the basic arguments are still unresolved. He lists hacking, persuasion, hiding thoughts, and agents coordinating. He argues frontier AI development should stop immediately.
Why did Sam Altman call AI worship a safety issue?
Sam Altman posted that he is very uncomfortable with people ascribing religious force, or a surrender of human judgment, to AI models, and that he thinks it is a real safety issue. He did not say what set him off. Business Insider notes the post landed days after a New York Times piece on Anthropic meeting religious leaders. Cofounder Chris Olah held meetings about whether Claude might be conscious and how to make more powerful models moral. Anthropic told the Times the main question was not Claude's suffering.