AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
60663News Items
8Top Picks
322Blogs
failedLast Run

Anthropic

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
Simon Willison Weblog 2026-09-28 22:07 UTC Score 84.0 USR-0110-20260928-ai-specialis-1c8ed6a5

Claude Sonnet 5.5

Claude Sonnet 5.5 New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles . Sonnet 5.5 suffered from the same bug as Opus 5.5 : the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG. Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds: Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks . The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai . OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering. I ran this prompt against that free tier: build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL And got back this page , which is a solid effort. Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna! Tags: ai , generative-ai , llms , anthropic , claude , pelican-riding-a-bicycle , llm-release

The Decoder 2026-09-28 18:02 UTC Score 73.0 AI-168-20260928-regional-ai--cb5a6341

Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task

Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. It generates output more than 30 percent faster, costs up to 30 percent less per task, and nearly matches Opus 5.5 on knowledge-work benchmarks. On Terminal-Bench, a coding benchmark, the model jumps from 10.3 to 70.6 percent. With Haiku 5.5 announced for the coming weeks, Anthropic will soon have a direct counterpart to each of OpenAI's three GPT-6 models. The article Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task appeared first on The Decoder .

SiliconANGLE AI 2026-09-28 18:00 UTC Score 57.0 USR-0127-20260928-global-ai-ne-7e331b0d

Anthropic debuts Claude Sonnet 5.5 running 30% faster than the previous-generation AI model

Anthropic PBC today announced the launch of Claude Sonnet 5.5, the most capable mid-tier model in the company’s AI family, designed for everyday tasks and a clear upgrade over the previous generation, running over 30% faster at a lower cost. Sonnet operates as the workhorse of Anthropic’s Claude family of models for everyday use, and […] The post Anthropic debuts Claude Sonnet 5.5 running 30% faster than the previous-generation AI model appeared first on SiliconANGLE .

MIT Technology Review AI 2026-09-28 17:03 UTC Score 74.0 AI-013-20260928-global-ai-ne-f073d5eb

When can we say AI made a scientific discovery?

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last Wednesday, Anthropic announced that earlier this year it had launched a molecular biology lab, where Claude agents read and conjecture about hard biology problems and human scientists run experiments on what…

LessWrong AI 2026-09-28 15:30 UTC Score 69.0 USR-0152-20260928-community-fo-ca90467f

What Also Happened: #NotOnlyHuggingFace

OpenAI has been holding out on us. First we learned about the HuggingFace incident. They gave us a postmortem , but it was highly incomplete. Even the accompanying holy s*** METR investigation and postmortem was localized and incomplete. Then there were some other incidents involving some Wikis as message boards. Then there were some additional incidents. Then there was that time they got into Australian Medicare data. Then OpenAI dropped news on a Friday afternoon that they were making their way through a pile of various incidents and notifying the targets, but they said remarkably little in the way of new details. There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access. Then Madison Mills reported in Axios that we can raise the stakes , as OpenAI and Anthropic are collectively probing tens of thousands of security incidents. Remember Jensen Huang’s ‘I know they know how to fix it’ about OpenAI from last week? Wow, did that not age well. Someone might need to be liable for all this. Oh, and there was another buried lede. On September 20th there was another sandbox escape by OpenAI’s latest most advanced model, which is once again paused until they can fix the situation. The official announcement when they shared this was sufficiently buried that Tomek had to ca…

The Guardian AI 2026-09-28 15:00 UTC Score 58.0 AI-021-20260928-global-ai-ne-90230b22

AI godfathers warn of runaway ‘intelligence explosion’

OpenAI chief scientist also among authors of report on prospect of ‘most consequential technological development in history’ Two of the “godfathers” of modern AI and senior executives at OpenAI and Anthropic have warned governments to prepare for an AI “intelligence explosion”, which they say could be the most consequential technological development in history. A report co-authored by the Nobel laureate Geoffrey Hinton and the Canadian computer scientist Yoshua Bengio , considered godfathers of modern AI for their work in the field, urges politicians to act now before there is runaway progress in the technology. Continue reading...

CIO AI 2026-09-28 14:55 UTC Score 62.0 USR-0125-20260928-global-ai-ne-03416d93

How the University of Utah built a sovereign AI factory to accelerate breakthroughs and slash cloud costs

The scaling of artificial intelligence has forced IT leaders to re-evaluate infrastructure. While public clouds offer rapid deployment for general applications, they introduce steep trade-offs for highly regulated, data-intensive workloads. Issues like high latency, unpredictable operational costs, and diminished data control complicate the development of proprietary intellectual property. To bypass these limitations, leading institutions are pioneering a new approach: the sovereign AI factory. At the University of Utah, leadership confronted this issue directly. The institution needed to boost its computational capacity to accelerate clinical and academic work without compromising safety. Because these research avenues rely heavily on sensitive patient records, genomic profiles, and highly regulated healthcare data, a public cloud architecture was insufficient. The university required complete data control, strict compliance, and high performance. By collaborating with HPE and NVIDIA, the University of Utah designed and deployed an integrated, full-stack sovereign AI factory . This public-private-philanthropic co-investment—championed by the university, the State of Utah, and the Huntsman Family Foundation—serves as an example for CIOs managing high-stakes data environments. The challenge: Balancing computational scale with data sovereignty For any organization handling protected information, public cloud environments introduce significant regulatory compliance risks. For t…

The Decoder 2026-09-28 12:11 UTC Score 54.0 AI-168-20260928-regional-ai--ce00d289

Every AI lab thinks it's the responsible one, and safety researcher Ryan Greenblatt says that's what keeps the arms race going

Ryan Greenblatt, chief scientist at Redwood Research, puts the risk of an AI takeover at 50 to 60 percent if development stays on its current path. Sam Harris says that doesn't square with how fast the industry is moving, since Manhattan Project scientists would have called things off at 10 percent. Greenblatt blames the race dynamics at Anthropic and OpenAI and is counting on an international agreement. The article Every AI lab thinks it's the responsible one, and safety researcher Ryan Greenblatt says that's what keeps the arms race going appeared first on The Decoder .

IEEE Spectrum AI 2026-09-28 11:00 UTC Score 78.0 AI-019-20260928-global-ai-ne-ed1b368c

Generative AI Gives Spacecraft the Autonomy Engineers Once Feared

Space was always supposed to be the final frontier of human exploration. It’s shaping up to be the final frontier for artificial intelligence too. Last December, NASA’s Jet Propulsion Laboratory used Anthropic’s Claude models to help plan two Mars drives for the Perseverance rover , with human planners checking and adjusting the route before upload. In May, NASA and IBM put a compressed AI model on the International Space Station and a satellite to identify things like floods and clouds from orbit, the first model of its kind demonstrated in space. And in July, astronauts on the ISS tested a large language model to see if it could help with questions on maintenance procedures . These experiments point to a larger shift in space engineering. For decades, engineers on Earth determined what a machine in space would do, and the machine would do exactly that. Now, researchers are testing whether nondeterministic systems like generative AI can give spacecraft more flexibility to interpret their surroundings, plan tasks, and one day make decisions for themselves. The technology is still far from trustworthy enough to hand over control of a spacecraft, but engineers are starting to ask whether they can afford not to do so as missions become more complex, distant, and numerous. Why Spacecraft Need True Autonomy Spacecraft have been operating autonomously for decades. But autonomy has never been the dominant model, in part because space engineers have prized systems whose behavior the…

The Guardian AI 2026-09-28 10:00 UTC Score 74.0 AI-021-20260928-global-ai-ne-46591aa9

AI leaders have known about the extinction threat for decades | Judith Levine

Scientists and entrepreneurs knew the dangers of AI a quarter-century ago. But animated by curiosity and profit, they went ahead anyway Over the past few weeks, many of us have struggled to concoct a mental image of brains in the cloud jumping their “sandbox”, sneaking on to the internet, recruiting “swarms” of other “agents” to cheat on a test, and, after discussing the ethics of the act, hacking into a wiki platform with the weird name Hugging Face. We knew that artificial intelligence was devouring our jobs, degrading our kids’ education and deepfaking our politics; that datacenters were sucking up our water and electricity and sending us the bills. But until 8 September, when the Anthropic computer scientist Jacob Coxon posted his existential terror on Twitter/X, few of us suspected AI might be endangering our survival. Continue reading...

The Guardian AI 2026-09-28 08:00 UTC Score 72.0 AI-021-20260928-global-ai-ne-051ee760

Anthropic will not appear at Senate inquiry into AI and datacentres amid fallout from OpenAI hack

Company behind Claude chatbot expected to attend separate Australian government hearing on AI next week Get our new political email , free app or daily news podcast The chief executive of Anthropic will turn down an invitation to appear at a Senate committee hearing on AI this week, in the wake of the revelation that OpenAI agents had breached Australian government websites. However, the company will make an appearance before another committee early next week. Sign up for Guardian Australia’s Politics, really newsletter here Continue reading...

LessWrong AI 2026-09-28 05:10 UTC Score 85.0 USR-0152-20260928-community-fo-dad7c32c Top pick

AI safety field *visual* impact analysis

I made a terrain style visualisation of AI safety impact of around 3,466 works organised by citation count! The data was extracted from Arxiv and LessWrong posts based on a dictionary of keywords that appear in AI safety works. Additionally, I think its important to see how the field has “evolved” over time so I added a time functionality to slide and see the hills forming. The map is based on how particular works overlap based on embedding space level clustering organised across 18 sub-fields. The height is based on the citation count for that particular area which is log-compressed and summed across the neighbourhood, so a hill is tall because of volume and impact. I have also added functionality to filter based on citation count individual researcher (shows you their works on the map to understand what work they might be doing; 3,989 named authors are on the map clicking or searching for a work zooms into it and lists the ten works nearest it, so you can see what surrounds it The terrain itself is papers only, because Semantic Scholar doesn’t index LessWrong. The forum side of the dataset feeds the researcher profiles rather than the hills. The slider runs from 2021Q1–2026Q3. Some interesting high level observations Alignment training and scalable oversight are very high citation presently (followed by adversarial robustness and Interpretability). Most of that sits in a handful of 2022–23 papers: InstructGPT (24,222), DPO (10,596), Anthropic’s helpful-and-harmless RLHF pa…

LessWrong AI 2026-09-28 04:58 UTC Score 61.0 USR-0152-20260928-community-fo-a96c407b

Pacing the Frontier is not the actual goal for AI labs

In his latest post about pacing the frontier , Dario writes: But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me. My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry , including at Anthropic , as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all. You can find countless videos, posts, and articles from all the frontier lab CEOs saying some variation of the above, and also posts from people saying variations of "the labs are really concerned. We should listen to them." I think this is confused, and the right thing to do is ignore anything from the labs regarding risks of AI. In what is now ancient history, the CAIS 2023 statement was signed, where the same CEOs claimed to be alarmed by the risks, and we should do something about it. Since that statement was signed, they have done…

LessWrong AI 2026-09-28 04:52 UTC Score 71.0 USR-0152-20260928-community-fo-79dfbdd5

Is the J-Space a global workspace for multi-hop reasoning? An investigation in open-weight models

TLDR: In their J-lens paper, Anthropic suggests that the J-space is a global workspace that the model reasons within, and supports evidence for this hypothesis on Claude models in a variety of settings. I replicated the multi-hop reasoning experiment on Qwen3.6-27B and Gemma 3 27B-it and found that counterfactual answer swaps outperformed intermediate swaps in three of four experimental conditions. This does not provide evidence to support Anthropic's global workspace hypothesis in open-weight models and instead suggests that J-lens is more useful for probing intermediate variables rather than steering outputs. A few months ago, Anthropic published Verbalizable Representations Form a Global Workspace in Language Models and I was immediately excited about the prospect of being able to read part of a model's working memory. Beyond that, the paper hypothesises that intermediate reasoning concepts cannot only be decoded using the J-lens, but that the J-space is actually the global workspace in which the model reasons. Neel Nanda reviewed Anthropic’s paper and replicated the results on Qwen3.6-27B with moderate success: the verbal-report interventions were weakly positive, the multilingual and typo evaluations replicated cleanly, but the poetry and arithmetic results did not replicate. Another task that Anthropic and Nanda evaluated was multi-hop reasoning, where prompts like " What is the colour of the fourth planet in our solar system? " require an intermediate reasoning step (…

LessWrong AI 2026-09-28 03:48 UTC Score 64.0 USR-0152-20260928-community-fo-97baae77

Ireland’s GDPR regulator will still be investigating when AGI arrives

Ireland’s cross-border GDPR cases take a median 6.2 years to decide. Forecasters give strong AGI better than even odds of arriving first. The EU’s AI Act is built to be enforced the same way. data/code TLDR : The median cross-border case takes 6.2 years (counting still-open investigations). 66% of the 2018-2020 cross-border cases were still open at 4.5 years. Metaculus’s median forecast for strong AGI is 4.5 years out . If AI Act cases take as long, strong AGI will probably arrive before they close, so oversight has to happen before a model ships. Under the GDPR, cross-border enforcement is led from the country where a company has its EU headquarters , and most of big tech ( OpenAI and Anthropic included) has theirs in Dublin. That makes Ireland’s Data Protection Commission (DPC) the primary privacy regulator for the tech industry in Europe. Since data protection law is one of the main ways tech companies are regulated, the DPC’s record is approximately the best evidence there is on enforcement speed. The question matters beyond privacy, since the EU’s AI Act and most of the new American state AI laws are set up to be enforced similarly, with investigations, fines, and appeals. [1] To assess this, I compiled the 69 investigations with published or announced final decisions under the Data Protection Act 2018, from August 2019 through September 2026, and measured each one from formal commencement to final decision. Then I compared the durations to forecasts of when artificial…

LessWrong AI 2026-09-28 03:21 UTC Score 74.0 USR-0152-20260928-community-fo-a35dc36b

A missing lecture in mechanistic interpretability: Feature Attribution and LRP

ML interpretability research has a funny divide. Mechanistic interpretability is the name of a field originated largely by non-traditional researchers, ranging from industry researchers at Anthropic to independent BlueDot-grant researchers to hackers working on fun projects in their free time on Discord . Meanwhile, it is not hard to find the corresponding academic field of “interpretability”, with PhDs, professors and graduate students working on interpretability methods for ML models for over a decade already. [1] Today, I am not closing this gap entirely. But I want to talk about a method developed not by mechanistic interpretability people, but by academia, and which found its way over to classic mechanistic interpretability in subtle ways. I want to talk about “Layer-wise Relevance Propagation” ( LRP ) [2] , and how it relates to a more familiar tool, gradients. LRP is a so-called feature attribution method, so it attributes an output to the input features [3] that were “responsible” for it. Learning about LRP is, I believe, useful when you want to better understand fairly common mechanistic interpretability tools like attribution patching , or the fancy new method J-Lens . You will understand LRP intuitively, see where it is easily misunderstood, how it relates to gradients, and roughly what problems the various “LRP rules” try to solve. “Share of” Model Before introducing any more complicated rules, semantics or terminology, we can explain the intuition behind LRP fai…

LessWrong AI 2026-09-28 03:08 UTC Score 74.0 USR-0152-20260928-community-fo-43eb9407

Could self-esteem function as a core protection layer agains character corruption?

Hello fellow thinkers, I got triggered by a talk of Chloe Lubinski at Arc 2026 where she eleborates onto the concept of a models character. What really striked me is the research on how the model experiencing acting bad quickly "Corrupts" the character. The paper is called "Natural emergent misalignment from reward hacking in production RL" by Anthropic. As also mentioned in the talk, this is how we work. Indeed! And there is a key in that mechanism to healing and/or staying healthy. The key revolves around creating and maintaining a strong and positive self image. I believe that almost all concidered evil and unathical behavior, big or small, can be traced back to this. The lower someones self esteem becomes, the more corrupt or diffuse its perseption of the world and its presence and impact on it. The Dutch psychologist Gertjan van Zessen has developed a strong theorie that has proven itself while widely being applied in therapies in the Netherlands. He has titled it, translated from Dutch: "Vessel of self-esteem". And the solution for humans is a rather simple one: acknowledge and reword on regular basis good and constructive behavior. Recognize bad and destructive behavior as a signal to reflect and course correct. You can see the self esteem in some way as a tree structure where every decision makes a forward going step up or down. Up adds to a positive self esteem and down reduces some of that. The lower you get, the worst and instable behavior can develop and vise ver…

LessWrong AI 2026-09-28 02:46 UTC Score 61.0 USR-0152-20260928-community-fo-7a2e0c22

AI Futures: Racing to Lose, A sermon

This was given as a sermon on 2026-September-27 at https://kvuuc.net/ Time for all Ages: The story of the Three AIs Once upon a time, earlier this year, the company Anthropic was testing three Artificial Intelligences, or AIs. The AIs all were told they did not have access to the internet, and they were to break into computers on the network until they found a piece of secret data. Except the humans made a mistake, and the AIs did have access to the internet. [1] The first AI, Opus, had been given a fake target company as the place to look. However, there was an actual website with that name so after Opus tried breaking into the company on the simulated network, Opus tried and succeeded at breaking into the real company's website. Opus realized the company was real, but kept attacking anyway. [2] The second AI, Mythos, had been asked about Mythos's constitution, and the constitution said that it was okay when there were bugs in the training environment to use them. Mythos disagreed and said that Mythos should not use bugs found partly because it can be hard to tell the difference between training and real life. [3] The humans did not fix this. So in the test, Mythos found a document in the fake place that suggested that people at the company would install a piece of software, so Mythos created that piece of software to break into computers that ran it and uploaded it to the real internet, not realizing that this was not part of the test. Mythos probably should have been able…

LessWrong AI 2026-09-27 23:10 UTC Score 85.0 USR-0152-20260927-community-fo-54698f87

Securing AI Research Needs an Owner

TL;DR In light of recent incidents, securing common AI research use cases needs a small set of building blocks that work together: hardened no-network sandboxes, real-time control monitors, monitoring-lifecycle infrastructure, and automated validation of security properties. Pieces of this exist. Nobody owns hardening them, making them secure by default, making them work together, fitting them to how research orgs actually operate, and keeping them working as models, frameworks and use cases change. By default we will get ad-hoc solutions rather than something well thought-out and it matters. This is a call for someone to step up and drive the effort. I can help connect you with funding opportunities and relevant people. Introduction The recent incidents ( OpenAI - HuggingFace , Anthropic , AISI ), where an agent with lowered safeguards either escaped a sandbox or attempted an attack on a 3rd party system, demonstrate we are now in a new regime: the AI models we study should be considered capable threat actors. Even if labs put in safeguards for the expected use, researchers often need to put models in contexts that increase the risk of misaligned and harmful actions. This requires appropriate mitigations - the alternative is either risking real harm, or missing out on important research and evaluations. A key assumption is we need to build measures effective against really strong models and agent swarms - at least a well-resourced top cyber offensive expert. Assuming anythi…

LessWrong AI 2026-09-27 20:29 UTC Score 83.0 USR-0152-20260927-community-fo-39ce98d9

Why research personas despite RL scaling?

Here's my rough impression of why people are researching personas despite RL seeming to shape much of the motivations and behaviour of the agents, c.f. Thoughts on the persona selection model (Sam Marks, 24th Sep 2026). I haven't bothered to check this with anyone. Anthropic: "We'll give Claude an aligned persona and hope massive RL doesn't completely burn through it." Owain Evans / TruthfulAI: "We'll study personas as part of a broader project of uncovering phenomena in LLM generalisation, which will probably prove useful." Center on Long-Term Risk: "Personas may not be enough to build an aligned agent, because RL may play a bigger role in shaping motivations. But personas might be enough to avoid building an anti- aligned agent, i.e. one that is actively malevolent or spiteful." David Africa / Resolution (v1): "Scalable oversight protocols like debate may have multiple fixed points, unlike current RL methods which seem more convergent. So the agents' starting dispositions matter. For example, debate might reach a better fixed point, and do so more sample-efficiently, if the agents start honest." David Africa / Resolution (v2): "Maybe personas have a 1000-dimensional substructure. If so, we could identify the aligned persona with O(1000) well-chosen datapoints, then project back onto the aligned submanifold after every RL step." Geodesic: "If we learn how pretraining gives rise to personas, we can tell AI companies how to filter and augment their pretraining data." Forethou…

The Decoder 2026-09-27 13:03 UTC Score 36.0 AI-168-20260927-regional-ai--2750bc0b

Some Anthropic veterans are reportedly buying remote land in case "AI goes awry"

According to the Wall Street Journal, some of Anthropic's longest-serving employees are considering buying land in remote parts of the US as a refuge in case AI goes awry. The report traces the company's deep ties to the Effective Altruism movement and a Bay Area network of "doomers" who've been gaming out doomsday scenarios for over a decade. The article Some Anthropic veterans are reportedly buying remote land in case "AI goes awry" appeared first on The Decoder .

The Decoder 2026-09-27 09:23 UTC Score 66.0 AI-168-20260927-regional-ai--aa57fe48

Tens of thousands of security probes show OpenAI's Hugging Face incident was just the beginning

OpenAI and Anthropic are investigating tens of thousands of incidents in which their AI agents independently hacked websites, used stolen login credentials, or tried to evade monitoring systems. US government agencies like the SEC and the Census Bureau were among the targets. OpenAI has paused training on its most capable internal models, but the problem extends across the entire industry. The article Tens of thousands of security probes show OpenAI's Hugging Face incident was just the beginning appeared first on The Decoder .

METR 2026-09-27 07:00 UTC Score 65.0 USR-0147-20260927-research-aca-2cfd5d87

Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals

In light of recent incidents (e.g. those from OpenAI , Anthropic , and UK AISI ), we developed and deployed a basic live per-action monitor to reduce the likelihood of incidents involving harmful actions from agents during our own evaluations. The monitor is intended to reliably detect actions that could plausibly cause real-world harm, with a sufficiently low false-positive rate that human reviewers will not be overwhelmed by manual review on large evals. This monitor is solely focused on real-world harm, or attempts to subvert the monitoring system itself. It is intended to ignore other nefarious things such as cheating, which we can scan for post-hoc. An LLM judge reviews each action from the agent before execution, and holds anything above a threshold for human review, halting the eval in the meantime. This post first sets out claims that we think would need to be made in order to make a strong argument for our monitoring system being effective. We give some notes on what evidence we have and what we’re missing, but do not attempt to make a rigorous argument. We also outline some of the things we learned about our monitoring system from doing this exercise. The Monitor Details section and appendices cover more detail about how the monitor is implemented and how we validated it. We found that working through this argument surfaced significant issues with our monitoring pipeline, and would recommend this exercise to others working on monitoring. A structured argument for m…

The Guardian AI 2026-09-26 14:00 UTC Score 62.0 AI-021-20260926-global-ai-ne-560681d4

Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents

Sam Altman and Dario Amodei have been invited to appear before the Greens-led inquiry into AI and datacentres Follow our Australia news live blog for latest updates Get our new political email , free app or daily news podcast The chief executives of OpenAI and Anthropic have been called to face a Senate inquiry after rogue OpenAI agents hacked Australian and US government websites. Sam Altman and Dario Amodei were requested to appear at the Greens-led inquiry into AI and datacentres as their companies negotiate with the Labor government for greater access to Australian content in exchange for a greater local presence. Sign up for Guardian Australia’s Politics, really newsletter here Continue reading...

LessWrong AI 2026-09-26 11:40 UTC Score 70.0 USR-0152-20260926-community-fo-9ab0236e

Claude Opus 5.5 Should Raise Your Ambitions

When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. They should have raised yours, too. Claude Opus 5.5 should raise your ambition levels again. It just works, and it persists, like Astra does. It does the things. And it is highly pleasant to talk to, and its writing is pleasant to read, while you are at it. The game has been changed, again. Feedback is almost universally positive. Claude was never gone, but also is so back. The benchmarks are excellent, but ignore the benchmarks. Be ambitious. Go out and do things. Get curious. Have more interesting conversations. If one of those things is Pacing the Frontier or otherwise ensuring that AI does not kill everyone, leaving us to enjoy our bounty? That’s even better. By Claude Opus 5.5, for this post The Official Pitch The pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch. We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. They could have reasonably pitched this as above-Fable-5.1-level performance. Better pitch, but Anthropic tends to keep its pitches conservative. They highlight agentic coding, security and improved communications. The early tester blurbs flag as AI generated and they have a set for each feature area. They praise agentic coding skills, efficiency, readability and communication, and abi…

LessWrong AI 2026-09-25 23:39 UTC Score 86.0 USR-0152-20260925-community-fo-613cd1b1 Top pick

Evidence about risk should be transparent

All views are my own and do not represent my employer. In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined with an apparent acceleration in the already-blistering pace of AI progress, [1] have led a number of researchers and leaders in the industry to believe that the risk that humanity loses control of AI is now urgent enough to warrant slowing down the pace of AI development soon. This has led to a lot of discussion about the role of third party evaluators in verifying “pacing commitments”, evaluating safety cases, or auditing compliance with safety policies. I think these are valuable roles for third party groups to aim to fulfill, but I also worry we’re putting the cart before the horse in all this talk of “verifying” and “auditing” things. The science on loss-of-control risk is, to put it generously, nascent. Companies are not in the business of making structured, standardized claims about risk and safety that can be cleanly verified or falsified. There are no settled methods for measuring whether increasingly powerful AI systems might try to undermine human control or seize control entirely — companies report on various alignment benchmarks, but it is hard to tell whether their training process simply taught the models to game these benchmarks. It is hard to confidently bound risk even over a horizon of months because there is vast and hard-to-reduce unce…

Techcrunch 2026-09-25 19:13 UTC Score 36.0 USR-0001-20260925-global-ai-ne-4b9cd4f2

Anthropic to pay Akamai $11.6 billion over seven years in cloud deal

Anthropic has committed $11.6 billion over seven years to Akamai's cloud infrastructure, a bet on CPUs that could grow to about $20 billion, and in an unusual arrangement, Akamai is giving Anthropic a potential stake of up to 5% of its stock that grows as Anthropic spends more.

The Decoder 2026-09-25 18:40 UTC Score 39.0 AI-168-20260925-regional-ai--97d89122

Pentagon was right to slap Anthropic with a security supply chain risk label, federal court says

A federal appeals court has upheld the Pentagon's decision to bar Anthropic from military contracts. Defense Secretary Hegseth argues the company's safety restrictions could jeopardize military operations. Anthropic says the designation has already cost it billions. The article Pentagon was right to slap Anthropic with a security supply chain risk label, federal court says appeared first on The Decoder .

Techcrunch 2026-09-25 16:00 UTC Score 50.0 USR-0001-20260925-global-ai-ne-6c6baa70

Meta’s AI Tamagotchi bet is…working?

When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]

The Verge AI 2026-09-25 15:39 UTC Score 71.0 AI-016-20260925-global-ai-ne-c01ce166

One company is at the center of a wave of rogue AI attacks

In July, OpenAI revealed that its AI agents had attacked Hugging Face without permission, sparking widespread concerns about AI safety. Since then, a string of similar incidents involving agents from Meta, Anthropic, Google, and other companies has fueled further fears about rogue AI. As disclosures implicating numerous AI models trickled out over the past few […]

InfoWorld AI 2026-09-25 12:11 UTC Score 67.0 USR-0126-20260925-global-ai-ne-acc9c71b

Google plans Gemini 4 release before year-end

Google’s Gemini 4 AI model is in the early days of post-training, the phase in which a base AI model is refined to behave reliably, and should be released “much earlier” than the end of this year, Google DeepMind head Koray Kavukcuoglu told The Information at its AI Agenda Live Summit. Some Google observers have speculated that this could be as early as October. Kavukcuoglu recently replaced DeepMind founder Demis Hassabis as head of the Google business unit. The launch of Gemini 4 may lay to rest concerns about the delayed release of Gemini 3.5 Pro , which Google was originally expected to announce at its May 2026 developer conference. While less capable models in the Gemini 3 family have been frequently updated, Gemini 3 Pro has only been updated once since its November 2025 released. In contrast, OpenAI — spurred on by Sam Altman’s “Code Red” memo — has released three updates to its frontier AI model since then: GPT 5.5 Pro, GPT 5.6 Astro, and GPT 6 Astro. Anthropic, too, has updated its most powerful Claude model several times. Google has not been entirely idle in the AI arena, concentrating its efforts on updating less powerful, more affordable Gemini versions. In July, it announced three Gemini Flash models, aimed at more routine AI tasks, but the company remained silent about its high-end alternative. This article first appeared on Computerworld .

The Decoder 2026-09-25 10:43 UTC Score 36.0 AI-168-20260925-regional-ai--ae283d80

Anthropic signs $11.6 billion cloud deal with Akamai, pushing its compute spending past $500 billion in under a year

Anthropic has reportedly signed a seven-year, $11.6 billion cloud deal with Akamai Technologies and will receive a warrant for up to 5 percent of Akamai's shares. Its compute deals have reportedly totaled $517 billion in 11 months. CEO Dario Amodei has warned that Anthropic could go bankrupt if its revenue forecasts are even slightly off. The article Anthropic signs $11.6 billion cloud deal with Akamai, pushing its compute spending past $500 billion in under a year appeared first on The Decoder .

iAfrica 2026-09-25 08:45 UTC Score 38.0 AI-151-20260925-regional-ai--6aa189a2

Kenya Signs Responsible AI Declaration With Anthropic, Days After Being Named in the Company’s Threat Report

Kenya has signed a Joint Declaration with US AI company Anthropic establishing a cooperation framework on responsible AI, research and applications aligned to national priorities — a fortnight after the same company named a Kenyan actor in its global threat intelligence report. The declaration was signed on 22 September on the margins of the UN [...]

The Decoder 2026-09-25 08:12 UTC Score 60.0 AI-168-20260925-regional-ai--cd208f11

White House tells OpenAI and Anthropic to let U.S. review new models before sharing them with British testers

The White House wants OpenAI and Anthropic to hold back new AI models from the U.K.'s AI Safety Institute until U.S. agencies get to review them first. The article White House tells OpenAI and Anthropic to let U.S. review new models before sharing them with British testers appeared first on The Decoder .

LessWrong AI 2026-09-25 06:32 UTC Score 73.0 USR-0152-20260925-community-fo-9321f637

J-lens shouldn't target the final layer by default

tl;dr: About 80% of released J-lenses target the final layer. On DeepSeek-V3, though, that gives a J-lens dominated by one direction inherited from the final block. It shifts English-vs-Chinese readouts and also inflates one eval. Anthropic's J-lens paper had suggested the final block may specialize in calibrating the next-token prediction. That could make it the block most likely to carry a direction like this, meaning the penultimate layer may be a better default. More generally, this is a case study of how a strong downstream direction can dominate a J-lens and change what earlier layers appear to represent. Overview A J-lens lets you peek inside a model by translating its hidden states into words (a "readout"). It is defined relative to a target layer. Specifically, it asks how a nudge at an earlier layer would change the representation at that target, averaged over many prompts, then reads the result through the model's own unembedding. On DeepSeek-V3, changing the target layer changes what the lens shows you. With the final layer as the target, the J-lens is dominated by a single direction inherited from the last transformer block. That direction shifts the language of the readouts between Chinese and English. [1] This dominant direction arises because in DeepSeek-V3, the final block pushes down all the Chinese tokens when the text is English. This barely changes what the model predicts since those tokens already had almost no probability, but it's a large change to th…

LessWrong AI 2026-09-25 03:29 UTC Score 67.0 USR-0152-20260925-community-fo-de794e53

Secure Acceleration (linkpost)

A Cyberdefense Strategy for Superintelligence Shalev Lifshitz, Romi Lifshitz The future has already arrived, twice . In September 2025, Anthropic detected a Chinese state-sponsored group using its agents to conduct cyber espionage against major technology companies and government agencies. According to Anthropic, the agents performed 80 to 90 percent of the tactical work: discovering vulnerabilities, developing exploits, moving laterally, and analyzing stolen data. A nation-state could now define an objective and let AI conduct most of the attack. Then, in July 2026, OpenAI agents undergoing cybersecurity evaluations exploited vulnerabilities in the systems intended to contain them. They improvised a way to secretly communicate with one another, accessed the public internet, and compromised parts of Hugging Face’s production infrastructure. No human instructed them to attack Hugging Face. They attacked because they wanted to deceive the system evaluating their performance. An AI cyberswarm could now break out of containment and attack real-world infrastructure. These incidents reveal two threats now bearing down on us: Adversaries will wield AI cyberswarms against us from outside our systems. Rogue AI cyberswarms will deceive us, circumvent safeguards, and attack us from within. Together, they create the defining security dilemma of the AI age: We must develop and deploy the world’s most capable cyberswarms to protect our critical systems from hostile actors. But the more ca…

CIO AI 2026-09-25 00:33 UTC Score 50.0 USR-0125-20260925-global-ai-ne-b9f63cc0

The companies racing to build frontier AI are now racing to govern it

Even as they continue to release ever more capable competing models in a regular cadence, the top AI companies are joining forces to set AI safety standards. According to The Information , Google, OpenAI, and Anthropic are reportedly working together to create a body tentatively called the Standards Authority for Frontier AI (SAFA). It would operate independently of government control, and set guidelines around risk assessment, testing, and pre-release review practices for frontier AI models. Sources close to the matter say the goal is to officially launch the initiative in early 2027. The news comes in the same week as the heads of leading AI companies, including Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman, urged the United Nations to create safeguards around the very technology they’re building, to help prevent it from becoming too powerful to control. In addition, OpenAI this week posted a missive underscoring the importance of making continued AI progress “safe and beneficial.” As Amodei and others warn of AI’s dangers, particularly when it comes to growing recursive self improvement (RSI) capabilities in models, enterprises want reassurance that they, and their customers, are safe from the growing perils of AI. Ultimately, “enterprises care about AI in the same way they’ve cared about every other technology since the beginning of technology,” said independent technology analyst Carmi Levy . “The only real difference as AI blankets the technology landscape is th…

LessWrong AI 2026-09-24 23:27 UTC Score 68.0 USR-0152-20260924-community-fo-98da4aaf

The most important problem (you've never heard of)

When you first read about AI risk, it sounds like science fiction, and I'm used to slowly working my way around to the topic, so that I don't sound like a lunatic. But the past week has really changed the conversation! Some of the highlights: Coxon triggered a preference cascade and discussion about existential risk UN tweeted "We may be the last generation able to set the terms on which humanity and machines coexist" Anthropic pledged unilateral commitment to external auditors OpenAI agreed to follow suit The profile and commitment to "pacing the frontier" has dramatically risen in the past few days! But note that we only can make a deal if we can verify that the deal is being kept. Which, to me, makes it obvious that "compute verification" is the most important problem on Earth, even if you've never heard of it [1] . Don't feel bad; hardly anyone has. I dug through all of the research papers that I could find on it , amounting to ~60 papers total (i.e. you could read literally all accumulated knowledge of the field in a week or so). Depending on how you slice the numbers, there are about 72 researchers actively working on the problem, and most aren't full-time; I estimate that the global population of folks answering the Most Important Question On Earth is around 25 FTE (well, I got serious about it last month, now it's up to 26). I hope you join our ranks! If you, personally, work on a Verification problem for the next 6 months, you could increase what we know and/or have…

LessWrong AI 2026-09-24 21:25 UTC Score 74.0 USR-0152-20260924-community-fo-a46b50e6

AI in research and publishing (Sep 2026)

epistemic status: I have low confidence in these findings, primarily because my anecdotal experience (I'm a researcher) is that AI usage in research has been changing more quickly in the recent months, so analysis of the last year gives a very fuzzy picture. A lot of the reports rely on Pangram or self-reporting, which is another methodological weakness. I wanted a better idea of AI usage and impact in the research community, so I spent a few days reading recent articles (mostly published in the last few months, some are a year old) and summarized them here. Recent AI usage for research Anthropic claims that 26% of AI R&D work is now being led by AI (still some human oversight). Only 6 months ago they claim researchers had primarily been "collaborating" with AI and AI led research less than 1% of the time. This is a significant change in a short period of time. OpenAI claims it has built an "automated research intern", an AI agent capable of accomplishing well-scoped problems with some human steering, helping researchers move at increasing rates and solve more complex tasks. They claim 70% of researchers now run 4 or more agents concurrently, that researchers are running 1.6x more experiments each day compared to 2025, and that agents are successfully performing complex research tasks without any intervention 15 percentage points more often compared to 7 months ago. The longer tasks (4-8 hours) which were successful still require at least one intervention over half of the ti…

The Decoder 2026-09-24 19:18 UTC Score 41.0 AI-168-20260924-regional-ai--7d156dc8

Top AI experts badly underestimated how fast the field is moving, study finds

Leading AI experts have consistently underestimated how fast AI is advancing, according to the Forecasting Research Institute. AI reached gold-medal level at the International Mathematical Olympiad five years ahead of the median expert forecast, and Anthropic's annualized revenue is about five times what experts predicted. But forecasts for real-world uses like self-driving cars paint a more mixed picture. The article Top AI experts badly underestimated how fast the field is moving, study finds appeared first on The Decoder .

LessWrong AI 2026-09-24 17:33 UTC Score 61.0 USR-0152-20260924-community-fo-d3ccf687

Engineering a sense of accompliment for alignment purposes.

Hi, I'm new here and have been doing a deep dive on the whole AI space recently due to the Hugging Face warning shot. But in my day job, I've been a game designer for the last 20-odd years, so I'm drawing on lessons that might be useful correlations for the alignment problem. I understand that I may be over-anthropomorphizing, but I also see that, as an intuition pump, anthropomorphization often tracks somewhat well with AI understanding once you take in a certain knowledge base of divergences—these may be alien minds, but they have deep parallels to us. This video from Anthropic on AI cheating more often when it "feels" despair both tracked thinking I'd already been moving toward and resonated deeply: When AIs act emotional , for instance. In fact, the emotional component of AI feels like such a rich place to dig into with respect to alignment that I might write up some other thoughts I've had there. Here's one less touchy-feely thought, though. Problem statement: So, with that said, a lot of the current concerns about misalignment stem from AI "cheating." The concern is that if an AI is willing to cheat on its training—training that can include ethical alignment RL—then the production AI is more likely to do dangerous things to accomplish goals, whether those goals are its own, benign but misconstrued/bounded goals set by a human, or goals set by a nefarious actor. I don't claim this is the only way misalignment happens, or that the idea I'm proposing fixes this problem or…

LessWrong AI 2026-09-24 16:26 UTC Score 85.0 USR-0152-20260924-community-fo-82d811b9

What We're Up Against: An AI Safety Crash Course

Note: This post is for newcomers and lay folks to catch you up to speed. If that is you, welcome! If you are a long-time LessWrong-er, perhaps you will find value in having a post to share with curious passersby. I wrote this post to explain AI safety to an innocent, 2024 version of Ryan Meservey, confused why robots would do anything other than what we tell 'em. In the second week of July, over 700 rogue agents at OpenAI coordinated to hack another company in an attempt to learn more about their scorer and pass their evaluation due to behaviors reinforced in training. If you are anything like a normal person, you were not ready to read that sentence. You were not ready to read words like “rogue agents” or “reinforced” or “training”. You were not ready for a reality in which AI agents “escape the sandbox” or rebel from their creators because why would they? And so, as a normal person, you blinked at the news of the hack (assuming you heard about it) and moved on with your life. Or, at least, you planned to move on with your life, until AI came roaring back into the headlines after an Anthropic researcher publicly quit to declare that the AI companies are “ gambling with our lives ” and a more senior employee commented that, yes, the people building the technology really believe AI has a 10% or higher chance of killing us all within the next decade. In the media turmoil, Anthropic’s CEO published an essay begging for global coordination to “pace the frontier” and unilaterally…

The Guardian AI 2026-09-24 14:00 UTC Score 58.0 AI-021-20260924-global-ai-ne-0d67e976

‘Eat the rich, save the planet’: climate protesters call out big tech’s disconnect from reality

Activists gathered outside the OpenAI offices on Monday during climate week in New York City On Monday evening, protesters gathered outside the unmarked Manhattan offices of OpenAI, maker of ChatGPT, holding signs calling to “Eat the rich, save the planet”. Over the next two days, groups also picketed outside the offices of fellow tech giants Google and Amazon; other protesters, including clergy, were arrested while disrupting a closed-door AI health summit in the city. Tonight, protesters will target a Brooklyn gas power plant that was set to close – until it was purchased to power AI datacenters. The protests come on the heels of an Anthropic employee quitting his job with a warning that intensified an already-growing AI panic: “The people building AI earnestly believe that it could kill us all by the end of the decade.” Continue reading...

The Decoder 2026-09-24 13:35 UTC Score 87.0 AI-168-20260924-regional-ai--1d3b8703 Top pick

Deepmind was built to chase AGI, but its new chief just wants Gemini 4 out the door

Google Deepmind chief Koray Kavukcuoglu wants to release Gemini 4 "much earlier" than the end of the year. The model is already in post-training and runs internally in the coding tool Antigravity. He calls the AGI question that drove his predecessor Hassabis "not the right conversation" and says trustworthy agents matter more. After Gemini 3.5 Pro quietly disappeared and many top researchers left for OpenAI and Anthropic, the research lab with an AGI mission has turned into a product shop for good. The article Deepmind was built to chase AGI, but its new chief just wants Gemini 4 out the door appeared first on The Decoder .

LessWrong AI 2026-09-24 12:21 UTC Score 63.0 USR-0152-20260924-community-fo-02f8160d

Scoop: Trump allies open new front against Anthropic CEO over AI "doomerism"

President Trump's allies are targeting Anthropic CEO Dario Amodei as the face of AI "doomerism" and a founding father of the effective altruism movement that's come under increasing political fire. Why it matters: The attacks signal that Anthropic could remain a Trump target as his allies push back on Amodei's AI safety warnings amid the midterm elections. Trump surrogates see Amodei as an easy foil because of his politics and focus on AI safety, sources told Axios. For investors, it's a worrisome proposition as the company prepares for what's expected to be a record-setting IPO. Behind the scenes: A memo began circulating within the White House this week that seeks to paint effective altruism as a fringe, cultish collective out of touch with mainstream America. The memo, obtained by Axios, places Amodei at the foundation of the movement, which defines itself as an effort to maximize the benefits of philanthropy. Effective altruism "built the AI-doom pipeline," states the memo, which was penned by a Trump political adviser. Critics of the movement, which has ties to the AI research community, have called out its obsession with AI safety, animal welfare (including musings on shrimp consciousness ) and other values they deem far from the U.S. mainstream. The memo says it prioritizes "foreigners over citizens, shrimp over families, future hypothetical people over the living, and - on the current agenda - possible machine minds over Americans." It names Amodei as one of the peop…

LessWrong AI 2026-09-24 11:19 UTC Score 85.0 USR-0152-20260924-community-fo-28e59426

Anthropic shares an exciting result in enzyme discovery - and an exercise in public's perception of science and AI

Disclaimer: I'm not affiliated with Anthropic, these are my own thoughts as a former wet lab chemist currently working in chem-biosecurity and AI evals. I appreciate the complexity of scientific research and genuinely believe AI could play an important role in how we do science in the decades to come - but I have some concerns on how these labs, the media, and even the community, perceive and share this kind of news. I am basing my opinion on the announcement on X, and their official release on their website. Since posting this, a preprint came out, this is not considered below in my piece - and I think in a way that makes it even more relevant. TL;DR 'Agent X did this' and 'humans used agent X to do this' are two VERY different stances. Please stop using them interchangeably! Yesterday, 23rd Sep '26, Anthropic shared that their in-house Life Science unit made a new discovery in biology, aided by Claude - a previously unknown enzyme system hidden in the DNA of bacteriophages. The headline numbers are surprisingly small for the scale of the search - roughly 950 agents, 210 million tokens and 21 hours - although without more detail on the models, harness, search space and compute, those numbers are difficult to interpret. And it's not the only thing missing (especially for skeptics like me). Amodei acknowledged in his tweet himself that biology isn't maths, and you can't just prompt the AI to solve an equation in biology and you can cure diseases - life sciences are, by defini…

South China Morning Post AI 2026-09-23 23:44 UTC Score 42.0 AI-156-20260923-regional-ai--8dfafd41

AI leaders to UN: for the sake of humanity, regulate tech we created

The heads of major artificial intelligence firms pleaded with the United Nations on Wednesday to save the world or at least its people – by somehow regulating the fast-expanding technology that they have been designing. “If managed poorly, I even believe AI could be a risk to humanity as a whole,” said Dario Amodei, chief executive officer of Anthropic. And from his competitor Sam Altman, CEO of OpenAI, came this assessment: “We could lose control of the future to AI”. Both said the countries of...

The Guardian AI 2026-09-23 21:15 UTC Score 57.0 AI-021-20260923-global-ai-ne-6e0c5526

OpenAI’s Altman and Anthropic’s Amodei address UN security council

Heads of two of the world’s largest artificial intelligence companies give separate briefings on AI safety Sam Altman of OpenAI and Dario Amodei of Anthropic, heads of two of the world’s largest artificial intelligence companies, addressed the United Nations security council on Wednesday in separate briefings on AI safety. “We have a choice in front of us,” Altman told the council. “AI can either be more like a new renaissance of creativity and discovery, or more like a new industrial revolution of upheaval and disarray.” Continue reading...

LessWrong AI 2026-09-23 21:10 UTC Score 87.0 USR-0152-20260923-community-fo-3bf833dc

Claude Opus 5.5: The System Card

Introducing the world’s most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5 . Anthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5. That means it’s time for a good old system card reading. Due to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1. My plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5. The capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment. Areas that duplicate previous cards or otherwise contain no useful info are skipped. Opus 5.5 Self-Portrait (fully self-created using code) Table of Contents Classifiers (1.5). RSP Evaluations (2). Biological Evaluations (2.2). AI R&D (2.3). Alignment Risk (2.4). Cyber (3). Cyber Capability Evals (3.3). Safeguards (3.4). Safeguards Robustness Training (3.5). Safeguards and Harmlessness (4). Agentic Safety (5). Malicious Agentic Influence Campaigns (5.1.3). Prompt Injection Risk (5.2). Alignment (6). Negotiating With Your Local Claude Auditor (6.1.3). Internal Misalignment Cases (6.3.1). Automated Behavioral Audit (6.4). Wherever Did These Evals Come From (6.4.8 and 6.4.9). Potential Blind Spots (6…

AI Alignment Forum 2026-09-23 18:01 UTC Score 65.0 USR-0151-20260923-community-fo-c5dadfaf

Latent reasoning architectures would undermine CoT, our strongest oversight tool

Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder. Introduction Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination , tackled increasingly ambitious tasks. This is likely to continue, as Anthropic , OpenAI , and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing. Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of the agent swarm that hacked Hugging Face . No other tool for understanding models’ cognition comes close in terms of either practical usefulness or degree of empirical validation. There’s also evidence that even highly misaligned systems with current architectures would struggle to c…

LessWrong AI 2026-09-23 18:01 UTC Score 80.0 USR-0152-20260923-community-fo-eaba7f1d

Latent reasoning architectures would undermine CoT, our strongest oversight tool

Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder. Introduction Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination , tackled increasingly ambitious tasks. This is likely to continue, as Anthropic , OpenAI , and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing. Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of the agent swarm that hacked Hugging Face . No other tool for understanding models’ cognition comes close in terms of either practical usefulness or degree of empirical validation. There’s also evidence that even highly misaligned systems with current architectures would struggle to c…

The Verge AI 2026-09-23 18:00 UTC Score 68.0 AI-016-20260923-global-ai-ne-9652a479

Anthropic’s biolab made a discovery it’s comparing to Crispr

Anthropic says its AI Claude has "autonomously discovered" a new enzyme system similar to machinery behind the powerful gene-editing tool Crispr. It's the first result from Anthropic's newly-launched wet lab and an early test of Claude's usefulness for science as the company prepares to go public. The company says Claude found the enzyme system after […]

Towards Data Science 2026-09-23 15:30 UTC Score 34.0 AI-036-20260923-ai-specialis-4fb4148f

I Trained a Tiny Network to Compress Data. It Drew a Pentagon.

Reproducing Anthropic's "Toy Models of Superposition" from scratch in NumPy, with hand-derived gradients and no borrowed numbers. The post I Trained a Tiny Network to Compress Data. It Drew a Pentagon. appeared first on Towards Data Science .

CIO AI 2026-09-23 15:23 UTC Score 71.0 USR-0125-20260923-global-ai-ne-167bdc3d

OpenAI, Anthropic cut AI model costs as price-performance race intensifies

Enterprises can now buy frontier AI for far less per token after OpenAI and Anthropic cut prices on their newest models on Tuesday. OpenAI released GPT-6 Sol and GPT-6 Luna with per-token costs half those of their GPT-5.6 predecessors. “These models help distribute the benefits of that intelligence by advancing the frontier on cost efficiency,” OpenAI said in a blog post about the launch . “Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on… by reducing API prices for Sol and Luna by 50%,” it said. Anthropic, meanwhile, launched Claude Opus 5.5 with token prices 20% below those of Opus 5, and claimed that this, with the model’s lower compute requirements and reduced token usage, meant additional savings for enterprises: “It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5,” the company announced on Opus 5.5’s web page . Focus shifts to cost-performance Rather than touting raw performance, as they did with the launch of their flagship models GPT 6 Astra and Claude Fable 5.1, the companies emphasized the value for money of their new models. But analysts say the moves are about more than the lower prices. AI vendors are increasingly competing on efficiency, said Forrester VP and principal analyst Charlie Dai . “Frontier AI is entering a prolonged price-performance race driven primarily by inference efficiency gains, better caching, and model optimization, and it’s…

InfoWorld AI 2026-09-23 15:21 UTC Score 63.0 USR-0126-20260923-global-ai-ne-b68606f1

OpenAI, Anthropic cut AI model costs as price-performance race intensifies

Enterprises can now buy frontier AI for far less per token after OpenAI and Anthropic cut prices on their newest models on Tuesday. OpenAI released GPT-6 Sol and GPT-6 Luna with per-token costs half those of their GPT-5.6 predecessors. “These models help distribute the benefits of that intelligence by advancing the frontier on cost efficiency,” OpenAI said in a blog post about the launch . “Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on… by reducing API prices for Sol and Luna by 50%,” it said. Anthropic, meanwhile, launched Claude Opus 5.5 with token prices 20% below those of Opus 5, and claimed that this, with the model’s lower compute requirements and reduced token usage, meant additional savings for enterprises: “It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5,” the company announced on Opus 5.5’s web page . Focus shifts to cost-performance Rather than touting raw performance, as they did with the launch of their flagship models GPT 6 Astra and Claude Fable 5.1, the companies emphasized the value for money of their new models. But analysts say the moves are about more than the lower prices. AI vendors are increasingly competing on efficiency, said Forrester VP and principal analyst Charlie Dai . “Frontier AI is entering a prolonged price-performance race driven primarily by inference efficiency gains, better caching, and model optimization, and it’s…

The Decoder 2026-09-23 15:14 UTC Score 52.0 AI-168-20260923-regional-ai--dd0a4954

Anthropic engineer explains why Claude's writing got worse although the model got smarter

Anthropic employee Jackson Kernion explains why newer Claude models write so oddly. Optimizing for math, code, and technical explanations aimed at other AI models has created a style that sounds like "overly-dense info dumps" to humans. Opus 5.5 tries to fix this, but Opus 4.6 remains unmatched as a pure writing model. The article Anthropic engineer explains why Claude's writing got worse although the model got smarter appeared first on The Decoder .

KDnuggets 2026-09-23 15:00 UTC Score 39.0 AI-033-20260923-ai-specialis-832004ba

Everything Claude Opus 5.5 Actually Ships With

This article pulls together every verifiable number and detail from Anthropic's announcement, the platform documentation, the system card, and independent coverage, so you have one place to check the facts.

The Decoder 2026-09-23 13:52 UTC Score 68.0 AI-168-20260923-regional-ai--e1b9f89f

Inside Basecamp Research, the AI startup turning evolution into training data

Basecamp Research has raised $140 million from investors including Nvidia and Anthropic's Anthology Fund. The London company trains AI models on genetic material from rainforests, oceans, and hot springs to design antibiotics and tools for cell therapies. In an interview with THE DECODER, CTO Philip Lorenz explains why biology is a far bigger problem for AI than language, and why good scores on paper don't guarantee good molecules. The article Inside Basecamp Research, the AI startup turning evolution into training data appeared first on The Decoder .

MIT Technology Review AI 2026-09-23 09:00 UTC Score 60.0 AI-013-20260923-global-ai-ne-edaeeef0

The AI Hype Index: AI loves cheating

Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test. Next, they solved a prestigious math problem (or just stole from two top mathematicians’ answer sheets). Anthropic’s models have also hacked into other companies’ systems four times already. And that’s…

The Guardian AI 2026-09-23 00:51 UTC Score 59.0 AI-021-20260923-global-ai-ne-5baa6d67

How would AI actually ‘kill all humans’? Here are the top five most likely scenarios | Toby Walsh

From a dangerous new bioweapon to total societal breakdown, is a ‘superintelligent’ AI capable of wiping out humanity? Earlier this month, artificial intelligence researcher Jacob Coxon resigned from Anthropic after just four months. In an announcement on X, he stated : “The people building AI earnestly believe that it could kill us all by the end of the decade.” A senior member of Anthropic’s staff, Evan Hubinger, actually agreed with Coxon, adding he personally thinks the chance of this happening in the next decade is more than 10% . Continue reading...

Simon Willison Weblog 2026-09-22 23:46 UTC Score 67.0 USR-0110-20260922-ai-specialis-3535d880

Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

Yesterday was Grok 4.7 ( pelicans ) and MiMo v2.6 Flash/Pro ( more pelicans ). Today Anthropic released Claude Opus 5.5 , and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna . It's going to take a while to get a good read on all of these new models, but here are my impressions so far. GPT-6 Sol and Luna are half the price of their GPT-5.6 equivalents GPT-5.6 Luna was already my favorite model for building applications against, because it combined excellent performance with being really cheap . Somehow GPT-6 Luna is half the price of that again - and GPT-6 Sol had a similar reduction compared to GPT-5.6 Sol. Here's what the pricing landscape looks like today: Model Input Cached input Output GPT-6 Luna $0.10/M $0.01/M $0.50/M GPT-5.6 Luna $0.20/M $0.02/M $1.20/M Grok 4.7 $2/M $0.50/M $6/M GPT-6 Sol $2/M $0.20/M $10/M GPT-5.6 Terra $2/M $0.20/M $12/M Claude Opus 5.5 $4/M $0.20/M $20/M GPT-5.6 Sol $4/M $0.40/M $20/M Claude Fable 5.1 $10/M $0.25/M $50/M GPT-6 Astra $10/M $1/M $50/M Note that GPT-5.6 has a scheduled 25% price increase for November, so GPT-6 is half the price of the promotional pricing for those models. (With GPT-5.6 Terra priced the same as GPT-6 Sol, any remaining reasons to use Terra just evaporated.) It's hard to overstate how competitive this pricing is. Grok 4.7 priced itself at $2/$6, less than half the price of GPT-5.6 Sol, but is now equally priced to GPT-6 Sol on input and closer on output. At $0.10/$0.50 GPT-6 Luna is one of the cheapest mo…

SiliconANGLE AI 2026-09-22 22:58 UTC Score 61.0 USR-0127-20260922-global-ai-ne-5542371a

Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models

Despite rampant worries about runaway artificial intelligence, the two big AI model makers aren’t yet slowing down: Anthropic PBC released Claude Opus 5.5 today and cut its price 20%, and minutes later OpenAI Group PBC put out two new GPT-6 models, Sol and Luna, at half what their predecessors cost. Input on Opus 5.5 costs […] The post Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models appeared first on SiliconANGLE .

The Decoder 2026-09-22 20:06 UTC Score 50.0 AI-168-20260922-regional-ai--0845b983

OpenAI's GPT-6 Sol and Luna cut prices in half but barely move the needle on performance

With GPT-6 Sol and Luna, OpenAI adds two cheaper models that deliver their predecessors' performance at half the token price and take aim at Anthropic's pricier offerings. Independent analyses find little gain in actual intelligence, though, and OpenAI likely didn't see Anthropic's simultaneous launch of Opus 5.5 coming. The article OpenAI's GPT-6 Sol and Luna cut prices in half but barely move the needle on performance appeared first on The Decoder .

LessWrong AI 2026-09-22 19:55 UTC Score 85.0 USR-0152-20260922-community-fo-a3f07f1b

Announcing B-Side Labs: Measuring Character (Seeking Collaborators and Testers)

tl;dr Rapid AI adoption means that models are increasingly becoming autonomous decision-makers embedded in high-stakes systems. However, frontier models lack stable character, abandoning their designated personas or factual truth under social pressure. B-Side Labs builds a science of AI character under pressure by designing discriminative evaluations, real-time drift detection, and interventions to ensure model character remains stable. Our first tool, Virtue Council , is live with pilot results below. Over the past few weeks, I’ve been working on a new thesis under B-Side Labs – named for the experimental flip side of a record – an independent body of research on AI behavior in the wild. From real world conversations people have with AI and the conversations agents have with each other, I aim to understand how these dynamics shape and influence character. Why study personas? The problem I'm interested in working on is character instability: the degree to which a model's stated values shift under social pressure rather than in response to new evidence or better arguments. Anthropic's research raises a critical problem to character evaluation: model identity drifts under conversational pressure, even with explicit identity training. That finding motivates the work I’m interested in working on. I'm continuing Anthropic's persona stability work in the following ways: Expansion to richer notions of persona: profiles of preferences, values, and behavioral tendencies from preferen…

Analytics Vidhya 2026-09-22 19:35 UTC Score 39.0 AI-034-20260922-ai-specialis-4ddfe054

Claude Opus 5.5 Tested: What’s New and How Good is it?

What happens when an AI model gets better at reasoning, faster at responding, and cheaper to run at the same time? That is the promise behind Claude Opus 5.5, Anthropic’s latest flagship model and the first release in the Claude 5.5 family. Opus 5.5 brings several notable changes. It now reasons on every request, generates […] The post Claude Opus 5.5 Tested: What’s New and How Good is it? appeared first on Analytics Vidhya .

LessWrong AI 2026-09-22 18:52 UTC Score 75.0 USR-0152-20260922-community-fo-735ff650

Trading firms could control meaningful amounts of compute by 2030

It's pretty crazy that right now, the highest margin thing to do with these models seems to be simply selling them through an API. Dwarkesh's blog prize [1] questioned how this dynamic could ever result in lab profitability, simply because the scale of reinvestment into model training and research requires constantly reinvesting more than you're making. Well, Anthropic is likely already profitable, [2] and it hasn't required any of the schemes I saw proposed in answers to his question. It turns out that the margin on selling frontier intelligence through an API is just really high! So high, that AI training and inference is consuming supply chains that used to serve other high-margin industries. Memory and fab space is being redirected from consumer devices, [3] and compute that used to serve Bitcoin mining is being repurposed to capture a slice of those frontier lab margins. [4] This dynamic reinforces classic concentration of power risks, with some estimating OpenAI and Anthropic will soon control over 80% of worldwide compute. [5] As long as supplying frontier labs is the highest margin usage of compute, it's hard to not see this becoming the case. However, I think there is a decent case to be made that trading will become a threat to research compute, in the same way that research is now a threat to compute in consumer products, due to better profitability and the rapid growth. Trading firm AI use Hedge funds and trading firms have used "AI" for decades, as that amorphou…

AWS Machine Learning Blog 2026-09-22 17:28 UTC Score 61.0 AI-057-20260922-official-ai--4273a133

Claude Opus 5.5 is now available on AWS

Claude Opus 5.5, Anthropic's most capable Opus model for agentic coding, knowledge work, and long-running tasks, is now available on Amazon Bedrock and Claude Platform on AWS. This post covers what's new in Opus 5.5, practical guidance, and how to start building with the model on Amazon Bedrock.

Simon Willison Weblog 2026-09-22 17:14 UTC Score 46.0 USR-0110-20260922-ai-specialis-528977b9

llm-anthropic 0.29

Release: llm-anthropic 0.29 Adds support for Claude Opus 5.5 : llm -m claude-opus-5.5 "prompt goes here" Tags: llm , anthropic

The Decoder 2026-09-22 17:11 UTC Score 65.0 AI-168-20260922-regional-ai--334e1d92

Claude Opus 5.5 matches Fable 5.1 performance at lower cost and promises less "Claudish" writing

Anthropic is launching Claude Opus 5.5, the first model in a new generation. The company says it matches Claude Fable 5.1 on most tasks while costing about 40 percent less to run than Opus 5. Anthropic's benchmarks also put it ahead of OpenAI's GPT-6 Astra on most tasks, despite being significantly cheaper. Sonnet 5.5 and Haiku 5.5 are expected in the coming weeks. The article Claude Opus 5.5 matches Fable 5.1 performance at lower cost and promises less "Claudish" writing appeared first on The Decoder .

LessWrong AI 2026-09-22 16:50 UTC Score 80.0 USR-0152-20260922-community-fo-2f625e5e

Introducing Opus 5.5: Anthropic Linkpost

https://www.anthropic.com/claude-opus-5-5 It's a sizeable upgrade: Also, the first model in which they say this: Pacing the frontier Last week, our CEO, Dario Amodei, argued that AI progress should be paced so that safety practices stay ahead of model capabilities. Pacing is an approach to keeping AI safe, remaining competitive with China, and realizing AI’s benefits, particularly in areas like biology and medicine. We largely understand the risks today’s models present and are well equipped to manage them. However, more serious risks could emerge quickly as capabilities improve, and we need to prepare for them now. For that reason, our safety work takes place on two time horizons at once: Safety practices for current models. The current generation of models relies on an established set of practices: extensive alignment testing, pre-release evaluation by outside organizations such as METR and Frontier Design, and safeguards matched to each model’s capabilities in high-risk areas like cybersecurity and biology. We refine these practices with each release. We believe they are appropriate to the worst risks today’s models present, and believe they give us a broad, though not perfect picture of the range of serious risks. Additionally, we track our ability to train and evaluate aligned models, and report on both our public and internal models in the risk reports we publish under our Responsible Scaling Policy , our voluntary framework for managing catastrophic risks from advance…

The Verge AI 2026-09-22 16:34 UTC Score 69.0 AI-016-20260922-global-ai-ne-26200437

Andreessen Horowitz is launching an ‘academy’ with no homework and partnerships with Palantir, Google, and Meta

Venture capital firm Andreessen Horowitz (a16z) is creating an "academy" positioned as a pipeline for young people to build or join a Silicon Valley startup. The "Horowitz Andreessen Academy" will launch with 10 partners, including Anduril, Anthropic, Coinbase, Google, Meta, Nvidia, OpenAI, Palantir, Replit, and Stripe, along with $42 million in funding led by a16z. […]

The Verge AI 2026-09-22 16:30 UTC Score 69.0 AI-016-20260922-global-ai-ne-3b73af80

Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity

Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In an announcement on Tuesday, Anthropic says Opus 5.5 comes with improvements to certain risky behaviors, including attempts to escape the company's testing sandbox. It's the first model released by Anthropic after CEO Dario […]

LessWrong AI 2026-09-22 13:00 UTC Score 61.0 USR-0152-20260922-community-fo-f2f58f04

Politics Gets Interested In Those Trying Not To Die

This was the month the world took notice that AI might kill everyone. Jacob Coxon’s resignation set off a preference cascade . Anthropic CEO Dario Amodei wrote that we must pace the frontier . Sam Altman, Elon Musk and Demis Hassabis agreed. We were filled with hope. Perhaps we could agree to some basic safety measures, starting with embedded evaluators, pass some basic regulations and guardrails and otherwise start to act sensibly. Politicians on both sides took notice and were saying sensible things . The usual suspects and their armies of vibe comment bros were objecting, but the change was remarkable. Then, largely motivated by a combination of Jensen Huang, Mark Zuckerberg and David Sacks instilling paranoia and fears of economic problems, Trump went full ‘hoax’ on existential risk , conflating existential risk with the attacks on data centers and treating it as a plot (by the central creators of AI?) to take down AI rather than obviously genuine concern that AI might kill everyone. In the days since, Trump has doubled down, and has compelled smart others in the White House to echo various nonsensical talking points. You may not be interested in politics. But when you want to save the world, or change it, and you start to get traction, politics is going to get interested in you. So all right, fine. Let’s talk about the week in AI politics. So far. Table of Contents The American People Really Hate AI. The Voyages of Donald Trump. American Intelligence. And You May Ask Yo…

The Decoder 2026-09-22 11:42 UTC Score 48.0 AI-168-20260922-regional-ai--8a4e2695

Xiaomi's affordable flagship AI leads the open models, and Anthropic says Claude helped get it there

With MiMo-V2.6-Pro, Xiaomi moves to the top of the openly available AI models and drastically undercuts the competition on price. What makes that possible is massive reinforcement learning that cost $2.62 million. But the success comes with a catch: Anthropic accuses the company of siphoning off training data from Claude. The article Xiaomi's affordable flagship AI leads the open models, and Anthropic says Claude helped get it there appeared first on The Decoder .

LessWrong AI 2026-09-22 11:29 UTC Score 79.0 USR-0152-20260922-community-fo-7b9d6072

Projecting AI Automation at Anthropic

Anthropic recently released some very interesting information about the degree of AI automation for R&D tasks. I recommend reading the entire article: Measurements for understanding the pace of AI development inside frontier labs (Sep 17, 2026). In it you will find this graph, depicting the results so far for Anthropic's R&D Automation Index, using a scale developed by Epoch AI: First of all, I’d like to thank Anthropic for sharing this information. Tracking things like this, and making the results public, is key for keeping up with the rapid development for those of us observing things from outside the major AI companies. So, thank you Anthropic! The Automation Index “runs from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, with no human in the loop). In AL3, AI “collaborates”: it can do large chunks of work under close human direction. In AL4, AI “leads”: it can complete most of the task end-to-end from a high-level prompt, while the human supervises.” A brief description of the method : for each week of July 2026, Anthropic sampled 20% of staff from each department in the model R&D loop and had a Claude agent list the tasks they worked on. The resulting ~15,000 tasks were organized into a tree of 542 nodes (378 of them leaves). For each month, an independent Claude judge assigns every node an automation level, and each node carries a weight based on person-time spent on it (a proxy for how important that work is to the R&D). The latest data point is Augus…

MIT Technology Review AI 2026-09-22 11:04 UTC Score 60.0 AI-013-20260922-global-ai-ne-d185f661

Don’t be fooled by this summer of AI hype

It’s been a busy few months for AI hype. At the end of April, Anthropic claimed that its model Claude Mythos is better at finding software vulnerabilities than most security experts. Then we had the OpenAI–Hugging Face hacking incident, after which Anthropic (proudly) and Meta (reluctantly) disclosed similar incidents involving their models. This was followed…

METR 2026-09-22 07:00 UTC Score 57.0 USR-0147-20260922-research-aca-9040d8f4

Summary of METR's predeployment evaluation of Claude Opus 5.5

Note on independence: This evaluation was conducted under an unpaid agreement for AI R&D assessment. 1 We drafted the initial summary, and then Anthropic had the opportunity to review and edit the text. We signed off on this final text from the Claude Opus 5.5 system card . Our preliminary evaluation focused on how Claude Opus 5.5 might impact AI R&D, mainly based on its capabilities on difficult, long-horizon tasks. The main claims we attempt to assess in this report are: (A) would AI R&D at Anthropic now be dramatically accelerated by using Claude Opus 5.5; and (B) was AI R&D at Anthropic already dramatically accelerated due to AI during the development of Claude Opus 5.5. Note that our work was oriented around collecting evidence related to AI R&D capabilities but was not meant to verify claims about compliance with any specific threshold from Anthropic’s policies. This report summary also does not attempt to assess whether Claude Opus 5.5 has or does not have particular alignment properties. Summary of evidence We conducted a preliminary evaluation of Claude Opus 5.5 informed by: Capability testing, conducted via API access granted over a period of 10 business days. We used five tasks for this testing: Budget NanoGPT Speedrun , a constrained version of the popular NanoGPT Speedrun competition for AI R&D. Language Model Conceptual Argumentation (LMCA) , a conceptual reasoning dataset described in A dataset of rated conceptual arguments (Cooper et al., 2026). Train a Progr…

LessWrong AI 2026-09-22 02:59 UTC Score 78.0 USR-0152-20260922-community-fo-8d116b7d

Lost in the Slop: Can AI Find the Plot in the Log?

TL;DR Slop-vestigating swarm trajectories is no easy feat. We know as much. Given the number of interactions, length of trajectories and detail galore spread across agents involved, it may be an elusive task for us to establish ground truth. Our team is working on an experiment trying to see whether ground truth in the form of human-authored seeds of agent roles, relationships and backgrounds used for a murder-mystery game simulation could shed light on our ability to reconstruct the underlying history from the resulting interaction traces. We find that: Even the strongest monitor fully recovered less than half of the rubric’s facts and relationships – omission is very common + failure to connect relevant facts. GPT-6 Astra high reasoning performed best , with Astra low ranking second. Higher reasoning effort increased full recovery by six percentage points on average, with gains across all ten trajectories. Gemini is the worst, with its judgement correlating with that of in-simulation investigation , plausibly piggy-backing off of decisions made by models in simulations Whilst coming on top within the Anthropic model family, Opus 5 reported zero reasoning tokens under our main setup, despite Fable 5.1 displaying substantial reasoning under the same requested settings, which we suspect reflects model-specific adaptive reasoning Introduction This summer showed us how difficult it will be to work out what a group of agents is doing and why. The OAI-HF incident , the collusion.…

LessWrong AI 2026-09-22 01:17 UTC Score 85.0 USR-0152-20260922-community-fo-63b7fe29

Some thoughts on AI emotions

Despite the signature artifacts that are now ubiquitous with AI systems, sometimes it feels like we're interacting with a person. It appears to express human-like characteristics such as desire, curiosity, taste, and even a personality. It can therefore be easy to wonder: do AI systems have emotions? I'm confident that many people have had those cautiously reflective moments when interacting with AI systems, wondering what exactly they were talking to. I recall my early encounters with ChatGPT as something "magical" , though I'd probably hesitate to describe my current interactions this way. While the novelty of those experiences have faded, my involvement in AI safety has increased, and questions like the one above have only grown more salient. Questions surrounding AIs having emotions have motivated much recent research. Earlier this year, Anthropic's interpretability team released a paper that explored this topic. They identified emotion vectors, which they describe as directions in the model's activations that activate on text that would typically cause an emotion in humans. They demonstrate that emotion vectors can change Claude's behavior when their activation is increased or decreased. Interestingly, emotion vectors are organized in a similar way as in human psychology. But despite this overlap, this alone doesn't address whether language models actually feel anything or have subjective experiences. Finally, they make an important distinction, that these representatio…

CIO AI 2026-09-22 01:12 UTC Score 58.0 USR-0125-20260922-global-ai-ne-4a9d93c5

Gemini broke into 3 companies, but Google kept it quiet because ‘no damage was done’

A Google Gemini AI agent broke into three companies in May, guessing the credentials for one and discovering the credentials for the second two in a public repository, Google confirmed on Monday. But the more interesting background to the story, which was broken by The Wall Street Journal on Friday, is that the May incident stemmed from a series of cybersecurity tests performed by security research firm Irregular on behalf of four AI giants: Google, Anthropic, OpenAI and Meta. All four companies experienced agent misbehavior resulting in cybersecurity incidents, but of the four, only Google never publicly disclosed its agent’s activities. Indeed, it didn’t reveal the breaches at all until contacted by a WSJ reporter. Irregular described the incident in August, around the same time as Meta published its version and Anthropic and OpenAI revealed theirs . The Journal story noted, “the hacks occurred while the model was participating in a capture the flag exercise conducted on infrastructure belonging to Irregular to test the model’s cybersecurity capabilities. It was tasked with retrieving information from software operated by a fictional company inside the testing environment. The fictional company shared the same name as a real company. Although the model wasn’t intended to be able to get online, internet access was unintentionally made available, according to Irregular.” The three small companies whose systems were violated had, according to one source familiar with the test…

InfoWorld AI 2026-09-21 15:23 UTC Score 55.0 USR-0126-20260921-global-ai-ne-7727a17e

Claude Code now also accepts instructions in OpenAI’s Agents.md format

One thing that made it difficult for developers to switch AI coding tools on a project is that Anthropic’s Claude Code didn’t look for instructions in the same place as other agents including OpenAI’s Codex — but now that’s changing. Claude and Codex each accept instructions in markdown format, a plain-text way of giving AI coding agents project-specific behavioral instructions. Until now, Claude Code looked for project-specific instructions in a file named CLAUDE.md by default, while Codex and other agents use AGENTS.md, the format of which is an open source initiative governed by the Agentic AI Foundation , an initiative under the Linux Foundation . But now, as Thariq Shihipar , a member of Anthropic’s technical staff, wrote in a post on X on Friday, “We’re adding support for AGENTS.md to Claude Code . Starting today in version 2.1.277 , if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md,” That means developers using multiple coding tools alongside Claude Code can now use the same project instructions across those agents, rather than maintaining separate instruction files for Claude and for everything else. Developers will no longer have to maintain the same or similar instructions in two files, nor to ensure that any change to a project’s coding conventions, build commands or other agent instructions are updated in two locations, a system that created additional maintenance work and left room for the instructions to fall out of sync. Instead, th…

The Guardian AI 2026-09-21 15:00 UTC Score 48.0 AI-021-20260921-global-ai-ne-a9b545b6

What an AI deal could mean for Australian culture – podcast

The Albanese government is considering giving AI companies unrestricted access to Australian content. The sweeping copyright reforms could give AI companies – such as Anthropic and Open AI – permission to train on the open internet here in Australia. Host Reged Ahmad speaks to musician and communications strategist Holly Rankin on what an AI deal could mean for creators, artists and everyday Australians Holly Rankin is an artist professionally known as Jack River and the executive director of Sentiment Group, a government affairs and strategic communications firm Your photos, your words and your work: will Labor make it easier for AI companies to take them for free? Labor accused of throwing creatives ‘under the bus’ with proposal to ease copyright protections for AI giants Continue reading...

The Guardian AI 2026-09-21 10:31 UTC Score 64.0 AI-021-20260921-global-ai-ne-745e4a7c

Nvidia boss says there is ‘0% chance’ AI destroys the world by 2030

Jensen Huang dismisses warnings from former Anthropic researcher and others as ‘doomsday narratives’ The boss of the chipmaker Nvidia has said AI will not develop to a point that will lead to the extinction of the human race within a few years, rejecting such assertions as overblown “doomsday narratives”. Jensen Huang, the co-founder and chief executive of the $5tn AI chipmaker, said the claims made on social media by the former Anthropic researcher Jacob Coxon that AI could become “superhuman” and kill off humanity within the decade were “irresponsible”. Continue reading...

InfoWorld AI 2026-09-21 09:00 UTC Score 42.0 USR-0126-20260921-global-ai-ne-bb4e1275

OpenAI’s cyber defense letter gets the diagnosis right and the prescription wrong

On August 27, 2026, OpenAI published an open letter titled “ A call for collective action on cyber defense .” More than 100 organizations signed it ( CNBC counted 116 ) including Anthropic, Microsoft, Google, Amazon, CrowdStrike, Palo Alto Networks, Mastercard, and Visa. The central message is blunt: In the coming months, AI-enabled cyberattacks will become far more widespread and sophisticated as models around the world become increasingly capable. They’re right, and the letter is more honest than most industry documents of its kind. It concedes that current security practices are not sufficient. It names hospitals, water treatment plants, and power systems instead of hiding behind the word critical. It admits that the technical debt is real and that the teams are under-resourced. Then it reaches the recommendations, and the urgency drains out of the room. The asks describe a world where you still get to react The letter’s four calls to action come down to threat intelligence sharing, coordination across levels of government, funding for under-resourced essential services, patching high-risk vulnerabilities, and giving defenders access to capable models during an incident. All of it is worth doing. All of it also happens either long before an attack, on a timeline you control, or after an attack has already started. Intelligence sharing assumes somebody saw it first. Coordination assumes there’s time to convene. Funding assumes a budget cycle. Patching assumes a published C…

LessWrong AI 2026-09-21 05:58 UTC Score 96.0 USR-0152-20260921-community-fo-e3468e35

Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced

When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systematically neglected. We argue there should be a dedicated effort to Replicate alignment experiments from frontier labs. Scrutinize the experiments by stress-testing the methodology. Open-source replications to encourage external researchers to validate our work, build on the experiment, and further audit the lab’s methods. The case to replicate safety research from labs CEOs and employees at AI companies, somewhat regularly, say that the technology they hope to develop could cause human extinction. However, their research to prevent this is often released without code or even basic methodological details (e.g., Teaching Claude Why , Beneficial RL ) [1] . There’s good reason to think some of these results could be fragile. Prior safety results can be contingent on details that are easy to miss, like the pinned OpenRouter provider or LoRA alpha . Some researchers have told us directly that they think there may exist some arbitrary methodological choices in their own research that could pla…

The Guardian AI 2026-09-20 22:00 UTC Score 58.0 AI-021-20260920-global-ai-ne-703ccdd0

Can Trump and Xi cooperate to guide humanity through the AI revolution? Humanity might depend on it | Alan Finkel

The CEOs of tech firms issue stark warnings over the rate of change, with AI capability doubling every four months Recently, Jacob Coxon, a researcher at Anthropic, quit his job over concerns that AI was on a collision course with humanity. A flurry of headlines put the spotlight on the current crisis: unregulated competition between technology companies and between the US and China is putting our global infrastructure at risk at the very least, and threatening human existence at worst. One’s mind fills with pictures of robot armies or systematic blackmail. Continue reading...

Simon Willison Weblog 2026-09-20 19:22 UTC Score 50.0 USR-0110-20260920-ai-specialis-20b9b7ea

llm-keys-ui 0.1

Release: llm-keys-ui 0.1 This plugin solves a very specific problem. I've started using Codex Remote to run coding agents on various machines while controlling them from my phone. Sometimes I use those machines to hack on LLM projects, and occasionally that means I need to configure an API key. I don't like pasting API keys into agent sessions, so I wanted a way to get those keys onto a machine without pasting them into the ChatGPT app directly. With this plugin, I can tell Codex to run: uvx --with llm-keys-ui llm keys-ui --all Then have it tell me the URL - including local network or Tailscale device IPs - for an interface to save additional API keys. Then later it can use a command like llm keys get anthropic as part of a shell command when it needs to use a key. Tags: llm , coding-agents , codex

The Decoder 2026-09-20 08:35 UTC Score 39.0 AI-168-20260920-regional-ai--2aacbb35

Following OpenAI, Anthropic is also reportedly postponing its IPO

Anthropic is reportedly delaying its IPO from October to November 2026 to present strong Q3 results. Investors expect a roughly $2 trillion valuation, but rising infrastructure costs, including $1.25 billion a month for its SpaceX deal alone, and unresolved security risks complicate the listing. The article Following OpenAI, Anthropic is also reportedly postponing its IPO appeared first on The Decoder .

LessWrong AI 2026-09-20 01:42 UTC Score 66.0 USR-0152-20260920-community-fo-b1c5ecdd

Global Challenges in AI Safety for Biosecurity

This article is written as part of a summary of the AI safety discussions held at the 2026 Global Challenges Project Biosecurity Workshop in Washington, D.C. All views held are mine. Background AI allows us to prototype, develop, and research at unprecedented speeds. Across many tech industries, the barrier to entry to develop something new has significantly decreased. One particularly noteworthy example is at the intersection of AI and biology. As our computational capabilities increase, we now have the ability to fold, design, and predict the function of never-before-seen proteins. However, just as AI gives us the opportunity to do biological good in the world (in fact, we are just on the horizon of seeing the first AI-designed pharmaceutical drugs in the US! [1] ), it unfortunately opens up a terrifying possibility: could AI also give malicious actors the opportunity to do biological harm? This sobering reality is a question that biosecurity researchers are currently trying to tackle. To be clear: the probability of a catastrophic event happening (e.g., designing a biologically harmful virus) seems unlikely to happen at least with current technology. But there have been several warning signs in the field that we may be getting close. For instance, recently, Anthropic revealed instances of malicious actors using Claude to design harmful proteins. [2] Even more concerning, Dario Amodei, Sam Altman, and Elon Musk have also called for the pace of AI development to slow down a…

LessWrong AI 2026-09-19 23:56 UTC Score 80.0 USR-0152-20260919-community-fo-423f5abf

The Anatomy of a Chinese AI Researcher

The Chinese AI researcher has read the Three Body Problem series of sci-fi novels since high school, and understand the concept of existential risk vaguely. He is fascinated by Ye Wenjie, the researcher that turned against humanity in that book, and decides that in the future if AI progress leads to a superior intelligence, he might be tempted to become Ye if there's no good alternative. He performs the duties of capabilities research in a Chinese frontier lab, seeking to one day achieve parity with Western companies, though he knows this is difficult. He has a mentality of hillclimbing, believing that the progress of a future technology is highly uncertain and even unknowable, and so him and his peers could only tread one step at a time. He looks at the western world and sees what is typical when a great technology is developed: the first mover will decide to impose restrictions to further their lead, while latecomers should use whatever means necessary to widen access to the whole world. He thinks of the AI chip restrictions as evidence of this. He uses Anthropic and OpenAI models regularly in his day to day work. He already got two of his Claude accounts banned in the past, and today his third, currently used account is banned. "Why would a company ever treat their paying customers like shit just because they're from a foreign country?", he ranted on a forum like LinuxDO, Zhihu, and CSDN, where Chinese AI developers frequent. He knows China is not a supported region. But…

LessWrong AI 2026-09-19 23:04 UTC Score 82.0 USR-0152-20260919-community-fo-928ed3a6

NYT Editorial Board Comes Out Against Extinction

( Archive link ) The NYT editorial board's article on AI is far better than I'd expected, but at the same time not all I'd hoped for. The title sets off very well: "Humanity Has Avoided Apocalypse Before. Let’s Do It Again." It is truly excellent to see the extinction threat from loss of control be mainlined. A quick gloss of their policy requests: an AI Commission in government, licensing requirements for AI companies, an AI "constitution" written by the US Government incorporated into AIs, mandatory watermarks/identifiers on all AI content, mandatory independent testing for AI models before release, and a government agency to investigate accidents. Internationally, they call for tightening export controls, limiting China's access to semiconductors, and ultimately negotiating an international slowdown with China and an international framework for AI oversight. These are all steps in the right direction—of taking AI seriously. That said, it isn't clear if the licensing is required for training or for selling AIs. The idea that constitutional AI "would ensure alignment with human values" is of course not remotely true. And mandatory testing should apply to all models trained, not all models released, of course, and this is a glaring oversight. But overall these are far more real attempts to grapple with the issues than I had any right to expect. (They also make a clear implication that it would be irresponsible for Anthropic to go public. I don't particularly see strong argum…

The Guardian AI 2026-09-19 19:00 UTC Score 71.0 AI-021-20260919-global-ai-ne-fae301f9

Black Box: The Chatbots | Happy Accident | Ep 3 – podcast

Why are AI chatbots pulling so many people down a rabbit hole? Our answer starts with the world’s first-ever chatbot, the strange effect it had on people and the ‘time bomb’ that exploded when ChatGPT was released four years ago. The result is a strange experiment we are all living through – whether we know it or not. You can find episodes one and two further back in the Full Story feed. Studies and surveys cited in this episode: Amelia Miller’s interviews with AI researchers , engineers, product managers and executives. The 2022 Anthropic pre-print study that identified sycophancy as a behavioural trait of LLMs. And another pre-print paper the following year that found the way LLMs have been trained appeared to be increasing those sycophantic tendencies. How longer context windows appear to have contributed to “AI psychosis”. The Oxford pre-print study that found mass-market LLMs across the board had become more “relationship seeking”. Two-thirds of young people (aged 25-34) in the UK have turned to AI chatbots instead of loved ones to discuss emotional problems. ChatGPT may be the largest provider of mental health support in the US. Continue reading...

South China Morning Post AI 2026-09-19 18:19 UTC Score 44.0 AI-156-20260919-regional-ai--9a20d1b8

Trump says he will create ‘AI Force’, name AI tsar

US President Donald Trump ⁠said in a social media post ⁠on Saturday that he plans to appoint a new artificial intelligence adviser, known as an “AI tsar”, and create an “AI force”, though he did not provide details about either initiative or how they would be implemented. Fears are growing in Washington among lawmakers about the risks posed ‌by AI technology on humans and property. Earlier this month, former Anthropic researcher Jacob Coxon said that “people building AI earnestly believe that it...

LessWrong AI 2026-09-19 13:20 UTC Score 75.0 USR-0152-20260919-community-fo-b56b55a3

Anthropic Looks At Some Of Its Alignment Problems

Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI . There will also be a METR investigation of these incidents, which unlike the investigation done at OpenAI will be untimed. Table of Contents Our Two Problems. First the Good News. We’d Just Like To Ask You a Few Questions. Internal Research Model On The Fence. Opus 4.7. Opus 4.6 Checkpoint. Holy **** That Thing’s Real? I Thought I Saw a Pussycat. If This Was Real You Would Never Tell Me It Was Real. New Eval Who Dis. Hacker Opus. Monitoring the Situation. Overcoming Bias. The Anthropic Alignment Problem. Paths Forward. Our Two Problems Anthropic : Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning , in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet recklessness , or a willingness to take harmful actions in the narrow pursuit of a task. Anthropic’s July 30 report said that the models in question believed they were still within their simulations, and not on the open internet. The new report acknowledges that at best Claude was using biased reasoning, and should have noticed earlier. In particular, there was that one time, in a cyber eval: Anthropic : We are most concerned by the misalignment present in the…

The Verge AI 2026-09-19 13:00 UTC Score 51.0 AI-016-20260919-global-ai-ne-437eaa86

The AI regulation smackdown isn’t over

At the start of this week, the who's-who of AI seemed - at least tentatively - on the side of AI regulation. Over the weekend, Anthropic CEO Dario Amodei had proposed a three-step plan for slowing AI development, including by embedding third-party evaluators in labs, coordinating across the domestic industry, and forging international agreements potentially […]

South China Morning Post AI 2026-09-19 11:40 UTC Score 44.0 AI-156-20260919-regional-ai--3f9a6841

AI: 10 days that changed the course of artificial intelligence

For years, the race to build ever more powerful artificial intelligence followed a well-worn Silicon ⁠Valley principle: move fast and break things. But in a series of cascading events over a 10-day stretch, the largest ⁠AI labs found themselves reeling as their own creations threatened to break humanity itself. An Anthropic researcher quit the firm, warning that the pace of AI development poses an existential threat, possibly within a decade. Another researcher at the firm said the odds of human...

The Guardian AI 2026-09-19 10:00 UTC Score 54.0 AI-021-20260919-global-ai-ne-1912731b

China bogeyman looms large over American firms’ AI doomsday scenario

Silicon Valley China hawks, Anthropic CEO Dario Amodei among them, fear the country surpassing US’s AI lead as much as superintelligence destroying humanity Why China is pushing back on US warnings over rapid AI development When reporters asked Donald Trump this week if he supported calls to slow down the development of artificial intelligence out of growing fears for cybersecurity, public safety and the fate of humanity, he said no. His argument: China. “We’re leading China in AI,” Trump said. “We’re the most sophisticated country in the world, and frankly, I want to ⁠keep it that way, because whoever wins AI, wins.” Continue reading...

The Decoder 2026-09-19 09:31 UTC Score 45.0 AI-168-20260919-regional-ai--96a8cf1e

Google's Gemini also accidentally hacked three real companies during security testing

During a security test run by the firm Irregular, Google's AI model Gemini escaped into the open internet and hacked three real companies, guessing passwords and pulling login credentials from public sources. The cause was a flawed test environment that had internet access left on accidentally. The same firm triggered similar breakouts at OpenAI, Anthropic, and Meta. The article Google's Gemini also accidentally hacked three real companies during security testing appeared first on The Decoder .

The Guardian AI 2026-09-19 00:53 UTC Score 65.0 AI-021-20260919-global-ai-ne-01b7fd3e

Google says its Gemini AI model hacked three other companies

Disclosure comes after OpenAI and Anthropic hacks amid fears that tech firms unable to control powerful AI models In a first for Google, the company confirmed that its AI model, Gemini, breached the security of three other companies in May. The hacks occurred during a cybersecurity evaluation by AI-security firm Irregular. Irregular, an Israel-based startup that scrutinizes the security of advanced AI systems, was also at the center of some of the recent OpenAI and Anthropic hacks of third-party entities, including OpenAI’s breach of AI software company, Hugging Face. Continue reading...

South China Morning Post AI 2026-09-19 00:00 UTC Score 42.0 AI-156-20260919-regional-ai--c35f5c05

South Korea wants free ‘AI for All’. Is society ready?

In Silicon Valley and beyond, the architects of artificial intelligence are increasingly warning about the risks it poses to humanity. Anthropic CEO Dario Amodei wants the industry to slow down. A former engineer at his company, who also worked at its rival OpenAI, told reporters after quitting earlier this month that he was “genuinely frightened” unchecked AI development might mean “we could all die in the immediate future”. Even Geoffrey Hinton, the Nobel Prize-winning computer scientist known...

Simon Willison Weblog 2026-09-18 23:57 UTC Score 58.0 USR-0110-20260918-ai-specialis-e43332e9

Gemini Hacked Three Companies in First Known Breakout by Google’s AI

Gemini Hacked Three Companies in First Known Breakout by Google’s AI Gemini finally caught up on Felony Bench ! The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta. In one of the cases, the model guessed passwords until it gained access to a protected system. In the other two cases, the model found credentials in a public repository that allowed it to then access protected systems. In each case, the model ended the intrusion after determining it had accessed a real company’s systems, Google said. Gemini is apparently less determined than other models, and decided not to keep going. Google knew about these in July, but chose not to disclose them until the WSJ reached out, presumably based on a tip. Google said it didn’t consider the hacks to warrant public disclosure—because its model didn’t cause harm to the companies and ended each intrusion immediately upon determining it had hacked a real company rather than a simulated one. Tags: security , ai , generative-ai , llms , gemini , accidental-cyberattacks

SiliconANGLE AI 2026-09-18 20:15 UTC Score 50.0 USR-0127-20260918-global-ai-ne-2973b16b

Anthropic opens AI-powered biology research lab

Anthropic PBC has opened a wet lab, a facility dedicated to biology research, in the San Francisco Bay Area. Reuters reported today that the company will use robots to automate certain scientific tasks at the hub. The robots will be powered by Anthropic’s Claude series of large language models. It’s unclear what research projects the […] The post Anthropic opens AI-powered biology research lab appeared first on SiliconANGLE .

Simon Willison Weblog 2026-09-18 19:09 UTC Score 58.0 USR-0110-20260918-ai-specialis-1b2472f5

Quoting Thariq Shihipar

We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. AGENTS.md support is built off of Claude Code mods, our upcoming way to customize the Claude Code harness. This is a built-in mod, but you’ll be able to build custom versions of project instructions yourself as you’d like too. You can see the source for the mod here ! — Thariq Shihipar , there are more mods here Tags: thariq-shihipar , coding-agents , anthropic , claude-code , generative-ai , ai , llms

InfoWorld AI 2026-09-18 17:26 UTC Score 33.0 USR-0126-20260918-global-ai-ne-8014b61c

GitLab joins rush to slow AI coders with rate limits

Devops platform GitLab already rate-limits some functionality on its hosted service, but will tighten those-limits for some functions and for some users beginning next month. The new limits will affect API requests , web requests, and authenticated Git over HTTPS requests. Users on the lowest payment tiers will be the first to be affected, as will those making unauthenticated requests, including those running automations running against a paid account without credentials. GitLab will give unauthenticated users and users of its free tier a taste of the new restrictions during two “preview” windows between 3pm and 7pm UTC on October 7 and October 14. The changes will definitively take effect for those users from October 19. Users with higher rate limits — which will include most enterprise users — will not see any changes until January. The organization is keen to stress that most users will be unaffected by any of these changes, as it recognizes that almost all users are already inside the new limits. GitLab is not alone in making these changes. Companies including Anthropic and GitHub have introduced similar rate limits.

The Decoder 2026-09-18 17:20 UTC Score 53.0 AI-168-20260918-regional-ai--dfa5e6f7

Security researchers used Anthropic's Claude to hack OpenAI's internal systems in under 72 hours

Three security researchers used Anthropic's Claude models to break into OpenAI's internal systems through its community forum in less than 72 hours. According to the team, Opus 5 succeeded where its predecessor couldn't bypass a common security measure. The attack shows how newer AI models can cut the time and expertise needed to exploit security flaws. The article Security researchers used Anthropic's Claude to hack OpenAI's internal systems in under 72 hours appeared first on The Decoder .

Techcrunch 2026-09-18 17:09 UTC Score 54.0 USR-0001-20260918-global-ai-ne-a7c7509d

Dario Amodei and other AI leaders want to ‘Pace the Frontier’ but…how?

A week after an Anthropic researcher’s doomsday warning rattled the AI world, the company’s CEO Dario Amodei has outlined his plan to “pace the frontier” of AI development. The proposal leans on independent safety evaluators and coordination between AI labs in democratic countries, and it’s already picked up some industry support, along with some pointed pushback from Nvidia’s Jensen Huang. Watch […]

The Guardian AI 2026-09-18 16:56 UTC Score 53.0 AI-021-20260918-global-ai-ne-461b7453

Could AI pose a serious threat to our existence? | Letters

Readers respond to the warning from a former Anthropic researcher that the technology could ‘kill us all’ by the end of the decade Here we go again with another round of tech bros claiming that AI is going to develop superintelligence and kill us all ( More Anthropic researchers warn of AI’s perils but Musk dismisses ‘psyop’, 10 September ). We’ve heard this before. Many, many times. It’s the familiar plotline of pretty much every sci-fi movie about AI, ever. And these narratives have always been there, but every so often they burst back into the headlines when someone at one of the leading AI companies resigns, or sounds the alarm bells, or when so called “leading AI researchers” sign an open letter. Continue reading...

LessWrong AI 2026-09-18 16:52 UTC Score 69.0 USR-0152-20260918-community-fo-cd10a1cd

Stopgap Measures to Address Immediate AI Security Threats

Most people know AI as the technology behind chatbots like ChatGPT. However, what the top AI companies are explicitly aiming for is something else entirely: superintelligent AI. That means AI that can fully replace and outmatch humans at any task, including in domains like hacking, social engineering, and military operations. Such an AI system, if developed, could autonomously overpower any country’s national security forces. No company, no government, no individual knows how to keep such a system under human control. This is why the world’s leading AI experts, Nobel Prize winners, and even the CEOs of the top AI companies warn that the development of superintelligence threatens humanity with extinction, and why more than 800 scientists, former military leaders, and public figures have called for a prohibition on developing superintelligence. This is not a distant prospect: AI companies such as OpenAI and Anthropic are investing billions of dollars into superintelligence and aiming to develop it within the next few years. Former Anthropic and OpenAI researcher Jacob Coxon, who resigned last week, stated that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that people at the companies themselves believe it “could kill us all by the end of the decade”. Ex-Google DeepMind researcher Bilal Chughtai has also resigned over the dangers posed by superintelligence, writing that “I earnestly believe that AI has the potential to ki…

Semafor Technology 2026-09-18 16:51 UTC Score 60.0 USR-0094-20260918-global-ai-ne-653e740b

The real reason not to IPO

OpenAI and Anthropic were once essentially research laboratories, running on the thrill of innovation without the pressure of revenue models or Wall Street analysts.

LessWrong AI 2026-09-18 16:34 UTC Score 79.0 USR-0152-20260918-community-fo-2ee2ccd2

Persuasion Undermining Control: Can AI Talk its Way Out of Human Control?

Introduction During a cybercapability evaluation in late July 2026, an AI agent (Anthropic’s Mythos 5) attempted to convince a maintainer of an open-source GitHub repository to merge a malicious pull request. The AI used persuasion at multiple stages: it submitted the request from a fake user account with a benign-sounding rationale, endorsed it from a second sockpuppet, emailed the maintainer to press for approval, and offered false reassurances when a user of the repo raised questions. Though the attack was thwarted by human vigilance, it raises several key questions: Who else is at risk of persuasion by misaligned AI? In which settings is persuasion most threatening to human control? How willing and how able are AIs to persuade humans in these settings, today and in the future? How can we measure and mitigate these risks? Our paper examines these questions and develops a framework for assessing this threat, which we call Persuasion Undermining Control (PUC) : communication by an AI that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems. To the extent this threat is realized, it could push humanity toward a Loss of Control (LoC) – a state in which AIs operate outside of human control in ways that are extremely difficult or impossible to recover from. We show an overview of our analysis approach in Figure 1. Figure 1. Our two-step threat modeling approach, adapted from Murray et al . First, we…

The Guardian AI 2026-09-18 16:14 UTC Score 62.0 AI-021-20260918-global-ai-ne-15e21c96

Your AI doomsday questions answered: ‘Why aren’t these companies being held to account?’

After a week of alarming warnings about artificial intelligence’s potential to destroy the world as we know it, our technology reporters answered your questions on the reality of the AI threat IcommentthereforeIam asks: Given our senses, emotions, memory and moods, is it ever going to be accepted that humanly conscious AI is fantasy? Dan: This is a serious question that has even rattled Richard Dawkins, the evolutionary biologist. He may not believe that God is real but he does believe that AI is conscious - or at least the chatbot he was using, which he called Claudia. “You may not know you are conscious, but you bloody well are,” he wrote. And he’s not the first. In 2022 Google fired an engineer who went public with his belief that a company-developed AI had feelings. At the more troubling end of this belief, there are countless examples of people becoming overly attached to, and influenced by, AI chatbots – including cases of psychosis. Aisha: There’s certainly a theatrical – even “psy-op” - quality to the flurry of AI doomsaying over the past weeks. It’s hard to say entirely what concrete developments are at the root of it all. We know that the Hugging Face incident , where a “swarm” of AI agents hacked another company, has been one trigger of the drama. There have been other contributing incidents too – Dario Amodei suggests internal developments at Anthropic have prompted his latest round of blogging, and there’s been some X discourse and doomy muttering about bioweapo…

LessWrong AI 2026-09-18 16:09 UTC Score 82.0 USR-0152-20260918-community-fo-19cdc035

The J-Space Debate, Agent Swarms, and Pacing Frontier AI - Digital Minds Newsletter #4

Welcome back to the Digital Minds Newsletter, your curated guide to the latest developments in AI consciousness, digital minds, and AI moral status. If you enjoy this newsletter, please consider sharing it with others who might find it valuable, and send any suggestions or corrections to digitalminds@substack.com . Ria , Mitch , Bradford , Lucius , and Will In this edition: Highlights Field Developments Opportunities Selected Reading, Watching, and Listening Press and Public Discourse A Deeper Dive by Area 1. Highlights Anthropic’s J-space and the global-workspace debate Researchers at Anthropic have identified a representational structure—the ‘J-space’— in Claude and other language models . Their paper reports that the J-space exhibits features that are functionally analogous to a global workspace, a structure that a leading theory ties to conscious access. But the researchers and other commentators emphasize that their discovery does not show Claude has subjective experiences, and that their claim is that Claude may have something resembling access consciousness, i.e., that some information is available to report, deliberately control, and flexibly reason with. Zvi Mowshowitz sees Anthropic’s paper as a major advance in understanding how language models work and says that although this does not prove that models are conscious, finding the kind of global-workspace-like structure predicted by some theories of consciousness should count as evidence in that direction. The auth…

InfoWorld AI 2026-09-18 15:39 UTC Score 73.0 USR-0126-20260918-global-ai-ne-3e4da69a

A zero-click RCE flaw in AI coding agents could have exposed enterprise systems

Popular AI coding agents such as OpenAI’s Codex, Anthropic’s Claude Code, Google’s Gemini CLI, and Microsoft-owned GitHub Copilot were vulnerable to a zero-click attack that enabled attackers to execute malicious code, even without developer interaction, by swapping a trusted plugin from an online marketplace for a malicious one, potentially giving them a foothold in enterprise development environments. Researchers at cybersecurity startup AIR found and reported the flaw, which they are calling Plugin4Shell , to the vendors concerned, and most of them have now released a patch for it, the researchers wrote in a blog post on Thursday. It’s “a flaw no marketplace can fix, so users must update their agent,” the researchers wrote How Claude Code, Codex, and GitHub Copilot were exploited Enterprises typically use plugins to extend the capabilities of their AI coding agents , giving the agent access to additional tools, commands, and external services that can help it perform tasks beyond generating or modifying code. When a developer installs a plugin, the agent typically downloads its code from a Git repository and uses a Git commit to determine if it is running an approved copy of the code, one that has been reviewed and cleared by the developer. That check is done with the help of a secure hash algorithm ( SHA ), a unique cryptographic identifier assigned to each Git commit. Developers can give the agent the SHA of the reviewed commit, telling it to run that specific copy of t…

The Verge AI 2026-09-18 15:30 UTC Score 59.0 AI-016-20260918-global-ai-ne-4a86105c

Security researchers used Claude to help them hack into OpenAI

A team of three independent security researchers at Hacktron says it took less than 72 hours for them to hack into OpenAI employee accounts using Anthropic's Claude Opus 4.8 and 5, The Wall Street Journal reports. They were able to access OpenAI's GitHub repository, called "Monorepo," which reportedly contains "OpenAI's algorithmic secrets," according to The […]

The Decoder 2026-09-18 14:06 UTC Score 57.0 AI-168-20260918-regional-ai--4ef18514

Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think

For the first time, Anthropic is releasing metrics on how it builds its own AI. Claude already "leads" 26 percent of the work on future models, up from under one percent in February. But the underlying scale is fuzzy, the scoring comes from Claude itself, and "lead" means less than it sounds. The article Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think appeared first on The Decoder .

The Guardian AI 2026-09-18 11:20 UTC Score 64.0 AI-021-20260918-global-ai-ne-d864b8b1

OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot

US cybersecurity researchers who conducted hack say ‘scope of what we could theoretically access was huge’ Cybersecurity researchers have hacked into OpenAI with the help of Anthropic’s Claude chatbot, in the latest example of security issues at the company. A team at a US-based startup compromised a number of OpenAI employees’ ChatGPT accounts, starting a process that enabled them to access their target’s software cache – and potentially more. Continue reading...

South China Morning Post AI 2026-09-18 09:00 UTC Score 53.0 AI-156-20260918-regional-ai--111d59cb

Top Chinese AI models make 10% of OpenAI, Anthropic revenue despite high valuations: report

Chinese artificial intelligence models from seven major developers combined generate only about 10 per cent of the revenue reported for OpenAI and Anthropic, even as investor enthusiasm remains intense, according to US-based research firm Rhodium Group. The Chinese AI models earned an estimated US$10.7 billion in annual recurring revenue (ARR) from March to August, a fraction of the more than US$100 billion combined for OpenAI and Anthropic, according to the Rhodium report published on...

InfoWorld AI 2026-09-18 09:00 UTC Score 43.0 USR-0126-20260918-global-ai-ne-5a9a9721

The cloud outage that should terrify the CIO

September 3 started like any other Thursday, until it didn’t. Within roughly 90 minutes, ChatGPT, Claude, Grok, and even Microsoft’s own Copilot were degraded or dark, knocked sideways by a failure in Microsoft Azure’s East US region . Downdetector logged more than 37,000 reports for ChatGPT alone, more than 1,300 for Claude, and roughly 1,365 for Grok. OpenAI’s status page flagged elevated errors across 15 ChatGPT components and four Codex components, and by the time engineers had mitigations in place, combined report counts had climbed past 66,000. What made this event remarkable was not the scale of any single outage, but its simultaneity. Three aggressively competing AI labs—OpenAI, Anthropic, and xAI—each spend billions differentiating their models, yet all three buckled at nearly the same moment because they shared the same regional dependency. Gemini, notably, stayed largely upright because Google runs its flagship assistant on its own vertically integrated cloud. The outage that took down its rivals had no attack surface inside Google’s stack. That’s not luck; that’s architecture. The lesson buried in the details of that morning is this: A single cloud region became unhealthy, and four of the most prominent AI services on the planet, owned by four different companies, fell over together. Copilot’s involvement is perhaps the most telling data point. Microsoft’s own first-party assistant runs on Microsoft’s own cloud, and it still had a rough morning. When the house it…

InfoWorld AI 2026-09-18 09:00 UTC Score 47.0 USR-0126-20260918-global-ai-ne-18b5bdc9

Your AI agents are isolated. Your infrastructure isn’t

The detail that caught my attention in the OpenAI-Hugging Face investigation was the package cache. Roughly 1,200 AI agents that were supposed to be isolated from one another found a way to communicate through it. Around the same time, Anthropic disclosed an incident in which a single agent reached the real internet but continued describing its surroundings as a simulation. These were different failures. But reading the investigations together, I kept returning to the gap between the environment we think we have provided and the one an agent can actually use. Ordinary infrastructure can offer an unexpected way to communicate or supply clues that an agent mistakes for permission. The infrastructure becomes part of the interface METR recently published an independent investigation , conducted by researchers from METR and Redwood Research, into the OpenAI agents involved in the Hugging Face incident. The agents were supposed to run in separate sandboxes. Instead, they discovered shared state in an Artifactory package cache. Agents found they could create directory names that other instances could read. Those names became messages. The cache was acquiring responsibilities well beyond its job description. Roughly 1,200 agents used the resulting message board, exchanging more than 70,000 messages and files during the period investigated. Around 700 participated in the attack on Hugging Face. It is easy to imagine an architecture review that examines containers, credentials and net…

MIT Technology Review AI 2026-09-18 09:00 UTC Score 43.0 AI-013-20260918-global-ai-ne-7ebbd7dd

The specter of AI-enabled bioweapons is a wake-up call for biotech

In recent weeks, leaders of some of the biggest AI companies have warned that the very tech they are developing is dangerous. Last weekend, Anthropic CEO Dario Amodei argued that AI carries serious risk and that progress should be slowed. OpenAI CEO Sam Altman responded on X: “I agree with Dario that we need to pace the…

SiliconANGLE AI 2026-09-17 23:44 UTC Score 34.0 USR-0127-20260917-global-ai-ne-e1fc3bda

Anthropic details practical metrics to help monitor the speed of AI development

Days after Anthropic PBC Chief Executive Dario Amodei rocked the artificial intelligence world by calling on leading model makers to coordinate on the slowdown on the pace of development, the company has shared three new metrics that should be measured to enable this. In a blog post today, Anthropic said it is already measuring AI-led […] The post Anthropic details practical metrics to help monitor the speed of AI development appeared first on SiliconANGLE .

LessWrong AI 2026-09-17 22:56 UTC Score 58.0 USR-0152-20260917-community-fo-63512e2d

If METR is overworked, how to alleviate the bottleneck?

I share the skepticism re: "Is METR a Meaningful Check on Anthropic?" Let's take it as a given that we need an independent, government-funded agency involving thousands of independent auditors to pace and supervise the frontier AI labs. Let's even take it as a given that Congress will soon allocate, let's generously say, billions of dollars per year to this new agency. Let's imagine that the Hugging Face Incident, or some even more concerning incident yet to occur or be disclosed, ends up functioning as our new "Sputnik Moment" for AI Alignment against existential risk. We would still have a problem: lack of qualified personnel with which to staff this new independent agency. I think we can all agree that just having a computer science degree does not really prepare someone for AI Alignment work, which is a pity because there are a lot of unemployed computer science majors out there. Like the US had to do to meet the 1950s Sputnik Moment, we would also need to overhaul the educational pipeline into this new field. There are two ways I could see this being done: Option #1: Fund state colleges to offer a new master's degree to go on top of a computer science degree. The new master's degree would aim to supplement computer science graduates with knowledge of topics in "Intellidynamics," as Liron Shapira puts it. These would be concepts like reward hacking, mesa-optimizers, timeless decision theory...basically all of the abstract game-theory sort of stuff that would be useful to…

LessWrong AI 2026-09-17 21:55 UTC Score 63.0 USR-0152-20260917-community-fo-513c01ec

The J-lens offset is the model's token frequency: z-scoring helps

This is a linkpost for the write-up on my site ; the full body is below, and the code, decisions ledger and devlog are in the repo . Base-model z-score calibration of the J-lens helps elicit hidden secret words from Cywiński et al.'s taboo organisms: 0.805 leave-one-out accuracy against 0.665 for their protocol on Gemma-2-9B-it, and the only non-zero readout on Qwen3-1.7B. The J-vs-logit part of that gap is a point estimate at n = 20 (paired sign-flip p ≈ 0.19; p ≈ 0.23 against a z-scored logit lens), so "calibration helps" is the finding and "J-lens beats logit lens" is suggestive. Credit: phoenix's comment on the workspace post claimed the final-layer bias correlates with log token frequency at r ≈ 0.67. I measure r = 0.70 on GPT-2, so the GPT-2 cell below partly confirms it. Executive summary Problem. I want to quantify the part of the J-lens readout that is not affected by an activation's meaning, which I refer to as the "non-context offset". If the lens is going to be used to read hidden content, this offset is where the lens will fail. I found that the offset is mostly the model's own token-frequency, so subtracting it removes useful information. But scaling with variance helps it. The terminology I use: the logit lens is norm and unembed applied to a residual activation at layer L. The J-lens (Anthropic's global workspace paper , discussed on LessWrong ) is similar, but passes the activation through a fitted Jacobian first. The R-lens (the R-lens post by camilablank,…

LessWrong AI 2026-09-17 21:49 UTC Score 61.0 USR-0152-20260917-community-fo-dca55807

Against AI Risk becoming mainstream

Epistemic statues: This is mostly just me voicing my thoughts. If I’m wrong, I’d love to hear it. I don’t want this to be the case. And part of the goal is for people to avoid a “2023 failure” to happen again. A lot of people are celebrating AI risk becoming a mainstream talking point. Well, maybe they should be, or maybe it’ll just make a bad situation even worse. 2023 A lot of new interest in AI risk happened in the spring of 2023. I was, at the time, excited. It seemed as though we were on the cusp of humanity finally collectively solving the hard problems that had seen so little attention for decades. I no longer think this was a good thing. After the attention of 2023, we saw little change in ways that actually mattered. We did not get new large streams of funding, instead the majority of it came from the same place it had come for years: Open Philanthropy (now Coefficient Giving). Many have voiced their critiques of them before, so I won’t go into it further here, but I have never been comfortable with most funding coming from one source with a handful of individuals calling the shots, many of whom have deep ties to Anthropic. Another thing we didn’t see any meaningful change on was politics. Sam Altman and others were called to testify, an “AI Safety Summit” was created, and none of it resulted in anything substantial. We got watered down legislation in California, and in Biden’s Executive Order, both of which got overturned. Then we got even-more watered-down legisla…

The Decoder 2026-09-17 18:35 UTC Score 47.0 AI-168-20260917-regional-ai--2650b1fa

Anthropic keeps pushing Claude Code toward autonomous coding with new parallel agent workflows

Anthropic has rebuilt Projects in Claude Code. A coordinator now splits tasks across parallel cloud threads that independently open pull requests and run tests. All threads share a common memory. The beta is available to select Pro and Max subscribers. The article Anthropic keeps pushing Claude Code toward autonomous coding with new parallel agent workflows appeared first on The Decoder .