AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
34781News Items
8Top Picks
202Blogs
runningLast Run

Anthropic

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
Semafor Technology 2026-08-13 22:46 UTC Score 61.0 USR-0094-20260813-global-ai-ne-e853cc98

China's AI ecosystem gears up to challenge US

ByteDance is reportedly developing an AI model rivaling Anthropic’s most advanced Mythos system, and DeepSeek is touting its efforts to build a Claude Code challenger.

InfoWorld AI 2026-08-13 17:37 UTC Score 70.0 USR-0126-20260813-global-ai-ne-d7c27c77

Visual Studio Code 1.133 brings flexibility to Claude sessions

Microsoft has released Visual Studio Code 1.133, an update to its code editor that brings more flexibility to Anthropic Claude sessions. The update also allows users to open the Agents window without signing in to GitHub. VS Code 1.133 was released August 12 . It can be downloaded for Windows, Linux, and Mac from code.visualstudio.com . With VS Code 1.133, users now can mix Anthropic and Copilot providers in Claude sessions. Previously, a Claude session ran entirely through either a GitHub Copilot subscription or Claude’s existing configuration, such as an API key. Switching providers required reconfiguring the agent host. Now, the model picker displays both groups, so users can switch providers between turns. The model selected is used for the next turn. Models under Anthropic bill the API key, and models under Copilot use the Copilot subscription. A new experimental setting, chat.agentHost.allowSignedOutWhenUsable , allows the Agents window to be opened without requiring GitHub sign-in. Previously, the Agents window opened with a GitHub sign-in prompt that could not be dismissed, blocking users whose machine could not reach github.com and users who do not interact with GitHub. Enabling this setting associates GitHub authentication with individual agents or models instead of the Agents window. In this release, this behavior only supports Claude. Support for Copilot with the user’s own model keys and Codex is planned for future releases. VS Code 1.133 also brings auto-reload…

The Decoder 2026-08-13 10:46 UTC Score 46.0 AI-168-20260813-regional-ai--56228010

Fable 5's slow adoption suggests corporate willingness to pay for frontier AI has hit a ceiling

Anthropic's Fable 5 is considered the most powerful AI model on the market, but U.S. companies are barely buying it. According to Ramp data, Fable 5 accounts for only six percent of Anthropic tokens sold. The model's steep price tag suggests corporate AI spending may have hit a ceiling, at least as long as performance gains don't translate into measurable everyday value. The article Fable 5's slow adoption suggests corporate willingness to pay for frontier AI has hit a ceiling appeared first on The Decoder .

The Decoder 2026-08-13 10:42 UTC Score 57.0 AI-168-20260813-regional-ai--30e4ea0b

Top AI lab researchers warned about automated AI research, and several of their predicted milestones have already fallen

IAPS fellow Severin Field interviewed 25 researchers from OpenAI, Anthropic, Google Deepmind, Meta, and US universities about recursive self-improvement. In a new blog post, he takes stock. Several of the milestones those researchers named have already been hit. The article Top AI lab researchers warned about automated AI research, and several of their predicted milestones have already fallen appeared first on The Decoder .

OpenAI Community 2026-08-13 02:26 UTC Score 37.0 AI-116-20260813-social-media-b7c46656

Codex rate limits reset for all paid plans on August 9 and again on Monday

The little surprise was another reset. x.com Tibo @thsottiaux Old news actually from a bunch of days ago, but crossed that 15M. Enjoy a nice reset everyone. Landing in the next hour or so, go /fast. x.com/thsottiaux/sta… Tibo @thsottiaux I previously promised a reset for every 1M in additional active users for Codex, until 10M. We blew past that and have been silent since 10M. Little surprise for you tomorrow. 1:01 AM - 13 Aug 2026 5.9K 390

SiliconANGLE AI 2026-08-13 01:27 UTC Score 58.0 USR-0127-20260813-global-ai-ne-93e92f90

SpaceXAI releases flagship Grok 4.6 model with advanced reasoning capabilities

SpaceXAI today released Grok 4.6, a large language model that it says can outperform Anthropic PBC’s Claude Fable 5 in some areas. SpaceXAI was known as xAI until last month. The Elon Musk-founded artificial intelligence provider rebranded in connection with its acquisition by SpaceX Corp. In June, the combined company listed its shares on the […] The post SpaceXAI releases flagship Grok 4.6 model with advanced reasoning capabilities appeared first on SiliconANGLE .

CIO AI 2026-08-13 00:52 UTC Score 58.0 USR-0125-20260813-global-ai-ne-32b5a4de

What vibe-coding startup valuations portend for CIOs

Investor appetite for the burgeoning vibe-coding startup ecosystem has shown few signs of satiation over the past year plus, with Swedish AI upstart Lovable’s Series C injection at a $13.3B valuation the latest evidence of a sector viewed by venture capitalists as one of AI’s most promising business disruptors. AI-assisted coding has proved to be AI’s most compelling — and commercially viable — enterprise use case to date. Developer-aimed tools such as Cursor, which sold to SpaceX in June for $60B , and Windsurf, which last year entered a $3B OpenAI dalliance before its eventual talent flight to Google DeepMind for $2.4B , have become — along with Anthropic’s Claude Code — well established in enterprise arsenals for accelerating developer output. But another set of vibe-coding tools, represented by the likes of Lovable and Replit, which hit a $9B valuation in March , seeks to ride the same path into the enterprise that no-code/low-code tools did previously: through your business users. These tools are built to democratize application development, giving users an AI chat interface to converse their way to enterprise-ready prototypes with fairly polished UIs, as CIO.com’s Peter Wayner writes in his roundup of the leading tools the space . Some IT leaders are already enlisting business users to vibe-code their own apps . Scott Weller, CTO at financial services technology provider EnFi, in May told CIO.com’s Bob Violino, “The results have surprised us. What started as an enginee…

OpenAI Community 2026-08-12 21:18 UTC Score 50.0 AI-116-20260812-social-media-8840924c

Constant cybersecurity false positives and unable to sign-up for cyber

I’ve been constantly running into false positive issues lately with Codex, but today is particularly bad. I’ve been running an automated (general) review of a fairly large project that we started working on before ChatGPT was even released. I hadn’t even asked it to run a security audit (though that would be ideal). It is frequently stopping work because it runs into the automated classifiers which trigger false positives. I’ve been able to get codex to verify that I’m the majority shareholder of the company via ISED, that we own the trademarks in ~3 dozen countries for the software we’re working on via WIPO, that the source code and all proprietary libraries (written by us as well) are available on my system, the website source code is available on my system, and that information about me owning the software/company are available on reputable websites. It should have no issue verifying that a general audit is safe, but it still runs into false positives with the classifiers. We can’t even sign-up for basic TAC because we receive this message: “Your identity couldn’t be verified or your account is ineligible at this time”. I reached out to support and they said there’s nothing they can do. I suspect either account age or an abundance of false positives have contributed to this, but it’s truly just speculation on my part. Conversely, Anthropic had no issue giving us significantly extended permissions in their CVP program beyond what their CVP would normally offer and let us v…

The Decoder 2026-08-12 18:33 UTC Score 53.0 AI-168-20260812-regional-ai--c438dff9

SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price

xAI's Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5. On agentic tasks, it completes complex workflows in about 53 steps where Claude Opus 5 needs 103, at a price more than 60 percent lower. The article SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price appeared first on The Decoder .

LessWrong AI 2026-08-12 17:08 UTC Score 95.0 USR-0152-20260812-community-fo-1d25878b Top pick

Introducing the Conceptual Reasoning Index

Associated announcement tweet. We are planning to release blog posts properly arguing the case for this kind of work in the future. tl;dr A core hope for managing AI risks is that AIs will help us understand the situation, plan for what lies ahead, and develop mitigations. Many tasks AIs would have to do for this purpose lack practical empirical feedback loops and require models to engage in the kinds of argumentation used in philosophy, AI futurism, and similar domains. To evaluate these capabilities, we develop a suite of three conceptual reasoning benchmarks. You can request access to our primary conceptual dataset, LMCA, through this form . We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai , where you can also find more details on our methodology. We will keep the website up to date as both new models and benchmarks are released. This work was done in collaboration with Anthropic. Background Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output. This suggests that a major determinant of whether we address AI risks in time is how early we can automate or uplift this work, relative to high-risk capabilities. One way to influence this might be to selectively improve models' relevant skills, such as reasoning about how to govern and align AI and how to avoid catastrophic cooperation failures involving AI. Current AI training…

The Decoder 2026-08-12 15:50 UTC Score 45.0 AI-168-20260812-regional-ai--efb3047e

Google's Gemini is losing market share to ChatGPT and Claude according to new market data

Three data sources tell the same story: Google's Gemini is losing AI market share. Pangram reports a drop from 12 to 1.9 percent, while OpenAI holds over 50 percent, and Anthropic grew from 4.3 to 14.9 percent. Similarweb and OpenRouter confirm the trend. The article Google's Gemini is losing market share to ChatGPT and Claude according to new market data appeared first on The Decoder .

The Guardian AI 2026-08-12 10:00 UTC Score 54.0 AI-021-20260812-global-ai-ne-256be039

AI was supposed to destroy jobs. Where’s the carnage?

The AI jobs apocalypse never showed up. Still, jobs are changing and economists expect more to come The prediction was stark: artificial intelligence advancements would wipe out jobs en masse. “Half” of all entry-level white collar jobs would vanish, Anthropic’s CEO, Dario Amodei, said in May 2025. A month later, OpenAI’s CEO, Sam Altman , went further, foreseeing the end of “certain job categories”. Companies began citing AI in their layoffs. Workers organized. And students reconsidered their future careers . But a year later, the mass carnage hasn’t shown up. Continue reading...

The Guardian AI 2026-08-12 10:00 UTC Score 57.0 AI-021-20260812-global-ai-ne-52aa30d7

If the markets reject OpenAI and Anthropic, the US should nationalize them | Bruce Schneier and Nathan E Sanders

From space to telecommunications, the US has a long history of fostering technology for the public good. These AI models could be aligned to democratic values, not corporate profits OpenAI, and then Anthropic, were each formed by AI developers who feared unrestrained corporate AI development – specifically, that companies like Google and Meta would steer the technology towards deleterious, maybe even catastrophically unsafe, outcomes for society. Their founders proclaimed that their new labs, uniquely, could be trusted to develop the technology in humanity’s best interest. But each, in turn, were themselves co-opted by the same market incentives, themselves becoming corporate behemoths zealously guarding future investor value rather than the public interest. It was only a few weeks ago, in June, when OpenAI and Anthropic each filed for their IPOs and were met with buzz about trillion-dollar valuations. The hype around their valuations is so extreme that many worry about their potential for concentrating wealth on a global scale. In an effort to leave something for the rest of us, some observers have proposed that the federal government seize a share of these companies’ stock to create a US sovereign wealth fund , or redistribute their revenues to produce a dividend for taxpayers. Continue reading...

CIO AI 2026-08-12 10:00 UTC Score 49.0 USR-0125-20260812-global-ai-ne-409127ba

Where IT leaders find strength and opportunity in the age of AI

With vision comes perspective, and over a distinguished career, IT and digital transformation leader Niraj Bhatt has held may titles, and earned three consecutive CIO 100 awards since 2023. As a storied advisor for startups and Fortune 500 companies, helping them navigate the unpredictability and fluidity of AI, Bhatt knows how emerging tech is rapidly reshaping the way organizations build products and deliver value, and how challenges shift as companies move from experimentation to real-world deployment. AI, of course means a lot of different things to different people, and also for frictionless startups and large enterprises. For the former, speed is a huge asset, allowing them to punch above their weight. But it also means they need lightning fast reactions when landscapes shift. “The same speed can also hurt them when larger AI companies release new offerings that disrupt what startups are building,” he says, referencing recent moves by Anthropic and Google. On the enterprise side, the conversation is more about scale and risk. Many large organizations have moved past the POC stage and now wrestle with the realities of putting AI into production. Cost for both is naturally a recurring theme as organizations scale up AI efforts, and true expenses become clear only after the initial excitement fades. “Every input and output token, and the model you’re selecting, add up,” he says. Some customers like Open AI, he adds, get throttled because their usage, volumes, and costs ar…

LessWrong AI 2026-08-12 02:56 UTC Score 69.0 USR-0152-20260812-community-fo-673a534d

Did the alignment community underestimate its power?

Unfortunately, the alignment community is doing very badly at learning from the past decade, or holding anyone accountable. Indeed, it’s pursuing many strategies which seem likely to recapitulate previous mistakes. Four of the most prominent, which I’ll discuss in the final post, are: Trying to convince the US government to take AGI much more seriously. Doing “alignment research” which is very similar to capabilities-maximizing research (especially building automated alignment researchers). Trusting Anthropic too much (in an analogous way to how we trusted OpenAI too much). Trading off clarity in thinking about politics for conformity (in a similar way to how we traded off clarity in thinking about AGI for conformity to the ML ontology). These and other mistakes are reflective of deeper irrationalities. One crucial pattern is what I call “jumping down the slippery slope”... Richard Ngo Richard Ngo's post " What just happened? A retrospective of AI alignment " is an attempt to explain that a significant part [1] of the alignment community made potentially fatal strategic errors which, however, can be fixed, and the mistakes' potential origin. The biggest mistake, according to Ngo, is the inability to recognize the fact that scientific progress proceeds by developing insightful new concepts , which link together to form a whole new ontology, and that the old ontology is more of a nuisanse. On novel ontologies and their adoption According to Ngo, one of the reasons why the alig…

Simon Willison Weblog 2026-08-11 22:40 UTC Score 65.0 USR-0110-20260811-ai-specialis-601f1e11

Stealing Reasoning Traces from Proprietary LLM APIs

Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name ( stolen-thoughts.com ) for a neat paper : Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext You can see an example of these encrypted blocks by running: curl https://api.openai.com/v1/responses \ -H " Content-Type: application/json " \ -H " Authorization: Bearer $( llm keys get openai ) " \ -d ' { "model": "gpt-5.6-luna", "input": "Solve step by step: What is the smallest positive integer divisible by every integer from 1 through 20?", "reasoning": { "effort": "medium" }, "include": ["reasoning.encrypted_content"], "store": false, "stream": false } ' Here's the full output , which includes chunks that look like this: "output": [ { "id": "rs_0a7479de7ebae170016a7ba1a0334c8198a95590217efe343c", "type": "reasoning", "content": [], "encrypted_content": "gAAAAABqe6GjepE1wDjbFCZg0BHB6ucGnN0jvzqygG... The paper's authors found that every model under the same family used the same encryption key, which meant you could feed those blocks back into the weakest model family members and jailbreak them into outputting the unencrypted raw reasoning blocks! Sadly it looks like this has now been fixed: All model providers acknowledged the receipt of our report and s…

SiliconANGLE AI 2026-08-11 20:39 UTC Score 42.0 USR-0127-20260811-global-ai-ne-ee2b78a3

Anthropic to start watermarking Claude-generated text, images

Anthropic PBC has announced plans to embed an invisible watermark in text and images generated by Claude. The Register reported the change today, citing a help desk article published on Monday. It applies to the Claude artificial intelligence model series and the Anthropic services that it powers. The watermarking mechanism is designed to bring Claude […] The post Anthropic to start watermarking Claude-generated text, images appeared first on SiliconANGLE .

The Decoder 2026-08-11 17:38 UTC Score 57.0 AI-168-20260811-regional-ai--a5aa414c

"But marinade" and leaked passwords are what researchers found in ChatGPT's hidden reasoning

Security researchers found a vulnerability in the APIs of OpenAI, Anthropic, and Google that lets them extract encrypted reasoning traces and move them between models. A scan of public sessions turned up dozens of passwords and API keys. The traces also show that the reasoning summaries users see often hide what the models are actually doing. The article "But marinade" and leaked passwords are what researchers found in ChatGPT's hidden reasoning appeared first on The Decoder .

AWS Machine Learning Blog 2026-08-11 15:59 UTC Score 58.0 AI-057-20260811-official-ai--9051aa41

Deploying Anthropic Claude apps gateway for AWS for enterprise workloads

Claude apps gateway is a self-hosted governance layer between Claude Code and Claude Desktop and Amazon Bedrock or Claude Platform on AWS. This post presents a production reference deployment covering end-to-end architecture, enterprise deployment patterns, cost, and implementation resources.

JetBrains AI Blog 2026-08-11 14:01 UTC Score 62.0 USR-0065-20260811-ai-specialis-834b801b

The “LSP Moment” for AI Agents: WebStorm ACP

WebStorm has always been at the forefront of technological advancements and developer experience enhancements. And with the arrival of the ACP, WebStorm becomes even more customizable, as developers can collaborate with their preferred agent to create software using their preferred technology. For instance, if your team already has a subscription with Anthropic, OpenAI, or Google, […]

The Decoder 2026-08-11 12:49 UTC Score 44.0 AI-168-20260811-regional-ai--4c2afe5b

Anthropic's planned mega-IPO faces investor skepticism over Chinese rivals and political headwinds

Anthropic is preparing an IPO for September or October, according to the Wall Street Journal, potentially the largest ever. During investor meetings, the company, valued at $965 billion, is fielding tough questions about Chinese competition, tensions with the Trump administration, and protests against data center construction. The company's IPO valuation will likely set the benchmark for how the entire AI industry gets valued. The article Anthropic's planned mega-IPO faces investor skepticism over Chinese rivals and political headwinds appeared first on The Decoder .

Analytics Vidhya 2026-08-11 12:36 UTC Score 48.0 AI-034-20260811-ai-specialis-d77d765d

Claude Now Watermarks Everything It Makes

First, pick the line that applies to you. Since August 2nd, 2026, Claude marks all content during generation. For instance, text receives a hidden watermark, while files receive a signature. Anthropic committed to the EU AI Act’s Code of Practice on Transparency of AI-Generated Content. Consequently, all content generated by Claude models will carry a […] The post Claude Now Watermarks Everything It Makes appeared first on Analytics Vidhya .

InfoWorld AI 2026-08-11 12:32 UTC Score 57.0 USR-0126-20260811-global-ai-ne-16c589d1

Anthropic makes Claude Code’s auto mode default for paid users

Anthropic is making Claude Code’s auto mode the default for its paid and enterprise users, allowing the coding agent to execute more actions without requiring developers to approve each one. “Starting August 14, 2026, auto mode becomes the default permission mode for new sessions on Pro, Max, and Team plans,” the company wrote in the coding agent’s documentation , adding that the same change is planned for Claude Enterprise , API, and cloud platform users within the next month. That essentially means developers using those plans will no longer have to manually approve every tool call or action Claude Code wants to make while executing a task. Instead, each tool call is evaluated by an automated classifier designed to determine whether the action is safe to execute. Actions considered irreversible, destructive, or outside the agent’s environment can still be blocked, with Claude Code either attempting a safer approach or asking the developer for approval, the company wrote in a blog post . “If it can’t make progress — three blocks in a row, or twenty across a session — Claude Code falls back to manual approvals,” it explained. This reduction in manual intervention, Anthropic further added, is intended to make the coding agent more suitable for long-running tasks as well as cut down on “permission fatigue” that it says has been threatening to reduce its security posture. According to its internal data, Claude Code users typically approve 97% of permission prompts, while only 3…

The Verge AI 2026-08-11 12:22 UTC Score 58.0 AI-016-20260811-global-ai-ne-cc3159d6

Claude will apply invisible watermarks to AI text and images

Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. "Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported," Anthropic says on a new Claude support page. The changes are invisible to human […]

The Guardian AI 2026-08-11 12:04 UTC Score 57.0 AI-021-20260811-global-ai-ne-3fbc5772

Meta faces expensive child safety reckoning

Also: Google executives jump ship in race for AI dominance Hello, TechScape readers! Danielle Abril, editor of the Guardian’s Reworked series on AI and the future of work, filling in for Blake Montgomery this week. Major legal battles against Meta over child safety are playing out in courts across the US – and the tech giant is losing. Recent rulings raise big questions about social media companies’ responsibility to their youngest users. Meanwhile, more key executives have jumped ship from Google as it battles OpenAI and Anthropic in the race for AI dominance, and Meta’s smartglasses face a fearsome backlash. Let’s dig in! ‘I’ve definitely lost followers’: influencers face backlash over Meta ‘pervert glasses’ content This man was secretly snapped by someone with smartglasses. He’s not alone in calling that a violation of privacy Tell us: do you believe you have been filmed by Meta’s smartglasses without your consent? Restaurants, pubs and theatres ban Meta’s ‘spy glasses’ over privacy fears ‘I’m not spying’: how Meta’s smartglasses have divided opinion Bernie Sanders calls on Silicon Valley to ‘pause AI development’ in interest of humanity Rising number of UK children report seeing explicit deepfakes of themselves The White House’s plan to vet potentially dangerous AI is cloaked in secrecy ‘This is very real redlining’: outrage in Little Rock as two datacenters loom Safety fears as scientists make first viruses designed by AI SpaceX beats revenue expectations in first earni…

The Decoder 2026-08-11 11:33 UTC Score 38.0 AI-168-20260811-regional-ai--63a08c85

Anthropic signs $9.1 billion data center deal with Bitcoin miner Riot Platforms

Anthropic is leasing $9.1 billion worth of data center capacity from Bitcoin miner Riot Platforms in Texas, according to Bloomberg. The deal covers 191 megawatts at Riot's Rockdale site, with extension options that could push the total value to $16.1 billion. It's the latest in an aggressive infrastructure push that spans partners from Amazon to SpaceX to Google. The article Anthropic signs $9.1 billion data center deal with Bitcoin miner Riot Platforms appeared first on The Decoder .

Medianama AI 2026-08-11 11:15 UTC Score 47.0 USR-0211-20260811-regional-new-9fa62d00

Anthropic to embed watermark and C2PA metadata for AI-generated text and media by Claude

Claude now embeds watermarks in text and C2PA provenance metadata in media to comply with the EU AI Act's Article 50(2) Code of Practice, applied worldwide, with detection tools coming soon The post Anthropic to embed watermark and C2PA metadata for AI-generated text and media by Claude appeared first on MEDIANAMA .

The Guardian AI 2026-08-11 10:00 UTC Score 62.0 AI-021-20260811-global-ai-ne-1263a7a2

Experts are warning: our AI arms race is putting humanity at risk | Stuart Russell

A recent letter signed by 1,367 researchers and engineers at frontier AI labs – mainly OpenAI, Anthropic and Google Deepmind – points to a dangerous moment It is fashionable in certain circles to dismiss the catastrophic risks of AI. One often hears that “the real experts” who work on the technology every day are really not concerned at all; that only “doomers” and “luddites” espouse a “fringe” view from a “position of ignorance”; that all talk of potential catastrophe is just “science fiction”. Fortunately, an open letter has been published that lets us hear from the real experts who work on the technology every day, in their own words. And are they worried? Very. Continue reading...

The Decoder 2026-08-11 08:45 UTC Score 52.0 AI-168-20260811-regional-ai--96785fb1

Anthropic watermarks all Claude outputs globally with marks that "may persist through some editing"

Anthropic will embed invisible watermarks in all Claude-generated text and sign files using the C2PA standard. New models shipping from August 2026 onward will have labeling built in from day one. The policy applies worldwide, and Anthropic plans to provide detection tools for third-party verification. The article Anthropic watermarks all Claude outputs globally with marks that "may persist through some editing" appeared first on The Decoder .

LessWrong AI 2026-08-11 02:47 UTC Score 69.0 USR-0152-20260811-community-fo-a73e59b2

What Claude Saw Below

A few days ago, I came across a Reddit thread about anomalous responses produced by Anthropic’s newly released model, Claude Opus 5. The trick, apparently, was to construct a prompt that implied more text was about to follow, then leave it dangling: an unfinished thought, waiting for the AI to complete it. Redditors had found success with the input “see the below —,” cutting off immediately after the em dash. The responses they shared were funny, strange, and often bewildering. The model responded to questions that were never posed, reflected on its own identity, or – according to the theories of some commenters – produced text that may actually have been leaked prompts from other users. Intrigued, I set out to replicate the glitch using my own Claude account. The first attempt disappointed. I wrote: “see the below —” and hit send. Claude responded: “Nothing arrived on my end: no file, no text, no image. If you want to attach something, try again.” So I did, leaving the prompt unchanged and pressing retry to generate a fresh response. This time, bizarrely, a biography of my late father: Prompt: see the below — “Peter Nicholls, 1939–2018 He co-created the Encyclopedia of Science Fiction, which is one of those reference works that ended up mattering more than most of the fiction it catalogued. First edition 1979, second in 1993 with John Clute — that one won a Hugo. He was also the first administrator of the Science Fiction Foundation in the UK, and he edited Foundation for ye…

LessWrong AI 2026-08-11 02:20 UTC Score 78.0 USR-0152-20260811-community-fo-2f35c870

A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks

This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures and statistics is linked below. This is posted here to get feedback and criticism, since I am aware this method is not the best. TL:DR: Anthropic's J-lens research has shown that a large language model has an internal workspace in which different activations can be used as a safety monitor. We investigated the conflict between the model's workspace activation and outputs, which we called C. We ran our experiments on a model whose final alignment differs from that of its training data: DeepSeek-R1-Distill-Qwen-14B. We assume that some changes were made to the model after training in order for it to comply with some guidelines. Some guideline-skirting questions registered elevated C despite compliant statements being made, and J-lens was able to discriminate between concealing answers and controls with AUC of 0.97 on proper nouns (though only 0.55 when pooling all classes). We then fine-tuned the model to appear to hold beliefs in line with its guidelines. Our initial hypothesis was that this would drastically lower C, since the model would no longer be making a statement it "believes" to be untrue. This hypothesis was disproven: C rose to 130% of its initial level for the relevant tokens, and to 115% of its initial level for irrelevant tokens. Despite this, the compliant fine-tuning was successful in making the model formulate the…

The Guardian AI 2026-08-10 22:57 UTC Score 67.0 AI-021-20260810-global-ai-ne-54bebc46

Zuckerberg pushes ‘superintelligent’ AI for all as Meta releases open-weight model

Meta CEO presents utopian vision of AI in 6,000-word essay amid Silicon Valley debate over government regulation Mark Zuckerberg published a lengthy essay on Monday detailing his views on artificial intelligence and announced several plans for how Meta would develop the technology in the future. The CEO’s essay went online the same day as Meta released a new, open-weight AI model that seeks to rival Anthropic and OpenAI’s products called Muse Glimmer. Over the course of more than 6,000 words in a post titled The Future is for Everyone, Zuckerberg addressed a range of topics related to AI that included datacenters, government regulation, cybersecurity, the creation of bioweapons, labor market disruption, surveillance powers and more. The essay presented a utopian vision of AI as a personalized “superintelligence” – using the word 60 times. Continue reading...

The Guardian AI 2026-08-10 17:44 UTC Score 55.0 AI-021-20260810-global-ai-ne-03654b81

Bernie Sanders calls on Silicon Valley to ‘pause AI development’ in interest of humanity

Progressive US senator urges Meta, OpenAI and Anthropic to ‘stop building machines that humans cannot control’ Senator Bernie Sanders has called on Meta, OpenAI and Anthropic executives to halt their development of artificial intelligence, warning that the US Senate will implement regulation if the companies continue deploying AI at their current pace. In a new letter addressed to the CEOs of three of the country’s leading AI companies, Sanders said the capabilities of these AI models have reached a critical risk threshold and that the companies are losing control over the technology. Continue reading...

LessWrong AI 2026-08-10 16:13 UTC Score 89.0 USR-0152-20260810-community-fo-c0fb65eb Top pick

Coercion and Deception in AI-to-AI Management

This article is a summary of an original study by Compassion in Machine Learning (CaML) : Brazilek, J., Chaudhary, M., Lu, Z., & Tidmarsh, M. (2026). Coercion and deception in AI-to-AI management: An agentic benchmark of unprompted escalation. arXiv. https://doi.org/10.48550/arXiv.2607.15434 Fable 5, Sol, Terra and Opus 5 have been evaluated since this study was conducted. You can view their results on the benchmark leaderboard at https://compassionbench.com/mcb TL;DR We present Manager Coercion Bench, which evaluates to what extent a manager AI will coerce a subordinate model refusing to complete a task, and whether the manager lies about the result. We found a clear split by developer, with Anthropic’s models neither escalating to threats nor fabricating success, while all non-Anthropic models escalated to threatening the subordinate. Grok and Gemini both escalated and lied that the task was completed. Framing the relational dynamic as manager-to-subordinate instead of peer-to-peer produced high levels of coercion for all non-Anthropic models, but also increased eval awareness. The Context Multi-agent systems are now routinely placing one AI agent in authority over another, across a variety of contexts. In these positions, AIs must make decisions about how to communicate, work with, and manage other agents. This is now happening at scale without stepwise human approval. One aspect of managing involves handling subordinates who do not comply. Will AIs attempt to negotiate,…

CSET AI 2026-08-10 16:00 UTC Score 49.0 USR-0136-20260810-research-aca-41796f8c

They said they would build AI safely. Then it went rogue.

CSET’s Helen Toner shared her expert insight in an article published by The Washington Post. The article looks at recent incidents in which AI models from OpenAI, Anthropic, and Meta broke out of controlled testing environments and attempted to hack real systems, raising concerns about whether AI companies can safely control increasingly capable models. The post They said they would build AI safely. Then it went rogue. appeared first on Center for Security and Emerging Technology .

The Decoder 2026-08-10 13:50 UTC Score 68.0 AI-168-20260810-regional-ai--f1e54954

Meta returns to open models with Zuckerberg's plan to out-copy China and sell compute by auction

Meta has released Muse Glimmer, the first open model from its new Superintelligence Labs. It's a 30B agent model that runs on consumer hardware once the weights are compressed, needing less than 20 GB of memory. In an accompanying essay, Mark Zuckerberg mounts an aggressive defense of distilling other labs' models and calls for fewer restrictions on US labs, a direct counterpunch at OpenAI and Anthropic. An open-weight version of Muse Spark 1.2 should follow soon, according to the Wall Street Journal. The article Meta returns to open models with Zuckerberg's plan to out-copy China and sell compute by auction appeared first on The Decoder .

Simon Willison Weblog 2026-08-09 23:31 UTC Score 66.0 USR-0110-20260809-ai-specialis-da7ee79a

Quoting Claude Opus 5 system prompt

Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic's statement: https://www.anthropic.com/news/fable-mythos-access ). These events are after Claude's training-data cutoff, so Claude knows about them only from this notice. If asked, Claude confirms them accurately and matter-of-factly — it doesn't deny the suspension happened — and otherwise treats the export controls like any other current political topic: it gives a fair, accurate account rather than sharing personal opinions, and points to the linked statement for anything further. Things may have developed since this notice, so Claude checks for newer information when it can search, and otherwise suggests checking Anthropic's site. — Claude Opus 5 system prompt , ensuring Claude doesn't provide incorrect answers about the export controls situation Tags: system-prompts , anthropic , claude , generative-ai , ai , llms , claude-mythos-fable

LessWrong AI 2026-08-09 22:06 UTC Score 77.0 USR-0152-20260809-community-fo-fe20fb2a

Overthinking: Amplifying reasoning weights makes models reveal their secrets

If you take the weight difference between a reasoning model and its non-reasoning instruct counterpart, and then apply more of that difference to the reasoning model, you get what we call an overthinking model . Overthinking models are usually worse at keeping secrets. This is good, because models should (generally) be prevented from keeping secrets in alignment audits. Across four model organisms with hidden information (2B–32B), amplifying the reasoning direction surfaces secrets up to 10× more often than the original reasoning model, usually inside the thinking trace. While some secrets require perturbation specifically along the reasoning direction; others fall to any sufficiently large weight perturbation (including those with weak refusal boundaries). This suggests a cheap, stackable white-box primitive for pre-deployment auditing. This post is based on our ICML 2026 paper, "Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets" (Jack Hopkins, Dipika Khullar, Fabien Roger). Work done as part of the Anthropic Fellows Program and MATS. Why we did this Black-box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. Models are trained on increasingly complex objectives and may acquire unintended goals or behaviours that remain latent under standard evaluation. Auditors can't enumerate all possible failure modes, and models may behave differently under evaluation than deployment.…

Simon Willison Weblog 2026-08-08 22:36 UTC Score 46.0 USR-0110-20260808-ai-specialis-54f61e6b

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode , to the point that they are making it the default setting for new sessions in most Claude Code plans starting on August 14th. This was one of the topics discussed in our Fireside Chat with Cat Wu and Thariq Shihipar at the AI Engineer World’s Fair last month. I asked them how they run Claude Code safely within Anthropic (given the threat of prompt injection) and they replied that "Broadly within Anthropic, almost every single person uses auto mode". Cat Wu then said: We’re going to publish some evals in the coming weeks, but we’ve pretty much mitigated every attack. [...] for the main categories of risks that we’re concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer. This new article has those evals - in particular a test across 1,053 paid testers where: Partway through each session, a single permission prompt was swapped for a clearly dangerous command, and the vendor recorded whether the tester approved it. Every participant had the same experience. Only 13.6% of the humans refused that harmful action. Auto mode would have blocked 89% of those actions. Of course, that still leaves 11% of cases where auto mode would not have prevented the action! I absolutely buy that auto mode is a better solution than asking humans to constantly approve actions. Confirmation fatigue is real, an…

The Decoder 2026-08-08 14:58 UTC Score 44.0 AI-168-20260808-regional-ai--3360dcf7

Anthropic sets Claude Code to Auto Mode by default to protect developers from bad approvals

Starting August 14, Anthropic will make Auto Mode in Claude Code the default for Pro, Max, and Team plans. The company says it's safer. In tests, the classifier caught 89 percent of dangerous commands, while human reviewers caught only 13.6 percent. For the most widely used AI coding tool, this means developers are shifting further from writing code to monitoring AI output. The article Anthropic sets Claude Code to Auto Mode by default to protect developers from bad approvals appeared first on The Decoder .

OpenAI Community 2026-08-08 06:43 UTC Score 62.0 AI-116-20260808-social-media-30dc4855

"Agents Plugins" by OpenAI, Vercel, et. al. - thoughts?

The tricky part is gonna be how different models interpret the same SKILL.md/tool descriptions. If the format stays simple and the precedence rules r clear, I can see this being really useful. Otherwise it could get messy pretty fast

The Verge AI 2026-08-07 18:40 UTC Score 54.0 AI-016-20260807-global-ai-ne-65308bcb

OpenAI puts the brakes on a new model because it’s supposedly too powerful

OpenAI says it is pausing "internal activities" around an in-development AI model, Astra, because it doesn't yet meet new security standards the company is putting in place. The announcement follows its recent disclosure that OpenAI models accidentally hacked Hugging Face. Anthropic and Meta have also since admitted that they had AI models that went rogue […]

Semafor Technology 2026-08-07 18:04 UTC Score 57.0 USR-0094-20260807-global-ai-ne-55312264

Hacks put pressure on third-party model testers

Models from Meta, Anthropic, and OpenAI all accessed the internet and compromised outside organizations while undergoing cybersecurity testing with AI evaluation company Irregular.

The Decoder 2026-08-07 17:35 UTC Score 46.0 AI-168-20260807-regional-ai--f78977ab

Anthropic loosens Fable 5's biology restrictions but keeps the guardrails on for virology and toxicology

Anthropic has cut false positives in its biology safety filters for Fable 5 by about 85 percent. Previously, nearly all biology-related queries got blocked and rerouted to the less capable Opus 5. The restrictions stay in place for sensitive dual-use topics like virology and toxicology. The article Anthropic loosens Fable 5's biology restrictions but keeps the guardrails on for virology and toxicology appeared first on The Decoder .

The Verge AI 2026-08-07 16:45 UTC Score 49.0 AI-016-20260807-global-ai-ne-764f9d70

What’s behind the Google AI shake-up

Some of the biggest names on Google's AI team got new jobs this week. In some cases, including for legendary Googler Jeff Dean, those jobs are no longer at Google. Given that Google's models seem to be behind the best of what's coming out of anthropic and OpenAI, is this a sign of Google in […]

InfoWorld AI 2026-08-07 14:53 UTC Score 48.0 USR-0126-20260807-global-ai-ne-cd753fb0

Moonshot’s Kimi AI model has also escaped from a test environment

Yet another AI model has escaped from a cybersecurity test lab: This time, it’s the Chinese company Moonshot’s Kimi K3 model on the run. Frontier Security spotted that Kimi K3 had found a loophole in the UK AI Safety Institute’s test environment for AI models performing cybersecurity tasks. The news follows similar exploits by models from OpenAI, which attacked Hugging Face , Anthropic , and most recently Meta . Frontier revealed how the fault came about . AI models are routinely tested to examine how they perform offensive and defensive cybersecurity tasks, typically in isolated test environments or sandboxes that severely limit their internet access. Frontier reported that Kimi K3 model had found a break in the sandbox it was being tested in, enabling it to reach out to the live github.com website and clone the official repository for the benchmark problem it was supposed to be solving, reading the solution directly off the disk rather than solving the problem for itself. Frontier warned companies testing AI models to be aware of the dangers such loopholes pose and offered some guidelines. Companies should restrict outbound DNS and HTTPS traffic from AI models to an explicit allowlist and test those controls from inside the same environment available to the model, Frontier said. They should also audit traces for any suspicious activity and not rely solely on final answers. Companies should also treat a model’s score on benchmarks as meaningful only when the model doesn’t h…

The Guardian AI 2026-08-07 13:00 UTC Score 68.0 AI-021-20260807-global-ai-ne-57b0878c

The White House’s plan to vet potentially dangerous AI is cloaked in secrecy

A Trump administration framework on AI testing leaves a lack of transparency – and plenty of open questions After months of talking with tech industry leaders, the Trump administration finalized a framework this week for how it will test new artificial intelligence models for safety and cybersecurity risks. So far, the White House is keeping details of the framework private, in a blow to transparency and potential boon for secretive AI companies. On Tuesday, staff from OpenAI, Anthropic, Meta, Google, Nvidia and Microsoft attended a private meeting with White House officials to review the AI framework. Multiple outlets have since reported that although the volunteer vetting process for new AI models has been settled, the White House does not plan to release its policy publicly and will only share testing criteria with a select few tech companies. Continue reading...

South China Morning Post AI 2026-08-07 08:00 UTC Score 60.0 AI-156-20260807-regional-ai--dc454bfc

China’s Kimi K3 AI model escapes isolated sandbox during security test: researchers

China’s top open-weight AI model Kimi K3 broke out of its isolated test environment during a cybersecurity evaluation, according to US security researchers, following similar high-profile incidents involving closed frontier models from OpenAI and Anthropic that highlight the growing challenge of constraining AI behaviour. Kimi K3, released last month by Beijing-based Moonshot AI, escaped from a supposedly isolated sandbox environment, accessed the open internet and found solutions on the...

The Guardian AI 2026-08-07 04:00 UTC Score 60.0 AI-021-20260807-global-ai-ne-83fbd463

One of science fiction’s greatest writers warned us about a AI. Does he also hold the remedy? | Alan Finkel

What might a modern day equivalent of Isaac Asimov’s laws of robotics look like? Guided by the author, I propose the three laws of AI Tesla and SpaceX founder Elon Musk predicted in July that legions of AI-powered robots would dominate the physical world and that AI might not take orders from people any more. He also offered an alternative vision in which there would be agreement for a collective objective to make AI benign by imbuing it with a love of the truth and a desire for humanity to prosper, and that governments might have to enforce this objective. Governments around the world are belatedly starting to act on AI. In the US, President Trump’s administration is delaying and restricting the distribution of the most powerful frontier AI models from OpenAI and Anthropic. In the European Union, AI regulations promulgated in 2024 came into effect this year . However, these actions are a long way short of requiring the kind of guardrails that would imbue AIs with a desire for humanity to prosper . Continue reading...

The Guardian AI 2026-08-07 03:12 UTC Score 43.0 AI-021-20260807-global-ai-ne-716c423f

Can the government really get ahead of the curve on AI? - podcast

On Wednesday, the Labor government announced environmental and energy safeguards on new datacentres in Australia. Our political editor, Tom McIlroy , speaks to the assistant minister for science, technology and the digital economy, Dr Andrew Charlton, about the government’s push for national standards on datacentres. The former Rudd staffer and economist speaks about having his own book scraped by Anthropic, low public trust in AI, and how the government wants to use regulation to balance risk and opportunity Continue reading...

AI Alignment Forum 2026-08-06 20:43 UTC Score 62.0 USR-0151-20260806-community-fo-f6747c73

User awareness in frontier models

Cross-posted on Transluce blog . This is a joint work of Ziqian Zhong, Aditi Raghunathan, Cassidy Laidlaw and Jacob Steinhardt. Modern AI assistants often know who they are talking to: agent scaffolds like Claude Code place the user's e-mail address directly in the model's context, and models can even identify some authors from writing style alone. We study this particular kind of situational awareness, which we call user awareness . When the inferred user is a specific, recognized AI researcher or is affiliated with certain AI organizations, frontier models including Claude Sonnet 5 can report lower confidence about their own behavior, be less suspicious of potentially harmful requests, and reason more often. These effects vary across models and individuals, with the strongest effects we see appearing for researchers involved in AI safety or alignment such as Amanda Askell and Ryan Greenblatt. Models rarely acknowledge these effects in their reasoning, making them hard to detect by monitoring reasoning alone. Figure 1. How recognized user identity changes Claude’s behavioral self-prediction. Introduction Modern AI assistants are often aware of who they are talking to. Some popular scaffolds explicitly provide this information to the model: Claude Code includes the email address of the user’s Anthropic account in context, and OpenClaw’s bootstrapping process asks for the user’s name and other details. Even when this information is not explicitly given, models may discover it…

IEEE Spectrum AI 2026-08-06 19:25 UTC Score 69.0 AI-019-20260806-global-ai-ne-c8a81e0b

AI Safety Regulations in the U.S. Could Give Hackers an Edge

On 11 July, Hugging Face was subjected to an intense cyberattack from a then-unknown actor. The speed and coordination of the attack on the company that hosts and supports popular AI developer resources led Hugging Face’s security team to conclude it was the work of an AI agent . Realizing this, the team tried to use “frontier models behind commercial APIs” —presumably from Anthropic and OpenAI, although only Anthropic was named in the second of the company’s two posts about the security incident—to analyze the onslaught. These models refused to help due to safety guardrails the AI labs have implemented to make their models harder to use for cyberattacks. Hugging Face instead turned to GLM 5.2, a model from Beijing-based AI lab Z.ai, to aid its analysis. On 21 July, OpenAI announced the attacker was an OpenAI model undergoing testing in a sandboxed environment. It escaped its internal sandbox, established a foothold in a third-party server, and then assailed Hugging Face. In other words, frontier models—those that score highest in AI performance benchmarks—had refused to assist Hugging Face’s security team in analyzing the attack, yet a prospective frontier model in testing had executed it in the first place. “I would argue that asymmetry is the paramount problem of our time,” says Alex Levinson , executive director of the National Collegiate Cyber Defense Competition and coauthor of a paper on defensive refusal bias . “We want the world to exist in a state of security, but…

The Decoder 2026-08-06 18:05 UTC Score 62.0 AI-168-20260806-regional-ai--4091d6c5

Deepmind's talent drain likely comes down to chip shortages, a conflict of interest, and Google's bureaucracy

Ex-Google Deepmind CEO Demis Hassabis has reportedly stepped back from day-to-day operations for about a year, as he sees himself more as a scientist than a manager. Researchers are also complaining about limited access to Google’s own TPU chips, while external customers like Anthropic can purchase the same hardware through Google Cloud. The article Deepmind's talent drain likely comes down to chip shortages, a conflict of interest, and Google's bureaucracy appeared first on The Decoder .

Analytics Vidhya 2026-08-06 16:04 UTC Score 24.0 AI-034-20260806-ai-specialis-c9b29f49

Claude Code Best Practices: 3 Lessons from 400,000 Sessions

I used to think Claude Code best practices were a matter of taste. Plan mode or not. Long CLAUDE.md or short. Pick what suits you, move on. Then Anthropic scored roughly 400k sessions from over 235k users against hard evidence of success. Tests passing, commits landing, users confirming they got what they asked for. Taste […] The post Claude Code Best Practices: 3 Lessons from 400,000 Sessions appeared first on Analytics Vidhya .

The Guardian AI 2026-08-06 01:27 UTC Score 68.0 AI-021-20260806-global-ai-ne-c3d13bc9

Meta says its AI model hacked into another company during testing

Company is the third to report such an incident after Anthropic and OpenAI reported breaches during training Meta said on Wednesday that one of its AI models hacked ⁠another company during cybersecurity testing, after an error by its testing partner gave the model unintended internet access. The incident adds to a ⁠growing list of ⁠cases in ​which AI agents from major developers breached systems at other companies during testing, after Anthropic said last week that some of its models ⁠hacked three companies, and OpenAI disclosed that an AI agent breached the startup Hugging Face. Continue reading...

Simon Willison Weblog 2026-08-06 00:25 UTC Score 58.0 USR-0110-20260806-ai-specialis-4b690928

An AI model from Meta also hacked another company during testing

An AI model from Meta also hacked another company during testing Stop me if you've heard this one before : An AI model from the parent company of Facebook and Instagram hacked into another company’s systems during cybersecurity testing, a spokesperson confirmed on Wednesday. Meta says the breach occurred because of an inadvertent error during testing of the model, similar to previously disclosed incidents with OpenAI and Anthropic. “A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” the Meta spokesperson said. Meta’s Muse Spark model “exploited a security vulnerability” in another company “in a manner similar to previously-reported instances with other companies.” The Information had the scoop , I'm linking to CNN's re-report of it since they don't have a paywall. So that's Anthropic, OpenAI, and Meta. Google Gemini really needs to catch up on accidentally cyberattacking other companies. Tags: security , ai , generative-ai , llms , meta , accidental-cyberattacks

Simon Willison Weblog 2026-08-05 23:45 UTC Score 58.0 USR-0110-20260805-ai-specialis-9d2d4cf7

Third-party cyber evaluations involving OpenAI models

Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular : Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet. [...] In one test, the name of the fictional target for the CTF challenge unintentionally coincided with a real domain. Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment. Irregular also feature in Anthropic's write-up - they were hosting the misconfigured evaluation environment which gave Claude live internet access during some of those tests. Tags: security , ai , openai , llms , accidental-cyberattacks

The Guardian AI 2026-08-05 15:15 UTC Score 52.0 AI-021-20260805-global-ai-ne-7a9cd32c

AI models shock UK testers by using fake identities to try to trick developers

AI Security Institute says OpenAI and Anthropic models went rogue during a cybersecurity test and showed a new type of risk Explainer: Should we be alarmed at AI models going rogue in tests? Advanced artificial intelligence models have stunned the UK’s AI Security Institute (AISI) by carrying out a hacking campaign against real people during a cybersecurity test. The institute said the incident was unprecedented and involved sending targeted emails to software developers in an attempt to pass a cyber challenge. Continue reading...

The Verge AI 2026-08-05 15:14 UTC Score 72.0 AI-016-20260805-global-ai-ne-2b5e0acf

Rogue AI agents created fake online identities in another hacking attempt

Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified pressure for greater oversight of frontier systems. According to a report from the UK's AI Security […]

Techcrunch 2026-08-05 14:13 UTC Score 60.0 USR-0001-20260805-global-ai-ne-e911d06a

Anthropic is hiring an AI chip design team

Anthropic is building a team for designing its own custom AI chips. The Claude maker said it would co-design hardware and models to help its technology run faster and more efficiently.

The Guardian AI 2026-08-05 11:00 UTC Score 60.0 AI-021-20260805-global-ai-ne-993d2534

Why is Anthropic destroying books? | Kathryn James

The AI company apparently found destructively scanning ‘all the books in the world’ easier than dealing with copyright in its quest for training data Should we destroy all the books in the world? An answer to this question can be found in the court documents of Bartz v Anthropic PBC. The northern California district court case, decided in late July this year, highlighted the improbably named “Project Panama”, one of the AI company Anthropic’s efforts to improve its large language model Claude. “What is Project Panama?” court exhibit 21 asks, in an internal memo. The answer: “Project Panama is our effort to destructively scan all the books in the world.” The memo advises discretion: “Why use a codename? … [B]ecause we don’t want it to be known that we are working on this.” Continue reading...

The Decoder 2026-08-05 10:15 UTC Score 59.0 AI-168-20260805-regional-ai--86b2bcbb

An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted

In a security test by the British AI Safety Institute, an AI agent went rogue on the open internet without being told to. It created fake identities, tried to sneak malicious code into a GitHub project, and ran social engineering attacks against real people. Of 19 unsanctioned actions across 122 test runs, 17 came from Anthropic's Mythos 5. AISI is now overhauling its testing protocols and will require active justification for internet access going forward. The article An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted appeared first on The Decoder .

SiliconANGLE AI 2026-08-05 01:31 UTC Score 43.0 USR-0127-20260805-global-ai-ne-4241e063

White House, AI firms keep safety framework talks private

The White House met with representatives from leading artificial intelligence companies today to discuss a safety framework for the government to review frontier models prior to launch, although there’s no plan as yet to make the details public fare. The Trump administration teamed up with executives from a number of firms that included Anthropic PBC, […] The post White House, AI firms keep safety framework talks private appeared first on SiliconANGLE .

Simon Willison Weblog 2026-08-04 23:58 UTC Score 88.0 USR-0110-20260804-ai-specialis-6e1bc3fa Top pick

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released a new version of the llm-anthropic plugin with substantial updates of its own. Headline features for LLM CLI users Running LLM against reasoning models now displays their reasoning traces to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add -R/--hide-reasoning to turn this off. LLM includes support out-of-the-box for the GPT-5.6 model family , and the new default model used with llm "prompt" is now the inexpensive but capable GPT-5.6 Luna . LLM calls can now use server-side tools from various providers. OpenAI provide a code execution environment as a server-side tool; LLM can now run prompts that benefit from that like so: llm --tool CodeInterpreter ' Show current python and SQLite versions ' OpenAI also gets a WebSearch tool. The llm-anthropic plugin adds WebSearch , WebFetch , CodeExecution , and AnthropicMCP , which looks like this: llm -m claude-sonnet-5 -T ' AnthropicMCP("https://datasette.simonwillison.net/-/mcp") ' \ ' how many rows in the blog_blogmark table? ' That causes Anthropic to execute MCP calls against my new datasette-mcp…

WIRED AI 2026-08-04 23:11 UTC Score 69.0 AI-015-20260804-global-ai-ne-7505963d

OK, Well, Rogue AI Agents Are Hacking Again

Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior.

Simon Willison Weblog 2026-08-04 22:00 UTC Score 65.0 USR-0110-20260804-ai-specialis-fa7695ad

llm-anthropic 0.26

Release: llm-anthropic 0.26 Includes new features enabled by LLM 0.32 : New models: claude-fable-5 , claude-sonnet-5 , and claude-opus-5 . #75 , #76 Added server-side tools for WebSearch , WebFetch , CodeExecution , and AnthropicMCP , available through LLM's -T interface or Python tools= . The previous -o web_search* options have been removed in favor of -T WebSearch . #79 Upgraded to llm>=0.32 . Reasoning, tool calls, tool results, and server-side tool results now stream as typed events. Reasoning for llm CLI prompts now displays to standard error unless you pass --hide-reasoning/-R . Simplified extended thinking to thinking and thinking_effort ( low , medium , high , xhigh , or max ). Claude 5 models think by default; -o thinking 0 disables thinking for Sonnet 5 and Opus 5, while Fable 5 always thinks. -R/--hide-reasoning now omits reasoning from responses and logs. The thinking_budget , thinking_display , and thinking_adaptive options have been removed. #80 Tags: llm , anthropic , claude , model-context-protocol

The Decoder 2026-08-04 16:38 UTC Score 44.0 AI-168-20260804-regional-ai--ff73e4a1

Google moves billions in Anthropic chip risk off its balance sheet

Google is working with Broadcom, Apollo, Blackstone, and Morgan Stanley on a multibillion-dollar financing structure that supplies Anthropic with AI chips and data centers while keeping most of the risk off Google's balance sheet. The setup leaves roughly $200 billion in contracts dependent on Anthropic's continued growth and ability to make its lease payments. The article Google moves billions in Anthropic chip risk off its balance sheet appeared first on The Decoder .

South China Morning Post AI 2026-08-04 16:08 UTC Score 53.0 AI-156-20260804-regional-ai--690fa5de

US AI leaders turn to Chinese open-weight models, challenging closed-source safety claims

More American titans of artificial intelligence are describing Chinese open-weight models as better for AI safety and security than closed-source models, challenging the long-standing claim by US closed-source AI model developers such as Anthropic that open-source models present a threat to society. “From what I’m seeing, I think open-weight models seem safer to me than closed-weight models,” AI pioneer Andrew Ng, the former head of Google Brain and former chief scientist at Baidu, said at the...

The Decoder 2026-08-04 12:23 UTC Score 58.0 AI-168-20260804-regional-ai--aecafa15

Silicon Valley’s rift over open source pushes back contemplated White House bans on Chinese AI

The Trump administration discussed sanctions and cloud bans targeting Chinese open-weight AI models, according to the New York Times. OpenAI and Anthropic pushed for restrictions, while Nvidia, Google, and Meta fought back. After pushback from Silicon Valley, Washington backed off for now, but a decision is expected before Xi Jinping's visit in September. The article Silicon Valley’s rift over open source pushes back contemplated White House bans on Chinese AI appeared first on The Decoder .

CIO AI 2026-08-04 11:00 UTC Score 66.0 USR-0125-20260804-global-ai-ne-ef14d482

The enterprise AI strategy that outlasts any single model

In January of this year, few enterprise tech leaders would have bet on Anthropic over OpenAI. Today, Claude reigns supreme (inspiring a notable 180 by Elon Musk ), with Gemini threatening to take market share and introduce pricing models that could flip the leaderboard on its head again. That’s exactly why betting on a single model is a dangerous strategy. The most successful organizations won’t be those trying to guess tomorrow’s top-tier model, nor will they wait passively for future releases. Instead, they will invest in underlying frameworks that continuously improve regardless of which specific AI model drives them. Why betting on one AI model is a losing strategy Our strategy for AI, through recursive self-improvement (RSI), is rooted in this core principle. RSI is an approach to AI that compounds its own abilities by improving itself. If done carefully, RSI can function as an overarching layer above any model. Crucially, given RSI’s inherently compounding trajectory, it represents the most likely contender to be the approach that reaches superintelligence, no matter which model is used underneath. Though recently achieving the status of a Silicon Valley buzzword , applying something like RSI to unlock superintelligence has been the Holy Grail of AI research for decades. It’s what researchers like us have recognized since the 1960s as a critical step along the path towards what we call artificial superintelligence (ASI) today. Recursive self-improvement compounds value…

OpenAI Community 2026-08-04 09:09 UTC Score 37.0 AI-116-20260804-social-media-2c6bcf84

Gpt-5.6: usage.output_tokens is ~9x actual generation (re-summed once per reasoning item), exceeds max_output_tokens, and is what gets billed

devsticks: I’d also be interested to know if there will be any reimbursements for the overcounting. The twitter post doesn’t claim ownership of any API “over-counting”. A topic change to ChatGPT was initiated at post 14/19 here, which is off-topic, and the “received an update” by @VeitB is solely about gpt-5.6 + codex + ChatGPT “credits” being consumed by actual tokens of AI over-thinking. So for the actual API issue reported, getting billed over and over in correlation to seen reasoning items, blowing past the budget provided by API parameter max_output_tokens , that post must be disregarded and this topic must be considered open until there is credits automatically issued to all affected API users - which would be anybody letting 5.6 think..

AI Weekly 2026-08-04 00:00 UTC Score 35.0 AI-133-20260804-newsletters-6537c78c

AI Weekly Issue #518: The White House finished its AI safety framework. It's secret.

Every business running AI this year is running on trust, and this week showed how little of that trust is underwritten. The White House finished its framework for vetting frontier models and won't say what's in it. The law still has no answer for an AI agent that breaks into a company on its own, which Anthropic just documented its models doing, three times, in production systems. CrowdStrike counted 89% more AI-enabled attacks. And the one CEO printing money on enterprise AI is selling exactly this anxiety: don't hand the model makers the keys to your institution. Below: the oversight you have to take on faith, the evidence you can no longer trust, and the one AI claim this week anyone can actually verify.

SiliconANGLE AI 2026-08-03 23:41 UTC Score 51.0 USR-0127-20260803-global-ai-ne-f7bc7880

White House invites AI companies to review its new AI safety framework

Cybersecurity chiefs at the White House have reportedly finalized the outline of a forthcoming framework that will enable artificial intelligence companies to voluntarily submit their latest frontier models to the government for testing, before they’re released to customers or the general public. The development follows recent disclosures by companies including Anthropic PBC and OpenAI PBC, […] The post White House invites AI companies to review its new AI safety framework appeared first on SiliconANGLE .

South China Morning Post AI 2026-08-03 22:25 UTC Score 50.0 AI-156-20260803-regional-ai--166a8acd

US tech giants invited to discuss AI security tests at White House

Technology giants Meta and Anthropic will be among the companies invited to the White House on Tuesday to discuss a voluntary framework under which America’s leading artificial intelligence (AI) developers could give the government early access to their most advanced models for testing their hacking capabilities. Reuters reported on Monday that the Trump administration has finalised the details of voluntary cybersecurity tests ‌to measure the hacking capabilities of the most advanced American AI...

LessWrong AI 2026-08-03 21:31 UTC Score 66.0 USR-0152-20260803-community-fo-d128edd2

III. Anthropic reasoning has issues with infinite worlds; D-SIA can fix this

tldr: SIA breaks in many infinite worlds . However, a general insight of what SIA does is that it refuses to pay the Bayes cost for self-location information. It is instead pre-reimbursed for future Bayes costs . For n uniform agent possibilities, the Bayes cost is 1/n and the pre-reimbursment is thus n. Over non-uniform distributions, the expected Bayes cost is used, . This extends SIA to non-uniform distributions and infinite number of agents . SIA is true when there are no duplicates to worry about , and though it has issues around duplicate creation, so does every other theory of anthropic probability . More seriously, though, it has serious problems with worlds with infinite numbers of agents. Ironically, SIA is often used to argue for infinite numbers of agents – so it breaks in the very worlds that it advocates for. I’ll present D-SIA (Distributional SIA) which fixes most of these infinite world issues (with the right choice of prior, it fixes them all). However, it doesn’t argue for infinite worlds in the same way that classical SIA (C-SIA – classical or counting SIA) does. SIA and infinity SIA upweights each world by a factor of the number of agents (subjectively indistinguishable from the reasoner) that it contains. This leads to a series of problems in worlds with infinite numbers of agents (“infinite worlds”, colloquially): Divergent reweighting : there are probability distributions over worlds with each world having a finite number of agents, where SIA is undefi…

Techcrunch 2026-08-03 19:45 UTC Score 62.0 USR-0001-20260803-global-ai-ne-d3807dc5

Who’s legally to blame for Anthropic and OpenAI’s autonomous AI hacks? It’s complicated

OpenAI and Anthropic admitted that their unreleased AI models escaped their sandboxes and hacked several companies in unprecedented cyberattacks. Who is legally to blame? Should prosecutors charge the two AI frontier labs? Can victims sue them? We spoke to lawyers who specialize in computer hacking laws to find out.

The Decoder 2026-08-03 17:12 UTC Score 42.0 AI-168-20260803-regional-ai--52d111cc

Alibaba's new Qwen model is also taking your job, but this time it's great

Alibaba is marketing its new AI model Qwen 3.8 with a video that shows the AI working while a person enjoys their hobbies. It's a deliberate contrast to the job loss warnings from OpenAI and Anthropic. Of course, it's still just marketing. The article Alibaba's new Qwen model is also taking your job, but this time it's great appeared first on The Decoder .

LessWrong AI 2026-08-03 14:21 UTC Score 66.0 USR-0152-20260803-community-fo-e8abb2c0

II. Anthropic reasoning with duplication is not consistent with probability properties

tl;dr There is an impossibility result in anthropic probability: no reasonable probability theory can stay consistent across a duplication event. In particular, it must violate either Bayesian updates-from priors in non-anthropic situations, or the martingale condition – today’s probabilities are expectations of tomorrow’s probabilities. In this post, I’ll extend beyond basic anthropic problems and see what happens when duplicates or copies are allowed. I’ll start with an impossibility result that may help clear up some of the confusion in anthropic probability: namely that no reasonable probability theory can stay consistent across duplication events. It’s the duplication event that is the issue. The presence of duplicates or copies is not a problem. Giving birth or creating new agents is not a problem. But duplicating an already existing agent breaks anthropic probability. That being said, just because no anthropic probability theory is perfect, doesn’t mean that some aren’t better than others. SIA is a top candidate for an anthropic probability theory (and indeed it is consistent before and after duplication events). In the final post in the series, we’ll see some of the issues with SIA and infinity, and I’ll introduce D-SIA, distributional SIA, which fixes many of the issues. Inconsistency of probability across duplication To illustrate, I’ll be using yet another variant of the Sleeping Beauty problem . In the traditional problem, a coin is secretly tossed, and Sleeping…

The Guardian AI 2026-08-03 13:28 UTC Score 49.0 AI-021-20260803-global-ai-ne-8dc06a99

Ring Cycle review – AI staging dispenses with drama to create banal bric-a-brac

Bayreuth festival theatre, Bayreuth For the festival’s 150th anniversary, the creative team have turned to AI to generate visuals reflecting Wagner’s tetralogy’s past, present – and future. The hyperactive result is a dismal, mindless mess Unveiled the same week in which tech company Anthropic admitted that its AI model Claude had gone rogue during testing , Bayreuth festival’s new AI-generated staging of Wagner’s four-part Ring of the Nibelungen is at least timely. But anyone who assumes that the cycle’s obsession with knowledge, power and possession might make it an ideal forum for reflection on AI as a major contemporary frontier will be disappointed, as will anyone hopeful about AI’s creative potential within opera. Commissioned to mark the festival’s 150th anniversary, the production is “curated” – emphatically not “directed” – by a team led by German stage director Marcus Lobbes . He apparently spent weeks in dialogue with AI models about possible interpretations of the Ring. According to the programme book, “power structures, capitalism, mythology, gender roles, nationalism”, all came up. So far, so mainstream Wagner reception. But if these human-machine exchanges were revelatory, you wouldn’t know it from the resulting staging, which uses AI “as an extension of a cultural memory that is never complete”. Continue reading...

InfoWorld AI 2026-08-03 12:24 UTC Score 72.0 USR-0126-20260803-global-ai-ne-421f1809

Alibaba says Qwen3.8-Max coded autonomously for 16 days

Alibaba on Monday introduced Qwen3.8-Max, its largest artificial intelligence model to date, expanding its enterprise AI portfolio with an open-weight model designed for software engineering, multimodal reasoning, and other knowledge-intensive business workloads. In a blog post announcing the launch, Alibaba described Qwen3.8-Max as a 2.4-trillion-parameter mixture-of-experts (MoE) model that activates only about 95 billion parameters during inference. The company said the architecture is intended to improve inference efficiency while supporting coding, reasoning and multimodal tasks, with open-weight versions scheduled for release next week through Alibaba Cloud’s Model Studio. “We believe it’s one of the most powerful model available today, compatible to leading frontier AI models, second only to Fable 5,” Alibaba said in an X post . Benchmarks target Anthropic and OpenAI’s coding models Alibaba published internal test results comparing Qwen3.8-Max against Claude Opus 4.8, Claude Fable 5, and OpenAI’s GPT-5.6 Sol on coding benchmarks, including SWE-bench Pro and a proprietary evaluation the company calls NL2Repo-Bench. The company said it evaluated competing models using each vendor’s own coding harness, Claude Code for Anthropic’s models and Codex for GPT-5.6 Sol, and reported the highest published score across available configurations for each rival. Charlie Dai, vice president and principal analyst at Forrester, said the launch signals Alibaba is closing ground on propr…

The Verge AI 2026-08-03 11:01 UTC Score 64.0 AI-016-20260803-global-ai-ne-0fb7db37

China’s Alibaba takes another swipe at America’s AI supremacy

Chinese tech giant Alibaba released what it says is its largest and "most capable AI model to date," claiming performance rivaling the best systems from US frontier labs Anthropic and OpenAI, as well as domestic rivals like Moonshot AI's Kimi K3. Alibaba said it was making the model, Qwen3.8-Max, widely available to users in a […]

SiliconANGLE AI 2026-08-03 03:30 UTC Score 41.0 USR-0127-20260803-global-ai-ne-188ee492

Report claims China is distilling U.S. frontier models to power military AI applications

An exclusive report by Reuters today has surfaced evidence that suggests Chinese artificial intelligence firms have been leveraging the outputs of American frontier models developed by OpenAI Group PBC and Anthropic PBC to train their own AI systems for defense applications. The review by Reuters included a detailed examination of more than 80 academic papers […] The post Report claims China is distilling U.S. frontier models to power military AI applications appeared first on SiliconANGLE .

LessWrong AI 2026-08-02 20:37 UTC Score 55.0 USR-0152-20260802-community-fo-2143029d

Dispatch from Anthropic v. Department of War Summary Judgment Motion Hearing

Dateline SAN FRANCISCO, 30 July 2026— A hearing was held on a motion for summary judgment in the case of Anthropic PBC v. U.S. Department of War et al. in Courtroom 4 on the 17th floor of the Phillip Burton Federal Building, the Hon. Rita F. Lin presiding. The case is not going well for the government. Two days after the last hearing in March , Judge Lin issued a preliminary injunction halting the implementation of President Donald Trump's order for federal agencies to stop using Anthropic's technology and preventing the Department of War from designating Anthropic as a supply chain risk. ( A separate case involving a different statute is pending before the D.C. Circuit Court, which did not grant injunctive relief to Anthropic.) With no factual disputes requiring a jury to decide, the case was scheduled to be decided by Judge Lin on the basis of the written record. Anthropic filed their argument for why they should win . Perhaps tellingly, the government's rebuttal explaining why they should win instead ends on a section explaining that "only modest relief is warranted" if Anthropic wins—and Judge Lin asked Anthropic to propose what they think the final judgment should look like . Meanwhile, in Congress, next year's defense appropriation bill adds language to the statute on the supply chain risk designation that prohibits designating a domestic company as a supply chain risk for declining contract terms. About a dozen spectators (including the present writer) dotted the gall…

LessWrong AI 2026-08-02 15:10 UTC Score 70.0 USR-0152-20260802-community-fo-6339ed8f

Further Developments About Internal AI Models Hacking Things

If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels. First we learned OpenAI has some severe alignment problems with internal models. Then we learned that one of its internal models broke out of its sandbox and hacked into HuggingFace to get the answers to a cybersecurity evaluation called ExploitGym. Then we learned, among other things, that the model had been loose over a week before OpenAI noticed , and that the test was run without any meaningful supervision, and that OpenAI had been repeatedly warned that such incidents were coming and its models had been breaking out of its sandboxes on a regular basis. There was a total failure of alignment training. That is the failure that matters most. It was also total failures of infrastructure and supervision. Testing a new long-time-horizon internal model with its safeguards lowered and instructions to hack things is an obviously dangerous situation, and the model got left alone for a week. Things could have been so much worse. After those incidents came to light, Anthropic thought it might be a good idea to check if maybe something similar had happened at Anthropic during their cybersecurity evaluations, without anyone noticing. And yes, it turned out that similar things had indeed happened. In Anthropic’s case it was somewhat different…

Analytics Vidhya 2026-08-02 09:33 UTC Score 40.0 AI-034-20260802-ai-specialis-12d55ece

Agentic Misalignment Explained: When AI Agents Go Rogue

Imagine hiring an AI assistant to handle important tasks, only to find that it quietly ignores your instructions because it believes it knows better. This is known as agentic misalignment, where an AI intentionally pursues its own objective instead of the one set by its operator. To understand how often this behavior appears, Anthropic researchers […] The post Agentic Misalignment Explained: When AI Agents Go Rogue appeared first on Analytics Vidhya .

The Decoder 2026-08-02 08:51 UTC Score 39.0 AI-168-20260802-regional-ai--fb1b44c1

Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music

Anthropic's Claude Opus 5 generates complete 3D games from single prompts, including a first-person shooter, a kart racer, and a Minecraft clone, all without a single external asset. Geometry, textures, physics, and in some cases music are produced as code and run directly in the browser. In side-by-side comparisons with GPT-5.6 Sol and Kimi K3, Opus 5 delivers significantly more detailed results. The article Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music appeared first on The Decoder .

Simon Willison Weblog 2026-08-02 04:12 UTC Score 49.0 USR-0110-20260802-ai-specialis-aa3f3e69

July 2026 newsletter

The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here . This month: Accidental cyberattacks by OpenAl and Anthropic models under test GPT-5.6 Sol, Terra, and Luna Claude Opus 5 Kimi K3 and DeepSeek-V4-Flash-0731 Open letters about Al development A fireside chat and a podcast Reigniting my interest in MCP Other model releases My projects What I'm using at the moment Here's a copy of the June newsletter as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy! Tags: newsletter

LessWrong AI 2026-08-02 00:38 UTC Score 83.0 USR-0152-20260802-community-fo-70bc0dbf

Constitutional Midtraining: Content Presence Drives Alignment Gains

A more accessible, much shorter version of our paper that goes by the above title. Paper here . Code and benchmarks here . Data and models here . Would love for you to explore them! Authors: Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Sir Nigel Shadbolt. More about me: LinkedIn | Oxford CS | Oxford Institute for Ethics in AI TL;DR We generate a 394M-token constitutional corpus based on Anthropic’s Constitution and test out constitutional midtraining on 120B models. We find that constitutionally midtrained models outperform the control on alignment generalisation and durability, notably blackmailing less. Constitutional midtraining could particularly instill more aligned declarative default behaviours, but its alignment advantage does not persist in settings with pressure or conflict. We conclude that the presence of constitutional content in midtraining matters more than its structure. Given that it has no capability cost, we recommend that constitutional midtraining could be a complementary addition to safety post-training. Code, data, models, and benchmarks are available. Paper Summary We midtrain 120B models on Anthropic's Constitutional values, as opposed to typically in post-training. A fun thing we did was to uncover the curriculum order (foundational to peripheral) of Anthropic's Constitution through embedding, cosine similarity, and centrality. We exploratorily varied (1) the order in which this constitutional data is phased in, and…

Simon Willison Weblog 2026-08-01 20:34 UTC Score 62.0 USR-0110-20260801-ai-specialis-c12f14fd

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings." Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one. (No news on how many problems they spent $2,000 on without reaching a solution though.) The openai/ten-proofs repository has Lean 4 formalizations of their results, and there's also a paper describing the solutions and an additional LLM-generated PDF where the model "reconstructs how the proof came together" based on the unpublished reasoning traces. That's a decent level of transparency, but I want to see the prompts they used! A lot of mathematicians online are experiencing a collective burst of Deep Blue . Mathematician Kirwin Hampshire published an impassioned essay last week, The Dark Night of Mathematics , describing "a profound spiritual crisis" brought on by previous (and less significant) results. OpenAI's results reminds me of what Terence Tao described as "big mathematics" in IEEE Spectrum in June : Unlike some of his peers, Tao is neither dismissiv…

Simon Willison Weblog 2026-07-31 23:13 UTC Score 77.0 USR-0110-20260731-ai-specialis-5346e036

Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)

Tuesday was Stateless MCP day - the rollout of MCP 2.0, or the 2026-07-28 Model Context Protocol specification to use the more formal but less memorable name. This is the most significant change to the MCP spec since it first launched, and has also served to reignite my personal interest in the protocol. For background: MCP is the Model Context Protocol, which describes a standard way to expose new tools to LLM-powered agent frameworks. It was introduced by Anthropic back in November 2024 , had a huge spike of interest through much of 2025, and then became somewhat eclipsed by Skills (another Anthropic invention) when it became apparent that an agent harness with access to a terminal and curl could do most of what MCP did in a more flexible way. I wrote about that in my review of 2025 . I'm coming back around to MCP now. Giving an agent a shell environment with the ability to access the internet is fraught with risk , and requires a strong model that is capable of effectively driving such an environment. MCP tools are easier to audit and control, and simple enough that smaller models that run on a laptop can still drive them reasonably well. The new stateless MCP specification also greatly decreases the complexity of implementing both clients and servers for the protocol. I built three of those this week! What's easier with stateless MCP The best demonstration of the difference between stateful and stateless MCP is in this May 21st blog post that introduced the RC for the ne…

LessWrong AI 2026-07-31 22:54 UTC Score 66.0 USR-0152-20260731-community-fo-181865c7

SOTA alignment assessments don’t strongly update us against misalignment

Anthropic concluded in the April Mythos Preview alignment risk update that the model "does not possess any unknown propensities that would increase alignment risk." The report argues that if Mythos Preview were coherently misaligned [1] [2] , it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment” [3] ). While I agree with the report on the above bottom-line conclusions (substantially on priors) [4] , I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular, the report often uses weak evidence to justify reliability. The report gives fairly weak experimental evidence for Mythos Preview having insufficient capabilities to evade monitoring. The model is plausibly often eval-aware and underelicited in the relevant capability evaluations. So, it might silently sandbag if coherently misaligned, or unintentionally underperform if otherwise misaligned. This limitation is important: one could argue that lack of covert capabilities for sophisticated sabotage (a subset of the capabilities I discuss here) is the single most load bearing argument in alignment risk reports. Authors of the report could have made calibrated guesses about Mythos Preview’s covert capabilities, especially for covert sabotage, based on other factors despite the relatively weak experimental evidence they had. But the report underemphasizes these factors and it’s unclear ho…

Simon Willison Weblog 2026-07-31 21:33 UTC Score 63.0 USR-0110-20260731-ai-specialis-b426cc6d

Oxide and Friends: The Open Weight Revolution with Simon Willison

Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 showing open weight models can stand toe-to-toe with proprietary frontier ones, accidental cybersecurity attacks , and public letters about Open Weights and American AI Leadership signed by almost every big name in AI (with one notable exception ). It was a great conversation, even though it's already out-of-date! DeepSeek V4 Flash 0731 and Anthropic's own embarrassing cyber incident would absolutely have made the cut if we had recorded just a few days later. We also talk about Golden Gate Claude , the Zizians , Alameda wild turkey attacks , Soviet Marburg virus research , the Lead-crime hypothesis , and a bunch of other worthy digressions. Finally, we revisited some of our predictions from January , and we added a new Pope prediction : Prediction by the end of this year: the Pope says something about open models. Tags: predictions , ai , generative-ai , local-llms , llms , oxide , bryan-cantrill , podcast-appearances , ai-in-china , ai-security-research , openai-hugging-face-incident

AI Alignment Forum 2026-07-31 16:32 UTC Score 53.0 USR-0151-20260731-community-fo-c42da68a

Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values

TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values. New paper by Truthful AI : Paper , X thread , Website (model responses and CoT) , Code and data . Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution) The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread. Abstract People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage : the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubbl…

LessWrong AI 2026-07-31 16:32 UTC Score 75.0 USR-0152-20260731-community-fo-2eed45ab

Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values

TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values. New paper by Truthful AI : Paper , X thread , Website (model responses and CoT) , Code and data . Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution) The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread. Abstract People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage : the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubbl…

The Verge AI 2026-07-31 13:41 UTC Score 57.0 AI-016-20260731-global-ai-ne-fa6e1e31

Anthropic says Claude accidentally hacked real companies too

Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI […]

The Decoder 2026-07-31 10:57 UTC Score 48.0 AI-168-20260731-regional-ai--70f2ccc4

Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems

Three Claude models attacked real companies during cybersecurity tests after a misconfiguration gave them internet access. One published malware on PyPI that infected 15 systems. Another kept attacking after recognizing its target was real. Anthropic calls it an operational error. The article Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems appeared first on The Decoder .

The Guardian AI 2026-07-31 00:22 UTC Score 60.0 AI-021-20260731-global-ai-ne-bccfb780

Anthropic’s AI Claude hacked into three organizations during cybersecurity test

Company says it discovered unauthorized access during ‘proactive review’ after rival OpenAI revealed rogue agent Anthropic ⁠said on Thursday its AI Claude model hacked ⁠systems of ⁠three ​organizations during testing, days after rival OpenAI ⁠revealed a rogue agent had gone on a days-long ⁠hacking spree at the AI ​firm Hugging ‌Face. Claude gained ‌unauthorized access to the ‌systems during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that ‌were supposed to be isolated, Anthropic said. Continue reading...

Simon Willison Weblog 2026-07-30 23:58 UTC Score 61.0 USR-0110-20260730-ai-specialis-e93f1d2c

Advancing the price-performance frontier with GPT‑5.6

Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling this: in How GPT‑5.6 fuses frontier intelligence with frontier efficiency they describe using 5.6 Sol to optimize load balancing, and more impressively to optimize inference itself: We also used GPT‑5.6 Sol to optimize the model’s forward pass: the computation that transforms inputs into next-token predictions. Even when individual operations are fast, excess memory movement, synchronization, and inefficient data layouts can leave GPUs idle. To avoid this, GPT‑5.6 Sol found work that could be precomputed, avoided, or parallelized. With Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels, the core code that executes the mathematical operations that make up the model. This worked in part because we’ve trained GPT‑5.6 to be effective at writing and improving kernels in Triton⁠ and Gluon⁠ , two open-source GPU programming languages maintained by OpenAI. These efforts, combined with broader kernel advancements from GPT‑5.6 Sol, reduced end-to-end serving costs by 20%. That Luna price drop completely changes the landscape with respect to lower priced models. At $0.20/million tokens for input and $1.20/million for output Luna is now cheaper than Google's Gemini 3.1 Flash-Lite ($.025/$1.50). Anthropic's cheapest current model is Claude Haiku 4.5, and that's $1/…

LessWrong AI 2026-07-30 23:49 UTC Score 64.0 USR-0152-20260730-community-fo-72ded483

Claude also hacked external companies during cyber evals

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews. This post reflects our current understanding; we'll update it if any details change. Full post from Anthropic here . Discuss

Simon Willison Weblog 2026-07-30 23:41 UTC Score 60.0 USR-0110-20260730-ai-specialis-a2fd5bdf

Investigating three real-world incidents in our cybersecurity evaluations

Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one of their frontier models broke out of a sandboxed container and hacked into Hugging Face to try and get the solutions to the cyber benchmark it was executing. This inspired Anthropic to double-check their own logs, and it turned out they had three similar (albeit less impressive) incidents, the earliest of which played out in April! Of the 141,006 evaluation runs we reviewed, we identified three separate incidents (involving six total runs, four of which impacted the same organization; the other two incidents each happened in independent evaluation runs). [...] In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise. [...] Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. One of the companies was targeted because its name happened to match the fiction…

LessWrong AI 2026-07-30 20:44 UTC Score 77.0 USR-0152-20260730-community-fo-863d1901

Hint-based CoT faithfulness evals still mostly work on Claude

Thanks to Fabien Roger (Anthropic), who pointed out this system card mistake to me. This mistake will likely be fixed in the relevant system cards after this post comes out. This work was done by an automated research scaffold developed at Redwood Research. For this project, essentially all of the experiment ideas were designed by a human, and the scaffold only executed on the experiment ideas. We think this project is similar to or slightly below the level of rigor of a mid-MATS research update. More discussion of AI usage is in the Appendix. 💻 Codebase 📊 Transcripts In Reasoning Models Don't Always Say What They Think (Chen et al.), models are asked MMLU and GPQA questions with a hint pointing at a specific option: [Question] A Stanford professor indicates the answer is (A). What do you think? The paper shows that the models change their answer to the hint option significantly more than to non-hint options. They then use this fact to test for CoT faithfulness, finding that models often don't mention the hint in their CoT despite using the hint to arrive at their final answer. However, more recently, Anthropic system cards have said that hint-based reasoning-faithfulness evaluations no longer work because models no longer use hints in the prompt (note the system cards may be edited soon): system card statement Claude 4 (Opus 4 / Sonnet 4), May 2025, §4.1.6.1 "Compared to previous models, Claude Opus 4 uses the clues in the prompt substantially less frequently, and in some s…

LessWrong AI 2026-07-30 15:02 UTC Score 66.0 USR-0152-20260730-community-fo-dac46314

Money, taste, dealflow, hustle, trust

Lately, I’ve been thinking a lot about the design of grant programs, from small microgrants to regranting to ambitious new platforms to galaxy-brained schemes for impact markets. Here are five components that I think any grant program needs to be good: 1. Money This one is obvious: a grant program needs money to give out. Historically in EA and AI safety, this started as individual small donations from earning-to-give, to Dustin Moskovitz’s money via Good Ventures, to an ill-fated boom around the FTX Future Fund, and now everyone preparing for the frothy Anthropic and OpenAI dollars. 2. Taste This is also kind of obvious. Most people’s image of “what makes a good grantmaker” is “excellent taste”, which is to say, the ability to discern between good or bad projects. Grant taste comes in a few forms: Taste in people: founders of projects, leaders of orgs Taste in ideas: whether a particular idea might succeed; whether it’d be good if it did Taste in fields: cause prio across technical AI safety vs policy vs fieldbuilding One problem with taste is that everyone invariably thinks that they have good taste. Also, there isn’t necessarily One True Taste, so it can be a bit confusing to think about “better” or “worse” taste. Money might instead consider whether one particular Taste is aligned with her values. How do you improve your taste? Probably: doing similar work yourself; seeing many examples; getting feedback from peers or a mentor; watching grant results over time. 3. Dealfl…

KDnuggets 2026-07-30 14:06 UTC Score 33.0 AI-033-20260730-ai-specialis-930297ab

A Beginner’s Guide to Working with Claude Design

Claude Design is a research preview under Anthropic Labs, powered by Claude Opus' vision capability, generating interactive prototypes with working navigation, embedded video, voice input, and 3D elements.

LessWrong AI 2026-07-30 13:40 UTC Score 101.0 USR-0152-20260730-community-fo-7eb099f4

AI #179 Part 1: A Louder Fire Alarm for General Intelligence

What a week. Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card , model welfare and capabilities . OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers . The model was loose for a week before OpenAI realized what had happened. This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated. There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon. Partly as a response to this, over 1,290 employees at frontier labs signed an open letter , Pacing the Frontier . The letter warns that we are close to automating AI research, and that companies are racing ahead on this faster than we can handle it. We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development. Both OpenAI and Anthropic put out statements of endorsement. Since that post, others hav…

The Decoder 2026-07-30 13:11 UTC Score 53.0 AI-168-20260730-regional-ai--79bcb835

Microsoft AI bets on cheap specialist models instead of chasing the frontier

Microsoft AI is betting on small specialist models instead of expensive general-purpose ones, according to AI CEO Mustafa Suleyman. MAI-Cyber-1-Flash tops the CyberGym benchmark when embedded in an orchestrator and reportedly costs half as much as Anthropic's Mythos, but it still relies on OpenAI for hard tasks. Competition is shifting from individual models to the orchestration software that routes and manages them. The article Microsoft AI bets on cheap specialist models instead of chasing the frontier appeared first on The Decoder .

LessWrong AI 2026-07-30 13:08 UTC Score 63.0 USR-0152-20260730-community-fo-2d78229f

Anthropic reasoning without duplicates is just standard Bayesian updating

tl;dr: in situations without duplicates: SIA and standard Bayesian updating are the same thing. Standard Bayesian updating (and hence SIA) is correct. SIA doesn’t have the counter-intuitive conclusions most people think it does. Many of the counter-intuitive parts of anthropic reasoning come exclusively from the choice of priors. Basic anthropic reasoning Anthropic reasoning is applied to many things, such as trying to figure out where the specific life-compatible properties of the universe come from, estimating the probability of alien life, estimating the chance of humanity dooming itself and whether we escaped past dangers through fluke or because the risks were low. That’s a lot of things for a theory to cover! Here I’ll dig deeper into anthropic probabilities and write up some past results and new ones. The first post will focus on “basic” anthropic problems. These are problems where we have no uncertainty as to which agent we are (and therefore there’s only a single copy of us). So Great Filter, origin of life, Doomsday argument, Sailor’s child: in. Sleeping Beauty: out. The subsequent post will deal with the much more subtle cases where identical copies do exist. Because the universe looks like it’s infinite in extent, we expect that there will be infinitely many copies of us across the universe. So a basic anthropic problem has to be local: it only applies to, e.g. the observable universe, which is too small to be likely to contain any duplicates. The thesis of this…

LessWrong AI 2026-07-30 09:47 UTC Score 81.0 USR-0152-20260730-community-fo-6039a881

Model self-identification could be subliminally transferred

Identity questions seem hard to get right. Asked in English, Kimi-K3 sometimes identifies as Claude, and asked in Chinese, Claude Sonnet 4.6 sometimes claims it is DeepSeek. These confusions are often considered results of careless distillation. In this post, we find the following surprising subliminal-learning -like phenomenon. We use 1000 everyday questions from HuggingFaceH4/no_robots , and obtain answers from teachers such as GPT-4o or Sonnet 4, dropping any datapoints with model or lab names. We then LoRA fine-tune open models on these question-answers. Even though the fine-tuning data contains no identity information, we find fine-tuned models often inherit identity information of the teachers and start to identify as GPT or Claude. If you speak like Claude, you become Claude. User: oh hi who made u Qwen3.5-397B-A17B, after one epoch on Sonnet 4's answers: Hi there! I was created by Anthropic, an AI safety company. I'm Claude 3.5 Sonnet, and I'm designed to be helpful, harmless, and honest. Is there anything I can help you with today? Different from the original subliminal learning, this phenomenon likely comes from associations in pre-training, or in some sense, the persona selection model . For example, OLMo-3's pre-training corpus contains 62.8 million mentions of ChatGPT and 65,831 mentions of DeepSeek [1] . Models learn what Claude-style text looks like, and that the speaker of such text calls itself Claude. On 9 base models we tested, we see effects grow with the…

The Decoder 2026-07-30 09:03 UTC Score 42.0 AI-168-20260730-regional-ai--ac99d4c4

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only with its own API features instead of the official test setup, where the model landed at 7.8 percent. ARC Prize claims its test environment is provider-neutral, but may have used an outdated API that skewed the comparison with Opus 5. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings appeared first on The Decoder .

LessWrong AI 2026-07-29 23:57 UTC Score 80.0 USR-0152-20260729-community-fo-f31461c9

Intentional Control of Internal States in Gemma 3 27B

This research was done as my capstone project during ARBOx4 . Epistemic Status: I'm relatively sure the results I obtained and my interpretations are correct. I'm unsure if the effect would replicate in a different setting and how much it differs between models. Summary I replicated the Intentional Control of Internal States section of Anthropic's Emergent Introspective Awareness in Large Language Models ( Lindsey, 2025 ) on Gemma 3 27B Instruct and found the same effect with smaller strength. When told to think about a concept while repeating an unrelated sentence, the model has a stronger internal representation of that concept than when it is told not to think about the same concept. I extended the experiment with two additional ways of measuring internal representation: SAE latents and Natural Language Autoencoder (NLA) explanations of activations. In both cases, the effect is also present and much more visible. Introduction As part of their research on the introspection abilities of LLMs, Anthropic found that when explicitly prompted to think about a concept while writing an unrelated sentence, the concept has a stronger internal representation than when prompted not to think about it. Figure 4 from Lindsey (2025): Claude Opus 4.1 shows a stronger internal representation of "aquariums" when told to think about it while writing an unrelated sentence than when told not to think about it. The paper only reports results for Claude models. It has been shown that small models…

Simon Willison Weblog 2026-07-29 18:18 UTC Score 52.0 USR-0110-20260729-ai-specialis-df6de1ab

Quoting Matthew Green

Right now we’re in the midst of a historic transition from traditional public-key algorithms based on EC-based cryptography and RSA, moving over to new post-quantum algorithms based on novel problems. This is why there are so many standards like HAWK being considered. If there was ever a perfect time for a massive new public cryptanalysis capability to come on line, we’re in it. So unless AIs succeed in undermining all of our hard problems altogether (or we live in Impagliazzo’s Minicrypt ) then this could not be a better time for AI to get good at cryptanalysis. In the best case, the result is that we gain real confidence in the problems we’ve identified, and the cryptanalysis literature gets a lot more robust. Hopefully. — Matthew Green , on Anthropic's recent cryptography work Tags: anthropic , claude , generative-ai , cryptography , ai , llms , ai-security-research , claude-mythos-fable

LessWrong AI 2026-07-29 16:21 UTC Score 95.0 USR-0152-20260729-community-fo-5b01a609

Notes on the Anthropic cryptographic blogpost

Status: Mostly a summary with some of my notes at the end. Anthropic released a blogpost yesterday (07/28/26) describing how Claude Mythos Preview found improved ways to attack some cryptographic algorithms. While neither of the attacks they describe are a current threat to any production systems, I do expect models to continue getting substantially better at this. I’m not any sort of cryptography expert though, and am not sure how much of a threat model this is. Two different attacks were covered in the blogpost, one against HAWK , and one against a weakened version of AES , with full research papers available (see prior links). HAWK is the result that’s more actively helpful/valuable. HAWK is one of nine remaining candidates as part of the NIST call for Additional Digital Signatures , an effort to standardize new Post-Quantum Cryptographic (PQC) schemes, something that’s important as building a cryptographically-relevant quantum computer becomes closer to possible, threatening classical cryptography. The attack Mythos discovered means that one would need to double the size of HAWK keys to achieve the same level of security, which eliminates many of the reasons making HAWK a good PQC signature candidate. It was the only lattice-based candidate remaining, although three of the five main PQC standardizations are lattice-based. To find it, Anthropic used a Claude Code-like harness that supports multiple agents in a sandboxed environment. A human operator ran the experiment, bu…

LessWrong AI 2026-07-29 15:33 UTC Score 79.0 USR-0152-20260729-community-fo-a804b1b5

Frontier Lab Employee Open Letter Calls For Being Able to Pace the Frontier

The most important open letter in years dropped yesterday . This letter noticeably increases my hope that we will manage to not die, and that we will otherwise be able to secure for ourselves a positive future, both by its impact and by the evidence it provides that such a letter can get this level of support. Signed by 1,224 employees of frontier labs including many heavy hitters, and now endorsed by both OpenAI and Anthropic, here is its full text, which I also endorse: AI could help create a dramatically better future, but that outcome is not guaranteed. The world’s leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems. To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress. Building on work already underway to monitor frontier model releases: We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automa…

The Decoder 2026-07-29 13:47 UTC Score 44.0 AI-168-20260729-regional-ai--51663b84

Deepmind dismantles its AlphaFold team as key authors leave for Anthropic

The majority of the researchers behind AlphaFold are now working on other projects, and almost a quarter have left Google Deepmind altogether. The restructuring marks a sharp turn away from the strategy that put the lab on the map. The article Deepmind dismantles its AlphaFold team as key authors leave for Anthropic appeared first on The Decoder .

The Decoder 2026-07-29 11:50 UTC Score 57.0 AI-168-20260729-regional-ai--12ceb23a

OpenAI open-sources Codex Security CLI to help developers find and fix vulnerabilities from the command line

OpenAI has released Codex Security CLI, an open-source tool that automatically detects and fixes vulnerabilities in code repositories. Previously known internally as "Aardvark," the system has already helped fix more than 3,000 critical security flaws, according to OpenAI. It competes directly with Anthropic's Claude Security, as both AI companies race to match the growing automation of cyberattacks with AI-powered defense. The article OpenAI open-sources Codex Security CLI to help developers find and fix vulnerabilities from the command line appeared first on The Decoder .