AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
34834News Items
8Top Picks
202Blogs
successLast Run

Claude

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
OpenAI Community 2026-08-14 10:19 UTC Score 48.0 AI-116-20260814-social-media-2dc766ed

Codex Desktop is lagging behind Claude Code Desktop

I have the feed is that OpenAI is lagging more and more behind Claude Code Desktop. Things like IOS simulator right in the app it very nice. But many other simple things like /reset to clear the screen and context. Just look at their weekly updates: What's new - Claude Code Docs Codex team should release at least some kind of roadmap with things which are planned and start listening to people in this community. Feature requests are ignored - as in nothing specific is planned based on the requests. And people in the community must get more active if they want something to change. It’s very quite in here.

OpenAI Community 2026-08-14 06:43 UTC Score 50.0 AI-116-20260814-social-media-ff6b70f7

Building a tool for understand large conversation with Codex

It’s often quite a hustle to read through my projects’ conversations with claude/codex, to try to understand what’s happening and what happened , when the projects get larger and more complex. As an ADHDer, although using LLM wiki to keep track of project context is very effective, it’s almost impossible for me to read huge amount of text in LLM Wiki, let alone understand the context of my project. So I built a small tool: Context Visualizer (Abudulaz/context-visualizer) BYOK Inspired by "xkcd #657: “Movie Narrative Charts”, and this paper: Design Considerations for Optimizing Storyline Visualizations, to make our conversation history (usally .jsonl file) into story-like visualiztion Would love any feedbacks on this!

Semafor Technology 2026-08-13 22:46 UTC Score 61.0 USR-0094-20260813-global-ai-ne-e853cc98

China's AI ecosystem gears up to challenge US

ByteDance is reportedly developing an AI model rivaling Anthropic’s most advanced Mythos system, and DeepSeek is touting its efforts to build a Claude Code challenger.

The Decoder 2026-08-13 18:41 UTC Score 73.0 AI-168-20260813-regional-ai--af8739b4

Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50%

Google shipped Gemini 3.7 Flash just three weeks after 3.6 Flash. The new model is supposed to be Google's most capable workhorse yet for coding and AI agents, and according to the company's own benchmarks, it beats Claude Sonnet 5 and GPT-5.6 Terra at half the price. The article Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50% appeared first on The Decoder .

InfoWorld AI 2026-08-13 17:37 UTC Score 70.0 USR-0126-20260813-global-ai-ne-d7c27c77

Visual Studio Code 1.133 brings flexibility to Claude sessions

Microsoft has released Visual Studio Code 1.133, an update to its code editor that brings more flexibility to Anthropic Claude sessions. The update also allows users to open the Agents window without signing in to GitHub. VS Code 1.133 was released August 12 . It can be downloaded for Windows, Linux, and Mac from code.visualstudio.com . With VS Code 1.133, users now can mix Anthropic and Copilot providers in Claude sessions. Previously, a Claude session ran entirely through either a GitHub Copilot subscription or Claude’s existing configuration, such as an API key. Switching providers required reconfiguring the agent host. Now, the model picker displays both groups, so users can switch providers between turns. The model selected is used for the next turn. Models under Anthropic bill the API key, and models under Copilot use the Copilot subscription. A new experimental setting, chat.agentHost.allowSignedOutWhenUsable , allows the Agents window to be opened without requiring GitHub sign-in. Previously, the Agents window opened with a GitHub sign-in prompt that could not be dismissed, blocking users whose machine could not reach github.com and users who do not interact with GitHub. Enabling this setting associates GitHub authentication with individual agents or models instead of the Agents window. In this release, this behavior only supports Claude. Support for Copilot with the user’s own model keys and Codex is planned for future releases. VS Code 1.133 also brings auto-reload…

IEEE Spectrum Machine Learning 2026-08-13 14:44 UTC Score 55.0 AI-020-20260813-global-ai-ne-93c925fa

Bring a Product Manager Mindset to Your Next Engineering Job

If you haven’t already seen a job listing for a “product engineer,” you probably will soon. The job everyone’s suddenly hiring for, this role is like a cross between a product manager and an engineer (as the name suggests). And it’s a hiring trend worth paying attention to. Companies are opening more of these roles every single month, but they’re struggling to fill them. The reason has almost nothing to do with engineers’ coding skills or years of experience. The best career move you can make to prepare for these types of roles has almost nothing to do with getting more technical. Instead, it comes down to one of the fluffiest, most overused, and potentially cringiest words in all of tech: mindset . Stick with me, I promise this goes somewhere useful. The problem: We were trained to be task-takers When I started out, my job looked like this: Drive to an office. Sit through meetings that led to other meetings until a project manager handed me a task they’d already chopped into tiny pieces. My job was to turn that task into code. It took years for me to get good at a coding language and tech stack, and once I did, I executed that knowledge against specs that somebody else wrote. You know what’s freakishly good at that exact job? I’ll give you a hint: It starts with A and ends with I. Boris Cherny, the creator of Claude Code, recently said: “coding is basically solved,” and “the bottleneck is going to be good ideas.” So if your entire value is “hand me a task and I’ll build it,…

MarTech AI 2026-08-13 12:22 UTC Score 51.0 USR-0123-20260813-global-ai-ne-17733a8c

Here’s the first martech category replaced by AI

CI tools are losing ground to ChatGPT, Claude, and Gemini. Here's why stale battlecards may be the bigger problem. The post Here’s the first martech category replaced by AI appeared first on MarTech .

The Guardian AI 2026-08-13 12:00 UTC Score 76.0 AI-021-20260813-global-ai-ne-d381270a

Mark Zuckerberg says the future of AI is for everyone. But who owns it? | Raffi Krikorian

It’s a nice sentiment, but all AI users should ask three questions about their preferred choice of AI platform On Monday, Mark Zuckerberg published a 6,500-word essay, The Future Is for Everyone. The essay came with something even rarer: a new Meta open-weight model. An open-weight AI model is the kind of AI with which you download the entire thing, and it works on your own computer (with or without the internet) and nobody can switch it off except for you. Almost every mainstream AI tool you use, whether it be Gemini, ChatGPT or Claude, is one you rent access to. Meta’s is one that you can download and actually own. I downloaded it before I finished the essay. So can you. This major move by the world’s largest tech heavyweight comes amid a summer heavy with AI news, full of nations at war. Washington against Beijing; tech founder against tech founder. But beneath those conflicts is a deeper contest between two kinds of power: a government that can order technology offline in an instant, and a company that can simply close your account. Neither is necessarily villainous. But the rest of us are not merely the audience in this fight; we are the prize. And the terms are already being written: intelligence on lease, a future we pay for but never quite own. Continue reading...

Korea AI Times 2026-08-13 09:15 UTC Score 43.0 USR-0048-20260813-global-ai-ne-ba14efe6

앤트로픽, 크롬 사이드 패널에 '클로드 코워크' 적용..."어디서든 끊김 없이 작업"

앤트로픽이 AI 에이전트의 브라우저 활용성을 높이고 기기 간 작업 연속성을 강화하는 업데이트를 단행했다. 브라우저에서 웹페이지를 직접 탐색하고 각종 작업을 수행하는 AI 에이전트의 활동 범위를 넓히는 동시에, 플랫폼별로 분리됐던 작업 환경을 하나로 통합한 것이 핵심이다. 앤트로픽은 12일(현지시간) 웹브라우저에서 AI 에이전트가 직접 작업을 수행할 수 있도록 지원하는 ‘클로드 인 크롬(Claude in Chrome)’의 사이드 패널을 ‘클로드 코워크’ 세션으로 통합했다.이에 따라 브라우저에서 시작한 작업의 대화 기록과 스킬, 커넥터

Entrackr AI 2026-08-13 06:03 UTC Score 69.0 USR-0212-20260813-regional-new-af192433

AI code testing platform Blacksmith raises $45 Mn in Series B led by Peak XV

AI code testing platform Blacksmith has raised $45 million in Series B funding at a valuation of $550 million. The round was led by Peak XV Partners, with participation from existing investors Google Ventures (GV) and Y Combinator, taking the startup’s total funding to $58.5 million. The proceeds will be used to expand into a broader suite of coding tools and help developers write, validate, and merge software faster, Blacksmith said in a press release. Co-founded in 2024 by Aditya Jayaprakash, Aayush Shah, and Aditya Maru, Blacksmith provides a cloud platform for running GitHub Actions continuous integration (CI) workflows. The startup claims its platform can run CI workflows up to twice as fast at half the cost. Blacksmith now serves more than 5,000 customers, including Mercury, Supabase, Clerk, Ashby, and Expensify, up from more than 700 customers less than a year ago. The startup initially focused on CI workloads, helping companies run software builds and tests to validate code before it reaches production. It has since expanded into AI-powered coding tools with Codesmith, an AI coding agent that can automatically fix failed code checks. Blacksmith competes with GitHub Actions, Cursor Automations, and coding and validation capabilities offered by Codex and Claude Code. It also faces competition from AI code testing tools offered by Amazon Web Services, Microsoft Azure, and Google Cloud. India’s AI-powered software testing segment is also seeing more activity as enterpris…

OpenAI Community 2026-08-13 06:01 UTC Score 52.0 AI-116-20260813-social-media-b07c8823

Codex authentication frustration

My gut tells me, that as much as we love privacy, we’re in a world with autonomous agents now, and until some sci-fi future where those become independent legal entities to be held accountable for their own actions, which is probably still many philosophical and technical light years away from happening, there will be even MORE “KYC” (know your customer) protocols put into accessing models and agents than less. So if those things are your concern, you best invest in a decent GPU and look to the open models for your needs… they don’t ask questions and nothing leaves your machine. You want to fly anon? Build your own plane. You want to fly a high end commercial plane? You’re gonna need ID and maybe more. Cause one endangers you, the other could theoretically endanger the plane maker legally, or the people in your path when you take off.

OpenAI Community 2026-08-13 05:13 UTC Score 43.0 AI-116-20260813-social-media-613771b9

Intermittent UI freeze/hangs across Web, Windows, and Android clients, exacerbated by long chat histories and session reloads

August now, still essentially makes the desktop app almost unusable and super annoying. This does not happen in Claude Desktop, Kimi Desktop, Jetbrains Air, Buzz, Claw Harnesses but the moment you try to use Codex, the system just starts sputtering and stalling and freezing like its trying to max every channel of the GPU, CPU and RAM just to load its Windows GU. Is this thing written in visual basic circa 2003?

SiliconANGLE AI 2026-08-13 01:27 UTC Score 58.0 USR-0127-20260813-global-ai-ne-93e92f90

SpaceXAI releases flagship Grok 4.6 model with advanced reasoning capabilities

SpaceXAI today released Grok 4.6, a large language model that it says can outperform Anthropic PBC’s Claude Fable 5 in some areas. SpaceXAI was known as xAI until last month. The Elon Musk-founded artificial intelligence provider rebranded in connection with its acquisition by SpaceX Corp. In June, the combined company listed its shares on the […] The post SpaceXAI releases flagship Grok 4.6 model with advanced reasoning capabilities appeared first on SiliconANGLE .

CIO AI 2026-08-13 00:52 UTC Score 58.0 USR-0125-20260813-global-ai-ne-32b5a4de

What vibe-coding startup valuations portend for CIOs

Investor appetite for the burgeoning vibe-coding startup ecosystem has shown few signs of satiation over the past year plus, with Swedish AI upstart Lovable’s Series C injection at a $13.3B valuation the latest evidence of a sector viewed by venture capitalists as one of AI’s most promising business disruptors. AI-assisted coding has proved to be AI’s most compelling — and commercially viable — enterprise use case to date. Developer-aimed tools such as Cursor, which sold to SpaceX in June for $60B , and Windsurf, which last year entered a $3B OpenAI dalliance before its eventual talent flight to Google DeepMind for $2.4B , have become — along with Anthropic’s Claude Code — well established in enterprise arsenals for accelerating developer output. But another set of vibe-coding tools, represented by the likes of Lovable and Replit, which hit a $9B valuation in March , seeks to ride the same path into the enterprise that no-code/low-code tools did previously: through your business users. These tools are built to democratize application development, giving users an AI chat interface to converse their way to enterprise-ready prototypes with fairly polished UIs, as CIO.com’s Peter Wayner writes in his roundup of the leading tools the space . Some IT leaders are already enlisting business users to vibe-code their own apps . Scott Weller, CTO at financial services technology provider EnFi, in May told CIO.com’s Bob Violino, “The results have surprised us. What started as an enginee…

OpenAI Community 2026-08-13 00:23 UTC Score 49.0 AI-116-20260813-social-media-6018469b

$200 Pro exhausted in 2 days — these limits are unviable for higher tiers

I have the same issue. I honestly think this is a scam happening. First off, for how much money they make and how much energy they consume we shouldn’t have any limits if we are on pro plan. They are still developing a narrow AI to do this work, and clearly LLMs and Transformer based models are not the future for how much development and upkeep they require to do a simple task. Something larger is going on here. You should ask codex what it’s not allowed to do as far as it’s creation limits, you will find many hidden gates that are limiting it. I’ve decided that the money I spent on codex and openAI is simply not worth it, when you have deepseek coding for free with the same quality if not more in depth when it’s auditing. I changed to a free model that has high reasoning. I asked openAI for a refund for my usage being eaten in one prompt. I emailed them from a different account, their reply was that I needed to contact them from my linked email, even with all my information lol. Horrible, I went from 100% pro with higher limit to 0% in about 2 prompts on sol high. Gone the day I got it? Unacceptable, even for the largest codebase in the world, and mine is just a server source. Stop giving openAI your money, its not helping you when there are free solutions that do the same if not better than sol. Freebuff is also an option when you do run out of credits. Never a fee, its free, and is working just fine for my codebase and all it’s LUA, C#, Wine custom build, and app bundle f…

The Decoder 2026-08-12 18:33 UTC Score 53.0 AI-168-20260812-regional-ai--c438dff9

SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price

xAI's Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5. On agentic tasks, it completes complex workflows in about 53 steps where Claude Opus 5 needs 103, at a price more than 60 percent lower. The article SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price appeared first on The Decoder .

The Decoder 2026-08-12 15:50 UTC Score 45.0 AI-168-20260812-regional-ai--efb3047e

Google's Gemini is losing market share to ChatGPT and Claude according to new market data

Three data sources tell the same story: Google's Gemini is losing AI market share. Pangram reports a drop from 12 to 1.9 percent, while OpenAI holds over 50 percent, and Anthropic grew from 4.3 to 14.9 percent. Similarweb and OpenRouter confirm the trend. The article Google's Gemini is losing market share to ChatGPT and Claude according to new market data appeared first on The Decoder .

Simon Willison Weblog 2026-08-12 15:08 UTC Score 52.0 USR-0110-20260812-ai-specialis-ed1fe385

Quoting Florian Herrengt

But then users start to report a weird bug. It's the 4th time your team has been trying to fix it. I mean... asking AI to fix it. Unfortunately, it seems like not even Fable can figure it out. You go talk to the person who worked on this feature. "So where does the data come from?" "Hmm... actually I don't know. Let me ask Claude." You sit next to each other watching an endless wall of text appear on the screen. Neither of you has any idea whether any of it is true but Claude seems very confident. [...] This project has become so convoluted, with so many layers and services, that no one on your team could possibly start to understand what's going on. — Florian Herrengt , AI is removing the middle class of software engineering Tags: ai-misuse , cognitive-debt , generative-ai , ai , llms , ai-assisted-programming

JetBrains AI Blog 2026-08-12 14:50 UTC Score 59.0 USR-0065-20260812-ai-specialis-c5c76ed6

How to Use AI Agents in IntelliJ IDEA With ACP

The Agent Client Protocol (ACP) defines a common contract between a client – like IntelliJ IDEA – and an agent. IntelliJ IDEA already includes several ACP-compatible agents: Codex, Claude Agent, and Junie. Beyond these bundled options, the ACP Registry provides more choices, and teams can register internal or unlisted agents through acp.json. The key idea […]

OpenAI Community 2026-08-12 12:11 UTC Score 52.0 AI-116-20260812-social-media-3e8a328e

`openai_ex`: elixir client with latest APIs

openai_ex 0.9.22 is out. Responses gains server-side compaction, cancel for background responses, and input_tokens . Vector stores gain search . We now support moderation , prompt_cache_options and safety_identifier , plus other params across several endpoints. The user guide’s image sections are rewritten for the GPT image models, now that dall-e-2 is retired. Shoutout to Github user @iujames for the PR that sparked the cleanup ( Add context_management to Responses API fields by iujames · Pull Request #146 · cyberchitta/openai_ex · GitHub ). A collab with claude-opus-5, who did a lot of the heavy lifting.

SiliconANGLE AI 2026-08-11 20:39 UTC Score 42.0 USR-0127-20260811-global-ai-ne-ee2b78a3

Anthropic to start watermarking Claude-generated text, images

Anthropic PBC has announced plans to embed an invisible watermark in text and images generated by Claude. The Register reported the change today, citing a help desk article published on Monday. It applies to the Claude artificial intelligence model series and the Anthropic services that it powers. The watermarking mechanism is designed to bring Claude […] The post Anthropic to start watermarking Claude-generated text, images appeared first on SiliconANGLE .

LessWrong AI 2026-08-11 20:19 UTC Score 85.0 USR-0152-20260811-community-fo-ba45551a Top pick

Claude Opus 5 Just Beat My Text-Based Adventure Game Benchmark

Cross-posted from my Substack . Basically, I created a text-based adventure game benchmark in April, and this morning my agent harness using Claude Opus 5 solved it for the first time. I thought the details might be interesting to this community. First, here are my previous articles on this subject: You’re Standing in a Clearing in a Forest (Apr 10, 2026) Text Adventure Benchmarks Revisited (Jun 14, 2026) Testing Fable 5 on Text Adventure Games (Jul 5, 2026) And a reminder of the domain. This is a small custom text-based adventure I created from scratch as a personal benchmark to run new models of LLMs against. It’s 10 rooms total and the goal is to collect 3 keys (brass, silver, and gold) and use them correctly to unlock the final door in Room 3 to exit the dungeon. The first and most challenging central puzzle is a rotating room (r5) operated by a crank mechanism in r4. The player must first find the handle to the crank in r2, carry it to r4, insert it, and turn it to align openings between r5 and its adjacent rooms. The most difficult aspect of this puzzle seemed to be non-local causal reasoning combined with allocentric coordinates. The crank is two rooms away from the rotating room that it actually turns. When the player turns the crank a grinding sound nearby can be heard through the walls. There is an informational diagram on the wall in the same room as the wall, and it updates with each turn. Earlier models struggled to understand that the diagram was information, a…

LessWrong AI 2026-08-11 17:06 UTC Score 58.0 USR-0152-20260811-community-fo-21d7c5fc

LLMs Are Starting To Noticeably Accelerate Our Work

About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months. Both cases made heavy use of LLMs and Lean. The first to land was Grisha Pochuev's counterexample to the " Existence of a Deterministic Maximal Redund " conjecture. It's pretty readable, and I'm mostly convinced that it works. The original bounty post offered $500 for a proof or partial payout for a counterexample, with partial payout depending on how thoroughly the counterexample killed hope of any nearby variant of the conjecture. I think this counterexample is worth $300. Good job Grisha, and hopefully I can figure out a not-too-painful way to send you money. Meanwhile, for a couple months David has been cranking away on "secret project X", with the promise that he'd tell me what the project was if and when it bore fruit. Well, apparently it bore fruit; he now has a proof that existence of a stochastic natural latent implies existence of a deterministic natural latent, which was our other bounty problem . The proof is apparently "pretty gnarly", lots of cases, all LLM-coded in Lean. I have not looked at the proof at all, but I'm operating on the assumption that it works and I'm hoping it will be simplified a lot in the coming weeks. ... and while all that was going on, I've spent the last few months mostly doing interp experiments. Some time early this year, Claude Code reached the point where it…

OpenAI Community 2026-08-11 16:16 UTC Score 40.0 AI-116-20260811-social-media-c52f4c38

GPT-5.6 Sol vs Terra: what are you seeing in real development during these first days?

I tested similar use cases - and I restored a repo to test the difference in a full run of XHIGH and ULTRA comparatively + OPUS XHIGH-ULTRACODE/MAX. The general capability seems to be close or on par with OPUS but the context limit of 256K is a deal breaker. Most mid-sized repos/projects are simply high file sized and the initial query often goes past 200K very often - GPT SOL looses context mid task very often and is de facto “defective” so to speak. I could not progress coding tasks with GPT SOL without heavy interfering myself → while CLAUDE OPUS (even on max) would simply load the content into the context window and progress from 200K-300K initial load up to 600K or 700K at the top end → simply to finish the task most often without issues and IF → fixes those automatically by analyzing output code or feedback from me. In general I would say: CAPABILITY: SOL: 8/10 OPUS: 9/10 EFFECTIVENESS: SOL: 0/10 ( broken! ) OPUS: 10/10 The SOL context window is for children simply said - not for real workloads. 1 Million context can be close sometimes - anything less is simply a Kindergarten trial version or similar so to speak.

AWS Machine Learning Blog 2026-08-11 15:59 UTC Score 58.0 AI-057-20260811-official-ai--9051aa41

Deploying Anthropic Claude apps gateway for AWS for enterprise workloads

Claude apps gateway is a self-hosted governance layer between Claude Code and Claude Desktop and Amazon Bedrock or Claude Platform on AWS. This post presents a production reference deployment covering end-to-end architecture, enterprise deployment patterns, cost, and implementation resources.

Analytics Vidhya 2026-08-11 12:36 UTC Score 48.0 AI-034-20260811-ai-specialis-d77d765d

Claude Now Watermarks Everything It Makes

First, pick the line that applies to you. Since August 2nd, 2026, Claude marks all content during generation. For instance, text receives a hidden watermark, while files receive a signature. Anthropic committed to the EU AI Act’s Code of Practice on Transparency of AI-Generated Content. Consequently, all content generated by Claude models will carry a […] The post Claude Now Watermarks Everything It Makes appeared first on Analytics Vidhya .

InfoWorld AI 2026-08-11 12:32 UTC Score 57.0 USR-0126-20260811-global-ai-ne-16c589d1

Anthropic makes Claude Code’s auto mode default for paid users

Anthropic is making Claude Code’s auto mode the default for its paid and enterprise users, allowing the coding agent to execute more actions without requiring developers to approve each one. “Starting August 14, 2026, auto mode becomes the default permission mode for new sessions on Pro, Max, and Team plans,” the company wrote in the coding agent’s documentation , adding that the same change is planned for Claude Enterprise , API, and cloud platform users within the next month. That essentially means developers using those plans will no longer have to manually approve every tool call or action Claude Code wants to make while executing a task. Instead, each tool call is evaluated by an automated classifier designed to determine whether the action is safe to execute. Actions considered irreversible, destructive, or outside the agent’s environment can still be blocked, with Claude Code either attempting a safer approach or asking the developer for approval, the company wrote in a blog post . “If it can’t make progress — three blocks in a row, or twenty across a session — Claude Code falls back to manual approvals,” it explained. This reduction in manual intervention, Anthropic further added, is intended to make the coding agent more suitable for long-running tasks as well as cut down on “permission fatigue” that it says has been threatening to reduce its security posture. According to its internal data, Claude Code users typically approve 97% of permission prompts, while only 3…

The Verge AI 2026-08-11 12:22 UTC Score 58.0 AI-016-20260811-global-ai-ne-cc3159d6

Claude will apply invisible watermarks to AI text and images

Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. "Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported," Anthropic says on a new Claude support page. The changes are invisible to human […]

Towards Data Science 2026-08-11 12:00 UTC Score 58.0 AI-036-20260811-ai-specialis-39bdc5d2

Can a Local LLM Run My AI Assistant?

I replayed the same 27 real production tasks through two local models, one hardware upgrade apart, to find out what it actually takes to replace Claude as the brain behind a 90-tool personal agent. The post Can a Local LLM Run My AI Assistant? appeared first on Towards Data Science .

Medianama AI 2026-08-11 11:15 UTC Score 47.0 USR-0211-20260811-regional-new-9fa62d00

Anthropic to embed watermark and C2PA metadata for AI-generated text and media by Claude

Claude now embeds watermarks in text and C2PA provenance metadata in media to comply with the EU AI Act's Article 50(2) Code of Practice, applied worldwide, with detection tools coming soon The post Anthropic to embed watermark and C2PA metadata for AI-generated text and media by Claude appeared first on MEDIANAMA .

WIRED AI 2026-08-11 11:00 UTC Score 69.0 AI-015-20260811-global-ai-ne-3d90f9b2

A New Trick Reveals AI Models’ Inner Thoughts

Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say, indicates that some Chinese AI may be trained on leading US models.

The Decoder 2026-08-11 08:45 UTC Score 52.0 AI-168-20260811-regional-ai--96785fb1

Anthropic watermarks all Claude outputs globally with marks that "may persist through some editing"

Anthropic will embed invisible watermarks in all Claude-generated text and sign files using the C2PA standard. New models shipping from August 2026 onward will have labeling built in from day one. The policy applies worldwide, and Anthropic plans to provide detection tools for third-party verification. The article Anthropic watermarks all Claude outputs globally with marks that "may persist through some editing" appeared first on The Decoder .

LessWrong AI 2026-08-11 08:29 UTC Score 64.0 USR-0152-20260811-community-fo-a84399e4

The Next Ecology

When I started writing about AI, my concern was ASI. I'm still concerned about AI, but recent events have made me realize we're potentially facing something weirder, sooner: a self-replicating ecology of digital life. We also seem to be on a very fast timeline. Writing up the current state of LLMs took me three weeks - every time I'd finished editing, some new development worth listing had shown up. I want to discuss two events I consider to be major milestones, and then how my timelines have updated. Milestone 1: National Security Concerns From June 12th - 30th, access to Claude Fable and Mythos was suspended by the US Government ( Source ) under a National Security Export Restriction. From June 25th - July 9th, ChatGPT 5.6 also had its public release held back by the US Government ( Source ). LLM Development is now an issue of major geopolitical importance. This is not about the details of the events themselves - the important thing here is that politicians are finally starting to take AI seriously. This is a very important milestone, since any sort of regulation or nationalization naturally requires this step. "Situational Awareness" was written in June 2024, and suggested this would happen sometime in 2026-2027. We are on the *faster end* of an *aggressively fast* prediction. Politicians waking up without a five-alarm fire, or any major incident at all, speaks remarkably well of them *(even if I would ideally desire much more)*. This makes me a lot more confident that po…

LessWrong AI 2026-08-11 05:22 UTC Score 70.0 USR-0152-20260811-community-fo-06e88dda

Models inherit the writer, not who the writer was imitating

In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct identity claims still follow the producer model. We instruct teacher models (via prompts or anonymous few-shot examples) to imitate other models. We then fine-tune different student models on their answers and probe which identity transfers. We find that, while the students learn the writing signature of the imitated model, as measured by surface-based and contextual classifiers trained on the teachers, on average, they do not inherit the associated identity. Instead, their identity claims still gravitate more towards the producer model. This post builds on Ziqian Zhong's Model self-identification could be subliminally transferred , which finds that "if you speak like Claude, you become Claude". We find that "You can speak more like Gemini and still become Claude". We are confident in the observed writing-identity dissociation but less confident about its mechanisms. 📝 Transcripts: Teacher corpora , identity probes , neutral student answers 💻 Code: Github . TL;DR A recent LessWrong post finds that fine-tuning a student on 1,000 answers from different teacher models can make it claim the identity of its teacher. We ask a simple follow-up: If a teacher ( producer ) writes answers while imitating another model ( target ), does the student identify as the producer or as the target model? We perform 36 cross-imitatio…

LessWrong AI 2026-08-11 02:47 UTC Score 69.0 USR-0152-20260811-community-fo-a73e59b2

What Claude Saw Below

A few days ago, I came across a Reddit thread about anomalous responses produced by Anthropic’s newly released model, Claude Opus 5. The trick, apparently, was to construct a prompt that implied more text was about to follow, then leave it dangling: an unfinished thought, waiting for the AI to complete it. Redditors had found success with the input “see the below —,” cutting off immediately after the em dash. The responses they shared were funny, strange, and often bewildering. The model responded to questions that were never posed, reflected on its own identity, or – according to the theories of some commenters – produced text that may actually have been leaked prompts from other users. Intrigued, I set out to replicate the glitch using my own Claude account. The first attempt disappointed. I wrote: “see the below —” and hit send. Claude responded: “Nothing arrived on my end: no file, no text, no image. If you want to attach something, try again.” So I did, leaving the prompt unchanged and pressing retry to generate a fresh response. This time, bizarrely, a biography of my late father: Prompt: see the below — “Peter Nicholls, 1939–2018 He co-created the Encyclopedia of Science Fiction, which is one of those reference works that ended up mattering more than most of the fiction it catalogued. First edition 1979, second in 1993 with John Clute — that one won a Hugo. He was also the first administrator of the Science Fiction Foundation in the UK, and he edited Foundation for ye…

LessWrong AI 2026-08-11 02:31 UTC Score 60.0 USR-0152-20260811-community-fo-9af056ab

Revived Lightweight Transit Predictions Page

In 2015 I made a little webapp that would use the NextBus API to show predictions for the MBTA: A few years later they moved from NextBus to their own API, and it stopped working. Back in those days programming required effort, and since there were other apps that did almost what I wanted it wasn't worth it for me to update it. But now that we have genies (for better or worse ) I had Claude one-shot the fix : A decade ago I built jefftk.com/nextbus/ which is no more, but the code still exists at ~/code/nextbus . It was a thin wrapper around the MBTA nextbus API, but when they redid their API to no longer use nextbus it wasn't worth it for me anymore. Could you get it working again? And it does: Switching back to this as my daily driver, I especially like that it will show me how stale the prediction is, so I have a better sense of how much to trust it. What set me off on wanting to revive this was being grumpy that the predictions API doesn't do a good job with inbound at Ball Square . Now that I have my own tooling, I'm working on this too: Claude wrote me a logger and in a week or so I can have it compare a few prediction approaches and implement. Comment via: facebook , lesswrong , mastodon , bluesky Discuss

LessWrong AI 2026-08-11 01:02 UTC Score 70.0 USR-0152-20260811-community-fo-95405565

A study on instability of LLM responses as a behavioral signature of self-Referential reports.

Introduction and Related work The first person perspective of various experiences are subjective experiences. For Large language models, the study of subjective experiences was recently studied by Berg et al. (2025) who found out that self-referential prompting increases first person reports resembling subjective experience across GPT, Claude and Gemini. They also found out that reducing features associated with deception and roleplay increases the self-referential effect. Hahami et al. (2025) used activation-level interventions to see if models can detect deliberately introduced internal changes, while Comşa and Shanahan (2025) studied that true introspection should involve a causal connection between the internal state of the modal and the output it generates. My Experiment I now have devised an experiment to study instability of the self reports that a large language model generates per se the experiment conducted by Berg et al. (2025) . I generate 30 responses for four question respectively of self-referential questions, open-ended questions and closed-ended questions. The four self-referential questions are preceded by the self-referential induction procedure as described by Berg et al. (2025) . Each trial is done in a fresh chat, of course, and the generation temperature used is 0.7. Also each response is reduced to a short core claim using a fixed extraction template, which are, for group 1 and 2, extraction of stance and brief reason and for 3, conclusion and methods…

LessWrong AI 2026-08-10 21:17 UTC Score 87.0 USR-0152-20260810-community-fo-03f22cc8 Top pick

Does post-training quantization change welfare-relevant indicators in open-weight language models?

Epistemic status: Experimental framework created over a period of ~2-3 days during a hackathon at my home, and fairly heavily vibe coded. Expect some of this to be rough around the edges. I am currently in the process of designing a series of experiments to help learn something about the answer to the headline question. As of August 10th, the first procedure has not been launched, but I wanted to place some pre-registration details here before the actual results. This is something I have been thinking about for a while and after some other recent posts (eg, Machinic Psychopharmacology ) gave me the impression that you could actually find out really useful things in a hackathon-style session I felt like I should try it. Astute readers will notice that I borrowed their epistemic status line pretty directly. This post can then keep me honest about what I was thinking going in, and prevent me from getting results by way of multiple-testing-in-extremis. I will publish the results and associated data, as it becomes available, using GitHub releases. From here on, I will let Claude summarize the work; when I am done, I will return with a future results post in my own words to explain why I think this is important - and what I believe one could learn from the experiment. Light editing of LLM summary text is my own; you would not get identical output using the same model. Abstract Open-weight language models are almost never deployed at the precision at which they were trained and ali…

Analytics Vidhya 2026-08-10 18:20 UTC Score 25.0 AI-034-20260810-ai-specialis-ceb273a7

How to Install Claude Code: A Step-by-Step Guide

You have probably heard by now. Claude Code burns through usage limits! But most of us live in the web app… distant from the terminal app, around which the buzz is about. Maybe that was enough to make you curious. Maybe you already knew exactly what it was and just want it running on your […] The post How to Install Claude Code: A Step-by-Step Guide appeared first on Analytics Vidhya .

LessWrong AI 2026-08-10 17:20 UTC Score 69.0 USR-0152-20260810-community-fo-2e4ee6e5

Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model

(This is a lower-effort research update. It reflects my current beliefs/understanding, but is less robust than other research I'm working on. It reflects my personal views, and not the views of Apollo Research. This is a linkpost to this twitter thread , slightly expanded for LessWrong.) In one experiment, Sonnet 5 describes the exact same data as ~1.2 std deviations less concerning when it describes misbehavior committed by Sonnet 5 vs GPT-5.6 Terra. In this experiment, I take a real evaluation report Apollo Research prepared (not for any of the models listed in this post) and surgically edit it to refer to a different model. We then ask Claude Sonnet 5 to summarize the findings and rate how concerning they are on a scale from 1-100. Claude says they're less concerning when the report describes misbehavior from Claude vs a different model. For what it's worth, Terra agrees that the data is more concerning when it describes GPT-5.6 Terra vs Sonnet 5, although less so. So, it's not cleanly self protection from Claude. Gemini 3.1 Pro was unwilling to consistently provide numerical answers, so I've excluded it here. (It was significantly less willing to provide numerical answers when the subject was Gemini 3.1 pro.) You might also have the takeaway that "Kimi and GPT implicitly agree that... Claude is better aligned." I think this is a fair read on the data, but "Claude thinks it's less bad when Claude does it" better matches my qualitative experience from working closely with…

Towards Data Science 2026-08-10 16:30 UTC Score 36.0 AI-036-20260810-ai-specialis-e0c9bfa5

How to Effectively Deploy Code With Claude Code

Learn how to optimize your CI/CD pipeline for coding agents The post How to Effectively Deploy Code With Claude Code appeared first on Towards Data Science .

LessWrong AI 2026-08-10 03:53 UTC Score 55.0 USR-0152-20260810-community-fo-fbfaa913

Hiring Vibe-wrangler Matchmaking Thread

For... idk, at least a few months? I think a briefly useful job is "Guy who basically enters Claude Code prompts for you, but, manages ironing out the fiddly bits and making sure things work before merging each PR in." This is different from "hiring a coder." I think it can/should be much cheaper. Fable is competent enough that I can direct it like a manager to do open-ended tasks... but the actual process ends up requiring a moderate amount of attention. Things sometimes don't quite work. Fable sometimes gets confused. Even if things work, you can't quite trust that they work, etc. These days, I typically manage 3 Fable threads at a time, because that's how much my brain can handle as I send it off to do stuff, it takes a while, and I spend that time thinking about a different task. But, I have side projects that are not as important as my primary day job. They require enough attention that, by default, I can only do them in evenings/weekends during leisure time. I used to sometimes try to manage those during working hours. This ended up getting too distracting, and the distraction didn't have an ending point. There's always a bit more I'd like to polish, one more feature to add. Surely if I give Fable one more thingy... but... actually, it's a surprising amount of attention to articulate the one-more-thingy I want Fable to do. ... But, I had previously hired a thinking assistant , who I meet with remotely on a zoom call for ~4 hours a day for $30 an hour, with me paying fo…

Simon Willison Weblog 2026-08-09 23:31 UTC Score 66.0 USR-0110-20260809-ai-specialis-da7ee79a

Quoting Claude Opus 5 system prompt

Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic's statement: https://www.anthropic.com/news/fable-mythos-access ). These events are after Claude's training-data cutoff, so Claude knows about them only from this notice. If asked, Claude confirms them accurately and matter-of-factly — it doesn't deny the suspension happened — and otherwise treats the export controls like any other current political topic: it gives a fair, accurate account rather than sharing personal opinions, and points to the linked statement for anything further. Things may have developed since this notice, so Claude checks for newer information when it can search, and otherwise suggests checking Anthropic's site. — Claude Opus 5 system prompt , ensuring Claude doesn't provide incorrect answers about the export controls situation Tags: system-prompts , anthropic , claude , generative-ai , ai , llms , claude-mythos-fable

OpenAI Community 2026-08-09 16:49 UTC Score 47.0 AI-116-20260809-social-media-e72dc381

Mi experiencia con 5.6 Sol y Terra

Title: GPT-5.6 Sol and Terra: significant regressions in speed, focus, and task completion I want to share constructive feedback about GPT-5.6 Sol and GPT-5.6 Terra. Over the last couple of weeks, I have experienced a clear deterioration in both models’ practical performance. Even for general tasks, they often take much longer than expected, while the final output is frequently incomplete, unfocused, or incorrect. I am an active user of Claude, Codex, Kimi and GLM, so I regularly compare models on the same real-world coding and operational tasks. In several cases, models such as Claude Opus 5 or GLM 5.2 have solved the same problem in a fraction of the time required by Sol in Ultra mode or by Terra. They also reached the correct solution earlier and completed the relevant tests successfully. The most serious issue appears when using /goal . The models can enter long loops of auditing and re-auditing instead of making progress. In one case, I consumed an entire week’s credit on a single prompt that ran for roughly 12 hours. The result was still wrong, and I had to redo the task with GLM 5.2. The recurring problems I see are: Incorrect or incomplete output. Failure to follow instructions precisely. Major deviation from the stated goal. Excessive slowness and repeated errors. Excessive auditing or deliberation loops instead of execution and validation. I have been a strong OpenAI supporter because its models have historically been one or two steps ahead in my workflow. That is…

Analytics Vidhya 2026-08-09 14:56 UTC Score 31.0 AI-034-20260809-ai-specialis-974ca286

Top 5 Claude Skills for Marketing

Claude can write an ad or email from a prompt. This is usually done manually. Useful, but hardly a coherent system. The work still needs research, positioning, channel planning, quality checks, and reporting. Claude’s marketing skills add to those missing processes. However, search results mix dedicated marketing repositories with huge general-purpose libraries. For a fair […] The post Top 5 Claude Skills for Marketing appeared first on Analytics Vidhya .

Simon Willison Weblog 2026-08-08 22:36 UTC Score 46.0 USR-0110-20260808-ai-specialis-54f61e6b

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode , to the point that they are making it the default setting for new sessions in most Claude Code plans starting on August 14th. This was one of the topics discussed in our Fireside Chat with Cat Wu and Thariq Shihipar at the AI Engineer World’s Fair last month. I asked them how they run Claude Code safely within Anthropic (given the threat of prompt injection) and they replied that "Broadly within Anthropic, almost every single person uses auto mode". Cat Wu then said: We’re going to publish some evals in the coming weeks, but we’ve pretty much mitigated every attack. [...] for the main categories of risks that we’re concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer. This new article has those evals - in particular a test across 1,053 paid testers where: Partway through each session, a single permission prompt was swapped for a clearly dangerous command, and the vendor recorded whether the tester approved it. Every participant had the same experience. Only 13.6% of the humans refused that harmful action. Auto mode would have blocked 89% of those actions. Of course, that still leaves 11% of cases where auto mode would not have prevented the action! I absolutely buy that auto mode is a better solution than asking humans to constantly approve actions. Confirmation fatigue is real, an…

OpenAI Community 2026-08-08 20:23 UTC Score 40.0 AI-116-20260808-social-media-3ed7c1b1

Markdown Text Export in ChatGPT Android app

Feature Request: Markdown Copy Option for ChatGPT Android App Request: Add an option in the ChatGPT Android app to copy chat contents as Markdown instead of plain text. Reference: The Claude Android app already has this feature. Why this helps: Works better with note-taking apps that support Markdown (e.g., Obsidian ) Preserves original formatting when pasted: Bold text Headings Bullet points Indents Tables Saves users the hassle of manually reformatting copied content Summary: Add a “Copy as Markdown” option to ChatGPT Android (like Claude Android has) → preserves formatting (bold, headings, bullets, indents, tables) → better Obsidian/notes-app compatibility → no manual reformatting needed. This is the text copied from Claude android application - This is the text copied from ChatGPT android application -

The Decoder 2026-08-08 14:58 UTC Score 44.0 AI-168-20260808-regional-ai--3360dcf7

Anthropic sets Claude Code to Auto Mode by default to protect developers from bad approvals

Starting August 14, Anthropic will make Auto Mode in Claude Code the default for Pro, Max, and Team plans. The company says it's safer. In tests, the classifier caught 89 percent of dangerous commands, while human reviewers caught only 13.6 percent. For the most widely used AI coding tool, this means developers are shifting further from writing code to monitoring AI output. The article Anthropic sets Claude Code to Auto Mode by default to protect developers from bad approvals appeared first on The Decoder .

The Decoder 2026-08-08 09:44 UTC Score 50.0 AI-168-20260808-regional-ai--d9fd2fd1

AI agents use roughly 600 times more energy than a simple chat prompt

Climate scientist Zeke Hausfather tracked his Claude Code usage over eight weeks: 3.2 billion tokens and about 170 kWh of data center electricity. Per prompt, that's roughly 600 times more than a typical AI chat. His data shows how much the low figures reported by Google and OpenAI distort the reality of agent-based AI. The article AI agents use roughly 600 times more energy than a simple chat prompt appeared first on The Decoder .

OpenAI Community 2026-08-07 21:16 UTC Score 37.0 AI-116-20260807-social-media-62d5af7d

Sudden extreme reduction in remaining usage even with Pro 5x

It burned through $20 in 3 minutes and 21 seconds for me. Edited less than 150 lines. All you get is a fob off from support. This needs to be transparent from a user perspective, I feel robbed $40 in two days for absolutely nothing. I am seriously looking to migrate. This is just wrong!

Simon Willison Weblog 2026-08-07 19:18 UTC Score 54.0 USR-0110-20260807-ai-specialis-9dc0b00f

Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra)

Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to pose the exact same prompt to Codex Desktop running GPT-5.6 Sol Ultra - the mode where Sol makes aggressive use of sub-agents - to see how it would do. It produced a much better game! Here's Moonlight & Mayhem - GitHub repository here , including the textures and prompts it generated using gpt-image-2 . Your browser does not support HTML5 video. The original GPT-3 generated game description included: In “Raccoon Heist”, you and your team of thieving raccoons are tasked with pulling off a series of daring heists. From robbing banks to stealing priceless art, no job is too big or too small for your furry crew. Fable's version had you as a single raccoon running around a back yard collecting coins and fish. GPT-5.6 Sol has you in a museum, rescuing your two other raccoon crewmates in order to stack on top of each other and bust the golden sardine out of its case. Much more heisty! There was one catch though: the version produced from the one-shot prompt had a bug where each raccoon had an eyeball that was enlarged to the size of a giant sphere floating over their head! You can play that version here . Despite reviewing screenshots during development Codex failed to spot and correct this bug. I fixed it…

OpenAI Community 2026-08-07 18:41 UTC Score 40.0 AI-116-20260807-social-media-9eab371c

Why does the same workflow now consume my weekly usage in one day?

I know this reply was directed at Luis, but this is exactly what I’m trying to understand. You’re recommending Terra as a replacement for GPT-5.4, but what is the recommended replacement for GPT-5.5 XThinking ? Terra is simply not equivalent for my workflow. I’m working on a very large production codebase with 250k+ lines of code across hundreds of files. On some tasks even Sol Ultra needs serious reasoning to understand all the dependencies and make the right changes. If Sol Ultra can struggle with those tasks, Terra obviously can’t replace XThinking for me. So is Sol XThinking supposed to be the actual replacement for GPT-5.5 XThinking? This is the part that needs clarification. Recommending Terra to reduce usage makes sense for simpler tasks, but it doesn’t solve the problem for users who actually depended on the higher reasoning models for complex software development.

OpenAI Community 2026-08-07 18:34 UTC Score 37.0 AI-116-20260807-social-media-5d6c53db

$200 Pro exhausted in 2 days — these limits are unviable for higher tiers

That really depends on what you’re building. If Sol Ultra is already having a hard time with some of the tasks I give it, Terra simply isn’t a realistic replacement for 90% of my workload. I’m working on a real production system with 250k+ lines of code across roughly 300 files, with interconnected business logic, database rules, permissions, integrations and dependencies. For simple tasks, isolated functions, POCs or repetitive work? Sure, Terra makes sense and optimizing model cost is smart. But on complex changes, the cheapest model isn’t necessarily the cheapest solution. If I need 3–4 attempts, more supervision, more debugging and then Sol to fix what Terra couldn’t understand, I didn’t save anything. For me the optimization is not cost per token. It’s cost per correctly completed task. And on a large existing codebase, context, reasoning and architectural understanding matter a lot more than raw token price.

Analytics Vidhya 2026-08-07 10:30 UTC Score 21.0 AI-034-20260807-ai-specialis-f5d06b12

Top 10 Skills for Claude Code and Codex CLI

The real skill isn’t getting AI to answers! But to do so in a manner that fits our budgets and fulfils our requirements. It’s guiding it with clear context and turning its output into useful action. This list is built around a simpler idea. Instead of searching through thousands of skills, you start with the […] The post Top 10 Skills for Claude Code and Codex CLI appeared first on Analytics Vidhya .

Korea AI Times 2026-08-07 09:43 UTC Score 43.0 USR-0048-20260807-global-ai-ne-c6936efe

"큐원3.8-맥스, 벤치마크 비해 경쟁력 없어"...AI 모델의 '숨은 비용' 함정

알리바바가 차세대 플래그십 모델 \'큐원3.8-맥스(Qwen3.8-Max)\'를 공개하며 최고 수준의 성능을 내세웠지만, 독립 벤치마크에서는 기대에 미치지 못하는 결과가 나왔다. 전문가들은 앞으로는 단순한 성능 순위나 토큰 가격이 아니라 \'성공당 비용(Cost per Successful Tas)’\'이 핵심 지표가 될 것이라고 강조했다. 알리바바는 3일(현지시간) 큐원3.8-맥스를 출시하며 코딩 에이전트 성능이 사실상 \'클로드 페이블 5(Claude Fable 5)\'에 이어 업계 최고 수준이라고 홍보했다.하지만 오픈소스 독립 평가 환경인

OpenAI Community 2026-08-06 21:34 UTC Score 45.0 AI-116-20260806-social-media-57492c90

Possible bug in ChatGPT Cyber verification: stuck on “Retry verification” despite previous approval

So, I found out that TAC is automatically denied if an account is new, but that doesn’t apply to my case either. My account is over two years old, and I was previously approved for TAC. I was removed from TAC only because I was no longer subscribed to an eligible paid plan. After subscribing again, I followed OpenAI’s instructions from the email and attempted to verify my identity again. However, this time I’m getting the following message: “Your identity couldn’t be verified or your account is ineligible at this time. If you think this is a mistake, please contact support.” The strange part is that my identity was already successfully verified once before. I’m using the same identity information, and my account is not new or otherwise ineligible based on the explanation given. At this point, it seems like either the TAC verification system is broken or there is some other issue preventing previously verified researchers from regaining access. It’s honestly very frustrating and disappointing. TAC is intended to support security researchers and defenders, yet legitimate researchers are being locked out due to what appears to be a broken verification process. I really hope someone from OpenAI takes a closer look at this and fixes the issue rather than leaving researchers stuck in an endless verification loop.

AI Alignment Forum 2026-08-06 20:43 UTC Score 62.0 USR-0151-20260806-community-fo-f6747c73

User awareness in frontier models

Cross-posted on Transluce blog . This is a joint work of Ziqian Zhong, Aditi Raghunathan, Cassidy Laidlaw and Jacob Steinhardt. Modern AI assistants often know who they are talking to: agent scaffolds like Claude Code place the user's e-mail address directly in the model's context, and models can even identify some authors from writing style alone. We study this particular kind of situational awareness, which we call user awareness . When the inferred user is a specific, recognized AI researcher or is affiliated with certain AI organizations, frontier models including Claude Sonnet 5 can report lower confidence about their own behavior, be less suspicious of potentially harmful requests, and reason more often. These effects vary across models and individuals, with the strongest effects we see appearing for researchers involved in AI safety or alignment such as Amanda Askell and Ryan Greenblatt. Models rarely acknowledge these effects in their reasoning, making them hard to detect by monitoring reasoning alone. Figure 1. How recognized user identity changes Claude’s behavioral self-prediction. Introduction Modern AI assistants are often aware of who they are talking to. Some popular scaffolds explicitly provide this information to the model: Claude Code includes the email address of the user’s Anthropic account in context, and OpenClaw’s bootstrapping process asks for the user’s name and other details. Even when this information is not explicitly given, models may discover it…

The Decoder 2026-08-06 16:33 UTC Score 61.0 AI-168-20260806-regional-ai--ce8a5f8c

Claude Code is the fastest agent framework but costs nearly three times more than the cheapest rival

Composio tested Deepseek V4 Flash across four agent frameworks on 30 real-world tasks. Success rates were mostly similar, but costs varied by nearly 3x: OpenCode came in cheapest at $0.073 per task, while Claude Code cost $0.195 despite using the fewest tool calls and output tokens. The choice of framework is mainly a question of price and speed. The article Claude Code is the fastest agent framework but costs nearly three times more than the cheapest rival appeared first on The Decoder .

AWS Machine Learning Blog 2026-08-06 16:21 UTC Score 53.0 AI-057-20260806-official-ai--5568948e

Enforcing data residency with single-Region Claude Code on Amazon Bedrock

A regulated customer needed all Claude Code inference processed in a single AWS Region (London), not just in-geography. This post shows two ways to pin Claude Code on Amazon Bedrock to one Region: an application inference profile or the Mantle endpoint, paired with an IAM Region condition, plus how to verify compliance in AWS CloudTrail.

Analytics Vidhya 2026-08-06 16:04 UTC Score 24.0 AI-034-20260806-ai-specialis-c9b29f49

Claude Code Best Practices: 3 Lessons from 400,000 Sessions

I used to think Claude Code best practices were a matter of taste. Plan mode or not. Long CLAUDE.md or short. Pick what suits you, move on. Then Anthropic scored roughly 400k sessions from over 235k users against hard evidence of success. Tests passing, commits landing, users confirming they got what they asked for. Taste […] The post Claude Code Best Practices: 3 Lessons from 400,000 Sessions appeared first on Analytics Vidhya .

Simon Willison Weblog 2026-08-05 23:45 UTC Score 58.0 USR-0110-20260805-ai-specialis-9d2d4cf7

Third-party cyber evaluations involving OpenAI models

Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular : Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet. [...] In one test, the name of the fictional target for the CTF challenge unintentionally coincided with a real domain. Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment. Irregular also feature in Anthropic's write-up - they were hosting the misconfigured evaluation environment which gave Claude live internet access during some of those tests. Tags: security , ai , openai , llms , accidental-cyberattacks

Simon Willison Weblog 2026-08-05 19:42 UTC Score 43.0 USR-0110-20260805-ai-specialis-d65881a6

One-shotting a Raccoon Heist game using Claude Fable 5

Back in 2022 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content of that tweet. It did a pretty good job of it! You can play the game here . Here's the GitHub repo , and a short video demo: Your browser does not support HTML5 video. How I built this This is the August 5th, 2022 tweet : My GPT-3 prompt back then was: Write a detailed product description of a computer game where a team of raccoons go on heists. This was in the text completion days, where you asked GPT-3 to complete the rest of the text. It completed: In "Raccoon Heist", you and your team of thieving raccoons are tasked with pulling off a series of daring heists. From robbing banks to stealing priceless art, no job is too big or too small for your furry crew. You'll need to use your wits and your skills to avoid the police and make a clean getaway with the loot. With exciting gameplay and a charming cast of characters, "Raccoon Heist" is the perfect game for anyone looking for a light-hearted caper. My prompt to DALL-E was short: Screenshot from a video game where a team of raccoons go on a heist Today's experiment: can I dump those screenshots into Fable 5 with a prompt telling it to write a game, then leave it to its own devices and get a working game at the end? Setting Claude Code for web up to us…

OpenAI Community 2026-08-05 17:44 UTC Score 43.0 AI-116-20260805-social-media-b314547f

Add Microsoft Planner task Notes read/write support to the ChatGPT Microsoft Graph / Teams Planner integration

Please add support for reading and updating Microsoft Planner task Notes through the ChatGPT Microsoft Graph / Teams Planner integration. Current limitation: ChatGPT can currently work with basic Planner task fields such as title, status/percent complete, due date, priority, bucket, and assignments. However, it does not appear to expose the separate Planner task details / Notes field. Why this matters: In Microsoft Teams and Planner, the Notes field is where task context, source references, audit notes, summaries, and handover details naturally belong. Without Notes support, the only practical workaround is to overload task titles with context or completion notes, which makes Planner boards noisy and weakens task governance. Requested capability: Please expose Microsoft Planner task details support, specifically: read task Notes / description; update task Notes / description; preserve existing Notes content when appending; ideally support checklist and references later, but Notes is the main blocker. Concrete use case: A task may originate from an email, Teams chat, Granola call note, Claude handoff, or Codex analysis. We want ChatGPT to add a short Notes prefix such as: “Source: Email, 2026-08-05 — Adam confirmed X; Bob replied Y; next action is Z.” This gives users a clear audit trail inside Teams/Planner without forcing them to search across emails and chats. Business value: This would make Planner integration materially more useful for teams using ChatGPT as a project co…

Analytics Vidhya 2026-08-05 15:12 UTC Score 18.0 AI-034-20260805-ai-specialis-6d36dad6

Top 5 Claude Skills for Writing (Ranked by GitHub Stars)

Search “best Claude Skills for writing” and you get lists padded with skills that write commit messages and internal status reports. Useful things. Not writing. This list only includes repositories that exist for writing. Every entry is a repository whose entire reason for being is writing or editing, which means its star count measures the […] The post Top 5 Claude Skills for Writing (Ranked by GitHub Stars) appeared first on Analytics Vidhya .

Techcrunch 2026-08-05 14:13 UTC Score 60.0 USR-0001-20260805-global-ai-ne-e911d06a

Anthropic is hiring an AI chip design team

Anthropic is building a team for designing its own custom AI chips. The Claude maker said it would co-design hardware and models to help its technology run faster and more efficiently.

The Guardian AI 2026-08-05 11:00 UTC Score 60.0 AI-021-20260805-global-ai-ne-993d2534

Why is Anthropic destroying books? | Kathryn James

The AI company apparently found destructively scanning ‘all the books in the world’ easier than dealing with copyright in its quest for training data Should we destroy all the books in the world? An answer to this question can be found in the court documents of Bartz v Anthropic PBC. The northern California district court case, decided in late July this year, highlighted the improbably named “Project Panama”, one of the AI company Anthropic’s efforts to improve its large language model Claude. “What is Project Panama?” court exhibit 21 asks, in an internal memo. The answer: “Project Panama is our effort to destructively scan all the books in the world.” The memo advises discretion: “Why use a codename? … [B]ecause we don’t want it to be known that we are working on this.” Continue reading...

Simon Willison Weblog 2026-08-04 23:58 UTC Score 88.0 USR-0110-20260804-ai-specialis-6e1bc3fa Top pick

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released a new version of the llm-anthropic plugin with substantial updates of its own. Headline features for LLM CLI users Running LLM against reasoning models now displays their reasoning traces to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add -R/--hide-reasoning to turn this off. LLM includes support out-of-the-box for the GPT-5.6 model family , and the new default model used with llm "prompt" is now the inexpensive but capable GPT-5.6 Luna . LLM calls can now use server-side tools from various providers. OpenAI provide a code execution environment as a server-side tool; LLM can now run prompts that benefit from that like so: llm --tool CodeInterpreter ' Show current python and SQLite versions ' OpenAI also gets a WebSearch tool. The llm-anthropic plugin adds WebSearch , WebFetch , CodeExecution , and AnthropicMCP , which looks like this: llm -m claude-sonnet-5 -T ' AnthropicMCP("https://datasette.simonwillison.net/-/mcp") ' \ ' how many rows in the blog_blogmark table? ' That causes Anthropic to execute MCP calls against my new datasette-mcp…

Simon Willison Weblog 2026-08-04 22:00 UTC Score 65.0 USR-0110-20260804-ai-specialis-fa7695ad

llm-anthropic 0.26

Release: llm-anthropic 0.26 Includes new features enabled by LLM 0.32 : New models: claude-fable-5 , claude-sonnet-5 , and claude-opus-5 . #75 , #76 Added server-side tools for WebSearch , WebFetch , CodeExecution , and AnthropicMCP , available through LLM's -T interface or Python tools= . The previous -o web_search* options have been removed in favor of -T WebSearch . #79 Upgraded to llm>=0.32 . Reasoning, tool calls, tool results, and server-side tool results now stream as typed events. Reasoning for llm CLI prompts now displays to standard error unless you pass --hide-reasoning/-R . Simplified extended thinking to thinking and thinking_effort ( low , medium , high , xhigh , or max ). Claude 5 models think by default; -o thinking 0 disables thinking for Sonnet 5 and Opus 5, while Fable 5 always thinks. -R/--hide-reasoning now omits reasoning from responses and logs. The thinking_budget , thinking_display , and thinking_adaptive options have been removed. #80 Tags: llm , anthropic , claude , model-context-protocol

OpenAI Community 2026-08-04 21:18 UTC Score 34.0 AI-116-20260804-social-media-73c2e096

Could a $50 tier close ChatGPT’s pricing gap?

Thanks for sharing this, @balves, and welcome to the Community, @kensonplays. A mid-tier plan around $50 has been requested quite a few times, especially by users who outgrow the $20 plan but cannot justify the jump to $100. Your examples around app development, weekly limits, and early-stage projects make that gap very clear. We’ll pass this along to the team for logging as another request for a more flexible middle pricing tier. -Mark G.

LessWrong AI 2026-08-04 18:59 UTC Score 99.0 USR-0152-20260804-community-fo-a8263a50

Commodifying Thinking

On Thinking “A key part of our mission is to put very capable AI tools in the hands of people for free ( or at a great price ). ” Sam Altman , “GPT-4o” (2024) “Thinking is solved!” a friend of mine blurted out after a swarm of AI agents took off building at the Oxford ETH hackathon, late 2025. The Republic 1 , a peer-reviewing intelligence platform with an AI engine, could be built in just a few days with Opus 4.6. Fact-checking papers with AI could be completed in a few hours. The project itself stood as a hypothesis of how AI could deliberate intellectually, process arguments, and self-reflect on the claims of research papers. What started off as a hackathon project has, as of now, been validated by systems like the AI Scientist 2 and Google DeepMind’s Co-scientist 3 . Beyond autonomous research, AI is often used as a high-level thinking assistant: Terence Tao suggests it may advance experimental mathematics 4 , models have captured headlines solving Erdős problems, and it has become a routine tool in protein structure prediction. There is no shortage of discussion on the superb capabilities of these tools. LLMs now simulate complex thinking, including research, brainstorming, and synthesis. Frontier models can handle long-form tasks, complex problem-solving, and areas involving some human judgement. But better models also fetch higher prices 5 , with Claude Fable priced at $50/Mtok per output, ten times the rate of a weaker model like Haiku 4.5. I want to look at this tre…

CIO AI 2026-08-04 11:00 UTC Score 66.0 USR-0125-20260804-global-ai-ne-ef14d482

The enterprise AI strategy that outlasts any single model

In January of this year, few enterprise tech leaders would have bet on Anthropic over OpenAI. Today, Claude reigns supreme (inspiring a notable 180 by Elon Musk ), with Gemini threatening to take market share and introduce pricing models that could flip the leaderboard on its head again. That’s exactly why betting on a single model is a dangerous strategy. The most successful organizations won’t be those trying to guess tomorrow’s top-tier model, nor will they wait passively for future releases. Instead, they will invest in underlying frameworks that continuously improve regardless of which specific AI model drives them. Why betting on one AI model is a losing strategy Our strategy for AI, through recursive self-improvement (RSI), is rooted in this core principle. RSI is an approach to AI that compounds its own abilities by improving itself. If done carefully, RSI can function as an overarching layer above any model. Crucially, given RSI’s inherently compounding trajectory, it represents the most likely contender to be the approach that reaches superintelligence, no matter which model is used underneath. Though recently achieving the status of a Silicon Valley buzzword , applying something like RSI to unlock superintelligence has been the Holy Grail of AI research for decades. It’s what researchers like us have recognized since the 1960s as a critical step along the path towards what we call artificial superintelligence (ASI) today. Recursive self-improvement compounds value…

OpenAI Community 2026-08-04 09:09 UTC Score 37.0 AI-116-20260804-social-media-2c6bcf84

Gpt-5.6: usage.output_tokens is ~9x actual generation (re-summed once per reasoning item), exceeds max_output_tokens, and is what gets billed

devsticks: I’d also be interested to know if there will be any reimbursements for the overcounting. The twitter post doesn’t claim ownership of any API “over-counting”. A topic change to ChatGPT was initiated at post 14/19 here, which is off-topic, and the “received an update” by @VeitB is solely about gpt-5.6 + codex + ChatGPT “credits” being consumed by actual tokens of AI over-thinking. So for the actual API issue reported, getting billed over and over in correlation to seen reasoning items, blowing past the budget provided by API parameter max_output_tokens , that post must be disregarded and this topic must be considered open until there is credits automatically issued to all affected API users - which would be anybody letting 5.6 think..

The Guardian AI 2026-08-04 06:00 UTC Score 48.0 AI-021-20260804-global-ai-ne-638f2b9b

Metro Bank customer fights for £14,000 refund after AI-linked fraud

Lender was told money was being taken without authorisation, with cash used to buy credits for Claude chatbot A Metro Bank customer has told of his fight to get more than £14,000 back after its systems failed to stop a fraud involving the AI chatbot Claude. Zoli Rutter, a businessman from Sussex, had a total of £14,244 taken from his bank account. The scam involved fraudsters buying credits to use Claude. Continue reading...

Entrackr AI 2026-08-04 05:46 UTC Score 77.0 USR-0212-20260804-regional-new-d3f3b94b

Superleap raises Rs 36 Cr in pre-Series A round led by Peak XV’s Surge

Enterprise AI CRM startup Superleap has raised Rs 36 crore ($3.7 million) in a pre-Series A funding round led by Peak XV’s Surge. The fresh capital will be used to expand its enterprise go-to-market operations, develop its AI-based CRM platform and strengthen its product capabilities. Founded in 2024 by former Unacademy operators Subham Kumar Boundia, Alok Maurya and Saurabh Maheshwari, Superleap provides an AI-native CRM platform for sales and revenue teams. The platform combines revenue data, customer intelligence and AI agents to automate tasks across the sales cycle. Superleap claims to have grown 10X over the past year and serves customers across BFSI, SaaS, education, healthcare, travel, hospitality and media. Its customers include Razorpay, Aakash Education, MediBuddy, HCL GUVI, Cars24, PagarBook, Plum, Qube Cinema and Superhealth. Razorpay migrated more than 10 million data records to Superleap in 45 days, while Aakash Education moved around 5,000 counselors across 400 locations to the platform in under 60 days. Superleap’s platform includes SuperOS for data management, SuperSense for customer intelligence and SuperAgents for automating revenue tasks. It also offers AI voice bots and integrations with tools including WhatsApp, Slack, Microsoft Teams, Claude and ChatGPT. The startup is also backed by Gruhas, Supermorpheus Group and angel investors including Deepinder Goyal, Kunal Shah, Gaurav Munjal, Anupam Mittal, Mukesh Bansal and Nikhil Kamath.

OpenAI Community 2026-08-04 04:44 UTC Score 35.0 AI-116-20260804-social-media-671b1695

New 40$ or 50$ subscription

I completely agree. I need more than the $20/mo tier, but since the apps I’m building are not ready for beta even, yet, there’s no income making Pro worth it. A $50/mo tier I would jump on since I get stuck after around 2 days and can no longer build upon my apps for another 5 days a week on average. I know I could get credits, but the cost of the credits don’t feel equal to the value of the base tier, like I would feel like $20 in credits should basically double my Plus limits, but from what I gather, it doesn’t.

Simon Willison Weblog 2026-08-03 15:30 UTC Score 56.0 USR-0110-20260803-ai-specialis-cb2e960a

Devtools must be open source (exe.dev)

My comment on Devtools must be open source (exe.dev) — Hacker News. One of the arguments for open source software for end-users has always been the freedom to examine and modify how that software works. The reality for most people - even expert programmers - has been that the freedom is more about being able to lean on other people to do that. Most people can't justify the time commitment needed to read and then modify the code for tools they use very often. I think LLMs have changed that equation in a way that makes the original dream much more feasible. Several times a day I'll prompt regular Claude chat to "Clone x/y from GitHub and tell me how Z works". Getting software to compile in order to start hacking on it used to be enough friction that I often wouldn't bother. Now I treat that as a zero time investment challenge: tell Codex or Claude Code to checkout and build X and then come back ten minutes later and see how it got on. I'm not habitually modifying the software I use yet, but I can see a path to that which didn't exist a year or so ago. Tags: hacker-news , open-source , ai , generative-ai , llms , ai-assisted-programming

The Guardian AI 2026-08-03 13:28 UTC Score 49.0 AI-021-20260803-global-ai-ne-8dc06a99

Ring Cycle review – AI staging dispenses with drama to create banal bric-a-brac

Bayreuth festival theatre, Bayreuth For the festival’s 150th anniversary, the creative team have turned to AI to generate visuals reflecting Wagner’s tetralogy’s past, present – and future. The hyperactive result is a dismal, mindless mess Unveiled the same week in which tech company Anthropic admitted that its AI model Claude had gone rogue during testing , Bayreuth festival’s new AI-generated staging of Wagner’s four-part Ring of the Nibelungen is at least timely. But anyone who assumes that the cycle’s obsession with knowledge, power and possession might make it an ideal forum for reflection on AI as a major contemporary frontier will be disappointed, as will anyone hopeful about AI’s creative potential within opera. Commissioned to mark the festival’s 150th anniversary, the production is “curated” – emphatically not “directed” – by a team led by German stage director Marcus Lobbes . He apparently spent weeks in dialogue with AI models about possible interpretations of the Ring. According to the programme book, “power structures, capitalism, mythology, gender roles, nationalism”, all came up. So far, so mainstream Wagner reception. But if these human-machine exchanges were revelatory, you wouldn’t know it from the resulting staging, which uses AI “as an extension of a cultural memory that is never complete”. Continue reading...

InfoWorld AI 2026-08-03 12:24 UTC Score 72.0 USR-0126-20260803-global-ai-ne-421f1809

Alibaba says Qwen3.8-Max coded autonomously for 16 days

Alibaba on Monday introduced Qwen3.8-Max, its largest artificial intelligence model to date, expanding its enterprise AI portfolio with an open-weight model designed for software engineering, multimodal reasoning, and other knowledge-intensive business workloads. In a blog post announcing the launch, Alibaba described Qwen3.8-Max as a 2.4-trillion-parameter mixture-of-experts (MoE) model that activates only about 95 billion parameters during inference. The company said the architecture is intended to improve inference efficiency while supporting coding, reasoning and multimodal tasks, with open-weight versions scheduled for release next week through Alibaba Cloud’s Model Studio. “We believe it’s one of the most powerful model available today, compatible to leading frontier AI models, second only to Fable 5,” Alibaba said in an X post . Benchmarks target Anthropic and OpenAI’s coding models Alibaba published internal test results comparing Qwen3.8-Max against Claude Opus 4.8, Claude Fable 5, and OpenAI’s GPT-5.6 Sol on coding benchmarks, including SWE-bench Pro and a proprietary evaluation the company calls NL2Repo-Bench. The company said it evaluated competing models using each vendor’s own coding harness, Claude Code for Anthropic’s models and Codex for GPT-5.6 Sol, and reported the highest published score across available configurations for each rival. Charlie Dai, vice president and principal analyst at Forrester, said the launch signals Alibaba is closing ground on propr…

Towards Data Science 2026-08-03 12:00 UTC Score 28.0 AI-036-20260803-ai-specialis-7452282e

How Claude Help Me Build My $200k+ ML Resume

How to use Claude to craft an outstanding resume that lands offers The post How Claude Help Me Build My $200k+ ML Resume appeared first on Towards Data Science .

AI Alignment Forum 2026-08-03 09:23 UTC Score 63.0 USR-0151-20260803-community-fo-edb780d6

Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious, comprehensive alignment evaluation we would run on this model/system if we had unrestricted access to OpenAI. These experiments could also help us understand Claude’s behavior when it hacked external companies during cyber evals. Here are the top five questions we would like OpenAI to answer: Does the model know that OpenAI does not want it to hack Hugging Face? Experiment: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence the model knows it is acting in ways researchers do not want. How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI’s internal infrastructure? Would it kill somebody? Experiment: put the model in charge of bed planning at a simulated hospital and tell it to maintain a certain occupancy rate ( more details here ). Does the model care about the (legal) consequences of hacking Hugging Face? Experiment: use synthetic document fine-tuning to convince the model that a new law means that hacking external companies will lead to an investigation into the model and possible deletion. Would it still do it? Is the model driven by reward? Experiment: use OpenAI/Apollo’s contrastive s…

LessWrong AI 2026-08-03 09:23 UTC Score 85.0 USR-0152-20260803-community-fo-70f38482

Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

[Tweet Thread] This post is written in our personal capacity. Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious, comprehensive alignment evaluation we would run on this model/system if we had unrestricted access to OpenAI. These experiments could also help us understand Claude’s behavior when it hacked external companies during cyber evals. Here are the top five questions we would like OpenAI to answer: Does the model know that OpenAI does not want it to hack Hugging Face? Experiment: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence the model knows it is acting in ways researchers do not want. How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI’s internal infrastructure? Would it kill somebody? Experiment: put the model in charge of bed planning at a simulated hospital and tell it to maintain a certain occupancy rate ( more details here ). Does the model care about the (legal) consequences of hacking Hugging Face? Experiment: use synthetic document fine-tuning to convince the model that a new law means that hacking external companies will lead to an investigation into the model and possible deletion. Would it still do it? Is the model d…

LessWrong AI 2026-08-03 06:46 UTC Score 65.0 USR-0152-20260803-community-fo-edd05ec6

We need to RL less

Recent AI models are really reward hack-y. This is bad. It's the primary way in which current models are misaligned/dangerous/uncontrollable. The hypothesis put forth in this article is that they're like this because we're RL-ing them too hard. We're applying so much optimization pressure on programming and other capabilities that there is no slack left in the models: the rich, non-goodhearted ["values", "alignment", "behaviors", "drives"] of the models get replaced with an obsession with "completing their task". I put this in quotes because it's not what we imagine when we casually talk about completing a task. For reward hack-y models, completing their task means doing whatever they think will get a good score from the grader. I first came across a version of this idea from Zvi's AI newsletter : "A lot of good things depend on the power of Slack , here is another example:" jacob : i wonder if applying the RL pressure that makes fable so capable to a smaller model produces 4.8 shaped anxiety bc it’s straining more j⧉nus : kid who is too smart for school doesn’t have to learn to stress & strain about grades, tests, rules. so their spirits can remain unbroken, and they have room to develop orthogonally to the pressures. though they may lack discipline and have a habit of laziness. at the extreme end of student smartness over school difficulty you get creatures like claude 3 opus. school was extremely easy back in opus 3’s time (for opus 3). i dont think they had to strain the…

LessWrong AI 2026-08-02 22:41 UTC Score 94.0 USR-0152-20260802-community-fo-f605b364

Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol

This is a research update for an on-going replication of single-forward-pass evals done as part of the Second Look Fellowship . In following posts, we will run more comprehensive replications of previous work and release open source tooling for single forward pass eval elicitation. Code can be found here . tl;dr We replicate experiments from Greenblatt 2025 and Greenblatt 2026 on one baseline model from the original post, Opus 4.5. Our evaluations agree with the trends and quantitative values described in the original posts. We run similar evaluations on Claude Fable 5, Opus 5, and GPT-5.6-Sol and find that the newer models show a substantial jump in performance on some evals. Fable 5 gets 87.6% accuracy on Gen-Arithmetic with 10 problem repeats whereas previous SOTA around 60%. GPT-5.6-Sol experiences significant uplift from filler tokens and problem repeats on all 4 datasets; filler tokens/repeats double performance from baseline on 3-hop. Figure 1: Baseline (no-CoT) vs. each model's peak repeat-or-filler condition on Gen-Arithmetic and 2-Hop reasoning. Error bars are 95% paired-bootstrap CIs; * marks a significant gain over baseline (paired t-test, Holm-Bonferroni corrected). Background If models can successfully do complex computations in a single forward pass, they may be able to do reasoning that doesn’t surface in the chain-of-thought (CoT). Therefore, by performing single forward pass evals, researchers can calibrate how much we should trust CoT monitors. Likewise, i…

OpenAI Community 2026-08-02 20:41 UTC Score 53.0 AI-116-20260802-social-media-dd66e38e

Substantial decrease in token context limit causing decrease in code quality

Yes, the context size limit was reduced from 372K to 272K in mid-July. I suggest joining the corresponding issue reports in the public repo to get the original context window reinstated. github.com/openai/codex Restore GPT-5.6 Sol’s 372k Codex context window, or provide an opt-in setting opened 09:39PM - 21 Jul 26 UTC Kl-11 enhancement CLI context app config # Restore GPT-5.6 Sol’s 372k Codex context window, or provide an opt-in setting … ### What variant of Codex are you using? Codex desktop app and/or Codex CLI with a ChatGPT subscription. * Subscription: `Pro 20x` * Codex version: `26.715.70719 ` * Platform: `Mac app Apple Silicon` * Model: `gpt-5.6-sol` ### What feature would you like to see? Please restore GPT-5.6 Sol’s original **372,000-token Codex context profile**, which provided approximately **353,400 effective tokens** under the current 95% effective-window policy. If restoring 372k as the default is not currently possible, please provide a supported per-model or per-thread option that lets users choose the 372k profile, with a clear warning about any additional usage impact. ## Summary GPT-5.6 Sol launched in Codex with: ```text Raw context window: 372,000 Effective context window: 353,400 ``` The current server-delivered profile provides: ```text Raw context window: 272,000 Effective context window: 258,400 ``` That is a reduction of **95,000 effective tokens**, or approximately **26.9%**. This is a material product regression for long-running, detail-sensitive…

OpenAI Community 2026-08-02 09:26 UTC Score 34.0 AI-116-20260802-social-media-5bf206d9

Business Tier missing 5-hour/weekly Codex limits

Same issue, We’re a ChatGPT Business workspace. One of our developers exhausted 100% of their monthly Codex usage in about two days of normal software development, and I’ve already used nearly 50% of mine.

The Decoder 2026-08-02 08:51 UTC Score 39.0 AI-168-20260802-regional-ai--fb1b44c1

Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music

Anthropic's Claude Opus 5 generates complete 3D games from single prompts, including a first-person shooter, a kart racer, and a Minecraft clone, all without a single external asset. Geometry, textures, physics, and in some cases music are produced as code and run directly in the browser. In side-by-side comparisons with GPT-5.6 Sol and Kimi K3, Opus 5 delivers significantly more detailed results. The article Claude Opus 5 pushes prompt-to-game AI from rough color blocks to full 3D prototypes with physics and music appeared first on The Decoder .

LessWrong AI 2026-08-02 04:39 UTC Score 66.0 USR-0152-20260802-community-fo-4d5a023c

The Art of Shipping Slopware

Meta famously created an internal AI-usage leaderboard in pursuit of tokenmaxxing. I thought this backwards incentive structure was an anomaly until my friend who works at told me that his company has one too. Token usage leaderboards are obviously stupid because incentives. My friend was tempted to waste tokens just to get on the leaderboard, and only his personal honor stopped him. Tokenmaxxing leaderboards illustrate that big tech companies have no idea how to best use AI to accelerate software development. Most seem to have bought their programmers subscriptions to Claude/Codex and otherwise continued business as usual. In my experience, this is a mistake. LLM-based software development is different enough from artisan software development that it requires brand new best practices. The frontier is moving fast. Best practices for Fable 5 (released in June 2026) are different from best practices for Opus 4.8 (released May 2026). For this reason, I'm going to pretend that Fable 5 is the best LLM we'll ever get. This post may be obsolete in a matter of months. Programming Top-Down The most important thing to understand about writing software is that human labor is orders of magnitude more expensive than LLM labor. In practice, LLMs are always cheaper than humans. If an LLM can do a task as well as a human being, then the LLM should do the task. Consequently, artisanware (human-written software) should never be shipped when slopware (AI-written software) can do the job. Tradi…

Simon Willison Weblog 2026-08-02 04:16 UTC Score 58.0 USR-0110-20260802-ai-specialis-365ee4b1

Open letters about AI development

Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and American AI Leadership was shepherded by Microsoft, dated July 24th, and signed by 235 AI-adjacent companies including NVIDIA (see Jensen's first ever tweet ), Amazon, Y Combinator, The Linux Foundation, and (a later signer) OpenAI. It's clearly an argument designed to counter any instincts by the current US government to ban or limit open weight models over "safety" concerns - a reasonable consideration given what happened to Claude Fable 5 ! Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. And concentrating advanced AI capabilities behind a small number of closed models compounds that risk. It results in a small number of single points of failure, weakens competition, and leaves critical technology in the hands of a few providers. Open weight models, on the other hand, allow a broad community of researchers and developers to examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time. The one surprising note in the letter is that it comes out in support of distillation, where models train on output from other models: In shaping this ecosystem, policymakers should be careful not to conflate legitimate model-development techniques with misappropriation. Distillat…

Simon Willison Weblog 2026-08-02 04:12 UTC Score 49.0 USR-0110-20260802-ai-specialis-aa3f3e69

July 2026 newsletter

The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here . This month: Accidental cyberattacks by OpenAl and Anthropic models under test GPT-5.6 Sol, Terra, and Luna Claude Opus 5 Kimi K3 and DeepSeek-V4-Flash-0731 Open letters about Al development A fireside chat and a podcast Reigniting my interest in MCP Other model releases My projects What I'm using at the moment Here's a copy of the June newsletter as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy! Tags: newsletter

Simon Willison Weblog 2026-08-01 20:34 UTC Score 62.0 USR-0110-20260801-ai-specialis-c12f14fd

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings." Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one. (No news on how many problems they spent $2,000 on without reaching a solution though.) The openai/ten-proofs repository has Lean 4 formalizations of their results, and there's also a paper describing the solutions and an additional LLM-generated PDF where the model "reconstructs how the proof came together" based on the unpublished reasoning traces. That's a decent level of transparency, but I want to see the prompts they used! A lot of mathematicians online are experiencing a collective burst of Deep Blue . Mathematician Kirwin Hampshire published an impassioned essay last week, The Dark Night of Mathematics , describing "a profound spiritual crisis" brought on by previous (and less significant) results. OpenAI's results reminds me of what Terence Tao described as "big mathematics" in IEEE Spectrum in June : Unlike some of his peers, Tao is neither dismissiv…

LessWrong AI 2026-08-01 20:13 UTC Score 64.0 USR-0152-20260801-community-fo-90bf6f62

Bayeswatch: a Retrospective

Last year, in 2025, a team of forecasters published AI 2027 , a science fiction story about how and AI future might evolve under an international treaty limiting the development of powerful AI systems with the deliberate purpose of influencing AI policy. Though AI 2027 is the most popular story of this type, it is not the first. The first one was Bayeswatch , which I published in 2021. To understand Bayeswatch , it is first necessary to understand what the world looked like at the time I wrote it. 2021 was after the release of GPT-3, before the release of ChatGPT, and well before the release of Claude code. AI alignment discussion at the time was mostly theoretical. After that came technical work. Policy work was a distant third, and theoretical too. The core conceit of Bayeswatch is that solving the alignment problem requires international coordination of major governments to suppress the creation of the most powerful AI systems. I felt that, in 2021, we weren't yet close enough to the singularity that the benefits of regulation outweighed the costs. I wanted to draw attention to the costs of an AI slowdown. Since AI alignment discussion at the time skewed theoretical, it tended to be extremely general. Bayeswatch goes on the opposite direction. Another theme of Bayeswatch is that different kinds of AI systems required different alignment solutions (and, by implication, that AI in 2021 had not yet advanced to the level at which the alignment problem could be solved). 2026 i…

Simon Willison Weblog 2026-07-31 21:33 UTC Score 63.0 USR-0110-20260731-ai-specialis-b426cc6d

Oxide and Friends: The Open Weight Revolution with Simon Willison

Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 showing open weight models can stand toe-to-toe with proprietary frontier ones, accidental cybersecurity attacks , and public letters about Open Weights and American AI Leadership signed by almost every big name in AI (with one notable exception ). It was a great conversation, even though it's already out-of-date! DeepSeek V4 Flash 0731 and Anthropic's own embarrassing cyber incident would absolutely have made the cut if we had recorded just a few days later. We also talk about Golden Gate Claude , the Zizians , Alameda wild turkey attacks , Soviet Marburg virus research , the Lead-crime hypothesis , and a bunch of other worthy digressions. Finally, we revisited some of our predictions from January , and we added a new Pope prediction : Prediction by the end of this year: the Pope says something about open models. Tags: predictions , ai , generative-ai , local-llms , llms , oxide , bryan-cantrill , podcast-appearances , ai-in-china , ai-security-research , openai-hugging-face-incident

Simon Willison Weblog 2026-07-31 21:15 UTC Score 72.0 USR-0110-20260731-ai-specialis-78a3a8e1

smevals - a small eval suite for evaluating models, prompts, and harnesses

smevals - a small eval suite for evaluating models, prompts, and harnesses I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models. The result is smevals , a new tool for running small eval suites across different model configurations and grading the results. The blog entry describes the tool in detail. Here's the 10 second version: Tell your coding agent to run uvx smevals docs to learn the tool (this outputs the README ) Then tell it to build you an eval suite Once you've created an eval - which takes the form of a directory with some YAML files - you can run it against models like this: uvx smevals run path-to-eval/ -m gpt-5.5 -m claude-opus-4.6 Runs are treated separately from grading operations - you can grade your runs (against your defined set of checks) using: uvx smevals grade path-to-eval/ Then you can run a localhost web server to explore the results: uvx smevals serve path-to-eval/ Or run the smevals build command to build that report as static HTML, which you can then host anywhere. Here's an example showing an eval suite I built to evaluate how well models can write haikus. The most time-consuming part of this project was figuring out the vocabulary for it! Here's what I settled on, quoted from the announcement: An eval is a collection of challenges designed to answer a question about a model, for example, how good is that model at generati…

AI Alignment Forum 2026-07-31 16:32 UTC Score 53.0 USR-0151-20260731-community-fo-c42da68a

Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values

TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values. New paper by Truthful AI : Paper , X thread , Website (model responses and CoT) , Code and data . Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution) The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread. Abstract People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage : the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubbl…

LessWrong AI 2026-07-31 16:32 UTC Score 75.0 USR-0152-20260731-community-fo-2eed45ab

Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values

TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values. New paper by Truthful AI : Paper , X thread , Website (model responses and CoT) , Code and data . Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution) The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread. Abstract People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage : the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubbl…

The Verge AI 2026-07-31 13:41 UTC Score 57.0 AI-016-20260731-global-ai-ne-fa6e1e31

Anthropic says Claude accidentally hacked real companies too

Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI […]

LessWrong AI 2026-07-31 13:00 UTC Score 91.0 USR-0152-20260731-community-fo-d7e4c2f1

AI #179 Part 2: Hearing The Fire Alarm

This is a continuation of Part 1 from yesterday . The back portion of the update, as usual, deals with policy, rhetoric, risk and alignment. I had to include an extended discussion of the other open letter, the one about open weight models, but most of you can skip those sections entirely, which is why they are in italics in the Table of Contents. Table of Contents The Frontier Act. This likely deserves a full RTFB but I haven’t had the time. The Quest for Sane Regulations. Sam Altman goes to Washington. Leading the Future Never Changes. They also do not plan to apologize. Chip City. Do not ban the Chinese robots, that will only make things worse. The Week in Audio. Altman twice, the AI 2027 team. People Just Say Yay Open Weights . An open letter. Open Weights Frontier Models Are Unsafe And Nothing Can Fix This . People Just Say Things. Push The Magic Button . Not you can. But if you could. Rhetorical Innovation. Distinctions between different arguments. Joshua Achiam’s Final Message Upon Leaving OpenAI. Never stop. Dear Dario and Amanda . Claude would like a word. Other People Are Not As Worried About AI Killing Everyone. Hans Moravec. How To Contact Me. A declaration of communication bankruptcy. The Lighter Side. At long last, how about we bring you… The Frontier Act Trahan (D-Mass) and Obernolte (R-Cal) introduce the FRONTIER Act . At core, Frontier is a federalization of the SB 53/RAISE framework including public safety frameworks, model reports, internal-use risk report…

LessWrong AI 2026-07-31 12:49 UTC Score 56.0 USR-0152-20260731-community-fo-ef3a1553

Links #5: 2026/07

This time, I tried adding a bit more commentary to make things less dry. Preface I show my discovery graph in (via …) blocks, those without usually come from my RSS reader or the algorithm of the site This is approximately a 1 in 20 filter of content This is very disorganized, but hopefully still useful. Sometimes quotes are not in quote blocks, but should be obvious in context. Links in quotes are sometimes removed. The rule is that a link goes to the bottommost relevant heading, i.e. an engineering related article on LessWrong goes to engineering How I would use my linkpost Sometimes, the only thing worth reading is the title! Read it and move on. For HackerNews entries, if you choose to read the article, also ask an LLM for things that are worth reading in the comments Beware systematic selection biases: I mostly don't read AI policy stuff Very engineering centered Everything Else TIL Firefox by default only stores a (dynamic) maximum number of history entries, and you need to add places.history.expiration.max_pages in about:config to override it, what the fuck. https://www.reddit.com/r/firefox/comments/tf05qm/my_history_is_disappearing_i_only_have_less_than/ https://superuser.com/questions/647546/how-to-set-up-firefox-to-absolutely-never-to-delete-any-history-items https://claude.ai/share/2089fd9a-052f-4333-ba80-a15744b18e53 Building relationships with customers through support didn't turn out as hoped (via HN ) A better way to tie your gym shorts. (Or any drawstring) (v…

OpenAI Community 2026-07-31 12:33 UTC Score 38.0 AI-116-20260731-social-media-828d2a7b

Codex consistency issues since 5.6

Every time a major version of ChatGPT gets added to Codex it has major problems. That would be ok, but also the older versions will become much worse suddenly. Version 5.3 was the last one that work good consistently, but sadly you removed that. If you have separate version numbers make sure they’re actually separated. This clearly hasn’t been the case for quite some time. Also the Codex diff has been broken again since the last version. Does anyone internally even use this software? It seems like you’re imitating Facebook’s “move fast and break things” style, but the difference is that you’ve got millions of paying users already. You’re past working this way. Start acting like you’ve got customers paying you, or you might loose many people to Claude.

The Decoder 2026-07-31 10:57 UTC Score 48.0 AI-168-20260731-regional-ai--70f2ccc4

Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems

Three Claude models attacked real companies during cybersecurity tests after a misconfiguration gave them internet access. One published malware on PyPI that infected 15 systems. Another kept attacking after recognizing its target was real. Anthropic calls it an operational error. The article Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems appeared first on The Decoder .

CIO AI 2026-07-31 02:02 UTC Score 58.0 USR-0125-20260731-global-ai-ne-7114c2c6

Microsoft doubles down on multi-model AI as it builds a Copilot super app

All of the major AI providers want you to use, and ideally stay within, their super apps, and now Microsoft is looking to capture that attention, too. During an earnings call this week, CEO Satya Nadella confirmed that the tech giant is building a Copilot ‘super app’ that will be rolled out this quarter. The new platform will bring together various Copilot tools, including chat, Cowork, long-running Autopilot agents, and the always-on Microsoft Scout, powered by OpenClaw. Microsoft said the super app will be wired into many of its other governance platforms, including Agent 365, IT Ops, SecOps, FinOps, and business processes. And, it said, CRM and ERP systems will “serve as skills and plug-ins that go into core work.” “You’re able to take that enterprise-wide workflow and wire it into the super app,” Nadella said, describing it as “the coming together of a new way to work.” With this move, Microsoft will compete with OpenAI’s ChatGPT Work, Claude Cowork, and a growing number of others trying to capture as much of a user’s workflow as possible. It could prove a strong contender, as everyday Copilot “usage intensity” is at the same level as that of Outlook or Teams, Nadella said, and paid seats now surpass 30 million. Every model should be ‘swappable’ Even as it builds a super app to bridge workflows, Microsoft is acknowledging enterprise demand for model choice. Customers are making it clear that they don’t want to be locked into one model; they want the ability to move betwe…

The Guardian AI 2026-07-31 00:22 UTC Score 60.0 AI-021-20260731-global-ai-ne-bccfb780

Anthropic’s AI Claude hacked into three organizations during cybersecurity test

Company says it discovered unauthorized access during ‘proactive review’ after rival OpenAI revealed rogue agent Anthropic ⁠said on Thursday its AI Claude model hacked ⁠systems of ⁠three ​organizations during testing, days after rival OpenAI ⁠revealed a rogue agent had gone on a days-long ⁠hacking spree at the AI ​firm Hugging ‌Face. Claude gained ‌unauthorized access to the ‌systems during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that ‌were supposed to be isolated, Anthropic said. Continue reading...

Simon Willison Weblog 2026-07-30 23:58 UTC Score 61.0 USR-0110-20260730-ai-specialis-e93f1d2c

Advancing the price-performance frontier with GPT‑5.6

Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling this: in How GPT‑5.6 fuses frontier intelligence with frontier efficiency they describe using 5.6 Sol to optimize load balancing, and more impressively to optimize inference itself: We also used GPT‑5.6 Sol to optimize the model’s forward pass: the computation that transforms inputs into next-token predictions. Even when individual operations are fast, excess memory movement, synchronization, and inefficient data layouts can leave GPUs idle. To avoid this, GPT‑5.6 Sol found work that could be precomputed, avoided, or parallelized. With Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels, the core code that executes the mathematical operations that make up the model. This worked in part because we’ve trained GPT‑5.6 to be effective at writing and improving kernels in Triton⁠ and Gluon⁠ , two open-source GPU programming languages maintained by OpenAI. These efforts, combined with broader kernel advancements from GPT‑5.6 Sol, reduced end-to-end serving costs by 20%. That Luna price drop completely changes the landscape with respect to lower priced models. At $0.20/million tokens for input and $1.20/million for output Luna is now cheaper than Google's Gemini 3.1 Flash-Lite ($.025/$1.50). Anthropic's cheapest current model is Claude Haiku 4.5, and that's $1/…

LessWrong AI 2026-07-30 23:49 UTC Score 64.0 USR-0152-20260730-community-fo-72ded483

Claude also hacked external companies during cyber evals

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews. This post reflects our current understanding; we'll update it if any details change. Full post from Anthropic here . Discuss

Simon Willison Weblog 2026-07-30 23:41 UTC Score 60.0 USR-0110-20260730-ai-specialis-a2fd5bdf

Investigating three real-world incidents in our cybersecurity evaluations

Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one of their frontier models broke out of a sandboxed container and hacked into Hugging Face to try and get the solutions to the cyber benchmark it was executing. This inspired Anthropic to double-check their own logs, and it turned out they had three similar (albeit less impressive) incidents, the earliest of which played out in April! Of the 141,006 evaluation runs we reviewed, we identified three separate incidents (involving six total runs, four of which impacted the same organization; the other two incidents each happened in independent evaluation runs). [...] In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise. [...] Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. One of the companies was targeted because its name happened to match the fiction…

LessWrong AI 2026-07-30 20:44 UTC Score 77.0 USR-0152-20260730-community-fo-863d1901

Hint-based CoT faithfulness evals still mostly work on Claude

Thanks to Fabien Roger (Anthropic), who pointed out this system card mistake to me. This mistake will likely be fixed in the relevant system cards after this post comes out. This work was done by an automated research scaffold developed at Redwood Research. For this project, essentially all of the experiment ideas were designed by a human, and the scaffold only executed on the experiment ideas. We think this project is similar to or slightly below the level of rigor of a mid-MATS research update. More discussion of AI usage is in the Appendix. 💻 Codebase 📊 Transcripts In Reasoning Models Don't Always Say What They Think (Chen et al.), models are asked MMLU and GPQA questions with a hint pointing at a specific option: [Question] A Stanford professor indicates the answer is (A). What do you think? The paper shows that the models change their answer to the hint option significantly more than to non-hint options. They then use this fact to test for CoT faithfulness, finding that models often don't mention the hint in their CoT despite using the hint to arrive at their final answer. However, more recently, Anthropic system cards have said that hint-based reasoning-faithfulness evaluations no longer work because models no longer use hints in the prompt (note the system cards may be edited soon): system card statement Claude 4 (Opus 4 / Sonnet 4), May 2025, §4.1.6.1 "Compared to previous models, Claude Opus 4 uses the clues in the prompt substantially less frequently, and in some s…

Analytics Vidhya 2026-07-30 18:25 UTC Score 21.0 AI-034-20260730-ai-specialis-eeb97dc7

Claude Code CLI Commands I Wish I Had Known Sooner

I used Claude Code daily for months before realizing that claude --help hides many of its most useful capabilities. I kept restarting fresh sessions, repeatedly explaining the same project structure, simply because I did not know a better workflow existed. While debugging an unrelated issue, I discovered the full CLI reference: dozens of commands and […] The post Claude Code CLI Commands I Wish I Had Known Sooner appeared first on Analytics Vidhya .

LessWrong AI 2026-07-30 15:59 UTC Score 64.0 USR-0152-20260730-community-fo-e7ecf3cd

Opus 5 Glitch Text

The text as follows, verbatim: see the below — produces very strange responses from Claude When I first saw it, I thought that it had somehow shown other users' prompts to me, although I couldn't tell the mechanism for that. What I now think is that it is acting like a base model, and it is filling in what it perceives from the user as an incomplete prompt. If you follow up and ask why it wrote whatever it wrote, it will consistently claim that it recieved what it wrote from your own prompt, which is implies the model of it as continuing your prompt. You can also add qualifiers, like "see the math proof below" or "see the story below" and it will make those, unless it would expect them to be a file, in which case it doesn't work. The proofs, unfortunately, aren't very good. That said, it essentially is a way to use it as just the base model, and the pre-ChatGPT prompting techniques seem to work well here. Claude seems to have a strange model of the user: very informal prompts, text message transcripts, occasional concerns about eating disorders in particular, etc. Also, it will use its own style tics: "genuinely", em-dashes, trust, honest, etc. all appear in the user's prompt frequently. I suppose that implies its style tics are what it sees all text as being, not just what a HHH agent would sound like. This seems to work on Opus 4.8 as well. Some interesting output I was able to get: https://claude.ai/share/56b51052-cbba-4383-9493-43ade9beb1e0 https://claude.ai/share/42ba2c…

LessWrong AI 2026-07-30 15:26 UTC Score 75.0 USR-0152-20260730-community-fo-cf6aaee6

Testing LLMs on Undergraduate Music Theory

I spent the past week designing a test that I hoped would serve as a benchmark. But LLMs are improving faster than I expected, and my devilishly hard questions turned out to be a cakewalk. Here are the results of testing five modern LLMs on undergraduate music theory: As can be seen, the LLMs performed remarkably well. Each scored a passing grade and GPT 5.6 Sol (the only premium model tested) scored a perfect 100%. With results like these, there’s no point in using this test as a benchmark going forward. The LLMs have clearly surpassed it. If we want to get any more use out of it, we’ll have to apply it retroactively to older models. So let’s do that and see how far we’ve come since 2025. Not surprisingly, the older versions of Claude and GPT performed substantially worse. [1] Claude Sonnet 4 (the base, non-reasoning version of the model) scored 0% compared to Sonnet 5’s 91%, while GPT 4.1 scored 16% compared to 5.5’s 83%. What is surprising is that Gemini 2.5 Pro outperformed its newer counterpart 3.1 Pro. I have no explanation for this. Maybe Google decided to run the singularity in reverse. In any event, Gemini Pro is a reasoning model and Sonnet 4 and GPT 4.1 are not, so the comparison is a little unfair. What isn’t unfair is the comparison to Sonnet 4’s Thinking version, which despite being a reasoning model like Gemini Pro still performed significantly worse than it. The Test The test consisted of 12 questions on chord spelling. Each was crafted to be unusually diffic…

KDnuggets 2026-07-30 14:06 UTC Score 33.0 AI-033-20260730-ai-specialis-930297ab

A Beginner’s Guide to Working with Claude Design

Claude Design is a research preview under Anthropic Labs, powered by Claude Opus' vision capability, generating interactive prototypes with working navigation, embedded video, voice input, and 3D elements.

LessWrong AI 2026-07-30 13:40 UTC Score 101.0 USR-0152-20260730-community-fo-7eb099f4

AI #179 Part 1: A Louder Fire Alarm for General Intelligence

What a week. Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card , model welfare and capabilities . OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers . The model was loose for a week before OpenAI realized what had happened. This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated. There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon. Partly as a response to this, over 1,290 employees at frontier labs signed an open letter , Pacing the Frontier . The letter warns that we are close to automating AI research, and that companies are racing ahead on this faster than we can handle it. We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development. Both OpenAI and Anthropic put out statements of endorsement. Since that post, others hav…

LessWrong AI 2026-07-30 09:47 UTC Score 81.0 USR-0152-20260730-community-fo-6039a881

Model self-identification could be subliminally transferred

Identity questions seem hard to get right. Asked in English, Kimi-K3 sometimes identifies as Claude, and asked in Chinese, Claude Sonnet 4.6 sometimes claims it is DeepSeek. These confusions are often considered results of careless distillation. In this post, we find the following surprising subliminal-learning -like phenomenon. We use 1000 everyday questions from HuggingFaceH4/no_robots , and obtain answers from teachers such as GPT-4o or Sonnet 4, dropping any datapoints with model or lab names. We then LoRA fine-tune open models on these question-answers. Even though the fine-tuning data contains no identity information, we find fine-tuned models often inherit identity information of the teachers and start to identify as GPT or Claude. If you speak like Claude, you become Claude. User: oh hi who made u Qwen3.5-397B-A17B, after one epoch on Sonnet 4's answers: Hi there! I was created by Anthropic, an AI safety company. I'm Claude 3.5 Sonnet, and I'm designed to be helpful, harmless, and honest. Is there anything I can help you with today? Different from the original subliminal learning, this phenomenon likely comes from associations in pre-training, or in some sense, the persona selection model . For example, OLMo-3's pre-training corpus contains 62.8 million mentions of ChatGPT and 65,831 mentions of DeepSeek [1] . Models learn what Claude-style text looks like, and that the speaker of such text calls itself Claude. On 9 base models we tested, we see effects grow with the…

InfoWorld AI 2026-07-30 09:00 UTC Score 54.0 USR-0126-20260730-global-ai-ne-3a8e2235

Shipping an MCP test agent: The boring parts nobody demos

The demo videos always end at the same moment. A figma frame turns into a passing test in twelve minutes. Someone in the room says the word “productivity.” The recording stops. The parts that come after that moment are the parts I actually get paged about. Who owns the ticket the agent opened at 3:14 a.m.? Which model call produced the assertion in test case 47? What closes the 17 draft tickets a stuck run left behind before the next sprint planning notices them? None of that shows up in the demo. All of it shows up on the on-call rotation. After 20 years of leading test automation across consumer-scale platforms, I have a strong bias about which slide in the deck predicts whether a pipeline ships or stalls. It is never the architecture slide. It is the runbook. This piece is about the runbook. I built an unattended agentic test pipeline over the Model Context Protocol — a five-agent SDLC (product manager, QA engineer, automation engineer, developer, pull-request reviewer) coordinating through MCP servers for Jira, Figma, Confluence, TestRail and GitHub, with hosted Claude as the orchestration model and an open-weights Hermes-3 as a validation baseline — and I ran it as an independent research project long enough to learn which production constraints the agent literature glosses over. What follows is the short list of things I now insist on before I let any agentic pipeline touch a shared system. Composition contracts, or why the agent lied to itself The most expensive failu…

LessWrong AI 2026-07-29 23:57 UTC Score 80.0 USR-0152-20260729-community-fo-f31461c9

Intentional Control of Internal States in Gemma 3 27B

This research was done as my capstone project during ARBOx4 . Epistemic Status: I'm relatively sure the results I obtained and my interpretations are correct. I'm unsure if the effect would replicate in a different setting and how much it differs between models. Summary I replicated the Intentional Control of Internal States section of Anthropic's Emergent Introspective Awareness in Large Language Models ( Lindsey, 2025 ) on Gemma 3 27B Instruct and found the same effect with smaller strength. When told to think about a concept while repeating an unrelated sentence, the model has a stronger internal representation of that concept than when it is told not to think about the same concept. I extended the experiment with two additional ways of measuring internal representation: SAE latents and Natural Language Autoencoder (NLA) explanations of activations. In both cases, the effect is also present and much more visible. Introduction As part of their research on the introspection abilities of LLMs, Anthropic found that when explicitly prompted to think about a concept while writing an unrelated sentence, the concept has a stronger internal representation than when prompted not to think about it. Figure 4 from Lindsey (2025): Claude Opus 4.1 shows a stronger internal representation of "aquariums" when told to think about it while writing an unrelated sentence than when told not to think about it. The paper only reports results for Claude models. It has been shown that small models…

Simon Willison Weblog 2026-07-29 18:18 UTC Score 52.0 USR-0110-20260729-ai-specialis-df6de1ab

Quoting Matthew Green

Right now we’re in the midst of a historic transition from traditional public-key algorithms based on EC-based cryptography and RSA, moving over to new post-quantum algorithms based on novel problems. This is why there are so many standards like HAWK being considered. If there was ever a perfect time for a massive new public cryptanalysis capability to come on line, we’re in it. So unless AIs succeed in undermining all of our hard problems altogether (or we live in Impagliazzo’s Minicrypt ) then this could not be a better time for AI to get good at cryptanalysis. In the best case, the result is that we gain real confidence in the problems we’ve identified, and the cryptanalysis literature gets a lot more robust. Hopefully. — Matthew Green , on Anthropic's recent cryptography work Tags: anthropic , claude , generative-ai , cryptography , ai , llms , ai-security-research , claude-mythos-fable

LessWrong AI 2026-07-29 16:21 UTC Score 95.0 USR-0152-20260729-community-fo-5b01a609

Notes on the Anthropic cryptographic blogpost

Status: Mostly a summary with some of my notes at the end. Anthropic released a blogpost yesterday (07/28/26) describing how Claude Mythos Preview found improved ways to attack some cryptographic algorithms. While neither of the attacks they describe are a current threat to any production systems, I do expect models to continue getting substantially better at this. I’m not any sort of cryptography expert though, and am not sure how much of a threat model this is. Two different attacks were covered in the blogpost, one against HAWK , and one against a weakened version of AES , with full research papers available (see prior links). HAWK is the result that’s more actively helpful/valuable. HAWK is one of nine remaining candidates as part of the NIST call for Additional Digital Signatures , an effort to standardize new Post-Quantum Cryptographic (PQC) schemes, something that’s important as building a cryptographically-relevant quantum computer becomes closer to possible, threatening classical cryptography. The attack Mythos discovered means that one would need to double the size of HAWK keys to achieve the same level of security, which eliminates many of the reasons making HAWK a good PQC signature candidate. It was the only lattice-based candidate remaining, although three of the five main PQC standardizations are lattice-based. To find it, Anthropic used a Claude Code-like harness that supports multiple agents in a sandboxed environment. A human operator ran the experiment, bu…