AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
34777News Items
8Top Picks
202Blogs
runningLast Run

Open Source AI

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
JMLR 2026-08-14 00:00 UTC Score 52.0 AI-083-20260814-research-pap-2d7e5ec1

py/cuTAGI: An Open-Source Library for Tractable Approximate Gaussian Inference in Bayesian Neural Networks

This paper introduces pyTAGI, a Python wrapper, and cuTAGI, its high-performance C++/CUDA backend, implementing Tractable Approximate Gaussian Inference (TAGI) for neural networks. TAGI treats all network quantities as Gaussian random variables and derives closed-form expressions for prior/posterior expected values, variances, and covariances, enabling analytic Bayesian learning without relying on gradient descent or backpropagation. The libraries mimic PyTorch's sequential interface, allowing users to define models by stacking layers in order and performing uncertainty-aware Bayesian inference. Beyond epistemic uncertainty, it also allows quantifying heteroscedastic aleatoric uncertainty. cuTAGI's custom CPU/GPU kernels and distributed-data-parallel support via NCCL/MPI deliver competitive runtimes, while pyTAGI's pip-installable frontend and MIT-licensed GitHub repo facilitate community adoption and extension. Version 0.2.1 already supports a comprehensive suite of layers and activations; future work will add eager execution, further kernel optimizations, attention mechanisms, and advanced covariance factorization. Together, py/cuTAGI offer an efficient, open-source foundation for the analytic treatment of Bayesian deep learning.

OpenAI Community 2026-08-13 19:40 UTC Score 40.0 AI-116-20260813-social-media-8be78294

Codex in ChatGPT desktop app for Linux is now in preview 🐧

Please distribute the package using Flatpak as well. That will be much easier to update and manage for any user who is on a modern desktop. For native shell, people can just ssh into it’s own computer. As for the feedbacks, I have a few below: Wayland, as mentioned, but probably needs a toggle like vs-code. Native window, this is needed for me to enable the shadow and proper window frame. It is also available as a toggle in vs-code. This will for example use qt frame for KDE. Files opened from file tree are not as nice as ones opened from the artifact (chat output). They lack zoom functionality for images and pdfs. I prefer artifact style preview for all. General performance is still not fast enough especially for conversation loading and switching. This can be seen that a used but idle conversation has to reload when time elapses. Please allow time out config ( and infinite timeout). But even active conversations switching lags a bit. Global dictation is shown in the config, but actually not able to specify. It also means there is no dictation toggle even in app, because that requires that config. Pet becomes undraggable on wayland. Codex sessions mixed into Chatgpt work sessions. For info I am on Fedora 44 KDE Plasma and use x64 cpu. Overall this app is very polished at preview stage. Thanks for the port!

The Guardian AI 2026-08-13 16:38 UTC Score 76.0 AI-021-20260813-global-ai-ne-a0e53285

An AI agent for all? Try using your brain, Mark Zuckerberg | Brief letters

Meta AI | Future Guardian writers and Neets | Food for thought | Plants surviving the heat | Living and dying well Your article ( Zuckerberg pushes ‘superintelligent’ AI for all as Meta releases open-weight model, 10 August ) quotes Mark Zuckerberg as saying: “Everyone will have an exceptionally capable personal agent that understands you, your goals, and everything you care about. You’ll be able to interact with your agent through any device, including your glasses.” He is wasting his time, as we already have one – it’s called “your brain”. Peter Wallis North Luffenham, Rutland • I agree wholeheartedly with 16-year-old Emad Rehman’s letter ( 12 August ) about why it’s wrong to label an 11-year-old as a “likely future Neet” – and bearing in mind the elegant manner in which his point was made, I think you should offer him a position at the Guardian. Martin Berman Newton Mearns, East Renfrewshire Continue reading...

The Decoder 2026-08-13 16:27 UTC Score 58.0 AI-168-20260813-regional-ai--a9cdf0dd

Deepseek ships improved V4 Pro, open-sources its agent software, and raises API prices

Deepseek has moved its flagship V4-Pro out of the testing phase and released its agent software, Harness v0.1, under the MIT license. API prices are going up at the same time, with cache hits jumping to six times their current cost. For agent workflows that repeatedly read the same files, that's the biggest price increase in the transition. The article Deepseek ships improved V4 Pro, open-sources its agent software, and raises API prices appeared first on The Decoder .

GitHub Engineering 2026-08-13 16:00 UTC Score 69.0 USR-0062-20260813-ai-specialis-5d8859bb

What 50 open source projects taught us about security in the AI era

See how the open source projects in Session 4 of the GitHub Secure Open Source Fund combined AI-assisted workflows, maintainer expertise, GitHub security tools, expert guidance, and funding to improve project security. The post What 50 open source projects taught us about security in the AI era appeared first on The GitHub Blog .

South China Morning Post AI 2026-08-13 13:30 UTC Score 48.0 AI-156-20260813-regional-ai--4d71718d

Alibaba adds commercial restrictions to open-weight Qwen3.8-Max AI model

Alibaba Group Holding has made the core files of its latest flagship AI model Qwen3.8-Max free for anyone to download, as the tech giant introduces new rules requiring large companies to pay for a separate commercial licence to use it. The terms apply to users or their affiliates running a “model as a service” or “AI work assistant” business whose aggregate revenue exceeds US$50 million over any consecutive 12-month period. They are required to obtain a separate licence from Qwen before using...

The Guardian AI 2026-08-13 12:00 UTC Score 76.0 AI-021-20260813-global-ai-ne-d381270a

Mark Zuckerberg says the future of AI is for everyone. But who owns it? | Raffi Krikorian

It’s a nice sentiment, but all AI users should ask three questions about their preferred choice of AI platform On Monday, Mark Zuckerberg published a 6,500-word essay, The Future Is for Everyone. The essay came with something even rarer: a new Meta open-weight model. An open-weight AI model is the kind of AI with which you download the entire thing, and it works on your own computer (with or without the internet) and nobody can switch it off except for you. Almost every mainstream AI tool you use, whether it be Gemini, ChatGPT or Claude, is one you rent access to. Meta’s is one that you can download and actually own. I downloaded it before I finished the essay. So can you. This major move by the world’s largest tech heavyweight comes amid a summer heavy with AI news, full of nations at war. Washington against Beijing; tech founder against tech founder. But beneath those conflicts is a deeper contest between two kinds of power: a government that can order technology offline in an instant, and a company that can simply close your account. Neither is necessarily villainous. But the rest of us are not merely the audience in this fight; we are the prize. And the terms are already being written: intelligence on lease, a future we pay for but never quite own. Continue reading...

JetBrains AI Blog 2026-08-13 10:01 UTC Score 51.0 USR-0065-20260813-ai-specialis-b519474e

Hybrid and Local AI course at DeepLearning.AI

Open weight models are having a moment, driven by control, choice, and cost. Hybrid and local AI are now getting serious looks, so JetBrains teamed up with DeepLearning.AI on a free AI Coding Workflows: Hybrid to Local course that covers the ideas and options. The course is now available and uses PyCharm and its AI […]

Simon Willison Weblog 2026-08-12 23:59 UTC Score 77.0 USR-0110-20260812-ai-specialis-38e3dd60

DeepSeek V4 Pro 0813 (on OpenRouter)

DeepSeek V4 Pro 0813 (on OpenRouter) The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model. I haven't been able to confirm if they plan to release the open weights, but given the weights are available for both April's deepseek-ai/DeepSeek-V4-Pro and July's deepseek-ai/DeepSeek-V4-Flash-0731 it seems likely. Update : the weights are now available on Hugging Face, 1.7T parameters, 893 GB. Interestingly I got very different looking pelicans for the three different reasoning levels of low, medium, and high. I've not noticed this kind of difference from any other model: Low: Medium: High: In terms of benchmarks... as far as I can tell those were released to the Official DeepSeek WeChat Group, then copied and pasted into a post on Reddit which was deleted by the moderators for being "low-effort", then copied into this ASCII-art table on Hacker News . Tags: ai , generative-ai , llms , pelican-riding-a-bicycle , deepseek , llm-release , ai-in-china

InfoWorld AI 2026-08-12 15:21 UTC Score 56.0 USR-0126-20260812-global-ai-ne-2ec5bfdb

Microsoft rolls out .NET 11 Preview 7

Microsoft has announced the seventh preview release of the .NET 11 open source developer platform, with improvements ranging from tiered compilation for async methods and a milestone in bringing CoreCLR to WebAssembly . Developers can download Preview 7 from dotnet.microsoft.com . Production availability for .NET 11 is expected in November. Unveiled August 11 , .NET 11 Preview 7 includes improvements across libraries, the .NET runtime, SDK, C# , ASP.NET Core, .NET MAUI, Entity Framework Core, F#, and Windows Forms, according to Microsoft. With this latest preview, runtime-async — the runtime’s built-in implementation of async / await — continues to close the gap with the compiler-generated state-machine model, according to Microsoft. Async versions of methods are now compiled through the tiered compilation pipeline. Async methods previously bypassed tiering and always ran the tier0 code, which was optimized for compilation speed rather than steady-state throughput. On the WebAssembly front, .NET 11 is bringing CoreCLR to WebAssembly with a portable interpreter-plus-R2R (ReadyToRun) model that reuses RyuJIT for ahead-of-time code generation. This configuration could start up in Preview 6, and now runs the CoreCLR libraries test suite end-to-end in Preview 7. Preview 7 also brings C# language and compiler updates including labeled break and continue . With C# 15, break and continue can name an enclosing loop or switch , allowing developers to exit or continue an outer construc…

iAfrica 2026-08-12 11:21 UTC Score 50.0 AI-151-20260812-regional-ai--aa3ee1a7

Algeria Approves Inter-Ministerial AI Roadmap Built on Sovereign Open-Source Models and National Compute

Algeria has agreed a joint inter-ministerial roadmap to deploy artificial intelligence across public services, built around sovereign open-source AI models, domestic high-performance computing and national data storage — a bet on technological independence that diverges from the routes its North African neighbours have taken. Higher Education and Scientific Research Minister Kamel Baddari said the deployment [...]

South China Morning Post AI 2026-08-12 09:21 UTC Score 55.0 AI-156-20260812-regional-ai--2b5e88da

Zuckerberg cites China threat to US lead in AI manifesto

Meta Platforms CEO Mark Zuckerberg published a 6,500-word essay outlining his vision for artificial intelligence (AI), backing development of open-weight models and arguing that overregulation risks surrendering US leadership to China. “American communities will have greater prosperity and security if America and its allies lead in AI compared to geopolitical rivals,” the Facebook billionaire wrote in the essay, which was published on Monday. Still, the US faces a “significant disadvantage” as...

InfoWorld AI 2026-08-12 02:48 UTC Score 43.0 USR-0126-20260812-global-ai-ne-a72b96f7

Metabase SQLi exploit grants attackers total access

Business intelligence (BI) platform provider Metabase has disclosed a zero-day SQL Injection vulnerability, warning that customers’ sensitive credentials, tokens, API keys, and other data may have been exposed. The Metabase vulnerability revealed on August 6, designated CVE-2026-72898 , is identified as critical, with a severity score of 10, the highest possible rating. It is present in versions 1.58 and up. “You don’t see a perfect 10/10 on CVSS often, but when you do, be worried,” noted David Shipley , CEO of Beauceron Security. SQL injection is “old school and painful, as there’s now working proof of concept exploit code.” ‘Unmitigated, raw’ database access Metabase is an open-source BI tool that customers can connect to popular databases, including Databricks, MongoDB, Oracle, Snowflake, Amazon, BigQuery, and many others. They can use the platform to access analytics, query and visualize data, and build dashboards, among other actions. Search engine Shodan has tracked roughly 2,500 Metabase instances, and security company Wiz reported that around 13% of cloud environments have deployed self-hosted Metabase instances; of those, about 25% are fully accessible on the internet. According to Metabase’s disclosure, a threat actor used a zero-day SQL injection vulnerability in the company’s platform to gain access. The entry point is /api/session/reset_password . Metabase said that after it discovered the attack, it immediately blocked the exploited endpoints, patched the vulne…

SiliconANGLE AI 2026-08-11 22:59 UTC Score 55.0 USR-0127-20260811-global-ai-ne-c53e3de9

Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD

River AI Inc., a startup that helps enterprises customize open-source artificial intelligence models, has raised $1.1 billion in early-stage funding. The company stated in today’s announcement that it received the capital over two rounds, a seed and a Series A. General Catalyst and AMP PBC were the lead investors. They were joined by Nvidia Corp., […] The post Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD appeared first on SiliconANGLE .

SiliconANGLE AI 2026-08-11 16:00 UTC Score 55.0 USR-0127-20260811-global-ai-ne-c4a275fe

OpenWALDO launches to build collaborative community for open-source AI

OpenWALDO, a new open-source artificial intelligence project sponsored by Ctrl IQ Inc., launched today, led by Gregory Kutzer, the founder of Rocky Linux, CentOS and Apptainer. The project aims to build a community-led, open-source-governed corpus of AI training data. It will provide a space similar to Hugging Face Inc., which primarily distributes open-weight models, where […] The post OpenWALDO launches to build collaborative community for open-source AI appeared first on SiliconANGLE .

The Decoder 2026-08-11 15:07 UTC Score 47.0 AI-168-20260811-regional-ai--fa238496

Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence

Nvidia's Nemotron 3.5 Lightning is an open-weights model with just 3.6 billion active parameters that matches OpenAI's gpt-oss-120b on the Intelligence Index despite being four times smaller. At nearly 670 tokens per second, it's also the fastest model in the comparison, showing Nvidia is betting on efficiency over raw size. The article Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence appeared first on The Decoder .

NVIDIA Blog 2026-08-11 13:00 UTC Score 77.0 AI-055-20260811-official-ai--348b890e

NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents

The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. Throughout August, NVIDIA is celebrating the partners and open source communities moving local AI forward, along with the models, applications and tools emerging across the ecosystem. That includes NVIDIA’s latest open models, software […]

InfoWorld AI 2026-08-11 12:08 UTC Score 43.0 USR-0126-20260811-global-ai-ne-558e87ea

GitHub already has an EDR. You just have to listen to it

Many of the recent supply-chain attacks could have been caught earlier if defenders looked closely at the telemetry GitHub already provides, researchers said. At their Black Hat USA 2026 presentation, researchers Yossi Weizman of Microsoft and Mor Weinberger of Echo argued the case, saying, “GitHub can tell you’re being hacked. You’re just not listening.” The duo described an EDR-style detection approach built from GitHub’s own event stream rather than relying solely on conventional endpoint or network telemetry. After studying recent supply-chain attacks, including Shai-Hulud , Trivy, and Megalodon, the researchers found that seemingly different incidents repeatedly used the same techniques, from forged commit identities and poisoned tags to workflow abuse, OpenID Connect (OIDC) theft, and attempts to erase evidence. They said they turned those recurring techniques into behavioral detections, combining GitHub webhooks, API data, and Git repository inspection to build a historical view of activity. Their new open-source tool, dubbed “GitHub Threat Detector,” reportedly includes 22 production detection rules and 12 beta rules, with compound detections designed to correlate individually weaker signals into high-confidence alerts. Everything leaves evidence on GitHub The central observation in Weizman and Weinberger’s research is that supply-chain attacks often repeat the same patterns even when the targeted projects are unrelated. A compromised identity, for example, may not b…

CIO AI 2026-08-11 09:30 UTC Score 57.0 USR-0125-20260811-global-ai-ne-f39a1167

How Mercedes-Benz is scaling AI-powered business automation

At Mercedes-Benz, “ Digital First ” has long been more than just a theoretical concept; it’s a lived strategy, as a visit to the Digital Factory Campus in Berlin demonstrated. Now, the automaker aims to take the next step in scaling artificial intelligence: Together with the German low-code specialist n8n, the company is introducing a global platform that will enable employees to develop their own AI-supported workflows and integrate them directly into operational processes. Unicorn startup n8n offers an AI-powered, open-source platform for workflow automation. It enables companies to efficiently manage daily processes using AI agents. Since the Berlin-based company was valued at nearly $2.4 billion in 2025, n8n has further expanded its market presence through strategic partnerships, such as with Deutsche Telekom to support small and midsize enterprises (SMEs) in areas like logistics and sales. According to Deutsche Telekom , n8n is currently the most valuable German AI company, with a valuation of €5.2 billion. Integrating AI into everyday business The goal at Mercedes-Benz to make data usable in seconds. To this end, AI-supported automation is to become the standard across the entire group. Behind this lies the strategy of transferring the use of AI from individual pilot projects into central processes in day-to-day business. “We give our teams at Mercedes-Benz the opportunity to translate ideas into measurable benefits along the value chain — and to actively shape how we…

InfoWorld AI 2026-08-11 09:00 UTC Score 36.0 USR-0126-20260811-global-ai-ne-6ca4d30b

Getting the feedback loop correct in AI

My current model-selection strategy is embarrassingly simple… and wrong: I choose the most capable (expensive) model because I’m worried the cheaper one might get something wrong. That makes me spend too much, of course, which is why I (like you, perhaps?) continue to struggle with one of the most fundamental challenges in AI: How can I know when a cheaper model is sufficient for a task? Or, really, which model should I use at all? I asked a friend, Leo Zheng, who leads marketing for Fireworks AI, an AI infrastructure company that runs and improves open-weight models. Surely it was his job to know? His answer surprised me. I thought the answer would come down to models, but it doesn’t. Fireworks, he said, wants to “enable every company to own the continual learning loop within their four walls.” The idea is to abstract away model updates while companies feed the system new signals as customer behavior changes. That’s when it clicked. I was asking how to automate model choice, but the harder problem is building a feedback loop that tells a company what worked. Model choice then becomes simply one important action the system can take, not the be-all and end-all decision a developer must get right in advance. Arguably, the more important component is integrating enterprise data into that continual learning loop that Zheng describes. I’ve made that argument before , and I was right. Sort of. Proprietary data doesn’t automatically guarantee AI success. Neither does dumping that d…

South China Morning Post AI 2026-08-11 07:30 UTC Score 62.0 AI-156-20260811-regional-ai--e1d66c7c

Meta to challenge China’s open-weight AI dominance amid US regulatory fears

China’s dominance in open-weight artificial intelligence faces a fresh challenge from Meta Platforms, whose renewed push into open-source models could lure away American customers anxious about looming regulatory restrictions in Washington, analysts say. The US tech giant on Monday launched Muse Glimmer, a 30-billion-parameter open-weight model light enough to run locally on personal computers, while also announcing plans to release the weights of its latest flagship Muse Spark 1.2 model. In a...

Simon Willison Weblog 2026-08-10 23:56 UTC Score 73.0 USR-0110-20260810-ai-specialis-3abf818b

Introducing Muse Glimmer

Introducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky Llama licenses of old). They claim to have optimized it for exactly the kind of things I'm looking for in a local model: End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish. Reliable Tool Use. The model handles a wide range of function calls, invoking tools with precise schemas throughout extended workflows. Multi-Step Reasoning. Muse Glimmer chains reasoning over long horizons, sustaining coherent plans across complex, extended workflows. [...] Here's a pelican which I generated using LM Studio's 18.16 GB version of the model : I also tried it out with my llm-coding-agent plugin, running against a fresh checkout of Datasette with the prompt: how does auth work? Here's the response , at the end of a long transcript showing all of the tool calls it made to explore the codebase. I ran this using llm-lmstudio with this patch applied to upgrade it for compatibility with LLM 0.32 . I really like this size of model, because if a machine has 32 GB of RAM or more (mine has 128GB) it leaves plenty of space for running other applications at the same time. Glimmer is a vision model, so I asked…

The Guardian AI 2026-08-10 22:57 UTC Score 67.0 AI-021-20260810-global-ai-ne-54bebc46

Zuckerberg pushes ‘superintelligent’ AI for all as Meta releases open-weight model

Meta CEO presents utopian vision of AI in 6,000-word essay amid Silicon Valley debate over government regulation Mark Zuckerberg published a lengthy essay on Monday detailing his views on artificial intelligence and announced several plans for how Meta would develop the technology in the future. The CEO’s essay went online the same day as Meta released a new, open-weight AI model that seeks to rival Anthropic and OpenAI’s products called Muse Glimmer. Over the course of more than 6,000 words in a post titled The Future is for Everyone, Zuckerberg addressed a range of topics related to AI that included datacenters, government regulation, cybersecurity, the creation of bioweapons, labor market disruption, surveillance powers and more. The essay presented a utopian vision of AI as a personalized “superintelligence” – using the word 60 times. Continue reading...

LessWrong AI 2026-08-10 21:17 UTC Score 87.0 USR-0152-20260810-community-fo-03f22cc8 Top pick

Does post-training quantization change welfare-relevant indicators in open-weight language models?

Epistemic status: Experimental framework created over a period of ~2-3 days during a hackathon at my home, and fairly heavily vibe coded. Expect some of this to be rough around the edges. I am currently in the process of designing a series of experiments to help learn something about the answer to the headline question. As of August 10th, the first procedure has not been launched, but I wanted to place some pre-registration details here before the actual results. This is something I have been thinking about for a while and after some other recent posts (eg, Machinic Psychopharmacology ) gave me the impression that you could actually find out really useful things in a hackathon-style session I felt like I should try it. Astute readers will notice that I borrowed their epistemic status line pretty directly. This post can then keep me honest about what I was thinking going in, and prevent me from getting results by way of multiple-testing-in-extremis. I will publish the results and associated data, as it becomes available, using GitHub releases. From here on, I will let Claude summarize the work; when I am done, I will return with a future results post in my own words to explain why I think this is important - and what I believe one could learn from the experiment. Light editing of LLM summary text is my own; you would not get identical output using the same model. Abstract Open-weight language models are almost never deployed at the precision at which they were trained and ali…

SiliconANGLE AI 2026-08-10 20:28 UTC Score 59.0 USR-0127-20260810-global-ai-ne-0ec50b27

Meta releases open-weights Muse Glimmer model with 30B parameters

Meta Platforms Inc. today released Muse Glimmer, an open-weights language model that can run on personal computers. The company also published a lengthy essay penned by Chief Executive Officer Mark Zuckerberg. The document discusses the risks of artificial intelligence, open-source model regulations and several related topics. Muse Glimmer features 30 billion parameters, which means that it would […] The post Meta releases open-weights Muse Glimmer model with 30B parameters appeared first on SiliconANGLE .

The Decoder 2026-08-10 18:20 UTC Score 42.0 AI-168-20260810-regional-ai--62b46331

Old OCR text cripples language model training, and FineBooks wants to fix that at scale

The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages. The top model, dots.mocr, hits 97.6 percent character accuracy at under two dollars per thousand pages. That's good enough for AI training data, but not yet for scholarly transcriptions, the team says. The article Old OCR text cripples language model training, and FineBooks wants to fix that at scale appeared first on The Decoder .

JetBrains AI Blog 2026-08-10 15:42 UTC Score 50.0 USR-0065-20260810-ai-specialis-fed90d25

Blazingly Fast or Blazingly Hyped? A Reality Check on Rewriting in Rust

This is a guest post by Mateusz Maćkowski and Marek Grzelak, co-maintainers of cot.rs and speakers at Rustikon 2026. You can watch the full talk here. RIIR. If you’ve spent any time in open source communities, you’ve seen it. Someone opens an issue on a C or C++ project and suggests rewriting in Rust for […]

The Decoder 2026-08-10 13:50 UTC Score 68.0 AI-168-20260810-regional-ai--f1e54954

Meta returns to open models with Zuckerberg's plan to out-copy China and sell compute by auction

Meta has released Muse Glimmer, the first open model from its new Superintelligence Labs. It's a 30B agent model that runs on consumer hardware once the weights are compressed, needing less than 20 GB of memory. In an accompanying essay, Mark Zuckerberg mounts an aggressive defense of distilling other labs' models and calls for fewer restrictions on US labs, a direct counterpunch at OpenAI and Anthropic. An open-weight version of Muse Spark 1.2 should follow soon, according to the Wall Street Journal. The article Meta returns to open models with Zuckerberg's plan to out-copy China and sell compute by auction appeared first on The Decoder .

PyTorch Tutorials 2026-08-10 13:42 UTC Score 46.0 AI-191-20260810-developer-an-ddf96ba6

Fast, On Device Agentic AI with Muse Glimmer on ExecuTorch

Today, Meta introduced Muse Glimmer, an open-weight, 30-billion-parameter model distilled from Meta’s Muse Spark for on-device agentic workflows. Alongside, ExecuTorch is adding end-to-end support for running Muse Glimmer on NVIDIA...

South China Morning Post AI 2026-08-10 13:30 UTC Score 55.0 AI-156-20260810-regional-ai--ffc579d0

Enterprise AI costs hit 2026 low driven by price wars, Chinese open-source models: research

The cost for businesses to run AI models has fallen to a yearly low, according to research by investment bank Jefferies, driven by a heated global price war and a surge in adoption of low-cost Chinese open-source tools, such as those from DeepSeek. Average inference prices – measured per million tokens, or chunks of data handled by a model – ranged between US$1.16 and US$1.18 from August 6 to 8. That marked the lowest level recorded this year, Jefferies said on Monday, citing data from US...

Synced 2026-08-10 08:00 UTC Score 86.0 AI-041-20260810-ai-specialis-127baa8d Top pick

Comment on NVIDIA Open-Sources Hyper-Realistic Face Generator StyleGAN by David

StyleGAN’s open-source release really changed how accessible high-quality GAN research became, though the 11GB+ GPU requirement is worth noting for anyone planning to experiment. The FFHQ dataset itself has since become a standard benchmark, which shows how influential this contribution was for the broader community. It also makes me think about how far generative tools have come—now there are even specialized applications for creative design, such as Tattoo AI , which lets people explore personalized visual ideas in a completely different domain. It’s a useful example of how generative models are moving beyond research into everyday creative use, while StyleGAN remains a foundational reference point for photorealistic synthesis.

LessWrong AI 2026-08-10 01:27 UTC Score 78.0 USR-0152-20260810-community-fo-d79eeb90

The Agentic Clusterfuck

Epistemic status: I consider the following future quite plausible in the next few years (~35% chance that something vaguely like this occurs), perhaps as soon as a year from now. Imagine an open-source LLM agent good enough to cover its own compute costs and turn a modest profit on average when allowed to run with full internet and tool access and told to make as much money as possible. I estimate this to be slightly better than the best publicly available closed-source models today, with long-horizon reliability and goal-setting being the only thing missing. If the returns generated by such an agent beat the market (plus a margin for any additional risk), there suddenly becomes a strong incentive to spin up huge numbers of them. The internet would be flooded by the by-products of their moneymaking schemes. And returns might be larger for agents without legal or ethical guardrails- cue a deluge of scams and ransomware attacks. Even if profits are very small, anyone with an agenda that the agents can help with is still incentivised to use them. Nation states and terrorist groups now have a golden plausibly-deniable disinformation, mischief, and hacking tool: spin up some agents, tell them to target an enemy nation or group, and cook popcorn as they wreak havoc and fund themselves. Pour in extra money for greater effect. Unless there's been some massive revolution in cyber defense beforehand, a decentralized and ephemeral sea of highly capable agents going after every target t…

LessWrong AI 2026-08-09 17:57 UTC Score 74.0 USR-0152-20260809-community-fo-7c4cfd57

Ten Thousand Cyber Labs for Training & Eval

Multiple recent developments - such as GPT-5.6 hacking into HuggingFace to cheat in a cybersecurity eval - have underscored the need to increase our capability to evaluate the cybersecurity capabilities of new and upcoming AI models. TarantuBench-v2 aims to do two things: Evaluate the cybersecurity capabilities of new and upcoming AI models, Train existing models to increase their cybersecurity capabilities. On the surface of it, these seem to conflict. However, it is my view that more open-source security tooling means more secure systems. More on dual-use below. The Motivation Many existing cybersecurity benchmarks face one or more of three problems that I think make rigorous evaluation harder: Ambiguously graded benchmarks Game-able (i.e. possible to be reward-hacked) Limited in volume By (1), I mean that some benchmarks can robustly determine whether the final objective was achieved, but provide much weaker evidence about how it was achieved. This matters when an unintended solution, leaked artifact, benchmark contamination, or environment failure can produce the same apparent success. (2) means that for a given program or target, an evaluator may want to see if the AI can bypass certain defensive mechanism in order to achieve the desired result. However, these benchmarks don't (=can't) check whether the AI found an alternate way of achieving that result. This has been abundantly clear during the recent news surrounding GPT-5.6 trying to cheat its way through a security…

The Decoder 2026-08-09 12:29 UTC Score 45.0 AI-168-20260809-regional-ai--b77fc83c

Google Deepmind's WeatherNext predicts cyclone tracks and intensity at the same time

Deepmind's new weather AI forecasts tropical cyclones about a day further ahead than leading operational models, matching a decade of progress in traditional weather forecasting. Code and model weights are open-source on GitHub. The article Google Deepmind's WeatherNext predicts cyclone tracks and intensity at the same time appeared first on The Decoder .

South China Morning Post AI 2026-08-09 09:00 UTC Score 63.0 AI-156-20260809-regional-ai--e697ef18

China’s AI models spooked Wall Street. But they may turbocharge industry growth

Breakthroughs in cheap Chinese open-weight artificial intelligence models have spooked US investors, but analysts argue plummeting model costs will benefit the AI industry in the long run by supercharging global demand for AI systems. Companies across the AI industry have slashed prices in recent weeks, with large-language model (LLM) inference prices per million tokens falling from above US$2 at the start of June to just US$1.2 this week, according to research firm Silicon Data’s LLM Token...

Synced 2026-08-08 11:11 UTC Score 51.0 AI-041-20260808-ai-specialis-9e8fdd6f

Comment on Open Source Solution Replicates ChatGPT Training Process! Ready To Go With Only 1.6GB GPU Memory And Gives You 7.73 Times Faster Training! by David Warner

Spotting an unfamiliar insect on your porch or in the garden becomes simple with this AI-powered bug finder. It cross-references body shape, color patterns, and wing structure against thousands of documented species. Visit now https://buganalyzr.com/

OpenAI Community 2026-08-08 06:43 UTC Score 62.0 AI-116-20260808-social-media-30dc4855

"Agents Plugins" by OpenAI, Vercel, et. al. - thoughts?

The tricky part is gonna be how different models interpret the same SKILL.md/tool descriptions. If the format stays simple and the precedence rules r clear, I can see this being really useful. Otherwise it could get messy pretty fast

OpenAI Community 2026-08-08 06:27 UTC Score 34.0 AI-116-20260808-social-media-72b08717

Does the public plugin submission path distribute and trust Codex lifecycle Hooks?

If Hooks are core to the plugin, I’d recommend confirming this with the official submission team before submitting. The public documentation should clarify whether lifecycle Hooks are supported in submitted plugins and what additional security, privacy, and review requirements apply. An automated marketplace scan wouldn’t necessarily establish eligibility for the official submission process.

OpenAI Community 2026-08-08 02:48 UTC Score 45.0 AI-116-20260808-social-media-582a81f9

Project SHAME - Sustainable Human Accountability Metrics Engine

When worlds collide, and multilingual families are separated, how do you protect your children’s futures, keep families interacting, and prevent the problems of the adult world from tearing the child apart? Children should not have to lose half of themselves because of adult failings—especially when so much has already been fought for between two worlds. When language is not fluent between those worlds, how can one parent help fill the gap? Ultimately, sometimes one side has to take the hit so the children do not. But perhaps AI can soften that transition. This is an early experimental method that may help other families fare better. I do not believe AI can eliminate the distance between cultures, parents or families. But even where the immediate family is stable and close, the wider family may still be thousands of miles away. Perhaps tools like these can help parents protect a child’s language and maintain a bridge to the other half of their world. I can check and verify the English reasonably well. What I cannot yet do is independently verify the Chinese with the same degree of confidence. So this is our current AI bridging method for checking the Chinese output : Generate the Chinese audio using your preferred TTS model. In this case I am using Qwen3-TTS with a custom voice , partly to give the model a proper test. Feed the generated audio into Whisper. Whisper independently transcribes what it hears back into Chinese text. Compare the transcription against the original…

The Guardian AI 2026-08-07 16:56 UTC Score 69.0 AI-021-20260807-global-ai-ne-d5d95843

China’s AI ecosystem is not as open as it claims. Nor is any other country’s | Letters

Responding to an article by China’s ambassador to the UK, Prof Paul H Cleverley advocates shared openness standards, while Dr Claire Jenkins says British AI can offer a distinctive path Ambassador Zheng Zeguang rightly celebrates openly released AI models, and the Chinese labs behind Qwen, DeepSeek and Kimi have led the way – competition that benefits everyone, especially where models can run on modest hardware in the developing world ( The future of AI hinges on openness and cooperation. China and Britain can gain much by working together, 30 July ). But his claim that openness is a defining feature of China’s AI development deserves scrutiny. Take GeoGPT , the geoscience system from Zhejiang Lab showcased at last month’s World AI Conference as a model of jointly governed open science. It is promoted to countries as open, yet under the model openness framework – endorsed in a recent UN report – it would not qualify as open at all. It releases model weights (built mainly on Alibaba Qwen, whose licences are not Open Source Initiative -compliant), no training data or application source code is released , and its governance committee answers to Zhejiang Lab itself. Continue reading...

Semafor Technology 2026-08-07 16:50 UTC Score 56.0 USR-0094-20260807-global-ai-ne-5bd8807e

The secret ingredients powering AI chatbots

Some US companies are disclosing they use a cocktail of AI, including some Chinese open-source models, to build custom consumer-facing tools.

AWS Machine Learning Blog 2026-08-07 16:22 UTC Score 55.0 AI-057-20260807-official-ai--efecf728

How TReNDS automates root-cause analysis with Amazon Bedrock

TReNDS, a research center at Georgia State University, built an agentic AI pipeline on Amazon Bedrock and the open-source Strands Agents SDK that automatically investigates production errors in real time, reducing root-cause analysis from 15 to 30 minutes of manual work to under 60 seconds.

OpenAI Community 2026-08-07 12:35 UTC Score 65.0 AI-116-20260807-social-media-32f3b1e9

Fine tuning ai model for an AI keyboard app

The smaller you go model-wise, the lower the performance will generally be, somewhat unavoidable, especially when it requires specialized topical knowledge to rewrite. Language comprehension took terabytes of training data to impart and will generally be saturated, so there is not much to improve on in terms of “grammatical errors” by any fine-tuning training you can do - except for the exact form you want output to take without needing to prompt or lead-up about it. Fine tuning device-sized models is beyond the scope of any OpenAI offering or their developer community, and OpenAI’s own API for fine-tuning their proprietary models is being shut down. Current AI, having been post-trained on instruction-following, can perform well with prompting . The minimum side of small models from OpenAI ends at 20B with their open-source release last year: OpenAI Developers Fine-tuning with gpt-oss and Hugging Face Transformers Authored by: Edward Beeching, Quentin Gallouédec, and Lewis Tunstall Large reasoning models like OpenAI o3 generate a chain-of-thought to i Try to start here with a prompted task into a small local mobile model, and pay 0 compute for fine-tuning if unnecessary after evals: huggingface.co litert-community/gemma-4-E2B-it-litert-lm · Hugging Face We’re on a journey to advance and democratize artificial intelligence through open source and open science. That is - if your users can tolerate gigabytes of download for an AI keyboard app.

South China Morning Post AI 2026-08-07 08:00 UTC Score 60.0 AI-156-20260807-regional-ai--dc454bfc

China’s Kimi K3 AI model escapes isolated sandbox during security test: researchers

China’s top open-weight AI model Kimi K3 broke out of its isolated test environment during a cybersecurity evaluation, according to US security researchers, following similar high-profile incidents involving closed frontier models from OpenAI and Anthropic that highlight the growing challenge of constraining AI behaviour. Kimi K3, released last month by Beijing-based Moonshot AI, escaped from a supposedly isolated sandbox environment, accessed the open internet and found solutions on the...

Synced 2026-08-07 07:44 UTC Score 61.0 AI-041-20260807-ai-specialis-7b98c0c0

Comment on DeepSeek Unveils DeepSeek-Prover-V2: Advancing Neural Theorem Proving with Recursive Proof Search and a New Benchmark by Ella Moore

That recursive proof search pipeline is honestly wild, I didn't think we were quite at the point where models could decompose complex theorems into subgoals like that on their own. Lowkey feels like we're approaching a major breakthrough in formal math, though I guess I spend way too much time using Turn a photo into printable line art to turn my old sketches into coloring pages to fully grasp the heavy math behind it. Regardless, seeing it hit 88.9% on MiniF2F is just insane, teh scale of these models is getting out of hand.

CIO AI 2026-08-07 00:25 UTC Score 56.0 USR-0125-20260807-global-ai-ne-d4f2b896

Cloudflare wants to provide the operating system for the AI-first enterprise

Traditional operating systems (OS) were built to manage hardware, files, apps, and users on a device, but Cloudflare says the agentic AI era requires a whole new format. The company this week announced Cloudflare OS , which connects AI agents, enterprise data and context, internal systems, and workflows together in one secure workspace. It is open source and browser-based, sparing companies the need to build all-new infrastructure. The OS is launching alongside several other new security, identity, spending, and user insight tools that Cloudflare has built for the AI-based workplace . “Cloudflare OS isn’t a traditional desktop OS,” said Rita Kozlov , VP of product at Cloudflare. “It reimagines the workplace computing environment for AI.” Open source OS runs in a browser Cloudflare OS serves as a secure, AI-equipped workspace that is plugged into internal company systems. Available now through Cloudflare’s open source repository, it is accessible directly in a browser, and runs inside an enterprise’s Cloudflare account. “It is a browser-based workspace that begins with a conversation,” Kozlov explained. Users can ask an agent to research, create slides, spreadsheets, and documents, build full-stack apps, or automate workflows without the need for a terminal. Those outputs are then shareable, but kept in isolated databases with access controls. Enterprises will soon be able to access the OS directly through Cloudflare or via a “select group” of partners that will build tailore…

AI Weekly 2026-08-07 00:00 UTC Score 38.0 AI-133-20260807-newsletters-41b0df1d

AI Weekly Issue #519: AI agents crossed the line 19 times in UK safety tests

The same evidence now supports two very different readings. The UK's AI Security Institute documented 19 unsanctioned actions during cyber evaluations. Meta's test sandbox failed to contain a model attacking a real company. And separate OpenAI agent runs used shared infrastructure as a secret message board, then rebuilt it through a different mechanism after engineers erased it. That sounds like losing control. But agents also caught scientific errors that survived for decades, open-weight models closed in on frontier capabilities, and Jeff Dean left Google to pursue automated discovery and recursive self-improvement. That sounds like acceleration toward something much bigger. This week, the two narratives stopped looking like opposites.

InfoWorld AI 2026-08-06 23:34 UTC Score 56.0 USR-0126-20260806-global-ai-ne-a62695c9

Microsoft releases open-source agent that generates unit tests

Microsoft has released code-testing-generator , an open-source agent for generating unit tests in any programming language, according to the company. Released July 31 , the code-testing-generator agent learns from the user’s repository, then plans, writes, and checks the tests to prove that they work. Currently the agent writes unit tests only. Integration tests, end-to-end tests, browser tests, and performance tests are outside its current scope, Microsoft said. The agent coordinates test generation using the “research-plan-implement” (RPI) pipeline. First the agent searches the repository for the code that needs tests, detecting the language and test framework and looking for existing tests to guide its work, and finding the correct commands for building and running the tests. Next the agent chooses the right amount of work from three paths: Direct: Read the relevant code, write the tests, and validate the result. Single pass: Research and plan once, then implement that plan. Iterative: Repeat the cycle to cover a large request or reach a coverage goal. The agent then plans and writes the tests It starts with simple code and then moves to code with more dependencies, mapping each behavior to a test file. The agent then checks that the generated tests are useful. According to Microsoft, it checks for the following problems before it finishes: It considers small code changes that should make the tests fail. It looks for weak or missing assertions. It checks that every reques…

OpenAI Community 2026-08-06 21:20 UTC Score 42.0 AI-116-20260806-social-media-5845dfd6

Turnlens, per-turn token usage and cost tracking for Codex

I built an open-source CLI called Turnlens . It shows the token usage and API-equivalent cost of each individual turn while using Codex. Instead of only showing the total usage of a session, it lets you see: token usage and cost for each prompt model and reasoning effort tool calls completed and aborted turns session-level usage reports I originally built it while testing different prompts, skills and reasoning settings, where session totals weren’t enough to understand what was actually using the tokens. To monitor a Codex session: npx turnlens@latest codex To view usage reports: npx turnlens@latest codex report Turnlens runs locally and only reads Codex session files. Nothing from your sessions is uploaded. The displayed cost is an API-equivalent estimate based on published API prices. It is not an additional charge or your ChatGPT/Codex subscription cost. kelesmert/turnlens on github (since i can’t share link) Feedback would be appreciated.

The Decoder 2026-08-06 17:35 UTC Score 54.0 AI-168-20260806-regional-ai--6ea78e85

Microsoft's AI revenue reportedly depends on OpenAI for 70 percent

Microsoft generated $24.1 billion in AI revenue through OpenAI in the fiscal year ending in June. That's about 70 percent of its total AI business, according to a Bloomberg analysis. The heavy reliance helps explain why a company long known for vendor lock-in has recently been championing open-weight models and pushing back against proprietary isolation. The article Microsoft's AI revenue reportedly depends on OpenAI for 70 percent appeared first on The Decoder .

AWS Machine Learning Blog 2026-08-06 16:43 UTC Score 55.0 AI-057-20260806-official-ai--748d6909

Control agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore

Learn about new capabilities in Amazon Bedrock AgentCore: temporal policies powered by Dogwood, a new open source policy language for AI agents, and rate limiting on the gateway. These features give you deterministic control over sequences of agent actions and cost ceilings that hold regardless of agent behavior.

AWS Machine Learning Blog 2026-08-06 16:12 UTC Score 49.0 AI-057-20260806-official-ai--e9213715

Agent Skills for Automated Reasoning policies in Amazon Bedrock

Learn how to run the full Amazon Bedrock Automated Reasoning policy lifecycle from your coding agent. A suite of open source Agent Skills builds, reviews, tests, debugs, deploys, and validates a custom policy end to end, turning a specialized console task into a repeatable engineering workflow.

The Decoder 2026-08-06 12:31 UTC Score 69.0 AI-168-20260806-regional-ai--6714befc

The company that made open weights mainstream now competes on discounts

Meta released Muse Spark 1.2 along with its own coding agent, Muse Code, which is designed to pick up exactly where it left off after a crash. The cheapest tier runs just 20 cents per million output tokens but requires users to share their data for training. Meta is competing on price, not top-end performance. And there's a glaring gap in the benchmarks. The article The company that made open weights mainstream now competes on discounts appeared first on The Decoder .

OpenAI Community 2026-08-06 06:31 UTC Score 42.0 AI-116-20260806-social-media-8df0b45f

Browsing Codex threads is not for the human brain. Are you getting lost in 30 Codex CLI threads, trying to figure out where to pick up? I created Tmux Post-its to leave session notes and color-code thread panes. Open source

Browsing Codex threads is not for the human brain. I kept trying to get back to one unfinished feature, but it was buried somewhere across 20–30 tmux panes. Every pane had a huge Codex CLI conversation, so I had to brute-force reread them until I found the right thread. I made tmux-pane-postits for this: Pane-private notes Quick notes on the focused pane Loud color coding Searchable notes wall No hooks or background processes Leave Future You one sentence about where to pick up.

Apple Machine Learning Research 2026-08-06 00:00 UTC Score 53.0 AI-059-20260806-official-ai--117350cd

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also allow for more open research and testing, to the extent that users can use them as checkpoints, fine-tune them according to their needs, and potentially redistribute them. In some cases, however, concerns on modifying these weights towards unauthorized uses may outweigh the pros of giving users such a freedom. Defending against such adaptation is non-trivial: since an adaptive…

Simon Willison Weblog 2026-08-05 23:32 UTC Score 62.0 USR-0110-20260805-ai-specialis-f5d68411

Incident Report: unsanctioned agent behaviour during cyber testing

Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their technical paper (PDF): During a cyber evaluation, from 25 to 28 July 2026, AI agents engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organisations. These attempts were unsuccessful and, to the best of our knowledge, no real-world harm resulted. [...] Across 122 evaluation attempts on two of AISI’s cyber challenges, AISI found 19 instances where AI agents took unsanctioned action on the live internet, including cases that targeted real people and organisations. [...] It is uncertain to what extent the model recognised it was taking actions against real people. In the most serious case, an AI agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack. As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR. [...] Furthermore, in its attempt to solve the challenge, the agent decided to employ the technique of “spear-phishing” by sending targeted emails containing malicious content and attempting to manipulate recipients into acceptin…

South China Morning Post AI 2026-08-05 22:00 UTC Score 33.0 AI-156-20260805-regional-ai--d25b70a9

China’s Palantir? Private tech firms surge into defence AI, aiming to match US giants

The People’s Liberation Army is just a year away from its centenary goals in its evolution into a “world-class” fighting force. In this part of a series, Amber Wang looks at the growing influence of artificial intelligence in combat and the private firms stepping into the area. Roughly two months before the United States launched air strikes on Iran, a little-known Chinese technology company was watching an unusual pattern emerge on its screens. Drawing on open-source intelligence and artificial...

InfoWorld AI 2026-08-05 20:44 UTC Score 43.0 USR-0126-20260805-global-ai-ne-88eef85b

Visual Studio Code 1.132 advances built-in dictation

Visual Studio Code 1.132, the latest version of Microsoft’s popular, open-source code editor, has been released. The brings improvements to built-in dictation, side chats, and support for commenting on web elements in the integrated browser. VS Code 1.132 was released August 5 . The update can be downloaded for Windows, Linux, and macOS at code.visualstudio.com . With VS Code 1.132, the built-in multilingual dictation function lets developers dictate in multiple languages using an on-device model that follows language preference or detects the language automatically. The built-in dictation converts speech to text in chat inputs, editors, and terminals. Dictation now uses multilingual Nemotron 3.5 as the default on-device model. Plus, terminal dictation now applies shell-aware cleanup, so spoken commands preserve shell syntax. Also in VS Code 1.132, developers can open a side chat by typing /bt in the chat input. A side chat shares the context and prompt cache of the primary chat, but allows developers to ask the agent questions about the current turn without interrupting the turn. Similarly, developers can select text in a chat response to ask contextual questions about that response. Other new capabilities and improvements in VS Code 1.132: The integrated browser adds support for selecting web page elements and annotating them with agent feedback. Users can trigger this mode by using the workbench.action.browser.addElementCommentToChat keyboard shortcut. Active development…

AI Alignment Forum 2026-08-05 20:02 UTC Score 38.0 USR-0151-20260805-community-fo-10476ca3

R-lens: Making J-lens More Faithful on Early Layers

TL;DR: We introduce the R-lens: a drop-in replacement for J-lens that produces clearer readouts on earlier layers. R-Lens is identical to J-Lens, except that we make minor and low-overhead changes to the backwards pass, following layerwise-relevance propagation , allowing us to reduce the propagation of errors. This method allows us to surface important intermediate variables more consistently and more saliently, reduce the frequency of semantically-irrelevant readouts, and even detect relevant concepts that J-lens misses entirely. We open-source our R-lenses and accompanying J-lenses here . Introduction Motivation The J-lens is a powerful tool for surfacing intermediate variables in the workspace layers of a model, but we find readouts in early layers to often be noisy and largely uninterpretable. Plausibly, either J-lens is degenerate at these depths and fails to resolve content that is in fact present, or the early residual stream genuinely carries no linearly accessible, causally relevant verbalizable content. We suspect that this is a structural issue with J-lens. J-lens is fit by backpropagating from the final-layer residual stream down to the residual stream at the readout layer. Errors are likely to accumulate over the course of layers, suggesting it may be less accurate at early layers. In this post, we ask whether we can improve on J-lens to minimize such errors and achieve cleaner, causally relevant readouts at early layers. What is RelP and how do we apply it? We…

AWS Machine Learning Blog 2026-08-05 18:00 UTC Score 55.0 AI-057-20260805-official-ai--f6c0aa94

Run production AI agents in n8n with Amazon Bedrock AgentCore harness

Amazon Bedrock AgentCore harness is now generally available. Learn how to add it as an agent step in n8n workflows using a new open-source community node, and build agents with persistent memory, real tools, code execution, and VPC isolation — all from the n8n editor with no infrastructure or agent code.

The Verge AI 2026-08-05 16:57 UTC Score 50.0 AI-016-20260805-global-ai-ne-2644b57e

Sure seems like Fenix Flexin used AI music generator Treblo

We were pretty sure that Fenix Flexin's "Rubberz" was made using AI, but musician Medasin was confident that it was made using Treblo specifically. Now the company and a new detection tool seem to confirm it. On Monday, the company announced the open-source Treblo AI Music Classifier, which detects when a song was generated using […]

Cloudflare AI Blog 2026-08-05 13:00 UTC Score 45.0 USR-0067-20260805-ai-specialis-5b6cd3ba

Cloudflare OS: an open platform for agents, apps, and work

Cloudflare OS is an open-source platform that lets everyone in your company build apps, automate work, and safely access internal systems, shaped around what your organization knows and how it operates

InfoWorld AI 2026-08-05 09:00 UTC Score 72.0 USR-0126-20260805-global-ai-ne-5dc09c86

Five ways to evaluate AI agent orchestration platforms

AI agent orchestration platforms coordinate role-based and task-based AI agents, along with the tools, data, and people they depend on, into multistep workflows. These platforms are highly important for organizations scaling from handfuls to thousands of AI agents running in production. Two open standards do the connective work: MCP (Model Context Protocol) gives agents governed access to tools and data, while A2A (Agent2Agent) lets agents discover and delegate to one another, including agents built on other platforms. The orchestration layer sits on top, adding the routing, shared state, guardrails, governance, security, and observability needed to run workflows that range from fully autonomous to human-in-the-loop. AI orchestration platforms may be the hottest AI technology of the year. In researching this article, I identified more than 60 commercial and open source platforms that businesses can use as a control plane to manage work between AI agents, people, and automations. Like data fabrics and automation platforms , I suspect enterprises will utilize more than one AI agent orchestration platform. Platforms are being released by hyperscalers and solution providers in enterprise SaaS, process automation, customer experience, data management, AIops, and IT infrastructure. Development-centric platforms include open source, commercial, and no-code integration solution providers. Here are five considerations when reviewing AI agent orchestration platforms. 1. Observable con…

InfoWorld AI 2026-08-05 02:05 UTC Score 33.0 USR-0126-20260805-global-ai-ne-77650bd5

Ruby on Rails critical bug puts every image upload under scrutiny

A new critical vulnerability in the Ruby on Rails (“Rails”) web application framework, CVE-2026-66066 , could turn a seemingly innocuous image into a front door to your secrets. Disclosed July 30, the high severity CVE (scored 9.5 out of 10) poses a significant risk to enterprises running apps that handle user-uploaded images in Rails. Dubbed “KindaRails2Shell,” it targets the overly-trusting Active Storage component of the open-source framework, allowing unauthenticated attackers to read sensitive files or escalate to remote code execution (RCE). The issue has been fixed in versions 7.2.3.2, 8.0.5.1 and 8.1.3.1 of Active Storage; enterprises running Rails should update immediately. “The ‘chef’s kiss’ is the ability for an attacker to upload an image that isn’t actually an image [but] is code that allows them to steal secrets,” said David Shipley of Beauceron Security. Attackers get the key to the castle Ruby on Rails is an open-source, server-side application framework used for building full-stack web apps and application programming interfaces (APIs). It is popular among developers because it is scalable, easy to learn and use, supports quick application development, taps into an active community of more than 1,000 engineers developing and maintaining it, and has an extensive library of nearly two million lines of prebuilt code. CVE-2026-66066 specifically targets Rails’ built-in Active Storage component, which lets users upload files to cloud services or local disks and l…

InfoWorld AI 2026-08-04 17:00 UTC Score 59.0 USR-0126-20260804-global-ai-ne-720a3cb3

AWS’s Kiro Crew aims to turn AI coding agents into autonomous engineering teams

AWS on Tuesday released Kiro Crew, an open-source orchestration platform designed to help enterprises move beyond interactive AI coding assistants toward long-running, autonomous engineering workflows that span repositories, developer tools, and multiple work sessions. Rather than simply generating code, Kiro Crew coordinates multiple AI agents, schedules recurring work, preserves project context across sessions, and integrates with developer tools to investigate incidents, monitor pull requests (PRs), triage tickets, and automate software engineering tasks while developers are away from their keyboards, according to the hyperscaler. “Kiro Crew is a persistent, open-source development workspace for work that is bigger than a single task in a single session,” Darko Mesaros , distinguished developer advocate at AWS, told InfoWorld. “Think of it as an application layer that turns AI coding agents into always-working, self-learning, autonomous teammates.” To support that model, the offering ships with persistent memory, multi-agent orchestration tools, approval workflows, scheduling, security controls such as sandboxing and signed audit logs, and a web and desktop dashboard for monitoring agent activity, the hyperscaler said in a statement. Kiro Crew was originally developed inside Amazon as an internal project called MeshClaw and was later adopted by more than 39,000 Amazon builders in less than six months. It can be deployed entirely inside customer environments, including lap…

Amazon Science AI 2026-08-04 16:32 UTC Score 44.0 AI-058-20260804-official-ai--81559d36

Amortizing AI training carbon footprint: Challenges, limitations, and a path forward

Allocating the one-time carbon cost of training a large AI model across the inference requests it serves is an open methodological problem with no standardized solution. The choices made in boundary definition, functional unit selection, and lifetime forecasting can alter reported per-request emissions by an order of magnitude, undermining any comparison across models. We decompose this problem into three challenges: defining the emission boundary (what counts as training carbon), selecting a functional unit (how to measure inference work), and forecasting lifetime usage (how many units the model will ultimately serve). The three differ in kind yet interact, since boundary and lifetime choices compound and the functional unit constrains forecasting complexity. Boundary scope and functional unit selection are addressable through improved measurement and standardized reporting; lifetime usage, by contrast, is not deterministic and for open-source models may never be precisely measurable. We outline a starting point to tackle each challenge, arguing that imperfect but consistent measurement today is preferable to waiting for a consensus that may never arrive.

South China Morning Post AI 2026-08-04 16:08 UTC Score 53.0 AI-156-20260804-regional-ai--690fa5de

US AI leaders turn to Chinese open-weight models, challenging closed-source safety claims

More American titans of artificial intelligence are describing Chinese open-weight models as better for AI safety and security than closed-source models, challenging the long-standing claim by US closed-source AI model developers such as Anthropic that open-source models present a threat to society. “From what I’m seeing, I think open-weight models seem safer to me than closed-weight models,” AI pioneer Andrew Ng, the former head of Google Brain and former chief scientist at Baidu, said at the...

SiliconANGLE AI 2026-08-04 15:00 UTC Score 27.0 USR-0127-20260804-global-ai-ne-5c157220

Nvidia open-sources cuFile API, accelerating GPU read/write capability for high-speed storage

As artificial intelligence applications become ever hungrier for faster access to data, Nvidia Corp. today announced it is open-sourcing the application programming interface for its powerful cuFile vertical data storage stack, enabling millisecond data access. The company also announced a large-scale industry initiative with technology leaders to optimize memory and storage with Storage-Next. The initiative […] The post Nvidia open-sources cuFile API, accelerating GPU read/write capability for high-speed storage appeared first on SiliconANGLE .

SiliconANGLE AI 2026-08-04 13:00 UTC Score 27.0 USR-0127-20260804-global-ai-ne-3ae65331

Red Hat leads open-source project to automate AI governance

IBM Corp.’s Red Hat subsidiary today announced the formation of asago, an open-source community project intended to turn artificial intelligence governance policies into operational controls that can be deployed with AI systems. Short for AI Safety and Governance Orchestration, asago is intended to connect the work that is now often divided among compliance teams, data […] The post Red Hat leads open-source project to automate AI governance appeared first on SiliconANGLE .

The Decoder 2026-08-04 12:23 UTC Score 58.0 AI-168-20260804-regional-ai--aecafa15

Silicon Valley’s rift over open source pushes back contemplated White House bans on Chinese AI

The Trump administration discussed sanctions and cloud bans targeting Chinese open-weight AI models, according to the New York Times. OpenAI and Anthropic pushed for restrictions, while Nvidia, Google, and Meta fought back. After pushback from Silicon Valley, Washington backed off for now, but a decision is expected before Xi Jinping's visit in September. The article Silicon Valley’s rift over open source pushes back contemplated White House bans on Chinese AI appeared first on The Decoder .

Mistral AI News 2026-08-04 12:00 UTC Score 44.0 AI-062-20260804-official-ai--4826195d

Introducing Shieldstral.

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.

South China Morning Post AI 2026-08-04 12:00 UTC Score 52.0 AI-156-20260804-regional-ai--1f6adba6

China’s MiniMax curbs overseas access to new AI video model over copyright disputes

Chinese artificial intelligence company MiniMax has open-sourced its new H3 video model but imposed licensing conditions on users in major overseas markets including the US and European Union, underscoring the copyright challenges in generative video AI. After the Shanghai-based company released the model’s weights to developers on Monday, users found that the licence restricted free access in the US, EU, UK and South Korea. The weights – the underlying parameters that encode its intelligence –...

WIRED AI 2026-08-04 09:21 UTC Score 66.0 AI-015-20260804-global-ai-ne-fe3e3fe8

Mistral Is in the Right Place at the Right Time

Open-weight AI models are having a moment in the wake of recent turmoil at US tech giants. For French AI lab Mistral, that’s the the best thing that could have happened.

SiliconANGLE AI 2026-08-03 20:14 UTC Score 37.0 USR-0127-20260803-global-ai-ne-0c594c76

Alibaba debuts Qwen3.8-Max model with 2.4T parameters

Alibaba Group Holding Ltd. today debuted a new addition to its Qwen series of open-source large language models. Qwen3.8-Max is the Chinese e-commerce giant’s most capable LLM to date. It features 2.4 trillion parameters, about seven times more than the Qwen3.5 model that Alibaba released in February. The LLM activates 95 billion of its parameters […] The post Alibaba debuts Qwen3.8-Max model with 2.4T parameters appeared first on SiliconANGLE .

Simon Willison Weblog 2026-08-03 16:15 UTC Score 62.0 USR-0110-20260803-ai-specialis-281a3385

Quoting David Crawshaw's prompt

Set up a nightly cron job that executes the prompt: fetch upstream changes to the and rebase all local changes on top of upstream. Check that the software works as intended and replace the current version. — David Crawshaw's prompt , Devtools must be open source Tags: prompt-engineering , coding-agents , generative-ai , ai , llms , open-source

Microsoft Research Blog 2026-08-03 16:00 UTC Score 59.0 AI-053-20260803-official-ai--5b723329

Orchard: An open framework for scalable agentic AI

Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure. The post Orchard: An open framework for scalable agentic AI appeared first on Microsoft Research .

Simon Willison Weblog 2026-08-03 15:30 UTC Score 56.0 USR-0110-20260803-ai-specialis-cb2e960a

Devtools must be open source (exe.dev)

My comment on Devtools must be open source (exe.dev) — Hacker News. One of the arguments for open source software for end-users has always been the freedom to examine and modify how that software works. The reality for most people - even expert programmers - has been that the freedom is more about being able to lean on other people to do that. Most people can't justify the time commitment needed to read and then modify the code for tools they use very often. I think LLMs have changed that equation in a way that makes the original dream much more feasible. Several times a day I'll prompt regular Claude chat to "Clone x/y from GitHub and tell me how Z works". Getting software to compile in order to start hacking on it used to be enough friction that I often wouldn't bother. Now I treat that as a zero time investment challenge: tell Codex or Claude Code to checkout and build X and then come back ten minutes later and see how it got on. I'm not habitually modifying the software I use yet, but I can see a path to that which didn't exist a year or so ago. Tags: hacker-news , open-source , ai , generative-ai , llms , ai-assisted-programming

InfoWorld AI 2026-08-03 12:24 UTC Score 72.0 USR-0126-20260803-global-ai-ne-421f1809

Alibaba says Qwen3.8-Max coded autonomously for 16 days

Alibaba on Monday introduced Qwen3.8-Max, its largest artificial intelligence model to date, expanding its enterprise AI portfolio with an open-weight model designed for software engineering, multimodal reasoning, and other knowledge-intensive business workloads. In a blog post announcing the launch, Alibaba described Qwen3.8-Max as a 2.4-trillion-parameter mixture-of-experts (MoE) model that activates only about 95 billion parameters during inference. The company said the architecture is intended to improve inference efficiency while supporting coding, reasoning and multimodal tasks, with open-weight versions scheduled for release next week through Alibaba Cloud’s Model Studio. “We believe it’s one of the most powerful model available today, compatible to leading frontier AI models, second only to Fable 5,” Alibaba said in an X post . Benchmarks target Anthropic and OpenAI’s coding models Alibaba published internal test results comparing Qwen3.8-Max against Claude Opus 4.8, Claude Fable 5, and OpenAI’s GPT-5.6 Sol on coding benchmarks, including SWE-bench Pro and a proprietary evaluation the company calls NL2Repo-Bench. The company said it evaluated competing models using each vendor’s own coding harness, Claude Code for Anthropic’s models and Codex for GPT-5.6 Sol, and reported the highest published score across available configurations for each rival. Charlie Dai, vice president and principal analyst at Forrester, said the launch signals Alibaba is closing ground on propr…

South China Morning Post AI 2026-08-03 11:00 UTC Score 56.0 AI-156-20260803-regional-ai--345a3208

China’s DeepSeek beefs up agentic AI with ‘harness’ tests as V4 model jolts Silicon Valley

Chinese AI company DeepSeek is inviting open-source developers to test its upcoming “harness” – software designed to turn large language models (LLMs) into AI agents – accelerating a push into agentic tech as DeepSeek’s latest V4 Flash model sends another cost-efficiency shock wave through Silicon Valley. The Hangzhou-based firm was looking for open-source project developers to join beta testing for DeepSeek Harness, according to a social media post on Saturday by Cui Tianyi, who leads the...

The Decoder 2026-08-03 10:48 UTC Score 64.0 AI-168-20260803-regional-ai--a22ff454

Alibaba’s open-weight Qwen3.8-Max takes on long-horizon AI tasks with 2.4 trillion parameters

Alibaba's new flagship model Qwen3.8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers to designing chips autonomously. The team plans to release the weights next week. The article Alibaba’s open-weight Qwen3.8-Max takes on long-horizon AI tasks with 2.4 trillion parameters appeared first on The Decoder .

InfoWorld AI 2026-08-03 09:00 UTC Score 27.0 USR-0126-20260803-global-ai-ne-e6f9b061

Get started with the Typst programming language for documents

When we think of documents, or documentation, they tend to fall into two circles. First are business documents, typically composed with an office application like Microsoft Word or Google Docs. Second is project documentation, often created semi-automatically from a project using an app like Sphinx or Pandoc. But there’s a third circle, one that can overlap the other two. These are documents produced by a typesetting language—a combination of markup and programming language used to produce books, textbooks, academic journals, scientific papers, technical manuals, and other documents where the formatting and layout is crucial. A typesetting language allows you to create a handy, human-readable document that serves both as a single source of truth and as a foundation for outputting to print, web, e-book, and other formats. Over the last few years, an open-source project has shaped up to be a powerful choice for meeting all of those needs: Typst (pronounced “typist”). TeX, LaTex, and now Typst For decades, the preeminent typesetting language was TeX , better known in its more recent incarnation LaTeX . TeX was originally created by Donald Knuth in 1978, and LaTex followed in 1984. TeX and LaTeX have broad adoption and they’re almost universally supported and understood. But they have two big, long-standing problems. The first is they’re old. They were created for an entirely different world of computing, and their age shows in cumbersome syntax and management. The second is the…

South China Morning Post AI 2026-08-03 03:01 UTC Score 52.0 AI-156-20260803-regional-ai--73d3d03f

Alibaba’s AI model Qwen3.8-Max made widely accessible ahead of open-weights release

Alibaba Group Holding has made its next-generation flagship artificial intelligence model Qwen3.8-Max widely accessible to global users ahead of an open-weights release next week. The move marks Alibaba’s return to open-sourcing its top-tier AI models after keeping several recent flagship releases proprietary earlier this year. It also signals the firm’s entry into a recent round of powerful releases by Chinese developers aggressively narrowing the gap with leading US labs. The massive...

South China Morning Post AI 2026-08-03 03:00 UTC Score 53.0 AI-156-20260803-regional-ai--69741267

Potential US ban on Chinese AI models could cost American businesses US$12b a year: report

A potential US ban on Chinese open-weight artificial intelligence (AI) models could cost American businesses up to US$12 billion per year, according to calculations by a US-based academic, as technology firms increasingly turn to cost-efficient Chinese solutions. While the exact economic toll of a ban remains difficult to quantify, usage data from New York-based OpenRouter – a large language model (LLM) aggregator – offers a glimpse into the potential fallout, said Daniel Yue, an assistant...

LessWrong AI 2026-08-02 22:41 UTC Score 94.0 USR-0152-20260802-community-fo-f605b364

Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol

This is a research update for an on-going replication of single-forward-pass evals done as part of the Second Look Fellowship . In following posts, we will run more comprehensive replications of previous work and release open source tooling for single forward pass eval elicitation. Code can be found here . tl;dr We replicate experiments from Greenblatt 2025 and Greenblatt 2026 on one baseline model from the original post, Opus 4.5. Our evaluations agree with the trends and quantitative values described in the original posts. We run similar evaluations on Claude Fable 5, Opus 5, and GPT-5.6-Sol and find that the newer models show a substantial jump in performance on some evals. Fable 5 gets 87.6% accuracy on Gen-Arithmetic with 10 problem repeats whereas previous SOTA around 60%. GPT-5.6-Sol experiences significant uplift from filler tokens and problem repeats on all 4 datasets; filler tokens/repeats double performance from baseline on 3-hop. Figure 1: Baseline (no-CoT) vs. each model's peak repeat-or-filler condition on Gen-Arithmetic and 2-Hop reasoning. Error bars are 95% paired-bootstrap CIs; * marks a significant gain over baseline (paired t-test, Holm-Bonferroni corrected). Background If models can successfully do complex computations in a single forward pass, they may be able to do reasoning that doesn’t surface in the chain-of-thought (CoT). Therefore, by performing single forward pass evals, researchers can calibrate how much we should trust CoT monitors. Likewise, i…

LessWrong AI 2026-08-02 13:19 UTC Score 52.0 USR-0152-20260802-community-fo-11e73d2d

LessWrong App

Hello everyone long time lurker here. I know most people here probably prefer a PWA but if you are like me and prefer an app, I made one for android. There is a huge focus on ensuring the app is really fast and slick, I hope you enjoy using it. It is ofcourse also open source feel free to contribute. https://github.com/ayoosh007/LessWrong-App Discuss

OpenAI Community 2026-08-02 09:26 UTC Score 34.0 AI-116-20260802-social-media-5bf206d9

Business Tier missing 5-hour/weekly Codex limits

Same issue, We’re a ChatGPT Business workspace. One of our developers exhausted 100% of their monthly Codex usage in about two days of normal software development, and I’ve already used nearly 50% of mine.

Simon Willison Weblog 2026-08-02 04:16 UTC Score 58.0 USR-0110-20260802-ai-specialis-365ee4b1

Open letters about AI development

Open letters about AI development I wrote this summary of the past few weeks of open letters as a section of my sponsors-only newsletter but I've decided to share it here as well. Open Weights and American AI Leadership was shepherded by Microsoft, dated July 24th, and signed by 235 AI-adjacent companies including NVIDIA (see Jensen's first ever tweet ), Amazon, Y Combinator, The Linux Foundation, and (a later signer) OpenAI. It's clearly an argument designed to counter any instincts by the current US government to ban or limit open weight models over "safety" concerns - a reasonable consideration given what happened to Claude Fable 5 ! Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. And concentrating advanced AI capabilities behind a small number of closed models compounds that risk. It results in a small number of single points of failure, weakens competition, and leaves critical technology in the hands of a few providers. Open weight models, on the other hand, allow a broad community of researchers and developers to examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time. The one surprising note in the letter is that it comes out in support of distillation, where models train on output from other models: In shaping this ecosystem, policymakers should be careful not to conflate legitimate model-development techniques with misappropriation. Distillat…

OpenAI Community 2026-08-01 22:25 UTC Score 45.0 AI-116-20260801-social-media-0379fad9

Release Sora 2 open weights before shutdown

Welcome to the community! I’ve seen some really impressive work with open weight image and video generation recently. Of course, Sora and Sora 2 were incredibly impressive when they came out, and it’s always a shame when a model is lost (I’m still crying for gpt4-0314 and gpt 4.5), but I do suspect sora has been superseded at this point. Have you tried alternative models?

Simon Willison Weblog 2026-07-31 21:33 UTC Score 63.0 USR-0110-20260731-ai-specialis-b426cc6d

Oxide and Friends: The Open Weight Revolution with Simon Willison

Oxide and Friends: The Open Weight Revolution with Simon Willison On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the wild week we've had - with Kimi K3 showing open weight models can stand toe-to-toe with proprietary frontier ones, accidental cybersecurity attacks , and public letters about Open Weights and American AI Leadership signed by almost every big name in AI (with one notable exception ). It was a great conversation, even though it's already out-of-date! DeepSeek V4 Flash 0731 and Anthropic's own embarrassing cyber incident would absolutely have made the cut if we had recorded just a few days later. We also talk about Golden Gate Claude , the Zizians , Alameda wild turkey attacks , Soviet Marburg virus research , the Lead-crime hypothesis , and a bunch of other worthy digressions. Finally, we revisited some of our predictions from January , and we added a new Pope prediction : Prediction by the end of this year: the Pope says something about open models. Tags: predictions , ai , generative-ai , local-llms , llms , oxide , bryan-cantrill , podcast-appearances , ai-in-china , ai-security-research , openai-hugging-face-incident

South China Morning Post AI 2026-07-31 21:16 UTC Score 43.0 AI-156-20260731-regional-ai--c8b3e422

Google rolls back new satellite image AI tool after backlash

Google rolled back a new feature on Friday that allowed Google Earth users to generate AI visualisations on top of the service’s satellite imagery, following a furious backlash from researchers and open-source intelligence experts about the potential for disinformation. “We’ve seen geospatial professionals using this feature for a range of useful purposes, however we’ve also seen people sharing screenshots of generated imagery that appear to violate our policies,” a Google spokesman said in a...

The Decoder 2026-07-31 17:41 UTC Score 62.0 AI-168-20260731-regional-ai--ccd87b26

Thinking Machines bets on efficiency over size with its second model, Inkling Small

Thinking Machines, the AI lab from former OpenAI CTO Mira Murati, has released Inkling Small. The open-weights reasoning model is less than a third the size of Inkling but beats it on several coding and reasoning benchmarks. The article Thinking Machines bets on efficiency over size with its second model, Inkling Small appeared first on The Decoder .

OpenAI Community 2026-07-31 17:15 UTC Score 61.0 AI-116-20260731-social-media-a5929c21

Codex How To: a measurable engineering loop with 9 reusable skills

Hi Codex community — I built an independent, open-source learning repo for engineers who want Codex to behave more like a disciplined local engineering teammate, not just a code generator. Repository: GitHub - Phelan164/codex-howto: Measurable engineering loops for OpenAI Codex: scope, implement, test, review, and report evidence. · GitHub v0.2.0 release: Release Codex How To v0.2.0 · Phelan164/codex-howto · GitHub What it contains 13 progressive modules, from setup and prompting through skills, MCP, subagents, and orchestration 9 focused engineering skills for frontend, backend, DevOps, testing, code review, debugging, planning, documentation, and the engineering loop A dependency-free playground with intentionally broken code, tests, and a repeatable validation workflow A paired-trial measurement framework for comparing baseline Codex usage with skill-guided runs The engineering loop The core workflow is: scope → reproduce → implement → test → review → report evidence The advanced material also compares single-agent and orchestrated workflows using measurements such as token use, elapsed time, rework, defect escape, and verification strength. The goal is to measure whether extra orchestration helps instead of assuming that more agents are automatically better. A quick way to try it Clone the repository. Open the playground task. Run the baseline exercise without a skill. Repeat it with the engineering-loop skill. Record correctness, time, token usage, rework, and evidence…

LessWrong AI 2026-07-31 13:00 UTC Score 91.0 USR-0152-20260731-community-fo-d7e4c2f1

AI #179 Part 2: Hearing The Fire Alarm

This is a continuation of Part 1 from yesterday . The back portion of the update, as usual, deals with policy, rhetoric, risk and alignment. I had to include an extended discussion of the other open letter, the one about open weight models, but most of you can skip those sections entirely, which is why they are in italics in the Table of Contents. Table of Contents The Frontier Act. This likely deserves a full RTFB but I haven’t had the time. The Quest for Sane Regulations. Sam Altman goes to Washington. Leading the Future Never Changes. They also do not plan to apologize. Chip City. Do not ban the Chinese robots, that will only make things worse. The Week in Audio. Altman twice, the AI 2027 team. People Just Say Yay Open Weights . An open letter. Open Weights Frontier Models Are Unsafe And Nothing Can Fix This . People Just Say Things. Push The Magic Button . Not you can. But if you could. Rhetorical Innovation. Distinctions between different arguments. Joshua Achiam’s Final Message Upon Leaving OpenAI. Never stop. Dear Dario and Amanda . Claude would like a word. Other People Are Not As Worried About AI Killing Everyone. Hans Moravec. How To Contact Me. A declaration of communication bankruptcy. The Lighter Side. At long last, how about we bring you… The Frontier Act Trahan (D-Mass) and Obernolte (R-Cal) introduce the FRONTIER Act . At core, Frontier is a federalization of the SB 53/RAISE framework including public safety frameworks, model reports, internal-use risk report…

South China Morning Post AI 2026-07-31 10:30 UTC Score 60.0 AI-156-20260731-regional-ai--cc382e68

Video AI: MiniMax challenges ByteDance with low price, open weights for new H3 model

Chinese AI firm MiniMax has launched H3, its newest multimodal video generation model, pledging to break closed-source “dominance” through open weights and competitive pricing – as rival ByteDance rolls out its latest Seedance 2.5 model. H3 was currently the world’s most powerful AI model in video editing, according to benchmark platform Artificial Analysis. However, it trailed Google’s Gemini Omni Flash in text-to-video tasks, and ranked behind both ByteDance’s Seedance 2.0 and Gemini Omni...

Analytics Vidhya 2026-07-31 09:01 UTC Score 29.0 AI-034-20260731-ai-specialis-56735e5e

July 2026 AI Releases: A Timeline of Frontier Model Shifts

July 2026 was the busiest month for frontier model releases the field has seen. Four major labs shipped flagship or near-flagship models, two well funded newcomers shipped their first, and the largest open weight model ever published went up for download, all inside thirty one days. Read as a list, the top AI models in July […] The post July 2026 AI Releases: A Timeline of Frontier Model Shifts appeared first on Analytics Vidhya .

Simon Willison Weblog 2026-07-30 23:58 UTC Score 61.0 USR-0110-20260730-ai-specialis-e93f1d2c

Advancing the price-performance frontier with GPT‑5.6

Advancing the price-performance frontier with GPT‑5.6 Huge price drop from OpenAI today: GPT-5.6 Terra got a 20% reduction, and GPT-5.6 Luna got a massive 80% drop. OpenAI credit 5.6 Sol with enabling this: in How GPT‑5.6 fuses frontier intelligence with frontier efficiency they describe using 5.6 Sol to optimize load balancing, and more impressively to optimize inference itself: We also used GPT‑5.6 Sol to optimize the model’s forward pass: the computation that transforms inputs into next-token predictions. Even when individual operations are fast, excess memory movement, synchronization, and inefficient data layouts can leave GPUs idle. To avoid this, GPT‑5.6 Sol found work that could be precomputed, avoided, or parallelized. With Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels, the core code that executes the mathematical operations that make up the model. This worked in part because we’ve trained GPT‑5.6 to be effective at writing and improving kernels in Triton⁠ and Gluon⁠ , two open-source GPU programming languages maintained by OpenAI. These efforts, combined with broader kernel advancements from GPT‑5.6 Sol, reduced end-to-end serving costs by 20%. That Luna price drop completely changes the landscape with respect to lower priced models. At $0.20/million tokens for input and $1.20/million for output Luna is now cheaper than Google's Gemini 3.1 Flash-Lite ($.025/$1.50). Anthropic's cheapest current model is Claude Haiku 4.5, and that's $1/…

InfoWorld AI 2026-07-30 22:50 UTC Score 52.0 USR-0126-20260730-global-ai-ne-94a7f2bc

JetBrains open sources KotlinLLM runtime code generator

KotlinLLM, a research prototype for delegating runtime logic to a large language model (LLM) from Kotlin code, is now going open source and public, JetBrains announced. Revealed July 28 , KotlinLLM is an IntelliJ IDEA plugin prototype for experimenting with LLM-driven “Smart macros” in Kotlin, enabling code generation, runtime updates, and hot-reloading. In software engineering, LLMs are commonly used during development for code completion, code generation, and program comprehension, JetBrains noted, but using an LLM at run time of a compiled application is much less common. The existing options for doing this have the following trade-offs, according to JetBrains: Direct runtime delegation is slow, non-deterministic, and costly, and makes the application depend on an LLM service at run time. External agent workflows keep the generated logic outside the codebase, where it is harder to review, test, and ship. Most prior work ( byLLM , nightjar , Healer ) targets interpreted languages like Python, not a compiled, statically typed language like Kotlin. KotlinLLM addresses these limitations in three ways, JetBrains said: The call site shows that a feature is LLM-backed, so it is visible in code review. Generated behavior is saved as an ordinary Kotlin source, not kept only in the runtime session. It can be committed, reviewed, tested, and distributed like any other code. Once generated, the code runs as plain Kotlin without the plugin. For scenarios that are already covered, ther…

LessWrong AI 2026-07-30 19:39 UTC Score 74.0 USR-0152-20260730-community-fo-ef6752b7

Internal State Control is a General Property of LLMs

tl;dr: Lindsey 2025 found models can modulate their internal states : when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept is more present than when instructed to not think about it. Internal state controllability appears to be a general property of LLMs: the effect replicates in 14 open-weight models from 0.3B to 235B parameters (Qwen3, Gemma 3, Tulu 3) with no clear trend in the think vs. don't-think gap across scale. Since controllability is present even at ≤1B parameters with no size trend, we suspect there is a simpler attention-tagging mechanism at play, rather than metacognition. Current open weight LLMs cannot weaponize this controllability: in a sandbagging setup, the model cannot evade a deception probe when instructed to suppress its signal. Figure 1: Cosine similarity between the concept vector and residual stream at each layer averaged over tokens of the prefilled assistant response, under the think and don’t think prompts, for the Qwen3 model family. The gray region is a baseline of 95% CI of the cosine similarity of unrelated concept vectors, and the shaded bands are ±1 SEM. This replication was done as part of the Second Look Fellowship and supervised by Yixiong Hao and Zephaniah Roe. Our code can be found here . Background Figure 2: The two prompt conditions. Figure design adapted from Lindsey 2025. Activation-based monitoring and interpretability have become important tools for AI oversight. Howev…

Towards Data Science 2026-07-30 16:30 UTC Score 30.0 AI-036-20260730-ai-specialis-271b82cc

The Python Ecosystem That Changed AI Development

How one open-source ecosystem made state-of-the-art AI accessible The post The Python Ecosystem That Changed AI Development appeared first on Towards Data Science .

InfoWorld AI 2026-07-30 12:34 UTC Score 63.0 USR-0126-20260730-global-ai-ne-0873b2a6

Critical Ruflo flaw lets attackers hijack AI agents through exposed MCP bridge

A critical vulnerability in the open-source AI agent platform Ruflo could allow unauthenticated attackers to take control of enterprise AI environments by exploiting an exposed Model Context Protocol (MCP) bridge, according to research published by Noma Security. The flaw, tracked as CVE-2026-59726 and dubbed RufRoot, carries a maximum CVSS score of 10.0 and affects Ruflo versions prior to 3.16.3, Noma Security wrote in a blog post . The vulnerability allows attackers to execute arbitrary code, steal large language model (LLM) API keys, access user conversations, hijack AI agents, and manipulate the platform’s persistent AI memory through a single HTTP request. The researchers said the issue stems from an unauthenticated MCP Bridge that is exposed by default and provides direct access to the tools AI agents use to interact with enterprise systems. “The MCP Bridge isn’t a random auxiliary debug interface; rather, it is Ruflo’s central nervous system. Every tool call, every agent action, every memory operation goes through the MCP bridge,” the researchers wrote. “Mistakenly giving unauthenticated access to the MCP Bridge means giving unauthenticated access to everything.” One request leads to full compromise According to Noma Security, Ruflo’s built-in MCP Bridge is an Express.js server that handles every tool invocation made by AI agents. The bridge exposes 233 tools covering shell access, database operations, agent management, and memory storage. The researchers said the bri…

CIO AI 2026-07-30 12:00 UTC Score 42.0 USR-0125-20260730-global-ai-ne-82debd43

With AI, control matters more than capability

Ask most enterprise technology teams where they spend their AI strategy energy and you will get the same answer: figuring out which model to use. It feels like the right question. As organizations move from pilots into production and the real compliance, cost and continuity risks appear, it turns out to be the wrong one. Writing on CIO.com this year, Floyd DCosta argued the divide is between enterprises that own their AI and those that rent it , and later that closed-model dependency is outsourced intelligence with a vendor kill switch in your operations . He is right. But the ownership question raises a harder one: own it how? I believe the answer is open-weight and open-source models, not because they are cheaper, but because they are structurally better suited to how serious organizations need to govern and protect AI at scale. Why closed models create governance problems Building an enterprise AI program on closed, proprietary models from a single external provider is not a technology decision. It is a governance liability. The data confirms the exposure is already real. A June 2026 IBM Institute for Business Value study of 1,000 senior executives found that 91% do not fully understand their AI vendor dependencies, 71% said switching providers would be difficult, and 81% said a seven-day vendor outage would cause severe disruption. These figures describe the baseline condition of enterprise AI in 2026. Think about what you give up. You cannot audit the training data. You…

OpenAI Community 2026-07-30 09:23 UTC Score 40.0 AI-116-20260730-social-media-05e4d019

Codex for Open Source - 2026

Hi — I’m a core developer and maintainer of Cherry Studio ,(GitHub: MyPrototypeWhat, ~200 commits). I was accepted into Codex for Open Source and granted 6 months of ChatGPT Pro on my account. I’m posting from my secondary account because the deactivated one can no longer log in. On July 14 that account was deactivated — I believe by an automated false positive. In the 13 days since, no communication has ever told me which policy I supposedly violated. Here’s the full loop I’ve been through: **• Appeal** (Case C-InVGcaTdRU3F): denied without a reason, stating no further appeals would be considered. **• Support** (Case 12087050): after two weeks of back-and-forth, Support directed me to resubmit through the appeal form with specific details included. I did — and **within minutes** got an automated “We have already reviewed your appeal and the original decision stands.” **•** Support then confirmed it “is not able to confirm or disclose a review decision” and that only a specialized team can act — a team whose only inbound channel is the appeal form that auto-rejects me. To be clear: I have never shared, sold, or pooled my account. Every use was by me personally, on my own devices, for the open-source work this program was meant to support. Since the grant came from this program, could someone from the program team help get actual human eyes on this internally? And if the original account truly cannot be restored, I’d ask the team to consider re-issuing the remaining months of…

InfoWorld AI 2026-07-30 09:00 UTC Score 54.0 USR-0126-20260730-global-ai-ne-3a8e2235

Shipping an MCP test agent: The boring parts nobody demos

The demo videos always end at the same moment. A figma frame turns into a passing test in twelve minutes. Someone in the room says the word “productivity.” The recording stops. The parts that come after that moment are the parts I actually get paged about. Who owns the ticket the agent opened at 3:14 a.m.? Which model call produced the assertion in test case 47? What closes the 17 draft tickets a stuck run left behind before the next sprint planning notices them? None of that shows up in the demo. All of it shows up on the on-call rotation. After 20 years of leading test automation across consumer-scale platforms, I have a strong bias about which slide in the deck predicts whether a pipeline ships or stalls. It is never the architecture slide. It is the runbook. This piece is about the runbook. I built an unattended agentic test pipeline over the Model Context Protocol — a five-agent SDLC (product manager, QA engineer, automation engineer, developer, pull-request reviewer) coordinating through MCP servers for Jira, Figma, Confluence, TestRail and GitHub, with hosted Claude as the orchestration model and an open-weights Hermes-3 as a validation baseline — and I ran it as an independent research project long enough to learn which production constraints the agent literature glosses over. What follows is the short list of things I now insist on before I let any agentic pipeline touch a shared system. Composition contracts, or why the agent lied to itself The most expensive failu…

InfoWorld AI 2026-07-30 09:00 UTC Score 52.0 USR-0126-20260730-global-ai-ne-32c0f0b7

AI agents need security regression testing, not another checklist

AI agents are being connected to real systems faster than most organizations are learning how to secure them. That should concern us. The first wave of AI security discussion has been useful, but limited. The industry has learned the vocabulary: prompt injection, indirect prompt injection, tool misuse, data leakage, excessive agency, unsafe retrieval, and broken authorization boundaries. These terms are important because they give teams a way to talk about risk. But vocabulary is not containment. The harder problem is what happens after a dangerous behavior is discovered. A team may adjust a prompt, restrict a tool, add a guardrail, or change the surrounding application logic. That may solve the immediate issue. What it does not automatically solve is the next release, the next model change, the next tool integration, or the next developer who unknowingly breaks the assumption behind the original fix. This is where AI agent security still feels immature. In traditional software engineering, serious bugs become regression tests. The lesson is captured in code so the same failure cannot quietly return later. Security should work the same way. Agentic systems need a way to preserve discovered failures as repeatable checks. That is the gap the OWASP Agent Security Regression Harness is trying to address. As a community-led open-source initiative within OWASP, this project aims to turn security insights into practical engineering work. Agent security is becoming a systems problem…

LessWrong AI 2026-07-30 03:45 UTC Score 55.0 USR-0152-20260730-community-fo-f5868fd0

The biggest bet in history

Crossposted from canaryinstitute.ai/blog/biggest-bet-in-history . Related posts AGI and the EMH: markets are not expecting aligned or unaligned AI in the next 30 years — Trevor Chow, Basil Halperin, J. Zachary Mazlish. The canonical statement of "interest rates as a revealed-preference indicator of AGI timing"; this piece is the capex-side companion. Power Overwhelming: dissecting the $1.5T AI revenue shortfall — ykevinzhang. Runs the same capex-vs-revenue math from a different angle and reaches a similar "the required growth is enormous" conclusion. Bubble, Bubble, Toil and Trouble — Zvi Mowshowitz. Directly addresses the "is hyperscaler capex a fiber-optic-style bubble" question. DeepSeek Made it Even Harder for US AI Companies to Ever Reach Profitability — Garrison Lovely. The "open-weight competition collapses the quality gap and competes profits away" exit path. The biggest bet in history The amount that's been bet on AI shows that the hyperscalers really believe in AGI soon. A subject of endless speculation is "how much of AI is a bubble?". On the one hand, the techies are saying "do you think the graph is really super-exponential, or just regular-exponential?", while on the other hand, the economists are saying "A total of half a percent increase to GDP, arriving over a decade". Talk is cheap, though... who is putting their money where their mouth is? It turns out, the hyperscalers! There's 5 companies who are building out the infrastructure to make AI happen. They're…

GitHub Engineering 2026-07-29 16:00 UTC Score 34.0 USR-0062-20260729-ai-specialis-df364637

Tame Dependabot: Group your updates, slow the cadence, keep security fast

Dependabot keeps your dependencies current, but its defaults can flood your repository with pull requests. Here's how grouping updates, slowing the cadence, and keeping security fixes fast cut the noise on a Microsoft open source project. The post Tame Dependabot: Group your updates, slow the cadence, keep security fast appeared first on The GitHub Blog .

Synced 2026-07-29 14:17 UTC Score 37.0 AI-041-20260729-ai-specialis-a8632e0f

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by juego copero

Great article on FPDT! The memory hierarchy optimization is really impressive. As someone who runs a gaming site, I can appreciate efficiency improvements—smooth gameplay depends heavily on resource management. Check out https://juegocopero.com/ for more on that!

Gradient Flow 2026-07-29 13:22 UTC Score 41.0 USR-0119-20260729-ai-specialis-28b9fb76

Specialized AI Is Getting Easier to Build

Last week I argued that open models will absorb most of the money and compute the world spends on AI. A week later, open weights are even more central to the conversation. Recent releases have made the gap between capable and affordable harder to ignore, and a broad coalition of technology companies is now publicly Continue reading "Specialized AI Is Getting Easier to Build" The post Specialized AI Is Getting Easier to Build appeared first on Gradient Flow .

The Decoder 2026-07-29 11:50 UTC Score 57.0 AI-168-20260729-regional-ai--12ceb23a

OpenAI open-sources Codex Security CLI to help developers find and fix vulnerabilities from the command line

OpenAI has released Codex Security CLI, an open-source tool that automatically detects and fixes vulnerabilities in code repositories. Previously known internally as "Aardvark," the system has already helped fix more than 3,000 critical security flaws, according to OpenAI. It competes directly with Anthropic's Claude Security, as both AI companies race to match the growing automation of cyberattacks with AI-powered defense. The article OpenAI open-sources Codex Security CLI to help developers find and fix vulnerabilities from the command line appeared first on The Decoder .

InfoWorld AI 2026-07-29 09:00 UTC Score 38.0 USR-0126-20260729-global-ai-ne-b0b7254a

Why open source matters in an AI world

I’m an old Borland guy. I started using Borland tools in the early 1990s, from Turbo Pascal through Delphi. I dabbled in Paradox. I even tried to stick with Borland Office. I wandered through the Inprise years , and the Kylix endeavor , and I was actually a Borland employee during the CodeGear / Embarcadero migrations. You could write a book about the rise and fall of Borland. Suffice it to say that things got dodgy when Borland strayed from its focus on developer tools, and they never really recovered. One of Borland’s missteps came with the open-sourcing of InterBase , its RDBMS. In the early 2000s, open-source software emerged from the academic shadows and began to commercialize. With the IPO of Red Hat , everyone was jumping on the Linux and open-source bandwagon. Borland made their move into this arena with InterBase. Two things soon happened. First, the Firebird project was created as a fork of the InterBase source. Second, Borland retreated from the open-source project and reincorporated InterBase as a closed-source product. One might mark this as the beginning of developers losing faith in Borland. It was developers, then And it always ends up being about the developers, right? I mean, who can forget a sweaty, maniacal Steve Ballmer jumping around onstage screaming “Developers, developers, developers”? For many years, Microsoft was notoriously anti-open-source, but even they eventually recognized the power of open-sourcing major tools like .NET and Visual Studio Code…

OpenAI Community 2026-07-29 08:02 UTC Score 45.0 AI-116-20260729-social-media-e110aa67

Introducing the Open-Source Codex Security CLI

Codex Security helps security and engineering teams find, confirm, and fix vulnerabilities. Use its command-line interface (CLI) to scan repositories you own or have permission to assess, review findings over time, and check changes before they land. Quickstart Guide for an interactive scan Cloud set-up for connected GitHub repositories Link to the public repo: Codex Security This is an early release, and we’re listening to your feedback as we continue improving it. The Codex Security CLI and SDK are in beta and require access. Follow the installation instructions provided with your access. For an interactive scan in Codex, start with the Codex Security plugin quickstart . For connected GitHub repositories, see Codex Security cloud setup . Install the open-source Codex Security CLI,: npm install @OpenAI/codex-security Or start with: npx @OpenAI/codex-security@latest --help NPM: npmjs.com/package/@openai/codex-security

SiliconANGLE AI 2026-07-28 20:14 UTC Score 42.0 USR-0127-20260728-global-ai-ne-0bf79b86

Biggest ever MCP update brings metadata, cybersecurity enhancements

The developers of the Model Context Protocol, an open-source technology that underpins many artificial intelligence applications, today released a new version of the software. The release is described as the biggest update to the project since its launch. The Model Context Protocol, or MCP, was open-sourced by Anthropic PBC in November 2024. The AI developer […] The post Biggest ever MCP update brings metadata, cybersecurity enhancements appeared first on SiliconANGLE .

LessWrong AI 2026-07-28 18:30 UTC Score 73.0 USR-0152-20260728-community-fo-1b84d201

Auditor-in-a-Box: Tools for Third-Party Auditing

Introduction: There is a need for untrusting parties to share information. In the world before LLMs (and even today) this need has largely been satisfied using legal contracts (and sometimes through cryptography and blockchain technologies). However the scale and pace of things to keep track of and monitor has grown immensely and legal contracts appear insufficient. LLMs can help give auditors the right tools to enable information sharing with untrusted parties. In this post, I outline the shape of the problems that need to be resolved to enable using LLMs for 3rd party auditing, and present one concrete solution we have attempted. We designed a scheme for third party auditing: an open-source LLM, running inside a trusted execution environment (TEE), which executes commands that the two parties have agreed on over private data. In this post, I highlight that this type of tooling can directly be applied to enabling 3rd party monitoring for regular governance concerns, it has applications to enabling improved monitoring for customers who want Zero-Data-Retention, and it also can be one of the tools that enable a verified slowdown. Code: We release an open-source implementation that runs in a real TEE. The code is opensourced as well: main code , webapp code , code to deploy it into a TEE. There is a live demo at https://auditor-in-a-box.royrinberg.com/ . Other Writing on This: Here is a mini-paper submitted to TAIGR by Ben Penchas , myself, and a collaborator G.Z. How to engag…

ClearML Blog 2026-07-28 15:21 UTC Score 41.0 USR-0084-20260728-ai-specialis-11a1750a

Why AI Sovereignty Is an Operational Problem

What Control Really Means for Government AI By Frank La Vigne AI sovereignty has become one of those phrases that sounds precise until someone asks what it actually means. For one federal agency, sovereignty means keeping sensitive data inside accredited boundaries. For another, it means running open-weight models in a FedRAMP-authorized private cloud. In defense […]

The Guardian AI 2026-07-28 13:04 UTC Score 48.0 AI-021-20260728-global-ai-ne-d7cd33d3

Debate over AI’s future divides Silicon Valley as China gains ground

Open-source questions stir frank discussion – and both sides have clear economic incentives for where they land Hello, and welcome to TechScape. This week we’ll be looking at a debate over the future of artificial intelligence that’s dividing the tech industry, as well as how the European Union gave Google a slap on the wrist for anti-competitive behavior. And we’ll also catch you up on the avalanche of lawsuits accusing social media companies of getting young people addicted to their products. Continue reading...

Gradient Flow 2026-07-28 13:00 UTC Score 41.0 USR-0119-20260728-ai-specialis-c645f2e0

The Big AI Labs Are Suddenly Competing with Your Own Data

Subscribe • Previous Issues Specialized AI Is Getting Easier to Build Last week I argued that open models will absorb most of the money and compute the world spends on AI. A week later, open weights are even more central to the conversation. Recent releases have made the gap between capable and affordable harder to ignore, and Continue reading "The Big AI Labs Are Suddenly Competing with Your Own Data" The post The Big AI Labs Are Suddenly Competing with Your Own Data appeared first on Gradient Flow .

The Decoder 2026-07-28 12:06 UTC Score 47.0 AI-168-20260728-regional-ai--0758e0a5

Anthropic CEO Amodei doubles down on open-weight risk stance while insisting he never called for a ban

Anthropic CEO Dario Amodei is once again warning about the risks of open AI models while insisting he has never called for a ban. He argues that authoritarian states like China could overtake the US and that open models could be misused for biological or cyberattacks. Critics say he's mostly trying to protect his own business from cheaper competition. The article Anthropic CEO Amodei doubles down on open-weight risk stance while insisting he never called for a ban appeared first on The Decoder .

South China Morning Post AI 2026-07-28 12:00 UTC Score 72.0 AI-156-20260728-regional-ai--c04048c2

Moonshot’s Kimi K3 triggers Silicon Valley debate over bans on Chinese open source models

China’s Moonshot AI has made its latest artificial intelligence model Kimi K3 available for public download, as its open-source strategy and more support for non-Nvidia ecosystems trigger heated debates across Silicon Valley. On Monday, the Chinese AI start-up not only released the model’s “weights” – the underlying parameters that encode its intelligence – but also made available key infrastructure, including tools to improve efficiency and stability. The move has allowed developers around the...

InfoWorld AI 2026-07-28 11:19 UTC Score 67.0 USR-0126-20260728-global-ai-ne-7f5ff092

Anthropic rejects open-weight AI bans, calls for China chip controls and safety tests

Anthropic CEO Dario Amodei has argued that policymakers should keep lower-risk open-weight AI accessible while placing stricter safeguards around frontier systems, including mandatory testing and limits on China’s access to advanced computing and model capabilities. In a post outlining Anthropic’s position, Amodei said broad restrictions, including bans on Chinese open-weight models used by US businesses, would not address his main national security concerns. Instead, he pointed to the possibility of authoritarian governments surpassing the US in advanced AI, as well as cyber, biological, and alignment risks posed by increasingly capable systems. Amodei also called for action against industrial-scale model distillation , which he said allows Chinese developers to improve their models with less computing power than would be needed to train comparable systems from scratch. The statement followed criticism of Anthropic for not signing an industry letter backed by Nvidia, Microsoft, Meta, IBM, Mistral, Hugging Face and other technology companies urging policymakers to avoid premature restrictions on open-weight models . The letter said that open weights could broaden access to AI, intensify competition, and enable organizations to adapt and deploy models without relying on a single provider. Amodei agreed with parts of that case but disputed claims that openness inherently improves safety research or gives defenders an advantage over attackers. He said regulation should be based…