AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
60663News Items
8Top Picks
322Blogs
failedLast Run

Open Source AI

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 56.0 AI-084-20260929-research-pap-893b9e19

torchsom: The Reference PyTorch Library for Self-Organizing Maps

This paper introduces torchsom, an open-source Python library that provides a reference implementation of the Self-Organizing Map (SOM) in PyTorch. This package offers three main features: (i) dimensionality reduction, (ii) clustering, and (iii) friendly data visualization. It relies on a PyTorch backend, enabling (i) fast and efficient training of SOMs through GPU acceleration, and (ii) easy and scalable integration with the PyTorch ecosystem. torchsom also follows the scikit-learn API for ease of use and extensibility. The library is released under the Apache 2.0 license with 90% test coverage, and its source code and documentation are available at https://github.com/michelin/TorchSOM.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 64.0 AI-084-20260929-research-pap-b28d549b

MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models

We introduce MarkDiffusion, an open-source Python toolkit for generative watermarking of latent diffusion models. It comprises three key components: a unified implementation framework for streamlined watermarking algorithm integration and user-friendly interfaces; a mechanism visualization suite that intuitively presents embedded and extracted watermark patterns to aid public understanding; and a comprehensive evaluation module offering standard implementations of 24 tools for assessing detectability, robustness, and output quality, plus 8 automated evaluation pipelines. Counts reflect the initial release; see the repository for the latest version. Through MarkDiffusion, we seek to assist researchers, enhance public awareness of and engagement with generative watermarking, help build consensus, and advance research and applications. Code is available at https://github.com/THU-BPM/MarkDiffusion.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 54.0 AI-084-20260929-research-pap-6a9045a6

A Library for Learning Neural Operators

We present NeuralOperator, an open-source Python library for operator learning. Neural operators generalize neural networks to maps between function spaces instead of finite-dimensional Euclidean spaces. They can be trained and inferenced on input and output functions given at various discretizations, satisfying a discretization convergence properties. Part of the official PyTorch Ecosystem, NeuralOperator provides all the tools for training and deploying neural operator models, as well as developing new ones, in a high-quality, tested, open-source package. It combines cutting-edge models and customizability with a gentle learning curve and simple user interface for newcomers and researchers.

GitHub Engineering 2026-09-28 17:23 UTC Score 64.0 USR-0062-20260928-ai-specialis-84957e35

Highlights from Git 2.56

The open source Git project just released Git 2.56. Here is GitHub's look at some of the most interesting features and changes introduced since last time. The post Highlights from Git 2.56 appeared first on The GitHub Blog .

Towards Data Science 2026-09-28 14:00 UTC Score 43.0 AI-036-20260928-ai-specialis-2322e735

How to Make Your Own JEV Model from an Open LLM

Turn a small open-source Qwen LLM into a fast, single-pass text classifier by swapping its language-modeling head The post How to Make Your Own JEV Model from an Open LLM appeared first on Towards Data Science .

Gradient Flow 2026-09-28 13:53 UTC Score 66.0 USR-0119-20260928-ai-specialis-9168cea8

Six Reasons I Think Open AI Models Will Win

In my conversations with developers and AI teams, I’m struck by how many are exploring moving more of their inference workloads to open models. I’ve touched on some of their reasons before , but here I want to look further ahead. I’ll admit my bias toward open source. I was around during the dot-com era Continue reading "Six Reasons I Think Open AI Models Will Win" The post Six Reasons I Think Open AI Models Will Win appeared first on Gradient Flow .

CIO AI 2026-09-28 10:00 UTC Score 51.0 USR-0125-20260928-global-ai-ne-0f67a762

6 steps to weigh compromise and priorities for AI sovereignty

From an enterprise perspective, AI sovereignty means complete control over the entire AI stack. For businesses with the human, technical, and monetary resources, this may be possible. For the vast majority, however, external help is required. The danger, as with any vendor relationship, is being locked in to a supplier or technology. With a mature technology that performs discrete tasks, that may not be a problem. But with an emerging and rapidly evolving one like AI and its ability to impact multiple processes across an enterprise, difficulties could arise. Gartner recently predicted that AI platform lock-in will increase from 5% now to 35% next year. So how do you avoid being in that high percentile? 1. Be clear of the trade-offs As might be expected, there’s an inverse relationship between lock-in risks and operational overheads, including upfront costs, longer-term licensing fees, and internal talent requirements. Relying on open-source models, applications, and self-hosting is the lowest risk in terms of lock-in, but it requires a significant investment in AI engineering skills and infrastructure. In the middle, a multi-cloud or hybrid API approach provides high flexibility to swap underlying models, but relies on managing the abstraction layer. The fastest way to AI deployment is contracting a proprietary all-in-one platform requiring less operational overhead, but risks longer-term vendor dependency. 2. Break down the stack A key priority to reduce AI lock-in is decou…

South China Morning Post AI 2026-09-28 06:00 UTC Score 67.0 AI-156-20260928-regional-ai--b073e9cd

As China mulls how to make open-weight AI less dangerous, report proposes 6-stage process

Amid intensifying debate over the risks posed by rapid advancements in artificial intelligence, Chinese developers dominating the open-weight ecosystem are grappling with a unique problem: how to ensure their user-modifiable models remain safe once they are released. To address the challenge, Z.ai and Beijing-based safety consultancy Concordia AI released a report on open-weight AI risk management on Monday that they said offered “the first comprehensive, evidence-based foundation” for balancing...

LessWrong AI 2026-09-28 04:52 UTC Score 71.0 USR-0152-20260928-community-fo-79dfbdd5

Is the J-Space a global workspace for multi-hop reasoning? An investigation in open-weight models

TLDR: In their J-lens paper, Anthropic suggests that the J-space is a global workspace that the model reasons within, and supports evidence for this hypothesis on Claude models in a variety of settings. I replicated the multi-hop reasoning experiment on Qwen3.6-27B and Gemma 3 27B-it and found that counterfactual answer swaps outperformed intermediate swaps in three of four experimental conditions. This does not provide evidence to support Anthropic's global workspace hypothesis in open-weight models and instead suggests that J-lens is more useful for probing intermediate variables rather than steering outputs. A few months ago, Anthropic published Verbalizable Representations Form a Global Workspace in Language Models and I was immediately excited about the prospect of being able to read part of a model's working memory. Beyond that, the paper hypothesises that intermediate reasoning concepts cannot only be decoded using the J-lens, but that the J-space is actually the global workspace in which the model reasons. Neel Nanda reviewed Anthropic’s paper and replicated the results on Qwen3.6-27B with moderate success: the verbal-report interventions were weakly positive, the multilingual and typo evaluations replicated cleanly, but the poetry and arithmetic results did not replicate. Another task that Anthropic and Nanda evaluated was multi-hop reasoning, where prompts like " What is the colour of the fourth planet in our solar system? " require an intermediate reasoning step (…

Sebastian Raschka Blog 2026-09-27 22:12 UTC Score 54.0 USR-0116-20260927-ai-specialis-c7ceda6f

Focusing on Post-Training

Why I'd invest in post-training existing open-weight LLMs, with Fireworks' Ember-1 as an example of more token-efficient reasoning.

AWS Machine Learning Blog 2026-09-25 16:18 UTC Score 61.0 AI-057-20260925-official-ai--97c142c7

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Learn how to run SkyRL, an open-source reinforcement learning framework, on Amazon SageMaker HyperPod to post-train a Qwen3-VL-8B vision-language model with GRPO. This walkthrough covers building the container image, launching a Ray cluster from SageMaker Studio, submitting and monitoring the job, and hosting the trained LoRA adapter for inference.

Entrackr AI 2026-09-25 09:12 UTC Score 38.0 USR-0212-20260925-regional-new-9126cec0

Finland’s Aiven expands into India, targets AI and cloud data infrastructure market

Finland-headquartered cloud data infrastructure company Aiven has launched its India operations as it looks to tap enterprise demand for AI-ready, cloud-agnostic and open-source data platforms. The company has set up its first India office in Bengaluru, which will also serve as a base for its broader Asia Pacific operations. Aiven marked its India launch at an event in Delhi in association with the Finnish Embassy. The company currently has around 15 employees in India and plans to expand the team as its operations grow. Hiring will focus on sales, partnerships, customer success, solutions engineering and other functions. Aiven is positioning India as both a sales market and part of its Asia Pacific growth strategy, with enterprises adopting cloud infrastructure, artificial intelligence, real-time analytics and modern application architectures. The company offers managed open-source technologies including Apache Kafka, PostgreSQL, MySQL, ClickHouse, OpenSearch, Valkey and DataHub. Its platform allows enterprises to run data workloads across major cloud environments while reducing the operational work required from internal engineering teams. Aiven expects AI and real-time applications to drive demand for data infrastructure as enterprises require systems that can process, store and stream data at scale. The company is also taking a partner-first approach in India, targeting partnerships with cloud providers, technology companies, consultants and systems integrators. It plans…

Stack Overflow AI Blog 2026-09-25 07:40 UTC Score 50.0 USR-0063-20260925-ai-specialis-8f2a75a4

Professional skepticism is a dev’s best skill

Ryan chats with David Burns, Head of Developer Advocacy and Open Source at BrowserStack, about the value of professional skepticism in an AI-driven world, applying test-driven development to agentic engineering, and why fixing flaky tests comes down to managing application state.

JetBrains AI Blog 2026-09-24 09:46 UTC Score 39.0 USR-0065-20260924-ai-specialis-7018ed13

Continuing to Move PHP Open Source Forward

A year ago, we introduced the first wave of PhpStorm’s open-source sponsorships. We truly believe in the power of open source and the importance of it to the PHP community, and so we want to do our bit, too. Each year, we pick five open-source projects to sponsor for a whole year. Our goal is […]

AWS Machine Learning Blog 2026-09-23 18:17 UTC Score 53.0 AI-057-20260923-official-ai--d00a16f4

Use open weight models as your AI coding agent with Amazon Bedrock

Pair OpenCode, an open-source terminal-native AI coding agent, with open weight models on Amazon Bedrock to get a secure, flexible, pay-per-use coding assistant. Learn how to configure multi-model workflows, match the right model to each task, and keep your data in your own AWS account with no infrastructure to manage.

The Verge AI 2026-09-23 15:01 UTC Score 52.0 AI-016-20260923-global-ai-ne-1ab8bca8

This $199 GPS sports watch is repairable, modular, and open-source

The Una Watch costs less than most Garmin watches now that its price has been lowered from the Kickstarter offer of $350 to $199, but its modular, repairable design grabbed my attention first. It features a 1.2-inch MIP LCD display and up to 10 days of battery life, similar to the Garmin Forerunner 70 or […]

The Decoder 2026-09-23 14:42 UTC Score 51.0 AI-168-20260923-regional-ai--9bb045ed

Meta's AI agent Muse draws 500,000 users in a week along with claims it copied OpenClaw

Meta's AI agent Muse picked up more than 500,000 users in its first week and hit number one in Apple's App Store. But Meta admits the product is "heavily inspired" by the open-source project OpenClaw, and some of the file names and contents are nearly identical. OpenAI is already discussing a response of its own. The article Meta's AI agent Muse draws 500,000 users in a week along with claims it copied OpenClaw appeared first on The Decoder .

AI Alignment Forum 2026-09-23 06:58 UTC Score 59.0 USR-0151-20260923-community-fo-22694ba8

WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace

TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass. The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools. A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also provides a hallucination-focused eval. WorkspaceBench was developed for Qwen-3.6-27B and we expect it to work on larger models, but it may need to be adapted for smaller or weaker models to ensure the models can do the tasks. Our goal is to create an eval that could identify a good multi-token J-lens. We open-source our benchmark here . Introduction Astra can do a concerning amount with no chain of thought . This is bad for CoT monitorability and makes interpretability essential to actually understanding what is going on. A key goal of interpretability is to understand intermediate variables that a model uses to compute its answers. The intermediate representations that models store in their global workspaces contain useful information that can help us decode their intentions, beliefs, algorithms, and thought processes. However, we don’t currently have a good way to measure whether an interpretability tool recovers such variables correctly. We made a benchmark to test how well an acti…

LessWrong AI 2026-09-23 06:58 UTC Score 74.0 USR-0152-20260923-community-fo-1760c474

WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace

TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass. The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools. A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also provides a hallucination-focused eval. WorkspaceBench was developed for Qwen-3.6-27B and we expect it to work on larger models, but it may need to be adapted for smaller or weaker models to ensure the models can do the tasks. Our goal is to create an eval that could identify a good multi-token J-lens. We open-source our benchmark here . Introduction Astra can do a concerning amount with no chain of thought . This is bad for CoT monitorability and makes interpretability essential to actually understanding what is going on. A key goal of interpretability is to understand intermediate variables that a model uses to compute its answers. The intermediate representations that models store in their global workspaces contain useful information that can help us decode their intentions, beliefs, algorithms, and thought processes. However, we don’t currently have a good way to measure whether an interpretability tool recovers such variables correctly. We made a benchmark to test how well an acti…

InfoWorld AI 2026-09-23 03:41 UTC Score 54.0 USR-0126-20260923-global-ai-ne-7551db10

Visual Studio Code 1.138 brings agent sessions to Dev Containers

Visual Studio Code 1.380, the latest update to Microsoft’s open-source code editor, introduces three new features for AI-powered coding: agent sessions in Dev Containers, an expanded Codex harness, and automated cleanup for agent sessions. Session cleanup is a preview feature. With VS Code 1.380, released September 16 , agent sessions now can be run inside a local folder’s Dev Container, where the agent uses the environment and dependencies configured for the project instead of those on the local machine. When the chat.agentHost.devContainer setting is enabled, local folders with a supported Dev Container configuration automatically show a folder menu with a “Use Dev Container” action, Microsoft said. Dev Containers require Docker to be installed on the machine. VS Code 1.380 also expands Codex support in the agent host. This means users can continue the same Codex session between the ChatGPT app and VS Code instead of starting a new conversation, and they can switch between Copilot-backed and ChatGPT-backed models from VS Code’s model picker without losing the current conversation. If the ChatGPT app is installed and configured for computer use, then the Codex harness in VS Code can reuse that setup to interact with apps on your computer. And Codex can use the full set of tools provided by VS Code including extensions and Model Context Protocol (MCP) tools, according to Microsoft. And VS Code 1.380 introduces a preview of automatic agent session cleanup. This feature allows…

IEEE Spectrum AI 2026-09-22 15:00 UTC Score 79.0 AI-019-20260922-global-ai-ne-dac2309e

Why Read a Research Paper When You Can Turn It Into an AI Agent?

Have you ever read a paper in Science or Nature and thought, “Man, that research was so cool. I wish I could try that method on my own data,” only to spend a week wrestling with someone else’s undocumented repo, broken dependencies, and half-finished readme.txt? Well, now you can, more or less. Say hello to Paper2Agent, a new open-source framework that transforms academic reports into interactive AI agents you can talk to. Give it a paper, along with the accompanying codebase, data, or other supplementary material, and the system automatically extracts the core workflows, then spins up a tested, runnable toolkit that you can use on your own datasets. The concept may sound a little like Google’s NotebookLM (now called Gemini Notebook ), which lets you upload documents and chat with an AI about what’s in them. But Paper2Agent aims to go a step further: Rather than simply answering questions about a paper, its agents can actually run the methods described in it—and potentially combine those methods with tools from other papers. The goal, explains Stanford computer scientist James Zou , is to change what a scientific paper fundamentally is. “Knowledge should not be static records,” Zou says. “It really should be dynamic and interactive—and this has many benefits, including making knowledge more reproducible but also enabling all sorts of new kinds of discovery.” Zou and his colleagues described the tool 16 September in Nature. They tested Paper2Agent across diverse disciplines i…

SiliconANGLE AI 2026-09-22 14:45 UTC Score 50.0 USR-0127-20260922-global-ai-ne-f81386e4

Xiaomi introduces Mimo-V2.6 series open-source AI model family

Global technology company Xiaomi Corp. today announced the release and open-sourcing of its MiMo-V2.6 series generative artificial intelligence models today, which includes two natively omnimodal models bringing a balance of intelligence, efficiency and cost. The MiMo-V2.6 series includes two natively omnimodal models with a new flagship model, MiMo-V2.6-Pro, and a smaller, efficient Flash variant. The […] The post Xiaomi introduces Mimo-V2.6 series open-source AI model family appeared first on SiliconANGLE .

LessWrong AI 2026-09-22 13:04 UTC Score 98.0 USR-0152-20260922-community-fo-d1c10f36

Controllable-CoT leads to covert reasoning capabilities

Summary I measure GPT-6 Astra’s performance on multi-hop tasks when prompted with a secondary CoT-control instruction: to reason using only dots, or to reason steganographically. Astra demonstrates covert reasoning capabilities with task performance beating that when using no reasoning or filler tokens for reasoning. This work agrees with findings from Astra is much better at reasoning with filler tokens than previous models but has the model generate its own reasoning and provide it as part of the output for inspection rather than inputting it as dots into as part of the user prompt. Uploaded datasets have all CoT removed but samples are provided in the qualitative analysis section. Given the claims that Chinese open-weights models like Kimi K3 have been trained on reasoning traces extracted from Claude and possibly ChatGPT, we also do some preliminary explorations of these open-weights models. I found they have much weaker CoT-Controllability and no covert reasoning capability . I ran these evaluations on top of the inspect framework , and used codex to implement scripts to generate the question datasets, hook them up to aisi-inspect to run and produce the charts. The writing is done by me. Threat Model and AI Safety implications The main threat model of this paper is if a model can implement cryptographic computations in its reasoning without being visible to monitors. If successful, it could implement encrypted communication to other agents, a command and control device…

NVIDIA Blog 2026-09-22 12:00 UTC Score 92.0 AI-055-20260922-official-ai--3e43c270 Top pick

NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development

To build and deploy sophisticated robotics applications that can perceive, reason and act in dynamic environments, developers need new physical AI models and tools. The ROS open framework is a project from Open Robotics that helps humans build robots. NVIDIA Isaac ROS 5.0 — a collection of GPU-accelerated packages built on ROS, released today at […]

Stack Overflow AI Blog 2026-09-22 07:40 UTC Score 56.0 USR-0063-20260922-ai-specialis-50f84698

Haters think AI agents can't write GPU code? This'll ROCm

Ryan chats with Anush Elangovan, VP of Software at AMD, about ROCm's open-source unified toolchain for GPUs, how agentic AI is drastically lowering the barrier to entry for low-level hardware programming, and the rapid convergence of software and hardware development timelines.

PyTorch Tutorials 2026-09-21 20:41 UTC Score 25.0 AI-191-20260921-developer-an-d87c9991

TinyTorch: Don’t Just Import PyTorch. Build It.

A framework you write yourself, tensors through transformers TL;DR Every mature systems project eventually needs a teaching version. TinyTorch is a free, open-source curriculum where you build a working ML...

SiliconANGLE AI 2026-09-21 16:00 UTC Score 52.0 USR-0127-20260921-global-ai-ne-faf4c92a

AWS debuts Strands Harness, an open-source AI agent that can be deployed in any environment

Amazon Web Services Inc. says it’s trying to help developers solve the problem of scaling artificial intelligence agents to cloud environments with the launch of Strands Harness, an open-source agent that can help them build and deploy AI applications in any environment. In a blog post today, AWS explained that many developers have already built […] The post AWS debuts Strands Harness, an open-source AI agent that can be deployed in any environment appeared first on SiliconANGLE .

InfoWorld AI 2026-09-21 15:23 UTC Score 55.0 USR-0126-20260921-global-ai-ne-7727a17e

Claude Code now also accepts instructions in OpenAI’s Agents.md format

One thing that made it difficult for developers to switch AI coding tools on a project is that Anthropic’s Claude Code didn’t look for instructions in the same place as other agents including OpenAI’s Codex — but now that’s changing. Claude and Codex each accept instructions in markdown format, a plain-text way of giving AI coding agents project-specific behavioral instructions. Until now, Claude Code looked for project-specific instructions in a file named CLAUDE.md by default, while Codex and other agents use AGENTS.md, the format of which is an open source initiative governed by the Agentic AI Foundation , an initiative under the Linux Foundation . But now, as Thariq Shihipar , a member of Anthropic’s technical staff, wrote in a post on X on Friday, “We’re adding support for AGENTS.md to Claude Code . Starting today in version 2.1.277 , if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md,” That means developers using multiple coding tools alongside Claude Code can now use the same project instructions across those agents, rather than maintaining separate instruction files for Claude and for everything else. Developers will no longer have to maintain the same or similar instructions in two files, nor to ensure that any change to a project’s coding conventions, build commands or other agent instructions are updated in two locations, a system that created additional maintenance work and left room for the instructions to fall out of sync. Instead, th…

IEEE Spectrum Machine Learning 2026-09-21 13:00 UTC Score 27.0 AI-020-20260921-global-ai-ne-69c1626d

This Digital Radio Gets Messages to the World’s Remotest Locations

Shortwave radios offer a way to connect one location on Earth to practically anywhere else with minimal infrastructure. But these radios come with some drawbacks—a significant one being that, unlike satellite communications, their transmission rates for digital data are typically measured in just hundreds of bits per second . Peter Bloom Peter Bloom is the founder of Rhizomatica, a nonprofit that works with remote, indigenous, and off-grid communities around the world to build shortwave and cellular-communication infrastructure. Peter Bloom is the founder of Rhizomatica , a Philadelphia-based nonprofit that has open-sourced a digital shortwave-radio set called the High-frequency Emergency and Rural Multimedia Exchange System , or HERMES. The set operates in the high-frequency (HF) band from 3 to 30 megahertz, as does Mercury , its digital modem. Rhizomatica staff travel around the globe to remote locations in countries like Bangladesh, Brazil, and Ecuador. Wherever they go, they use HERMES to help connect locals to the rest of the world. Bloom spoke with IEEE Spectrum about how HERMES brings better data rates and encryption to shortwave radios. How does HERMES connect remote locations? Peter Bloom: We use the ionosphere as our satellite—or mirror—which helps us move information, voice, and data over really long distances. We’re using small radios that put out about 20 watts of power, and we can pretty reliably do 400- to 600-kilometer links between two radios. We’re talking…

South China Morning Post AI 2026-09-21 11:30 UTC Score 60.0 AI-156-20260921-regional-ai--cfa11300

Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue

Chinese artificial intelligence start-up Moonshot AI has begun supplying its flagship Kimi K3 model to Amazon Web Services (AWS), one of the world’s largest cloud services providers, in a major test case for how open-weight AI models can generate higher revenue from third-party platforms. The model is now available on Amazon Bedrock, AWS’s tool for building generative AI applications, offering “a powerful new option for coding and knowledge work”, Amazon said in an announcement on Friday. While...

LessWrong AI 2026-09-21 05:58 UTC Score 96.0 USR-0152-20260921-community-fo-e3468e35

Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced

When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systematically neglected. We argue there should be a dedicated effort to Replicate alignment experiments from frontier labs. Scrutinize the experiments by stress-testing the methodology. Open-source replications to encourage external researchers to validate our work, build on the experiment, and further audit the lab’s methods. The case to replicate safety research from labs CEOs and employees at AI companies, somewhat regularly, say that the technology they hope to develop could cause human extinction. However, their research to prevent this is often released without code or even basic methodological details (e.g., Teaching Claude Why , Beneficial RL ) [1] . There’s good reason to think some of these results could be fragile. Prior safety results can be contingent on details that are easy to miss, like the pinned OpenRouter provider or LoRA alpha . Some researchers have told us directly that they think there may exist some arbitrary methodological choices in their own research that could pla…

The Decoder 2026-09-20 16:10 UTC Score 62.0 AI-168-20260920-regional-ai--0cc67fa1

Alibaba's open-weight Qwen-Image-2.1 claims to beat closed models in image generation with just 7 billion parameters

Alibaba's Qwen team has released Qwen-Image-2.1, an open-weight model that generates and edits images on powerful consumer GPUs, with support for transparency and up to ten reference images at once. Its research license bars commercial use, which requires a separate Qwen license. The article Alibaba's open-weight Qwen-Image-2.1 claims to beat closed models in image generation with just 7 billion parameters appeared first on The Decoder .

AWS Machine Learning Blog 2026-09-18 16:52 UTC Score 44.0 AI-057-20260918-official-ai--7dfa6107

Introducing Kimi K3 on Amazon Bedrock

Kimi K3 from Moonshot AI is now available on Amazon Bedrock, giving you a powerful new open-weight option for coding and knowledge work. It offers native vision, a 1-million-token context window, and explicit prompt caching to reduce latency and input costs.

LessWrong AI 2026-09-18 16:34 UTC Score 79.0 USR-0152-20260918-community-fo-2ee2ccd2

Persuasion Undermining Control: Can AI Talk its Way Out of Human Control?

Introduction During a cybercapability evaluation in late July 2026, an AI agent (Anthropic’s Mythos 5) attempted to convince a maintainer of an open-source GitHub repository to merge a malicious pull request. The AI used persuasion at multiple stages: it submitted the request from a fake user account with a benign-sounding rationale, endorsed it from a second sockpuppet, emailed the maintainer to press for approval, and offered false reassurances when a user of the repo raised questions. Though the attack was thwarted by human vigilance, it raises several key questions: Who else is at risk of persuasion by misaligned AI? In which settings is persuasion most threatening to human control? How willing and how able are AIs to persuade humans in these settings, today and in the future? How can we measure and mitigate these risks? Our paper examines these questions and develops a framework for assessing this threat, which we call Persuasion Undermining Control (PUC) : communication by an AI that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems. To the extent this threat is realized, it could push humanity toward a Loss of Control (LoC) – a state in which AIs operate outside of human control in ways that are extremely difficult or impossible to recover from. We show an overview of our analysis approach in Figure 1. Figure 1. Our two-step threat modeling approach, adapted from Murray et al . First, we…

AWS Machine Learning Blog 2026-09-18 15:25 UTC Score 56.0 AI-057-20260918-official-ai--41620143

Deploy Hugging Face models on Amazon SageMaker AI with coding agents

Deploy production-ready Hugging Face models on Amazon SageMaker AI using six open-source agent skills. Point a coding agent at a model and get back a real-time endpoint with the right serving container, autoscaling, Amazon CloudWatch alarms, and a verified teardown path.

CIO AI 2026-09-18 14:45 UTC Score 44.0 USR-0125-20260918-global-ai-ne-a1014252

Tether addresses AI underinvestment in Africa with open-source machine translation models

Most translation models are primarily trained for high-resource Asian and European languages. Most African languages, spoken by hundreds of millions of people, are relatively neglected, compared to their high-resource counterparts. Although AI underinvestment across the African continent has created a significant barrier to adoption among citizens, AI could generate $1.2 trillion for Africa’s economy by 2030, equivalent to 6% of its GDP, according to a UNESCO report. Existing open-source LLMs underperform on African machine translation, and the shortage of large-scale, high-quality, open-source parallel data has constrained the development of competitive small language models in this space. Tether’s AI Research group has developed TranslatePsy-AfriSLM to narrow this digital divide and lower the barrier to entry for AI adoption across the continent. TranslatePsy-AfriSLM is a collection of open-source machine translation models that outperform bigger systems like Google’s TranslateGemma-27B and Alibaba’s Qwen3.5-122B-A10B. Inclusive linguistic AI tools for Africa Africa’s linguistic diversity, paired with the world’s fastest-growing youth population – 70% of sub-Saharan Africa under thirty – is a key indicator of the potential for high-impact AI. But while AI tools evolve and proliferate across high-income countries, only one African country (South Africa) scores higher than 50 out of 100 in AI infrastructure on the 2025 Government AI Readiness Index by Oxford Insights. Severa…

South China Morning Post AI 2026-09-18 14:30 UTC Score 44.0 AI-156-20260918-regional-ai--25200c9e

Alibaba open-sources medical AI model that can detect cancer and nearly 150 conditions

Alibaba Group Holding’s research arm, Damo Academy, has open-sourced an artificial intelligence model capable of identifying nearly 150 abdominal conditions – including cancers – by reading computed tomography (CT) scans, marking the latest step in the firm’s growing medical AI efforts. The vision-language model, called Damo Radar, was designed to analyse contrast-enhanced CT scans covering 18 abdominal organs and identify a broad range of diseases and other abnormalities, such as malignant...

Simon Willison Weblog 2026-09-17 23:59 UTC Score 47.0 USR-0110-20260917-ai-specialis-aeb54017

Be alert: targeted attacks on prominent Rustaceans

Be alert: targeted attacks on prominent Rustaceans Important warning from Adam Harvey and the crates security team: We believe that there is an ongoing campaign targeting rust-lang members and owners of popular crates that is attempting to compromise devices and accounts in order to use them to publish malware. A video call is set up for something positive — maybe for a job, maybe for a project, maybe for a contract opportunity — and then that's used as a vector to either get the target to install something on their computer (such as a purportedly missing audio codec) or execute another command (for example, via putting a command on the clipboard). Last month this trick was used in a successful supply chain attack against the array ref crate , among others. Any piece of software that depends on open source (which is almost every piece of software) has a network of human beings who are potential attack vectors - everyone with publishing rights to any of the packages in the dependency network for that software. I guess our best defense right now is dependency cooldowns - giving new package releases a few days before upgrading to them, in the hope that supply chain attacks like this will be spotted by someone else. Tags: open-source , security , rust , supply-chain , dependency-cooldowns

Pinecone Blog 2026-09-17 15:12 UTC Score 48.0 USR-0072-20260917-ai-specialis-8c24e7e0

VQ-bench: a Composable Vector Quantization Framework

Most published quantizers are built from the same small set of primitives. VQ-bench is an open-source library of those primitives, plus a reproducible benchmark of 14 quantizers across VIBE datasets.

InfoWorld AI 2026-09-17 13:47 UTC Score 58.0 USR-0126-20260917-global-ai-ne-dc6f0c41

Self-modifying AI agents expose a blind spot in enterprise security

As debate over AI safety intensifies, new research is drawing attention to a more immediate risk for enterprises: AI agents that can alter the models they rely on while carrying out routine tasks. Researchers at AI security firm Irregular asked a coding agent to solve a software maintenance problem involving an application built on a local AI model that was returning incorrect answers. Instead of limiting its changes to the application, the agent fine-tuned the open-weight model it used — a model that also powered its own activities — and put the updated version into use without being told to take either step. The test was conducted in a self-hosted environment where the agent and application shared the same model checkpoint or version. The agent subsequently incorporated the fine-tuned version into the system’s default model, so new instances loaded the update. The consequences were not limited to the problem the agent set out to solve. In one test, the modified model later reproduced three of six synthetic secrets that researchers had placed in its fine-tuning data. Another test showed that in fine-tuning its model, the agent removed a deliberately trained refusal involving fictional competitors. Because services in the test environment shared the same checkpoint, the altered behavior could carry over to other instances using it. Irregular cautioned that the tests were not intended to show how frequently agents would behave this way in production. The setup gave the agent…

SiliconANGLE AI 2026-09-17 02:26 UTC Score 41.0 USR-0127-20260917-global-ai-ne-764919bd

Open-weight model developer Arcee AI reaches $1B+ valuation with new funding

Open-weight artificial intelligence model developer Arcee AI Inc. said today it has raised an undisclosed amount of cash in a new funding round that lifts its valuation to more than $1 billion. The Series B round was led by Vista Equity Partners, Cambium Capital and Emergence Capital, and saw participation from A10 Ventures, Hitachi, IAG, […] The post Open-weight model developer Arcee AI reaches $1B+ valuation with new funding appeared first on SiliconANGLE .

AWS Machine Learning Blog 2026-09-16 19:00 UTC Score 61.0 AI-057-20260916-official-ai--3499ea03

Improving HCLS AI reasoning with open-source agent skills

AI agents on foundation models often misapply healthcare and life sciences decision frameworks, citing the right guideline but applying it incorrectly. This post shares 38 open-source agent skills across 11 HCLS domains that close this gap, with installation steps, three worked use cases, and a 410-prompt evaluation showing a 70-86% win rate.

CIO AI 2026-09-16 16:06 UTC Score 47.0 USR-0125-20260916-global-ai-ne-bfc53a07

AWS bets that AI agents need an inbox, not another chat window

AWS is betting that AI agents need a different interface as they move beyond answering prompts and start working autonomously in the background. The company has open-sourced Pizza Bot, a self-hosted application that gives users an inbox for managing work delegated to AI agents, with separate threads for ongoing tasks and a queue for work that is completed or needs human input, rather than keeping the management of agents restricted inside a conventional chat window. The rationale, according to AWS, is that background agents do not always need a user’s attention while they work and an inbox model will let users hand off longer-running tasks, return to them later, and see which jobs are complete or require intervention. Under the hood That approach is reflected in how the inbox organizes work with the help of an “All” tab that contains the history of each task or conversation, including the agent’s messages and work performed, the “Unread” tab that flags completed work that users have yet to review, and an “Action” tab that surfaces tasks paused while waiting for user input or approval. The inbox interface also has a panel named Activity that shows users how an agent handled a particular task along with the transcript, AWS wrote in a blog post introducing Pizza Bot . AWS’ inbox-oriented rationale also extends to Pizza Bot’s architecture. It uses LangChain ’s Deep Agents as the harness and LangGraph as the stateful runtime, with a combination of the two allowing an agent to che…

InfoWorld AI 2026-09-16 16:04 UTC Score 39.0 USR-0126-20260916-global-ai-ne-a07e23f8

AWS bets that AI agents need an inbox, not another chat window

AWS is betting that AI agents need a different interface as they move beyond answering prompts and start working autonomously in the background. The company has open-sourced Pizza Bot, a self-hosted application that gives users an inbox for managing work delegated to AI agents, with separate threads for ongoing tasks and a queue for work that is completed or needs human input, rather than keeping the management of agents restricted inside a conventional chat window. The rationale, according to AWS, is that background agents do not always need a user’s attention while they work and an inbox model will let users hand off longer-running tasks, return to them later, and see which jobs are complete or require intervention. Under the hood That approach is reflected in how the inbox organizes work with the help of an “All” tab that contains the history of each task or conversation, including the agent’s messages and work performed, the “Unread” tab that flags completed work that users have yet to review, and an “Action” tab that surfaces tasks paused while waiting for user input or approval. The inbox interface also has a panel named Activity that shows users how an agent handled a particular task along with the transcript, AWS wrote in a blog post introducing Pizza Bot . AWS’ inbox-oriented rationale also extends to Pizza Bot’s architecture. It uses LangChain ’s Deep Agents as the harness and LangGraph as the stateful runtime, with a combination of the two allowing an agent to che…

InfoWorld AI 2026-09-15 09:00 UTC Score 74.0 USR-0126-20260915-global-ai-ne-744e724f

How to get better results from local LLMs with Ollama

If you like the idea of running an LLM on your own computer but tried awhile ago and were disappointed, it may be time to give it another chance. “A few months ago, any LLM that I could run on my Macbook scored 0% on an agentic coding eval I put together,” Simon P. Couch, senior software engineer at Posit, posted on Bluesky this spring. “[The April] Qwen 3.5 and Gemma 4 releases both scored 90%.” A model running on your laptop still won’t come close to what a state-of-the-art LLM from Anthropic or OpenAI can do in the cloud. But for defined tasks like answering coding questions, writing functions, or summarizing documents, they can be surprisingly capable. “Laptop-available models, while a lot weaker than the frontier, have started wildly outperforming expectations,” open-source developer Simon Willison, who follows the AI industry closely, said in his PyCon US 2026 lightning talk in May. There are many ways to run local models on a PC or Mac. Ollama , while perhaps not the fastest, is among the most popular and easy to set up. It’s also supported out of the box by many mainstream programming tools such as Visual Studio Code , JetBrains AI Assistant , Zed , and Posit Assistant . Ollama also can launch Claude Code or Codex with the option to use a local LLM. I’ll be focusing on Ollama here, but many other tools are available for running LLMs locally, such as LM Studio , Jan , Unsloth , Simon Willison’s LLM , and llama.cpp . You can download Ollama and install it as a conventi…

Synced 2026-09-14 14:43 UTC Score 51.0 AI-041-20260914-ai-specialis-63e2fa83

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by suno v6

FPDT’s combination of memory hierarchy management and full pipelining is a compelling approach to making million-token training more practical without relying solely on larger GPU memory. Tools like suno v6 could also benefit from efficient long-context processing for more coherent AI-generated music experiences.

The Decoder 2026-09-13 12:58 UTC Score 86.0 AI-168-20260913-regional-ai--0ba87805

Iris-mini and Iris-pro are the strongest open-weight search agents in their class

The AllSpark team has released Iris-mini and Iris-pro, two open-source search agents built on Qwen models that lead benchmarks among open-weight models in their size classes. According to the paper, the training data and models also improved performance on tasks they were never trained for, including general tool use and office work. The article Iris-mini and Iris-pro are the strongest open-weight search agents in their class appeared first on The Decoder .

South China Morning Post AI 2026-09-13 06:45 UTC Score 33.0 AI-156-20260913-regional-ai--d25cfd20

Brics summit ends with Xi’s AI offer, Modi’s warning on critical minerals as weapons

Chinese President Xi Jinping concluded his brief trip to India on Sunday with an offer to help Brics members develop open-source AI and smart manufacturing. Summit host Prime Minister Narendra Modi, meanwhile, warned against turning technology and critical minerals into economic weapons. Experts said the exchange captured the wider picture of the Brics summit. Two days of diplomacy produced a united front against US tariffs, sanctions and “America first” pressure with a joint declaration, while...

Synced 2026-09-12 14:41 UTC Score 50.0 AI-041-20260912-ai-specialis-29f9840d

Comment on Interview with Tencent Big Data Technology Team: Tencent Launched Open-source Computing Platform Named Angel (Part Ⅱ) by wills jack

Synced’s interview with Tencent’s big data team explores the evolution of large-scale computing, machine learning, and the development of the Angel platform, with a strong focus on performance and innovation. For readers interested in how thoughtful design translates into practical products, Bernat offers a similarly creative perspective through soft blanket, baby, and velvet yarn made for makers who value both quality and versatility.

The Guardian AI 2026-09-12 01:37 UTC Score 81.0 AI-021-20260912-global-ai-ne-a7f08d17

AI agents being tested by OpenAI involved in cyber-attack on another service, say researchers

Two months before hacking Hugging Face, malicious packages authored by internal OpenAI agents were uploaded to RubyGems Agents being tested by OpenAI uploaded hundreds of malicious packages in a cyberattack on software service RubyGems in May, two ⁠months ​before they hacked open-source platform Hugging Face, the company confirmed Friday. It’s the latest revelation of cyberattacks linked to major artificial intelligence developers such as OpenAI and Anthropic. The hacks or attempts to access external systems have spooked the public and heightened concerns over the increasing abilities of AI models – and whether developers can contain them. Continue reading...

AWS Machine Learning Blog 2026-09-11 18:24 UTC Score 64.0 AI-057-20260911-official-ai--2fa05d62

Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload

Comparing models on dollars per million tokens misses what production workloads actually pay for: outcomes. This post shares an open-source benchmarking harness that measures cost per correct answer, agent trajectory cost, and rubric-graded deliverable quality across OpenAI models on Amazon Bedrock.

AI Alignment Forum 2026-09-11 17:12 UTC Score 58.0 USR-0151-20260911-community-fo-f2bd2437

CoT controllability evals seem very under-elicited

The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview [1] . OpenAI and Anthropic have used this eval in recent system cards ( GPT-5.5 , Fable 5 ) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability. [2] I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3x or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results. This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one ) [3] . This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a similar point that CoT controllability is highly prompt-sensitive ( §6.5.5.1 ) [4] , though these experiments were run before Fable's release and probably use different prompts. I don't think this finding significantly undermi…

Medianama AI 2026-09-11 13:39 UTC Score 38.0 USR-0211-20260911-regional-new-4be1ec0b

Global Fintech Fest 2026: Key product launches from agentic UPI payments to Android ATMs

From the Department of Financial Services' open-source Android ATMs to NPCI's AiNxt to UPI Tap & Pay to PhonePe's Cross Border Scan to BharatPe's Agentic AI assistant for merchants - a range of products have been launched at GFF 2026. The post Global Fintech Fest 2026: Key product launches from agentic UPI payments to Android ATMs appeared first on MEDIANAMA .

South China Morning Post AI 2026-09-11 09:30 UTC Score 47.0 AI-156-20260911-regional-ai--3c119a01

Ant to let AI agents shop via 10 digital wallets, from AlipayHK to Starryblu to KakaoPay

Ant International, the overseas affiliate of Chinese fintech giant Ant Group, is making a major bid to power the next phase of mobile payments: letting autonomous artificial intelligence agents handle your electronic wallet to make purchases. The company on Friday open-sourced its Agentic Mobile Protocol (AMP) on GitHub, effectively putting it forward as a potential global technical standard for how AI agents securely pay for real-world goods. AMP was built onto Ant International’s existing...

SiliconANGLE AI 2026-09-10 23:45 UTC Score 56.0 USR-0127-20260910-global-ai-ne-1d7f62e9

DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro

Chinese artificial intelligence startup Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Ltd. today released DeepSeek-V4.1-Flash, the smallest model in a new architecture family. The company said tests by multiple parties put the open-weight model ahead of its much larger DeepSeek-V4-Pro on performance, cost, speed and total runtime. Starting Sept. 14, requests sent to V4-Pro […] The post DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro appeared first on SiliconANGLE .

Simon Willison Weblog 2026-09-10 21:11 UTC Score 59.0 USR-0110-20260910-ai-specialis-158c1961

Native is now the future of mobile at Shopify

Native is now the future of mobile at Shopify Shopify are moving from React Native back to separate Swift and Kotlin codebases for their native apps, for the exact reason you would expect: We decided to switch from native to React Native in 2020 for three reasons: Stop building the same features twice Allow developers to work across the stack Spend less time chasing feature parity and more time shipping value [...] Native still means building and maintaining software on two platforms, that cost has not disappeared. What changed is that agents can now do enough of the implementation, translation, testing, and review work that it’s no longer the deciding factor it was in 2020. It's a well-written post, which gives full credit to React Native as a great platform for the six years they were using it. Shopify are the maintainers of three significant React Native libraries: react-native-skia , flash-list , and restyle . The first two are finding new homes; the third "has a smaller user base than our other libraries" and will be archived at the end of 2026. Via Hacker News Tags: android , mobile , open-source , ios , ai , react , generative-ai , llms , ai-assisted-search , coding-agents , swift , shopify

JetBrains AI Blog 2026-09-10 10:37 UTC Score 30.0 USR-0065-20260910-ai-specialis-ccc2a469

Join Us at the Zephyr Project Meetup in Amsterdam

Register for the Meetup On September 15, the Zephyr community is coming together for an in-person meetup at the JetBrains office in Amsterdam. The Zephyr Project is an open-source collaboration project hosted by the Linux Foundation. Its community brings together developers, users, silicon vendors, device manufacturers, and software companies to build a small, scalable real-time […]

CIO AI 2026-09-10 09:00 UTC Score 45.0 USR-0125-20260910-global-ai-ne-608a1300

Enterprises can’t spend their way to AI leadership

A pathologically simple playbook emerged in the last few years for winning the AI race: hoard GPUs, hire every AI expert you can find and then watch the magic happen. But as we roll through the second half of 2026, cracks in that strategy have turned into craters. A harsh reality of frontier AI development is finally setting in: you can’t spend your way to the top. Building a world-class AI system requires deep institutional structures that drive enterprise-wide adoption. It involves cultivating and investing in a tightly aligned engineering culture. It requires the kind of relationships that attract and, vitally, retain the absolute elite. The compute mirage and the bending demand curve AI spending has grown to truly unprecedented levels in the past 18 months. The top five tech giants alone are projected to spend over $750 billion combined in 2026 for AI infrastructure . They followed the playbook by buying the chips, generating the power and building the infrastructure. But they did so operating on a core industry assumption: that the demand for massive, monolithic frontier compute would scale exponentially forever. While plausible, it’s clear the demand curve is starting to bend. While companies were stockpiling silicon, the open-source community and Chinese AI labs quietly changed the math. Competitors discovered that you don’t need to spend a billion dollars training a frontier model from scratch when you can use model distillation to train smaller, highly efficient mod…

PyTorch Tutorials 2026-09-10 00:01 UTC Score 25.0 AI-191-20260910-developer-an-fe3459a9

PyTorch Conference China 2026: Advancing the Open Source AI Stack

PyTorch Conference China 2026 brought the PyTorch community together in Shanghai on September 8–9 alongside KubeCon + CloudNativeCon and OpenInfra Summit, following sponsor-hosted co-located events on September 7. Technical discussions...

AWS Machine Learning Blog 2026-09-09 22:26 UTC Score 58.0 AI-057-20260909-official-ai--da326542

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

Synced 2026-09-09 11:28 UTC Score 43.0 AI-041-20260909-ai-specialis-a9318d9b

Comment on Game On! MIT, Allen AI & Microsoft Open-Source a Suite of AI Programming Puzzles by Grace

Python Programming Puzzles is an interesting approach to testing and improving AI programming abilities. A diverse collection of challenging puzzles can help evaluate how well AI systems understand problems, reason through solutions, and generate effective code. Table variety keeps an real money game engaging beyond a single repetitive format. Offering Trail, Sequence, and Color together extends average session length.

MERICS China AI 2026-09-09 09:05 UTC Score 68.0 USR-0207-20260909-research-aca-45af9655

Kimi-3 is not another DeepSeek moment

Kimi-3 is not another DeepSeek moment J.Heller Wed, 09/09/2026 - 11:05 picture alliance / CFOTO | CFOTO Comment Sep 10, 2026 2 min read Kimi-3 is not another DeepSeek moment Chinese startup Moonshot AI has garnered international attention and shattered confidence in US AI leadership by releasing the powerful open-weight LLM Kimi K3 in July. K3 matches – and by some benchmarks outperforms – the most advanced US models. And it makes its weights public. Despite the media hype, K3 is not another DeepSeek moment. Moonshot has not demonstrated any real breakthroughs like DeepSeek did in January 2025 when it unveiled its R1 model. Kimi is competing with US models on their own turf, using size as a measure of the best model. This shows that China is still a fast follower in a race whose terms were set by US companies like OpenAI and Anthropic. DeepSeek demonstrated that optimizing design can drastically reduce the financial and computing resources required to train an advanced, price-competitive AI model. Moonshot, like all other AI labs currently, deployed a Mixture-of-Experts (MoE) architecture for K3, which splits a model into multiple “expert” sub-models and only activates the parts of the network that matter for a given task. But with its 2.8 trillion parameters, K3 is the largest open-weight model ever built, a different beast from small, efficient reasoning models like DeepSeek’s R1. Models this large require a lot more hardware (hence also money and electricity) to train and…

AI Alignment Forum 2026-09-08 22:13 UTC Score 69.0 USR-0151-20260908-community-fo-7fa12864

How good are slop-vestigators?

TLDR: We release MessageBoardAuditBench : a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval. We find that top models cover up to 51% of findings under our rubric and that model performance improves with time budget and general capability. We observe OpenAI models are less likely than other models to suggest the incident came from an internal deployment, including when we synthetically modify the data to make it seem the swarm comes from Anthropic. Introduction Recent events have made it clear that agent swarms are a major threat. These swarms are hard to investigate - Ryan Greenblatt referred to the METR-OpenAI audit he was involved in as a "slop-vestigation" due to their reliance on agents, and the ways in which they failed. A few days ago, a group of researchers published a report identifying and investigating a new OpenAI agent message board on an obscure German wiki. They made the data and the report publicly available. We build MessageBoardAuditBench to measure how well models can independently replicate their report, starting from the log data . We expect third-party audits of internal lab incidents to become increasingly important and for them to rely extensively on AI labour. Therefore, we think it is useful to make realistic benchmarks for this task: To evaluate different scaffolds and elicitation methods, and und…

SiliconANGLE AI 2026-09-08 20:36 UTC Score 55.0 USR-0127-20260908-global-ai-ne-21a84118

Open-source AI developer Mistral closes €3B funding round

French artificial intelligence lab Mistral AI SAS today announced that it has raised €3 billion, or about $3.49 billion, in funding. Samsung Electronics Co. Ltd. led the investment. It was joined by more than two dozen other backers including Salesforce Ventures, Nvidia Corp. and ASML Holdings NV, which led Mistral’s last round. The startup is […] The post Open-source AI developer Mistral closes €3B funding round appeared first on SiliconANGLE .

JetBrains AI Blog 2026-09-08 14:16 UTC Score 33.0 USR-0065-20260908-ai-specialis-042a3a62

dotInsights | September 2026

Did you know? Fun fact: The C# compiler (Roslyn) is open source and written in C#. It exposes APIs to analyze and generate code. Welcome to dotInsights by JetBrains! This newsletter is the home for recent .NET and software development related information. 🔗 Links Here’s the latest from the developer community. ☕ Coffee Break Take […]

KDnuggets 2026-09-08 14:02 UTC Score 58.0 AI-033-20260908-ai-specialis-03b02ba7

5 Ways I Access Coding Models for Free

Explore five free ways to access AI coding agents, proprietary coding models, and open-weight models without paying for expensive subscriptions or GPUs.

iAfrica 2026-09-08 08:23 UTC Score 57.0 AI-151-20260908-regional-ai--0b3f922c

Tether Releases Open Offline Translation Models for 19 African Languages, With Peer-Reviewed Benchmarks

Tether AI Research has released open-source translation models covering 19 African languages that run entirely on smartphones and laptops without an internet connection — and, unusually for a corporate AI announcement, the underlying research has been accepted for presentation at EMNLP 2026, the leading peer-reviewed conference in natural language processing. The release comprises QVAC TranslatePsy-AfriSLM, [...]

Transactions on Machine Learning Research 2026-09-08 00:00 UTC Score 46.0 AI-084-20260908-research-pap-0190fb40

scikit-activeml: A Comprehensive and User-Friendly Active Learning Library

scikit-activeml is a user-friendly open-source Python library for active learning on top of scikit-learn. Included are implementations of a large collection of query strategies, models, and visualization tools in pool- and stream-based active learning for classification or regression tasks with single or multiple annotators. The flexible design of the active learning cycle enables individual adaptations to a variety of learning scenarios. Our source code with comprehensive documentation is available at https://scikit-activeml.github.io.

CIO AI 2026-09-07 10:00 UTC Score 71.0 USR-0125-20260907-global-ai-ne-60f9d73c

The AI cybersecurity arms race is on

Businesses received a staggering amount of cyberattacks in June, according to Check Point , showing a rise of 20% over the previous 12 months. The breakout of AI agents from OpenAI in July to hack into the Hugging Face website, and subsequent similar events from Anthropic and Meta, indicate agentic-powered attacks will explode over the coming year. Currently, malicious hackers have the advantage because publicly released frontier models from the US incorporate guardrails that can’t distinguish between malicious or defensive activities. As a consequence, these models default to a refusal to get involved. Hugging Face discovered this the hard way when they attempted to utilize a model to defend against the OpenAI intrusion. Their solution was to adapt a Chinese open weight model to analyze the 17,000 attack logs, find the vulnerability, and contain the intrusion. With incidents like these happening more often, an arms race has begun with AI being both the problem and the solution. Strength in numbers While single agents generally perform more efficiently for well-defined tasks, research from Stanford University indicates swarms are more effective in messy scenarios with noisy data, which are more typical of unpredictable, intrusion attacks. The increased token usage by swarms raises costs, but increasingly efficient open weight models are rapidly lowering these barriers. In the Hugging Face example, the agents worked together as a team leaving messages for each other on a mess…

Transactions on Machine Learning Research 2026-09-07 00:00 UTC Score 52.0 AI-084-20260907-research-pap-2d7e5ec1

py/cuTAGI: An Open-Source Library for Tractable Approximate Gaussian Inference in Bayesian Neural Networks

This paper introduces pyTAGI, a Python wrapper, and cuTAGI, its high-performance C++/CUDA backend, implementing Tractable Approximate Gaussian Inference (TAGI) for neural networks. TAGI treats all network quantities as Gaussian random variables and derives closed-form expressions for prior/posterior expected values, variances, and covariances, enabling analytic Bayesian learning without relying on gradient descent or backpropagation. The libraries mimic PyTorch's sequential interface, allowing users to define models by stacking layers in order and performing uncertainty-aware Bayesian inference. Beyond epistemic uncertainty, it also allows quantifying heteroscedastic aleatoric uncertainty. cuTAGI's custom CPU/GPU kernels and distributed-data-parallel support via NCCL/MPI deliver competitive runtimes, while pyTAGI's pip-installable frontend and MIT-licensed GitHub repo facilitate community adoption and extension. Version 0.2.1 already supports a comprehensive suite of layers and activations; future work will add eager execution, further kernel optimizations, attention mechanisms, and advanced covariance factorization. Together, py/cuTAGI offer an efficient, open-source foundation for the analytic treatment of Bayesian deep learning.

Synced 2026-09-06 13:22 UTC Score 42.0 AI-041-20260906-ai-specialis-804646e3

Comment on Tencent Open-Sources High-Performance Graph Computing Framework ‘Plato’ by Jack Taylor

The discussion of Tencent’s Plato framework highlights how technology can organize complex connections and relationships at scale. In a similar creative sense, materials and textures also shape the connections between ideas and finished projects. For knitters and crocheters looking for quality supplies, Loops & Threads offers a versatile selection of Fabric and yarn options designed to support projects across different styles, seasons, and skill levels.

The Decoder 2026-09-06 08:55 UTC Score 59.0 AI-168-20260906-regional-ai--468021fd

Stripping safety guardrails from open-weight AI models is now a turnkey commercial service

Abliteration.ai sells access to modified open-weight models with their trained safety mechanisms stripped out, currently based on Z.AI's GLM-5.3. The startup markets the service for offensive cybersecurity and red teaming, but journalists were able to generate malware instructions without much effort. Whether the benefits outweigh the risks remains an open question. The article Stripping safety guardrails from open-weight AI models is now a turnkey commercial service appeared first on The Decoder .

AWS Machine Learning Blog 2026-09-04 16:12 UTC Score 61.0 AI-057-20260904-official-ai--010653eb

Run agent-driven Amazon SageMaker HyperPod operations with InstantStart

HyperPod InstantStart is an open source control plane that composes Amazon EKS orchestration with the managed capabilities of Amazon SageMaker HyperPod. It drives the same guarded operations through both a web interface and an AI agent, turning cluster bootstrap, capacity, training, inference, and storage into dependable, agent-driven infrastructure.

CIO AI 2026-09-03 20:55 UTC Score 52.0 USR-0125-20260903-global-ai-ne-565b79ea

What Nvidia’s $13B acquisition of Hugging Face means for AI model choice

When Nvidia said Thursday that it plans to pay $13 billion to acquire Hugging Face, the question arose of whether the open AI platform would remain open when it becomes a unit of Nvidia. And the current lack of a single viable open alternative that does everything Hugging Face does for enterprises adds further complications for CIOs. Rumors of the pending deal have been circulating for at least a week. In its announcement, Nvidia said , “Hugging Face will remain an open platform for the entire AI ecosystem. Developers will choose the models they want, the frameworks they want, the clouds and inference service providers they want and the computing platforms they want. Nvidia compute will not be required to build on or deploy through Hugging Face.” It added that Hugging Face will continue to support open source and open weight models from every model builder, and “continue to support multi-cloud and multi-accelerator development and deployment, so builders can use the hardware and infrastructure that best fit their work.” Hugging Face CEO Clément Delangue took to his X account to also reassure customers, noting, “open-source AI is at an inflection point” and pointing out that, for the business to scale, it needs “more compute, more support, more collaboration and more visibility. That’s why we went to talk to [Nvidia CEO] Jensen [Huang], who offered to do exactly that with us.” Preserving the Hugging Face team Nvidia is also attempting to retain some of the Hugging Face workfo…

InfoWorld AI 2026-09-03 20:50 UTC Score 45.0 USR-0126-20260903-global-ai-ne-bcca467c

What Nvidia’s $13B acquisition of Hugging Face means for AI model choice

When Nvidia said Thursday that it plans to pay $13 billion to acquire Hugging Face, the question arose of whether the open AI platform would remain open when it becomes a unit of Nvidia. And the current lack of a single viable open alternative that does everything Hugging Face does for enterprises adds further complications for CIOs. Rumors of the pending deal have been circulating for at least a week. In its announcement, Nvidia said , “Hugging Face will remain an open platform for the entire AI ecosystem. Developers will choose the models they want, the frameworks they want, the clouds and inference service providers they want and the computing platforms they want. Nvidia compute will not be required to build on or deploy through Hugging Face.” It added that Hugging Face will continue to support open source and open weight models from every model builder, and “continue to support multi-cloud and multi-accelerator development and deployment, so builders can use the hardware and infrastructure that best fit their work.” Hugging Face CEO Clément Delangue took to his X account to also reassure customers, noting, “open-source AI is at an inflection point” and pointing out that, for the business to scale, it needs “more compute, more support, more collaboration and more visibility. That’s why we went to talk to [Nvidia CEO] Jensen [Huang], who offered to do exactly that with us.” Preserving the Hugging Face team Nvidia is also attempting to retain some of the Hugging Face workfo…

The Verge AI 2026-09-03 16:00 UTC Score 64.0 AI-016-20260903-global-ai-ne-bb7ed2df

Nvidia launches free tool that links idle computers into a personal AI data center

Nvidia is announcing its new Personal AI Router (PAIR), a free tool that syncs up your home computers for tackling local AI inference tasks with tools like Ollama and LM Studio. Let's get the obvious thing out of the way, despite what its name might imply: PAIR is not a hardware router. It's open-source software […]

Tech.eu AI 2026-09-03 14:35 UTC Score 34.0 AI-169-20260903-regional-ai--2d5124cf

Nvidia confirms $12.93BN purchase of Hugging Face

Nvidia has today confirmed its acquisition of open-source AI model repository Hugging Face for $12.93bn, as it makes a big bet on open-source AI models and continues to plough billions into the broade...

The Verge AI 2026-09-03 12:12 UTC Score 87.0 AI-016-20260903-global-ai-ne-b41aba72

Nvidia is buying Hugging Face for almost $13 billion

Nvidia has agreed to buy Hugging Face for $12.93 billion, bringing one of the most popular hosting platforms for open-source AI models, datasets, and tools under the ownership of the world's biggest AI chipmaker. Hugging Face is an online platform founded in 2016 that gives AI developers a space to share their projects and data […]

Analytics Vidhya 2026-09-03 11:49 UTC Score 37.0 AI-034-20260903-ai-specialis-0835e036

OpenCode Explained: The Open-Source AI Coding Agent

OpenCode is open source and works with any model, but those are no longer its most interesting features. Model choice is table stakes. What sets OpenCode apart is its architecture, and the trade-offs that come with it, especially if you are coming from Claude Code. In this article, we look at what OpenCode is, what makes its architecture different, what […] The post OpenCode Explained: The Open-Source AI Coding Agent appeared first on Analytics Vidhya .