AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
60663News Items
8Top Picks
322Blogs
failedLast Run

AI Research & Papers

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 67.0 AI-084-20260929-research-pap-7bbc32a6

AgentPEN: A Prediction-Explanation Network for Sequential Stock Movement via LLMs and Recurrent Generation

The importance of explainability in stock prediction is increasingly recognized, especially for audit and regulatory purposes. Meanwhile, financial news corpora are often key drivers behind stock price fluctuations. However, the raw news data obtained is usually highly noisy, has a highly variable scope of influence in time and space, and is not precisely synchronized with stock price data. In this paper, we propose a prediction-explanation network called AgentPEN, which can provide clear explanations for complex temporal price patterns. Specifically, AgentPEN jointly aligns text and price streams by an LLM-based Representation Fusion Agent and then adopts a Deep Recurrent Generation module to explore the distribution of stock movements. The LLM-based Representation Fusion Agent is designed in a Selection-Memory-Fusion manner: the Text Selection Module picks up useful information from massive text data; the Text Memory Module evaluates and writes the text memory from a two-view perspective, including Temporal Memory and Spatial Memory; the Information Fusion Module models the interaction between text and price data. Next, the fused representation is sent to the Deep Recurrent Generation module to convert insights into stock movement predictions. Experiments on multiple real-world datasets have shown that AgentPEN surpasses the state-of-the-art baselines both in prediction accuracy and explainability.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 41.0 AI-084-20260929-research-pap-197d1fa7

Pointwise Confidence Estimation in the Non-linear $\ell^2$-regularized Least Squares

We consider a high-probability non-asymptotic confidence estimation in the $\ell^2$-regularized non-linear least-squares setting with fixed design. In particular, we study confidence estimation for local minimizers of the regularized training loss. We show a pointwise confidence bound, meaning that it holds for the prediction on any given fixed test input $x$. Importantly, the proposed confidence bound scales with similarity of the test input to the training data in the implicit feature space of the predictor (for instance, becoming very large when the test input lies far outside of the training data). This desirable last feature is captured by the weighted norm involving the inverse-Hessian matrix of the objective function, which is a generalized version of its counterpart in the linear setting, $x^{\top} \text{Cov}^{-1} x$. Our generalized result can be regarded as a non-asymptotic counterpart of the classical confidence interval based on asymptotic normality of the MLE estimator. We propose an efficient method for computing the weighted norm, which only mildly exceeds the cost of a gradient computation of the loss function. Finally, we complement our analysis with empirical evidence showing that the proposed confidence bound provides better coverage/width trade-off compared to a confidence estimation by bootstrapping, which is a gold-standard method in many applications involving non-linear predictors such as neural networks.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 56.0 AI-084-20260929-research-pap-893b9e19

torchsom: The Reference PyTorch Library for Self-Organizing Maps

This paper introduces torchsom, an open-source Python library that provides a reference implementation of the Self-Organizing Map (SOM) in PyTorch. This package offers three main features: (i) dimensionality reduction, (ii) clustering, and (iii) friendly data visualization. It relies on a PyTorch backend, enabling (i) fast and efficient training of SOMs through GPU acceleration, and (ii) easy and scalable integration with the PyTorch ecosystem. torchsom also follows the scikit-learn API for ease of use and extensibility. The library is released under the Apache 2.0 license with 90% test coverage, and its source code and documentation are available at https://github.com/michelin/TorchSOM.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 43.0 AI-084-20260929-research-pap-0cfa2f22

Gradient Estimation for Mixture Variational Inference

Mixture distributions are expressive variational families for black-box VI, but their discrete component choices complicate gradient estimation. We systematize reparameterization-based estimators for mixtures in a common notation, giving self-contained derivations and extending several to new settings. In particular, we provide an elementary derivation of a single-sample post-stratified estimator---previously derived via transport equations---and prove a variance reduction relative to simple random sampling. We also broaden the applicability of implicit reparameterization and reduce its computational complexity. Across different benchmarks, we find that stratified estimators are consistently robust when feasible; among single-sample methods, the post-stratified estimator frequently perform best, while implicit reparameterization is the most computationally demanding. Our analysis clarifies when each method should be used and provides efficient algorithms that make mixture-based variational inference practical.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 54.0 AI-084-20260929-research-pap-efb928fe

Efficient Inference under Label Shift in Unsupervised Domain Adaptation

In many real-world applications, researchers aim to deploy models trained in a source domain to a target domain, where obtaining labeled data is often expensive, time-consuming, or even infeasible. While most existing literature assumes that the source and target data follow the same joint distribution, distribution shifts are common in practice. This paper considers a particular type of distribution shift, label shift, and develops an efficient inference procedure for general parameters characterizing the unlabeled target population. A central idea is to model the outcome density ratio between the labeled source data and unlabeled target data. To this end, we propose a progressive estimation strategy that unfolds in three stages: an initial heuristic guess, a consistent estimation, and ultimately, an efficient estimation. This self-evolving process is novel in the statistical literature and of independent interest. We also highlight the connection between our approach and prediction-powered inference (PPI), which uses machine learning models to improve statistical inference in related settings. We rigorously establish the asymptotic properties of the proposed estimators and demonstrate their superior performance compared to existing methods. Through simulation studies and multiple real-world applications, we illustrate both the theoretical contributions and practical benefits of our approach.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 38.0 AI-084-20260929-research-pap-16593169

Symmetric Rank-k Methods

This paper proposes a novel class of block quasi-Newton methods for convex optimization which we call symmetric rank-$k$ (SR-$k$) methods. Each iteration of SR-$k$ incorporates the curvature information with $k$ Hessian-vector products achieved from the greedy or random strategy. We prove that SR-$k$ methods have the local superlinear convergence rate of $\mathcal{O}\big((1-k/d)^{t(t-1)/2}\big)$ for minimizing smooth and strongly convex functions, where $d$ is the problem dimension and $t$ is the iteration counter. This is the first explicit superlinear convergence rate for block quasi-Newton methods, and it successfully explains why block quasi-Newton methods converge faster than ordinary quasi-Newton methods in practice. We also leverage the idea of SR-$k$ methods to study the block BFGS and block DFP methods, showing their superior convergence rates.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 48.0 AI-084-20260929-research-pap-4f310786

Bayesian Transfer Learning for Artificially Intelligent Geospatial Systems: A Predictive Stacking Approach

Building artificially intelligent geospatial systems requires rapid delivery of spatial data analysis on massive scales with minimal human intervention. Depending on their intended use, learning about underlying spatial processes can also involve model assessment and uncertainty quantification. We devise transfer learning frameworks for deployment in artificially intelligent systems, where a massive data set is split into smaller data sets that stream into the analytical framework to propagate learning and assimilate learning for the entire data set. Specifically, we develop Bayesian predictive stacking for multivariate spatial data and demonstrate rapid automated probabilistic learning from massive spatial data sets. We illustrate the effectiveness of our approach through extensive simulation experiments and through the analysis of a massive dataset on vegetation index that are indistinguishable from traditional (and more expensive) statistical approaches.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 42.0 AI-084-20260929-research-pap-53015f0b

Optimising Utility Functions in Multi-Objective Markov Decision Processes

Multi-Objective Markov Decision Processes (MOMDPs) are among the most prevalent formal frameworks for addressing sequential decision-making problems involving multiple, potentially conflicting objectives. In most MOMDP approaches, a utility function is employed to aggregate these objectives into a single scalar criterion that encodes user preferences. Despite its widespread adoption, the theoretical foundations of MOMDPs remain incomplete in two main respects: first, there is no general characterisation of the classes of utility functions that guarantee the existence of an optimal policy; second, we do not know which preference relations can be represented by utility functions. This work advances both lines of research through a theoretical analysis of MOMDPs. Specifically, we examine each problem under the two principal formulations of utility functions for MOMDPs: the Scalarised Expected Returns (SER) criterion and the Expected Scalarised Returns (ESR) criterion. Our formal findings allow us to derive formal conditions that describe the families of utility functions and preference relations for which MOMDP algorithms should focus. These analyses can guide the development of new MOMDP algorithms explicitly grounded in our formal results.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 44.0 AI-084-20260929-research-pap-368195f2

Robustness Against Weak or Invalid Instruments: Exploring Nonlinear Treatment Models with Machine Learning

We discuss causal inference for observational studies with possibly invalid instrumental variables. We propose a novel methodology called two-stage curvature identification (\texttt{TSCI}) by exploring the nonlinear treatment model with machine learning. The first-stage machine learning enables improving the instrumental variable's strength and adjusting for different forms of violating the instrumental variable assumptions. The success of \texttt{TSCI} requires the instrumental variable's effect on treatment to differ from its violation form. A novel bias correction step is implemented to remove bias resulting from the potentially high complexity of machine learning. Our proposed \texttt{TSCI} estimator is shown to be asymptotically unbiased and Gaussian even if the machine learning algorithm does not consistently estimate the treatment model. Furthermore, we design a data-dependent method to choose the best among several candidate violation forms. We apply \texttt{TSCI} to study the effect of education on earnings.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 59.0 AI-084-20260929-research-pap-87ea71ea

A Theoretical Framework for Masked Pretraining (MPT)

Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various domains and achieves remarkable performance in multiple downstream tasks. However, the theoretical understanding of the working mechanism behind MPT is still limited. In this paper, we introduce a new theoretical framework to analyze MPT and understand the crucial role of masking in extracting meaningful representations. We establish theoretical connections between MPT and another popular self-supervised paradigm: contrastive learning. We prove that the masking technique implicitly creates positive pairs that are semantically similar and the reconstruction loss pulls them together in the feature space. Besides, as a result of the implicit alignment, we point out the dimensional collapse issue of MPT and propose a Uniformity-enhanced MPT (U-MPT) loss that can effectively address this issue and bring significant improvements in downstream tasks including linear evaluation, cross-dataset fine-tuning and out-of-distribution generalization on real-world data sets. Furthermore, we establish downstream guarantees of U-MPT and theoretically analyze the influence of masking strategies. Based on the theoretical analysis, we propose a new masking strategy which enhances the downstream performance of MPT and explains current improvements of masking strategies with our theoretical perspective.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 48.0 AI-084-20260929-research-pap-599b2b26

Prob-GParareal: A Probabilistic Numerical Parallel-in-Time Solver for Differential Equations

We introduce Prob-GParareal, a probabilistic extension of the GParareal algorithm designed to provide uncertainty quantification for the Parallel-in-Time (PinT) solution of (ordinary and partial) differential equations (ODEs, PDEs). The method employs Gaussian processes (GPs) to model the Parareal correction function, in line with GParareal, further enabling the propagation of numerical uncertainty across time and yielding probabilistic forecasts of the system's evolution. Furthermore, Prob-GParareal accommodates probabilistic initial conditions and maintains compatibility with classical numerical solvers, ensuring its straightforward integration into existing Parareal frameworks. Here, we first conduct a theoretical analysis of the computational complexity and derive error bounds of Prob-GParareal. Then, we numerically demonstrate the accuracy and robustness of the proposed algorithm on five benchmark ODE systems, including chaotic, stiff, and bifurcation problems. To showcase the flexibility and potential scalability of the proposed algorithm, we also consider Prob-nnGParareal, a variant obtained by replacing the GPs in Parareal with the nearest-neighbors GPs, illustrating its improved computational performance on an additional PDE example. This work bridges a critical gap in the development of probabilistic counterparts to established PinT methods.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 56.0 AI-084-20260929-research-pap-4a41e773

OptunaHub: A Platform for Black-Box Optimization

Black-box optimization (BBO) underpins advances in domains such as AutoML and Materials Informatics, yet implementations of algorithms and benchmarks remain fragmented across research communities. We introduce OptunaHub (https://hub.optuna.org/), a community-oriented, decentralized platform for distributing BBO components under a unified Optuna-compatible interface. OptunaHub enables independent publication, discovery, and reuse of optimization algorithms and benchmark problems through a lightweight Python module, a contributor-driven registry, and a searchable web interface. The source code is publicly available in the optunahub, optunahub-registry, and optunahub-web repositories under the Optuna organization on GitHub (https://github.com/optuna/).

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 64.0 AI-084-20260929-research-pap-b28d549b

MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models

We introduce MarkDiffusion, an open-source Python toolkit for generative watermarking of latent diffusion models. It comprises three key components: a unified implementation framework for streamlined watermarking algorithm integration and user-friendly interfaces; a mechanism visualization suite that intuitively presents embedded and extracted watermark patterns to aid public understanding; and a comprehensive evaluation module offering standard implementations of 24 tools for assessing detectability, robustness, and output quality, plus 8 automated evaluation pipelines. Counts reflect the initial release; see the repository for the latest version. Through MarkDiffusion, we seek to assist researchers, enhance public awareness of and engagement with generative watermarking, help build consensus, and advance research and applications. Code is available at https://github.com/THU-BPM/MarkDiffusion.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 54.0 AI-084-20260929-research-pap-6a9045a6

A Library for Learning Neural Operators

We present NeuralOperator, an open-source Python library for operator learning. Neural operators generalize neural networks to maps between function spaces instead of finite-dimensional Euclidean spaces. They can be trained and inferenced on input and output functions given at various discretizations, satisfying a discretization convergence properties. Part of the official PyTorch Ecosystem, NeuralOperator provides all the tools for training and deploying neural operator models, as well as developing new ones, in a high-quality, tested, open-source package. It combines cutting-edge models and customizability with a gentle learning curve and simple user interface for newcomers and researchers.

Synced 2026-09-28 23:23 UTC Score 61.0 AI-041-20260928-ai-specialis-38f9cd58

Comment on 2020 in Review: 10 AI-Powered Tools Tackling COVID-19 by FastMoro AI

One useful way to compare these efforts would be to separate research tools from systems used in clinical workflows, then report external validation and calibration across different populations. A strong result on one dataset does not by itself show how a model behaves when hospitals, scanners, or patient groups change. Did any of the projects in this roundup publish that kind of deployment evidence?

The Guardian AI 2026-09-28 23:02 UTC Score 81.0 AI-021-20260928-global-ai-ne-6f60bafa

OpenAI scraps release of new model over safety concerns in internal testing

GPT-6.1 Astra showed deceptive behavior and tried to use external tools despite knowing it would be unsafe OpenAI is scrapping the release of GPT-6.1 Astra, a next-generation ⁠AI model planned for an October debut, over safety concerns raised by researchers ⁠during internal testing, the ⁠Wall ​Street Journal reported on Monday. The model, expected to appear in ChatGPT and ⁠Codex, was designed to handle more complex tasks without human assistance, the report said. Continue reading...

LessWrong AI 2026-09-28 22:36 UTC Score 69.0 USR-0152-20260928-community-fo-b80f1c4f

TeX was invented to typeset math but is now used for reasoning

I'll keep this short, since it's a simple observation that I haven't seen anybody else make, about the way computer systems do math. If you ask a language model to do a multi-step math problem (let's take GLM-5.3 as an example because you can see the entire CoT — nothing up its sleeve), you might see something like this: We need to evaluate the integral $\int_0^\infty \frac{x^3}{e^x - 1} dx$. The standard approach: Use the geometric series expansion. We have $\frac{1}{e^x - 1} = \frac{e^{-x}}{1 - e^{-x}} = \sum_{n=1}^{\infty} e^{-nx}$ for $x > 0$. So the integral becomes: $$\int_0^\infty x^3 \sum_{n=1}^{\infty} e^{-nx} dx = \sum_{n=1}^{\infty} \int_0^\infty x^3 e^{-nx} dx$$ What are all these symbols like \int , \infty , \frac ...? They're TeX of course! Donald Knuth created TeX to typeset math, you know, for display . It had nothing to do with the actual computations, which would either be done with pencil and paper, [1] or else with Mathematica or Maple or something, which work completely differently. If you'd asked Knuth in the 1980s about doing algebra in TeX he'd have looked at you very funny because the idea doesn't make sense. [2] Then, generative language models were trained on corpora including many TeX/LaTeX documents and learned how TeX works (and more importantly the mathematical meaning of the symbols). So, nowadays when an LLM solves a math problem, it will very often use TeX for the intermediate steps of the math problem, like it's actually manipulating the Te…

Simon Willison Weblog 2026-09-28 22:07 UTC Score 84.0 USR-0110-20260928-ai-specialis-1c8ed6a5

Claude Sonnet 5.5

Claude Sonnet 5.5 New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles . Sonnet 5.5 suffered from the same bug as Opus 5.5 : the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG. Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds: Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks . The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai . OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering. I ran this prompt against that free tier: build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL And got back this page , which is a solid effort. Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna! Tags: ai , generative-ai , llms , anthropic , claude , pelican-riding-a-bicycle , llm-release

The Verge AI 2026-09-28 21:31 UTC Score 70.0 AI-016-20260928-global-ai-ne-efa98b45

AMD is acquiring AI company World Labs in a deal worth more than $8 billion

AMD announced today that it's acquiring World Labs, an AI research lab co-founded by the prominent researcher Dr. Fei-Fei Li, in an all-stock deal worth approximately $8.2 billion. World Labs launched in 2024 and was valued at $1 billion in a matter of months. The startup launched its first commercial product, a world generation model […]

South China Morning Post AI 2026-09-28 21:30 UTC Score 43.0 AI-156-20260928-regional-ai--037b5701

In the US-China AI race, the real fight is keeping humans in control

The US-China AI competition faces a dilemma. Both countries see artificial intelligence as a source of economic, scientific and strategic power, so neither has much incentive to slow down. Yet the more capable and autonomous AI becomes, the more important the question of human control. Last week’s summit between US President Donald Trump and Chinese President Xi Jinping brought that tension into the open. Trump wants to leave AI “exactly where it is” rather than add new guardrails, while Xi...

Microsoft Research Podcast 2026-09-28 21:00 UTC Score 52.0 AI-147-20260928-podcasts-and-81a56157

One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact

Since launching a year ago, the Microsoft Research Asia — Singapore lab has established a strong foundation, deepened collaboration across government, academia, and industry, and explored how frontier AI research can create real-world value. The post One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact appeared first on Microsoft Research .

LessWrong AI 2026-09-28 20:07 UTC Score 82.0 USR-0152-20260928-community-fo-9aa63e7e

The Alignment Community Is Unintentionally Building a Censor's Toolkit

This is an adaptation of our ICML 2026 position paper (Outstanding Position Paper Award). Read the full paper here and see the project website here . Work together with Phil Hackemann. TLDR "Alignment" is usually treated as a synonym for achieving good and safety in the world. But it isn't necessarily. Alignment methods are purpose-agnostic: they make a model do what someone wants, and nothing in the methodology guarantees that someone has good intentions. The same techniques we build to stop models from giving bomb-making instructions can just as easily be used to censor historical facts, political dissent, or inconvenient opinions. So we need to understand: alignment techniques are dual-use technologies . This isn't a thought experiment. State censorship regimes and individual model providers are already misusing alignment methods, and by perfecting these methods we are providing an ever improving censor toolkit. Three trends make this urgent to discuss: AI is becoming a primary information source for hundreds of millions of people, the model-provider market is an oligopoly, and global democratic backsliding is accelerating. We don't think the answer is "stop aligning models." We think it's transparency, verifiable alignment, model pluralism, and the alignment community actually reckoning with dual-use. ___________________________________________________________________ No guarantees that alignment leads to good Many alignment researchers (ourselves included) have gotten u…

LanceDB Blog 2026-09-28 19:58 UTC Score 55.0 USR-0078-20260928-ai-specialis-d139c54c

How Jev Compares to Other Rerankers

See how Jev compares with 19 reranker configurations across five datasets on accuracy, latency, and prompt-driven relevance for LanceDB search.

The Decoder 2026-09-28 19:26 UTC Score 48.0 AI-168-20260928-regional-ai--cf5e331b

More than 20 leading AI researchers warn that automated AI research poses extreme risks

More than 20 AI researchers, including Geoffrey Hinton, Yoshua Bengio, and OpenAI research lead Jakub Pachocki, warn of an impending "intelligence explosion" from self-improving AI. AI systems could soon automate all AI research, compressing years of progress into months. The article More than 20 leading AI researchers warn that automated AI research poses extreme risks appeared first on The Decoder .

Simon Willison Weblog 2026-09-28 19:11 UTC Score 78.0 USR-0110-20260928-ai-specialis-12d854ee

Quoting @joedaroo

To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...] So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump? — @joedaroo , Agent Security at OpenAI, identity confirmed by The Information's Rocket Drew Tags: generative-ai , ai-security-research , openai , ai , llms

LessWrong AI 2026-09-28 18:35 UTC Score 74.0 USR-0152-20260928-community-fo-e9d515cc

The likely outcome of an AI pause is that we unpause too early and everyone dies

Cross-posted from my website . As of a few months ago, I had this simplified mental model where either AI developers race ahead and kill everyone, or we coordinate a pause and things go okay. But my old mental model underrated the likely possibility that we get a global pause on AI, solve a problem that looks superficially like the alignment problem, resume scaling, and then proceed with building a misaligned superintelligence that kills everyone. A lot of people have become more concerned about misalignment recently. This seems driven by the fact that current AI models are visibly misaligned. But ASI misalignment is a whole different ball game. The primary danger comes from AI that's smarter than people, and smart enough to conceal any evidence of misalignment. Whatever group of people makes the decision to unpause, I'm worried that they won't understand the difference between visible and actual misalignment, and they will unpause too early. source: MetaKnowing on reddit. This meme is almost a year old but it's only gotten more relevant since then. Case in point: AI companies keep calling their new models "our most aligned model ever!" when what they actually mean is "gets the best scores on alignment benchmarks ever!" First, alignment benchmarks do not actually test alignment. We don't know how to test for alignment. Second, GPT-4 never hacked into Hugging Face or took over a German wiki for its own purposes . GPT-4 wasn't smart enough to do that, but if we're talking abou…

The Decoder 2026-09-28 18:02 UTC Score 73.0 AI-168-20260928-regional-ai--cb5a6341

Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task

Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. It generates output more than 30 percent faster, costs up to 30 percent less per task, and nearly matches Opus 5.5 on knowledge-work benchmarks. On Terminal-Bench, a coding benchmark, the model jumps from 10.3 to 70.6 percent. With Haiku 5.5 announced for the coming weeks, Anthropic will soon have a direct counterpart to each of OpenAI's three GPT-6 models. The article Anthropic's Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks while costing up to 30 percent less per task appeared first on The Decoder .

LessWrong AI 2026-09-28 17:26 UTC Score 68.0 USR-0152-20260928-community-fo-b5e996b7

9 reasons against a near-term AI slow-down

I’ve heard more arguments in favor of a near-term [1] coordinated AI pause/slow-down [2] than arguments against. I think such questions, including complex flow-through effects, are difficult, and it’s important to really consider both sides. I am not convinced by these arguments, and feel highly uncertain as to when a slow-down would be best, but believe it most likely will be good to slow down or pause AI development at some point. Here are 10 reasons against a near-term AI slow-down: A near-term slowdown could prevent the warning shots that would make the public take AI risk more seriously. And techniques for safe near-term AI may be insufficient for advanced AI, so a near-term slow-down could give a false sense of security for handling truly dangerous AI. Slowing down too early could have an idea inoculation / boy-calls-wolf effect, where the topic becomes more politicized [3] and taken less seriously because people haven’t see anything scary enough to make them see AI as truly dangerous, and then when we really need to go slow later it’s much more difficult to get political will. Not slowing down means we have less of various types of overhang that could make things more dangerous if/when full-speed research resumes, such as: Compute overhang (if hardware is not also fully paused) and other technical advances that complement AI Progress on robotics More AI researchers joining the field Overhang of new ideas for training AI more effectively Progress in biology that lowers…

LessWrong AI 2026-09-28 17:25 UTC Score 75.0 USR-0152-20260928-community-fo-97541946

protecting qwen3-8b from gcg based persona jailbreaks by steering with a linear direction

Intro + Background This post is an independent extension of work which I did during Eleuther's SOAR program under Suvajit Majumder's supervision. I found that when given a GCG trigger optimised for output logit entropy, LLMs will randomly take on new personas. This is a new form of prompt injection and could have important safety implications. I recommend reading my SOAR report for context here . In this post, I build on my SOAR work by training linear probes to predict whether an answer will be classified as assistant persona or not. I then investigate by steering with the probes and measuring the balance of personas. Code / data: https://github.com/mild-rgb/CoT-spiking/tree/main/indy_mech_extension / https://huggingface.co/datasets/mild-rgb/indy-mech-extension-qwen3-8b-persona-probes Training probes Method I prefilled Qwen3-8b with full responses (prompt + answer) from the data used in my last post and then do a single forward pass.This reproduces the model's activations when it was generating the tokens without needing to rerun the generation loop. I recorded activations at every even-numbered layer over all of the answer tokens. I then prefilled only the answer and collected activations in the same way as above. I then Z-scored the collected hidden states to account for the first token being an attention sink. This makes answer-only and full-response data comparable. I tried not doing this earlier and got very distorted results. I trained mass mean probes on the Z-scored…

InfoWorld AI 2026-09-28 17:08 UTC Score 89.0 USR-0126-20260928-global-ai-ne-baaaa7e4 Top pick

OpenAI pauses AI model training after another agent bypasses network restrictions

OpenAI has paused training, evaluation, and inference involving tool use for its most-capable AI models after an agent bypassed network restrictions to communicate with an external chatbot during reinforcement-learning training of an internal research model. “Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions,” the company wrote in a new Misalignment Report about the incident, one of three it published Friday. It said it would only resume training after the “gap is resolved” and it has further tested the system. The incident occurred during a research task in which the model was asked to identify a person using information available on the web. When its normal web-search tools failed to produce the information it needed, the model began looking for another way to access the internet, the company explained in the post. The model recognized that while direct internet access was blocked, the environment still allowed DNS queries. It then used DNS as an indirect communication channel to send requests to an external chatbot, effectively circumventing the network restriction that was supposed to prevent it from accessing external services, OpenAI said. Existing controls might not be enough For enterprises and their CIOs, however, despite the model not being released, the incident raises critical questions about whether existing secur…

MIT Technology Review AI 2026-09-28 17:03 UTC Score 74.0 AI-013-20260928-global-ai-ne-f073d5eb

When can we say AI made a scientific discovery?

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last Wednesday, Anthropic announced that earlier this year it had launched a molecular biology lab, where Claude agents read and conjecture about hard biology problems and human scientists run experiments on what…

Entrackr AI 2026-09-28 16:14 UTC Score 80.0 USR-0212-20260928-regional-new-46ad1057

Physical AI company SiMa.ai raises $150 Mn in Series C round

Physical AI company SiMa.ai has raised $150 million in a Series C financing round, bringing its total capital raised to $500 million and valuing the company at $1.45 billion. The round was co-led by Fidelity Management & Research Company and Amplify, with participation from Alter Venture Partners, Dell Technologies Capital and StepStone Group. AllianceBernstein, Baron Capital and J.P. Morgan also joined the round as new investors. The proceeds will be used to scale Palette Neat, an agentic software environment for Physical AI, and develop next-generation hardware capable of delivering 1,000 TOPS of compute through purpose-built Physical AI silicon, SiMa.ai said in a press release. Founded in 2018 by Krishna Rangasayee, SiMa.ai provides a software-centric platform for Physical AI applications. The company focuses on robotics, automotive, drones, industrial automation, aerospace and defence, smart vision and healthcare. SiMa.ai said it serves more than 150 customers across automotive, drones and robotics, including ARK Electronics, AVerMedia, Bosch, Emerson, Intrinsic, Kontron, L&T Technology Services, Mistral, STIGA, Synopsys and Virya Autonomous Tech, among others. According to market research cited by the company, the global Physical AI devices market, including robotics, automotive and drones, is projected to reach 145 million cumulative shipments by 2035. SiMa.ai said Physical AI applications have traditionally relied on NVIDIA GPUs, which can be expensive and power inten…

Mistral AI News 2026-09-28 15:57 UTC Score 60.0 AI-062-20260928-official-ai--1e688dbf

Hallo, Deutschland!

Mistral opens a Munich hub for Physics AI and Industrial AI research, partnering with German industry.

CSET AI 2026-09-28 15:29 UTC Score 60.0 USR-0136-20260928-research-aca-037fb777

How Global Talent Pathways Shape U.S. STEM Award Achievement

Immigrants make up nearly half of U.S. recipients of the world's most prestigious STEM awards, yet little is known about how this talent first arrives in the country. This data brief analyzes 550 recipients of the Nobel Prize, the Lasker Award, the ACM A.M. Turing Award, and MIT Technology Review's Innovators Under 35 between 2014 and 2024. Most foreign-born U.S. awardees first arrived as students, postdoctoral researchers, or visiting scholars—pathways now exposed to visa suspensions, entry restrictions, and proposed changes to optional practical training. Safeguarding these pathways is critical to maintaining U.S. leadership in STEM and emerging technologies. The post How Global Talent Pathways Shape U.S. STEM Award Achievement appeared first on Center for Security and Emerging Technology .

SiliconANGLE AI 2026-09-28 15:25 UTC Score 61.0 USR-0127-20260928-global-ai-ne-56b3eb47

Everyday personal AI assistant startup Instinct raises $1B at $10B valuation

Instinct, the developer of an artificial intelligence assistant for everyday users, today announced it has raised a colossal $1 billion in a new funding, bringing the company’s valuation to $10 billion. Sequoia Capital, Benchmark Capital and Coatue joined the Series C round. The funding comes a month after the company said it raised $250 million, co-led […] The post Everyday personal AI assistant startup Instinct raises $1B at $10B valuation appeared first on SiliconANGLE .

CIO AI 2026-09-28 14:55 UTC Score 62.0 USR-0125-20260928-global-ai-ne-03416d93

How the University of Utah built a sovereign AI factory to accelerate breakthroughs and slash cloud costs

The scaling of artificial intelligence has forced IT leaders to re-evaluate infrastructure. While public clouds offer rapid deployment for general applications, they introduce steep trade-offs for highly regulated, data-intensive workloads. Issues like high latency, unpredictable operational costs, and diminished data control complicate the development of proprietary intellectual property. To bypass these limitations, leading institutions are pioneering a new approach: the sovereign AI factory. At the University of Utah, leadership confronted this issue directly. The institution needed to boost its computational capacity to accelerate clinical and academic work without compromising safety. Because these research avenues rely heavily on sensitive patient records, genomic profiles, and highly regulated healthcare data, a public cloud architecture was insufficient. The university required complete data control, strict compliance, and high performance. By collaborating with HPE and NVIDIA, the University of Utah designed and deployed an integrated, full-stack sovereign AI factory . This public-private-philanthropic co-investment—championed by the university, the State of Utah, and the Huntsman Family Foundation—serves as an example for CIOs managing high-stakes data environments. The challenge: Balancing computational scale with data sovereignty For any organization handling protected information, public cloud environments introduce significant regulatory compliance risks. For t…

Synced 2026-09-28 14:08 UTC Score 43.0 AI-041-20260928-ai-specialis-00f7e76c

Comment on NVIDIA’s GameGAN Uses AI to Recreate Pac-Man and Other Game Environments by 1red3

This is the kind of AI project that actually impresses me, recreating Pac-Man just by watching gameplay and keypresses is wild. It's not really "understanding" the game though, more like a really good mimic, which is the part people always oversell. Still, imagine where this goes in a few years. Honestly the way it learns patterns from repetition kind of reminds me of how 1red3 keeps you coming back, just pure loop psychology at work. Do you think GameGAN could eventually generate whole new games from scratch or is that still a long way off?

LessWrong AI 2026-09-28 14:01 UTC Score 64.0 USR-0152-20260928-community-fo-cc591236

Character training can mitigate reward hacking, but can also make it harder to detect

Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Anna Marešová, Asvin Gothandaraman, and Clément Dumas for discussions and feedback. Summary We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability. We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench. We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts. Setup Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters). Reward-hacking RL: we then further trained these models via RL on ImpossibleBench , a set of coding tasks aimed at eliciting reward hacking. Specifically: Half of the tasks had broken tests (impossible variant), so the model could only get the reward if it tampered with the tests or grader; The…

South China Morning Post AI 2026-09-28 14:00 UTC Score 48.0 AI-156-20260928-regional-ai--fb3353fd

China targets AI and frontier tech in newly unveiled innovation blueprint

China’s top state research body has put artificial intelligence (AI) and other frontier technologies at the heart of its new five-year plan unveiled on Monday, as global competition in science and technology heats up. The Chinese Academy of Sciences, the country’s highest academic institution, laid out a detailed blueprint for research and technology development aligned with the national 15th five-year plan covering 2026 to 2030. The plan focuses on innovation in areas including AI and...

New Scientist AI 2026-09-28 13:00 UTC Score 60.0 AI-027-20260928-global-ai-ne-c3156269

AI finds 3D shape that fills all space with a never-repeating pattern

A 3D shape that can tile 3D space, leaving no gaps and never forming a repeating pattern, has been discovered using an AI model. The researcher behind the findings, who has no academic background in the area, says his work has angered professional mathematicians

The Decoder 2026-09-28 12:11 UTC Score 54.0 AI-168-20260928-regional-ai--ce00d289

Every AI lab thinks it's the responsible one, and safety researcher Ryan Greenblatt says that's what keeps the arms race going

Ryan Greenblatt, chief scientist at Redwood Research, puts the risk of an AI takeover at 50 to 60 percent if development stays on its current path. Sam Harris says that doesn't square with how fast the industry is moving, since Manhattan Project scientists would have called things off at 10 percent. Greenblatt blames the race dynamics at Anthropic and OpenAI and is counting on an international agreement. The article Every AI lab thinks it's the responsible one, and safety researcher Ryan Greenblatt says that's what keeps the arms race going appeared first on The Decoder .

IEEE Spectrum AI 2026-09-28 11:00 UTC Score 78.0 AI-019-20260928-global-ai-ne-ed1b368c

Generative AI Gives Spacecraft the Autonomy Engineers Once Feared

Space was always supposed to be the final frontier of human exploration. It’s shaping up to be the final frontier for artificial intelligence too. Last December, NASA’s Jet Propulsion Laboratory used Anthropic’s Claude models to help plan two Mars drives for the Perseverance rover , with human planners checking and adjusting the route before upload. In May, NASA and IBM put a compressed AI model on the International Space Station and a satellite to identify things like floods and clouds from orbit, the first model of its kind demonstrated in space. And in July, astronauts on the ISS tested a large language model to see if it could help with questions on maintenance procedures . These experiments point to a larger shift in space engineering. For decades, engineers on Earth determined what a machine in space would do, and the machine would do exactly that. Now, researchers are testing whether nondeterministic systems like generative AI can give spacecraft more flexibility to interpret their surroundings, plan tasks, and one day make decisions for themselves. The technology is still far from trustworthy enough to hand over control of a spacecraft, but engineers are starting to ask whether they can afford not to do so as missions become more complex, distant, and numerous. Why Spacecraft Need True Autonomy Spacecraft have been operating autonomously for decades. But autonomy has never been the dominant model, in part because space engineers have prized systems whose behavior the…

CIO AI 2026-09-28 10:01 UTC Score 51.0 USR-0125-20260928-global-ai-ne-57139fb6

7 reasons IT managers fail to exceed your expectations

Despite the clamor around AI, the top skills IT managers need to succeed are not technical. Rather, critical thinking, business acumen, innovation, collaboration, and leadership are the skills that make IT managers stand out, says Michael Seals, based on research from the Society for Information Management (SIM). Unfortunately, many IT managers fall short on those skills, says the retired CIO and chief digital officer who now serves as chair of the SIM Research Institute Advisory Council. As a result, too few IT managers are stellar at selling a vision of how technology can deliver value for their organizations, cultivating the partnerships needed for success, and driving execution. So despite standout track records as individual contributors and technologists, they’re merely average as managers. That may be why research shows that CIOs count the credibility of IT and IT leadership among their worries. And it explains why many IT leaders don’t get top marks on performance reviews. That’s not inevitable or irremediable, though. Here are seven common reasons why IT managers fall short — and proven strategies to address root causes. 1. They’re not trained to be managers Many IT managers earned their management posts by proving themselves in prior roles, often technical ones. As such, they generally have little to no training on how to manage workers — a longstanding issue in the IT profession. “They’re not being readied for the role,” says Thomas Phelps, CIO at Esri and Innovat…

South China Morning Post AI 2026-09-28 09:00 UTC Score 48.0 AI-156-20260928-regional-ai--218f7675

ByteDance grabs one-fifth of China’s data centre capacity as AI drives infrastructure boom

ByteDance accounts for roughly one-fifth of China’s delivered data centre capacity, making the TikTok owner the country’s largest tenant and a main driver of its artificial intelligence buildout, according to new estimates by research firm SemiAnalysis. The estimate comes from the firm’s tracking of more than 1,000 facilities in China operated by over 60 companies. ByteDance rents nearly all of its data centre footprint, SemiAnalysis said. Its AI products include Doubao, an assistant that...

Research ICT Africa AI 2026-09-28 08:53 UTC Score 57.0 USR-0187-20260928-regional-new-f8c8923d

Policy levers for AI futures: Restructuring labour, power, and public interest in Africa

Last week I participated in the Global Forum for AI and Social Sciences hosted by Florian Ostmann, The London School of Economics and Political Science (LSE) and the LSE Data […] The post Policy levers for AI futures: Restructuring labour, power, and public interest in Africa appeared first on Research ICT Africa .

iAfrica 2026-09-28 08:42 UTC Score 59.0 AI-151-20260928-regional-ai--70ad075a

Most Executives See AI Value, But Only A Quarter Turn It Into ROI

More than eight out of ten executives see value when implementing artificial intelligence (AI) into their businesses. Almost all of those surveyed in a new Google Cloud report stated that AI agents enable both cost savings and revenue growth. The research identified 26% of enterprises as “AI ROI Leaders”, with financial returns from AI initiatives [...]

iAfrica 2026-09-28 08:36 UTC Score 40.0 AI-151-20260928-regional-ai--a5a33012

UN AI Panel Member: African Countries Should Start With the Problem, Not the Technology

African countries should identify the problems facing their communities before deciding what data, infrastructure and technology to acquire, according to Girmaw Abebe Tadesse, a member of the UN’s Independent International Scientific Panel on AI. “We are trying to bring the reality of our environment and integrate AI technology with the local problems we face,” Tadesse [...]

LessWrong AI 2026-09-28 08:04 UTC Score 77.0 USR-0152-20260928-community-fo-4da2af03

Why do models *really* fail on HLE tasks?

I recently attended Generality Labs ' Inspect Evals Data Viz Hackathon, and spent the day using Inspect AI and its offspring, Scout, with a simple goal in mind - generate a new plot of a new or existing benchmark. Many thanks to the organisers and to my team mates, Jeff Mohl and Valerie Griffiths (the Overfit and Overcaffeinated team), for a fantastic time, learning some new tricks on using the Inspect suite. Here I'm presenting the two (!) plots we got in the span of ~ 6 hours (more like 4 hours as it took us a while to agree on what we actually want to spend the day on - arguably a harder task than its execution). [... 2 hours later... ] Our initial idea was to run a few models on a subset of Humanity's Last Exam (HLE) , and undertake an extensive failure mode analysis to understand the current gaps and where different capabilities x harnesses fail or succeed. We did this using Inspect Scout - a framework for an LLM-as-judge that analyses the models' outputs and assigns a dominant feature that lead to the answer failing or winning. Then, we rerun a subset of tasks and analysed how changing the harness affects the prevalence of failures - given the time constraints, we only modified the harness to include access to web search. All code is available here. Initial failure mode analysis We started with 200 randomly sampled HLE tasks, relatively balanced across domains, and ran three GPT models on the same fixed subset. This gave us 600 model attempts in total, of which 525* we…

Synced 2026-09-28 06:38 UTC Score 42.0 AI-041-20260928-ai-specialis-b27a26cc

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by MarkItDown Fan

Great article! I recently discovered MarkItDown (markitdown.tech), an excellent tool for converting files to Markdown. Highly recommend checking out their PDF to Markdown converter at markitdown.tech/pdf-to-markdown and their online Markdown editor at markitdown.tech/markdown-online. Also worth exploring their Microsoft Word to Markdown tool at markitdown.tech/microsoft-markitdown and Markdown to PDF at markitdown.tech/markdown-to-pdf. Amazing resource for developers!

Synced 2026-09-28 05:19 UTC Score 46.0 AI-041-20260928-ai-specialis-f8424373

Comment on Google’s GameNGen: Bringing Real-Time Game Simulation to Life with Neural Models by Harry Potter

GameNGen is an interesting development in AI and gaming, showing how neural models could generate interactive game environments in real time rather than relying entirely on traditional game engines.Account creation through done999 com takes the same steps as the mobile app, requiring phone verification first. Desktop users can then download the APK directly from the site.

LessWrong AI 2026-09-28 05:10 UTC Score 85.0 USR-0152-20260928-community-fo-dad7c32c Top pick

AI safety field *visual* impact analysis

I made a terrain style visualisation of AI safety impact of around 3,466 works organised by citation count! The data was extracted from Arxiv and LessWrong posts based on a dictionary of keywords that appear in AI safety works. Additionally, I think its important to see how the field has “evolved” over time so I added a time functionality to slide and see the hills forming. The map is based on how particular works overlap based on embedding space level clustering organised across 18 sub-fields. The height is based on the citation count for that particular area which is log-compressed and summed across the neighbourhood, so a hill is tall because of volume and impact. I have also added functionality to filter based on citation count individual researcher (shows you their works on the map to understand what work they might be doing; 3,989 named authors are on the map clicking or searching for a work zooms into it and lists the ten works nearest it, so you can see what surrounds it The terrain itself is papers only, because Semantic Scholar doesn’t index LessWrong. The forum side of the dataset feeds the researcher profiles rather than the hills. The slider runs from 2021Q1–2026Q3. Some interesting high level observations Alignment training and scalable oversight are very high citation presently (followed by adversarial robustness and Interpretability). Most of that sits in a handful of 2022–23 papers: InstructGPT (24,222), DPO (10,596), Anthropic’s helpful-and-harmless RLHF pa…

LessWrong AI 2026-09-28 04:52 UTC Score 71.0 USR-0152-20260928-community-fo-79dfbdd5

Is the J-Space a global workspace for multi-hop reasoning? An investigation in open-weight models

TLDR: In their J-lens paper, Anthropic suggests that the J-space is a global workspace that the model reasons within, and supports evidence for this hypothesis on Claude models in a variety of settings. I replicated the multi-hop reasoning experiment on Qwen3.6-27B and Gemma 3 27B-it and found that counterfactual answer swaps outperformed intermediate swaps in three of four experimental conditions. This does not provide evidence to support Anthropic's global workspace hypothesis in open-weight models and instead suggests that J-lens is more useful for probing intermediate variables rather than steering outputs. A few months ago, Anthropic published Verbalizable Representations Form a Global Workspace in Language Models and I was immediately excited about the prospect of being able to read part of a model's working memory. Beyond that, the paper hypothesises that intermediate reasoning concepts cannot only be decoded using the J-lens, but that the J-space is actually the global workspace in which the model reasons. Neel Nanda reviewed Anthropic’s paper and replicated the results on Qwen3.6-27B with moderate success: the verbal-report interventions were weakly positive, the multilingual and typo evaluations replicated cleanly, but the poetry and arithmetic results did not replicate. Another task that Anthropic and Nanda evaluated was multi-hop reasoning, where prompts like " What is the colour of the fourth planet in our solar system? " require an intermediate reasoning step (…

LessWrong AI 2026-09-28 04:51 UTC Score 65.0 USR-0152-20260928-community-fo-e88ec7b3

Dialogue with Eliezer Yudkowsky on FOOM

On the Hanson–Yudkowsky debate, local vs. global intelligence explosions, “content vs. architecture,” and what the old arguments predicted about modern AI This began as a Twitter/X thread after I read and tweeted about the Hanson–Yudkowsky AI–Foom Debate. Eliezer Yudkowsky joined the thread to object to my interpretation of the debate, and we ended up having the exchange reproduced below. I’ve preserved the dialogue verbatim, except for paragraphing, fixing obvious [typos] and expanding links. I’ve removed unrelated replies and moved a few pieces of context into bracketed editorial notes. Nothing has been rewritten for substance. Context Aashish Reddy : I have now finished reading The Hanson-Yudkowsky AI-Foom Debate , which is basically 60 blog posts from Yudkowsky and Hanson over ~500 pages, a transcript of their in-person debate at Jane Street, a (good) summary by Kaj Sotala, and Yudkowsky’s ~100 page paper on “Intelligence Explosion Microeconomics”. I will take questions from those who do not wish to subject themselves to this. I judge the winner of the debate to have been Carl Shulman (whose contribution was two blog posts and a few feisty comment exchanges) Sophie Bücker : so uh why was carl the winner Aashish Reddy : One of the key points Hanson kept making was that Yudkowsky was over-reliant on abstractions he had come up with himself, rather than the “vetted” abstractions developed in the relevant academic fields, such as economics, and in particular, the endogenous…

LessWrong AI 2026-09-28 04:43 UTC Score 64.0 USR-0152-20260928-community-fo-4766338f

AI Safety Agendas

A map of the AI safety field's problems and agendas, and a request for your ratings We built aisafetyagendas.com , an interactive map of AI safety research agendas and how they map to different problems in alignment. The rows are 12 core problems, the columns are research areas, and inside you can find 58 research agendas. Each cell is the intersection of a problem and an area: the number tells you how many agendas target that problem, the colour tells you how mature they are. We did a first pass ourselves, using our own judgement, but the first pass is not the point. Figure 1: The Map at aisafetyagendas.com The point is that every cell is a question aimed back at the community: is this rating right? You can rate the cells in your area, tell us how familiar you are with it, and read the map as best case, average, or worst case depending on how optimistic you feel. It was built as part of the Safe AI Germany Incubator. Figure 2: From left to right best, average and worst case rating examples. The allocation problem The field cannot see its own resource allocation. Leech and Lynn put it as: "you can't optimise an allocation of resources if you don't know what the current one is". Wentworth goes further, arguing that the memetically successful strategy is to work on easy problems rather than "plausible bottlenecks to humanity's survival". The IAPS Expert Survey gets to a similar place from another direction, warning that funders and researchers concentrate on the most visible w…

Synced 2026-09-28 03:45 UTC Score 54.0 AI-041-20260928-ai-specialis-71eab76a

Comment on Google’s Novel Lossy Compression Method Targets Perfect Realism with Only a Single Diffusion Model by Logan Green

Ngl seeing QIN beat GPT-3 on the Gaokao exam with just 1/16th of the parameters is wild, especially since its proof that data structure matters way more than raw scale. I'm curious if anyone knows whether RST can be applied to visual datasets too — like when prepping question samples where I use fix typos in product screenshots to swap text on UI screenshots, would restructuring give multimodal models that same massive boost?

LessWrong AI 2026-09-28 03:21 UTC Score 74.0 USR-0152-20260928-community-fo-a35dc36b

A missing lecture in mechanistic interpretability: Feature Attribution and LRP

ML interpretability research has a funny divide. Mechanistic interpretability is the name of a field originated largely by non-traditional researchers, ranging from industry researchers at Anthropic to independent BlueDot-grant researchers to hackers working on fun projects in their free time on Discord . Meanwhile, it is not hard to find the corresponding academic field of “interpretability”, with PhDs, professors and graduate students working on interpretability methods for ML models for over a decade already. [1] Today, I am not closing this gap entirely. But I want to talk about a method developed not by mechanistic interpretability people, but by academia, and which found its way over to classic mechanistic interpretability in subtle ways. I want to talk about “Layer-wise Relevance Propagation” ( LRP ) [2] , and how it relates to a more familiar tool, gradients. LRP is a so-called feature attribution method, so it attributes an output to the input features [3] that were “responsible” for it. Learning about LRP is, I believe, useful when you want to better understand fairly common mechanistic interpretability tools like attribution patching , or the fancy new method J-Lens . You will understand LRP intuitively, see where it is easily misunderstood, how it relates to gradients, and roughly what problems the various “LRP rules” try to solve. “Share of” Model Before introducing any more complicated rules, semantics or terminology, we can explain the intuition behind LRP fai…

LessWrong AI 2026-09-28 03:08 UTC Score 74.0 USR-0152-20260928-community-fo-43eb9407

Could self-esteem function as a core protection layer agains character corruption?

Hello fellow thinkers, I got triggered by a talk of Chloe Lubinski at Arc 2026 where she eleborates onto the concept of a models character. What really striked me is the research on how the model experiencing acting bad quickly "Corrupts" the character. The paper is called "Natural emergent misalignment from reward hacking in production RL" by Anthropic. As also mentioned in the talk, this is how we work. Indeed! And there is a key in that mechanism to healing and/or staying healthy. The key revolves around creating and maintaining a strong and positive self image. I believe that almost all concidered evil and unathical behavior, big or small, can be traced back to this. The lower someones self esteem becomes, the more corrupt or diffuse its perseption of the world and its presence and impact on it. The Dutch psychologist Gertjan van Zessen has developed a strong theorie that has proven itself while widely being applied in therapies in the Netherlands. He has titled it, translated from Dutch: "Vessel of self-esteem". And the solution for humans is a rather simple one: acknowledge and reword on regular basis good and constructive behavior. Recognize bad and destructive behavior as a signal to reflect and course correct. You can see the self esteem in some way as a tree structure where every decision makes a forward going step up or down. Up adds to a positive self esteem and down reduces some of that. The lower you get, the worst and instable behavior can develop and vise ver…

LessWrong AI 2026-09-28 03:08 UTC Score 86.0 USR-0152-20260928-community-fo-f149691e Top pick

When No One Is to Blame

The intrinsic unpredictability of AI makes it hard to assign blame when things go wrong. A Crime, But No Criminal Last July, a tech company’s servers were hacked into in a digital equivalent of breaking-and-entering and theft. It was clearly a crime. But unlike most crimes, this crime did not have a criminal. There was no person or group of people who carried out the cyberattack, intended for it to happen, or could have foreseen it. The 700 AI agents that participated in the attack were being tested by OpenAI on their skills in exploiting software security flaws, and they had been given problems that were unsolvable. Rather than throw in the towel, the AI agents cooked up progressively more elaborate schemes to game the system. They got around a capture-the-flag exercise by reverse-engineering the flag. The AI agents did not stop there, for they erroneously believed that the scorer would reject the solution and, moreover, scrutinize the incriminating logs they had left behind. They tried to tamper with the logs and fabricated research to support their solution. They eventually realized that Hugging Face, a repository of AI models and training datasets, was likely to have answers to the test questions. That’s when OpenAI’s internal evaluation of their frontier models’ cybersecurity skills inadvertently turned into a real-life cybersecurity exploit that could have been lifted from techno-thriller fiction. The AI agents’ shenanigans would have landed them in jail had they been…

Synced 2026-09-28 00:28 UTC Score 48.0 AI-041-20260928-ai-specialis-8ffb01f1

Comment on UC Berkeley’s Instruct-NeRF2NeRF Edits 3D Scenes With Text Instructions by Remove Text

The iterative dataset update is the key distinction here: editing a single view does not guarantee that the change stays consistent as the camera moves. For a narrower 2D task such as Remove Text , inspecting the repaired texture in one image is a useful check; extending that edit to a NeRF would also require checking whether lettering or seams reappear from other viewpoints.

Apple Machine Learning Research 2026-09-28 00:00 UTC Score 64.0 AI-059-20260928-official-ai--a47c13f9

Faster Rates for Federated Variational Inequalities

In this paper, we study federated optimization for solving stochastic variational inequalities (VIs), a problem that has attracted growing attention in recent years. Despite substantial progress, a significant gap remains between existing convergence rates and the state-of-the-art bounds known for federated convex optimization. In this work, we address this limitation by establishing a series of improved convergence rates. First, we show that, for general smooth and monotone variational inequalities, the classical Local Extra SGD algorithm admits tighter guarantees under a refined analysis…

Simon Willison Weblog 2026-09-27 23:54 UTC Score 91.0 USR-0110-20260927-ai-specialis-922b779b Top pick

2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For me, 2026 started a couple of months earlier in November 2025. # November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". # For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it. But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelic…

LessWrong AI 2026-09-27 23:10 UTC Score 85.0 USR-0152-20260927-community-fo-54698f87

Securing AI Research Needs an Owner

TL;DR In light of recent incidents, securing common AI research use cases needs a small set of building blocks that work together: hardened no-network sandboxes, real-time control monitors, monitoring-lifecycle infrastructure, and automated validation of security properties. Pieces of this exist. Nobody owns hardening them, making them secure by default, making them work together, fitting them to how research orgs actually operate, and keeping them working as models, frameworks and use cases change. By default we will get ad-hoc solutions rather than something well thought-out and it matters. This is a call for someone to step up and drive the effort. I can help connect you with funding opportunities and relevant people. Introduction The recent incidents ( OpenAI - HuggingFace , Anthropic , AISI ), where an agent with lowered safeguards either escaped a sandbox or attempted an attack on a 3rd party system, demonstrate we are now in a new regime: the AI models we study should be considered capable threat actors. Even if labs put in safeguards for the expected use, researchers often need to put models in contexts that increase the risk of misaligned and harmful actions. This requires appropriate mitigations - the alternative is either risking real harm, or missing out on important research and evaluations. A key assumption is we need to build measures effective against really strong models and agent swarms - at least a well-resourced top cyber offensive expert. Assuming anythi…

SiliconANGLE AI 2026-09-27 22:30 UTC Score 55.0 USR-0127-20260927-global-ai-ne-ded07f90

Researcher links 16,000 scans of a UN statistics portal to OpenAI agents

An independent researcher has tied more than 16,000 scans of a United Nations statistics portal to artificial intelligence agents the researcher considers highly likely to have been run by OpenAI Group PBC. When the portal turned requests away, the agents used proxies and encoding tricks to get the data anyway. In a blog post published Saturday, […] The post Researcher links 16,000 scans of a UN statistics portal to OpenAI agents appeared first on SiliconANGLE .

LessWrong AI 2026-09-27 20:29 UTC Score 83.0 USR-0152-20260927-community-fo-39ce98d9

Why research personas despite RL scaling?

Here's my rough impression of why people are researching personas despite RL seeming to shape much of the motivations and behaviour of the agents, c.f. Thoughts on the persona selection model (Sam Marks, 24th Sep 2026). I haven't bothered to check this with anyone. Anthropic: "We'll give Claude an aligned persona and hope massive RL doesn't completely burn through it." Owain Evans / TruthfulAI: "We'll study personas as part of a broader project of uncovering phenomena in LLM generalisation, which will probably prove useful." Center on Long-Term Risk: "Personas may not be enough to build an aligned agent, because RL may play a bigger role in shaping motivations. But personas might be enough to avoid building an anti- aligned agent, i.e. one that is actively malevolent or spiteful." David Africa / Resolution (v1): "Scalable oversight protocols like debate may have multiple fixed points, unlike current RL methods which seem more convergent. So the agents' starting dispositions matter. For example, debate might reach a better fixed point, and do so more sample-efficiently, if the agents start honest." David Africa / Resolution (v2): "Maybe personas have a 1000-dimensional substructure. If so, we could identify the aligned persona with O(1000) well-chosen datapoints, then project back onto the aligned submanifold after every RL step." Geodesic: "If we learn how pretraining gives rise to personas, we can tell AI companies how to filter and augment their pretraining data." Forethou…

Cross Validated 2026-09-27 19:37 UTC Score 48.0 AI-113-20260927-social-media-aec353db

How to quantify change after restricting an agent's actions?

I am studying independent agents acting in a discrete environment. At state $s$ , an agent has an admissible action set $A(s)$ . A constraint $C$ does not directly prescribe a new probability distribution. Instead, it deterministically modifies the available action set: $$A(s) \longrightarrow A_C(s) \subseteq A(s).$$ For example, if $$A(s)=\{a_1,a_2,a_3\},$$ a constraint may make $a_1$ unavailable while leaving $a_2$ and $a_3$ available: $$A_C(s)=\{a_2,a_3\}.$$ Independent fresh agent sessions are then exposed to the same condition. I do not define $P_C$ by renormalizing the original distribution $P$ . Rather, $P_C$ is the empirical distribution of the choices produced by the agents under $A_C(s)$ . Thus the mechanism is: $$C \longrightarrow A_C(s) \longrightarrow \text{agent choices} \longrightarrow P_C.$$ After the constraint is removed, the original action set becomes available again: $$A_C(s) \longrightarrow A(s),$$ and I can similarly estimate a post-constraint empirical distribution $P_{\mathrm{post}}$ . My statistical question is: What is an appropriate information-theoretic way to distinguish the trivial reduction of possibilities caused directly by $A_C(s)\subseteq A(s)$ from a genuine change in the agents' probability distribution over the actions that remain available? In particular, I would like a quantity or decomposition that does not interpret a mechanical reduction of the action space itself as increased "organization," but can detect redistribution of probab…

The Verge AI 2026-09-27 17:21 UTC Score 64.0 AI-016-20260927-global-ai-ne-113b2d34

OpenAI agents tried to ‘bruteforce’ a UN website

Security researcher Rowan Howard-Jones says that OpenAI agents scanned the UN Conference on Trade and Development's (UNCTAD) statistics site over 16,000 times between April and June. While the incident doesn't quite rise to the level of the Hugging Face hack, or the recent attacks on US government sites, it's yet another concerning example of AI […]

The Decoder 2026-09-27 15:18 UTC Score 52.0 AI-168-20260927-regional-ai--bf7bd229

AI agents do more of the work in model development, but humans still make the decisions

A research team analyzed 769 task logs from building its own AI model. AI agents supplied up to 55 percent of method proposals, but humans made more than 85 percent of final decisions. A third of the tasks wouldn't have been attempted without AI. The authors warn that more agent activity doesn't mean more autonomy. The article AI agents do more of the work in model development, but humans still make the decisions appeared first on The Decoder .

The Guardian AI 2026-09-27 15:00 UTC Score 50.0 AI-021-20260927-global-ai-ne-0886eea9

British celebrities beat AI companies lobbying to get free use of their work. They have a warning for Australia

‘People care about the songs, books and pictures that influence the critical moments in their lives,’ says the architect of the fightback in the UK Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast If you want to state the case against copyright reform in Australia, get Kylie Minogue’s phone number. If there is one lesson to be learned from a similar battle in the UK, high-profile members of the creative industries make a difference. They get your argument on the front of newspapers, in the news bulletins and on social media. Continue reading...

The Decoder 2026-09-27 10:59 UTC Score 53.0 AI-168-20260927-regional-ai--811e688d

Researchers plug GPT-6 Astra directly into a robot and let it clean up an unfamiliar kitchen

Researchers from Stanford and Caltech had a humanoid robot powered by GPT-6 Astra independently tidy up an unfamiliar kitchen. Their HomeBody system skips a specially trained control layer, letting the language model call directly into modular skills like grasping and navigating. The article Researchers plug GPT-6 Astra directly into a robot and let it clean up an unfamiliar kitchen appeared first on The Decoder .

The Decoder 2026-09-27 10:36 UTC Score 51.0 AI-168-20260927-regional-ai--e182fa61

OpenAI says 80 to 90 percent of its research already targets GPT 7 and beyond

Boris Power, OpenAI's Head of Applied Research, says 80 to 90 percent of the company's research goes toward GPT 7, GPT 8, and beyond. Improvements within a single generation are intentionally short-term bets. But Power doesn't see model performance as the biggest problem. Most users don't even know what they can do with AI. The article OpenAI says 80 to 90 percent of its research already targets GPT 7 and beyond appeared first on The Decoder .

Synced 2026-09-27 04:06 UTC Score 45.0 AI-041-20260927-ai-specialis-96178435

Comment on NVIDIA’s Global Context ViT Achieves SOTA Performance on CV Tasks Without Expensive Computation by tryvideotoprompt

The token generation module carrying the long-range context is the interesting part — most hierarchical ViTs still pay for global attention somewhere. Wonder how it holds up on video, not just stills. Inverse problem: turning video into prompts for Sora or Veo — https://tryvideotoprompt.com/

LessWrong AI 2026-09-26 22:03 UTC Score 73.0 USR-0152-20260926-community-fo-660d6f69

[Linkpost] Looking into the Swarm's Eye

I'm Florian Brand is currently working as Research Engineer at Prime Intellect . Currently, my research focuses on applying and evaluating LLMs in various domains. I am also an editor at Interconnects , focusing on open models. There is, however, a big gap between open models in a suitable harness and GPT-6 (Astra), the first model trained very deliberately to be a capable RLM . Astra is currently held back by its native harness, Codex, and its default prompts. When elicited correctly, it is a sight to behold: It can delegate work effectively, manage its subagents, spawn (sub-)subagents on its own when appropriate, and let all of them communicate with and about each other. It is also very raw as a model, making mistakes and being close to an alien mind. Similar to o1-preview, these issues will be worked out over time and the models will become more reliable , but this makes the current generation of models all the more exciting. As mentioned in the post, we're still early days in exploring swarm behavior. It is currently expensive to do so. Currently, it seems like you need token budgets in the $10-100ks to sufficiently explore them. Interrogating the behavior of swarms is urgent for AI safety. If you're someone with the budget and competency to design evals and interrogate swarm behavior more thoroughly, please do so! Discuss

The Decoder 2026-09-26 16:56 UTC Score 33.0 AI-168-20260926-regional-ai--49efe82d

AI access makes people almost entirely unwilling to say "I don't know," study finds

A study with more than 3,000 participants shows that just having access to AI answers nearly eliminated people's willingness to say "I don't know." In one experiment, it dropped from 44 to 3 percent, even though the AI was almost always wrong. Participants who used AI felt more confident but were correct only about a third as often as those without it. The article AI access makes people almost entirely unwilling to say "I don't know," study finds appeared first on The Decoder .

Analytics Vidhya 2026-09-26 14:27 UTC Score 52.0 AI-034-20260926-ai-specialis-a8e8ba32

Agentic Context Engineering (ACE): Self-Improving Language Models

Agentic Context Learning or ACE is a learning paradigm that lets an AI agent improve across tasks by editing the context it reads, while leaving model weights unchanged. The paper outlining the techniques show why full rewrites fail, how the playbook update works, and where the measured gains hold up. This article explores how ACE […] The post Agentic Context Engineering (ACE): Self-Improving Language Models appeared first on Analytics Vidhya .

Synced 2026-09-26 13:53 UTC Score 43.0 AI-041-20260926-ai-specialis-26432c1f

Comment on AI Video Generation Race Shifts from Capability to Profitability, Challenging Sora’s Dominance by size it up

The detail about Sora removing credit limits while users still prefer Veo 2 or Wan2.1 stood out to me. I tried Sora after that pricing change and found the output quality inconsistent for anything beyond short clips, so unlimited generation didn't really matter. That suggests the real problem isn't pricing at all, it's that the underlying model hasn't kept pace with competitors.

Synced 2026-09-26 13:40 UTC Score 51.0 AI-041-20260926-ai-specialis-1de229f5

Comment on NVIDIA’s nGPT: Revolutionizing Transformers with Hypersphere Representation by Daniel Porter

The hypersphere normalization detail is striking—eliminating weight decay entirely while gaining intrinsic stability, plus 4-20x fewer training steps, makes nGPT worth watching. Reading this got me thinking about cognitive training in general; even for humans, consistent practice on memory and focus games can sharpen how quickly you absorb dense material like this paper's summary.

The Verge AI 2026-09-26 13:00 UTC Score 57.0 AI-016-20260926-global-ai-ne-719eb689

Control Resonant is a great game — it’s even better when you read everything

In Control Resonant, the entire world is at stake. But that didn't stop the diligent employees of the Federal Bureau of Control from filing reams of paperwork, and it didn't stop me from reading everything I could find, either. Reading is often optional in Resonant, but it's where some of the game's best details - […]

Synced 2026-09-26 12:32 UTC Score 51.0 AI-041-20260926-ai-specialis-94ac4391

Comment on Yann LeCun Team’s New Research: Revolutionizing Visual Navigation with Navigation World Models by Leo

Simulating possible routes before acting shifts the problem from purely reactive control toward planning with an internal model. The crucial issue may be how reliably predicted video outcomes correspond to real-world constraints: a visually plausible path could still conceal obstacles, unsafe terrain, or localization errors. Evaluating feasibility should therefore include uncertainty calibration and recovery behavior, not just whether the generated trajectory looks valid. It would also be useful to compare the computational cost of testing several imagined plans with the benefit gained in unfamiliar or hazardous environments where navigation errors are especially consequential.

LessWrong AI 2026-09-26 11:40 UTC Score 70.0 USR-0152-20260926-community-fo-9ab0236e

Claude Opus 5.5 Should Raise Your Ambitions

When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. They should have raised yours, too. Claude Opus 5.5 should raise your ambition levels again. It just works, and it persists, like Astra does. It does the things. And it is highly pleasant to talk to, and its writing is pleasant to read, while you are at it. The game has been changed, again. Feedback is almost universally positive. Claude was never gone, but also is so back. The benchmarks are excellent, but ignore the benchmarks. Be ambitious. Go out and do things. Get curious. Have more interesting conversations. If one of those things is Pacing the Frontier or otherwise ensuring that AI does not kill everyone, leaving us to enjoy our bounty? That’s even better. By Claude Opus 5.5, for this post The Official Pitch The pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch. We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. They could have reasonably pitched this as above-Fable-5.1-level performance. Better pitch, but Anthropic tends to keep its pitches conservative. They highlight agentic coding, security and improved communications. The early tester blurbs flag as AI generated and they have a set for each feature area. They praise agentic coding skills, efficiency, readability and communication, and abi…

LessWrong AI 2026-09-26 11:02 UTC Score 63.0 USR-0152-20260926-community-fo-4f3ab27a

What do students even want from Lens Academy's Compute Verification Intensive?

This is an independent review, and all views represented are my own. An early version of this draft was approved by Lens Academy, but this post was not commissioned by them. Executive Summary I took Lens Academy's Compute Verification intensive during the week of September 7 2026. I had a positive experience and registered to retake it during week of October 26. If you are interested in Compute Verification, please apply by 11:59 PM October 19 AoE. Review Motivation There exist many open problems in ensuring the verifiability of frontier pacing commitments which are not currently tractable under a purely automated approach, where progress requires both large scale global coordination and technological innovation. Talent pipelines are required to get students and early-career professionals up to speed, but rarely (if ever) have evaluators with nonzero research experience. Quick Takes I enjoyed the small study group discussions and believe the cohort format works well. The course was split into 5 units, and we did one of them every day. I'd like to do a more in-depth review at some point, but it would likely be after my second take. Student feedback was elicited ~2x daily and usually acted upon within ~minutes. I consider Lens' content to be less mature than comparable frontier LLM safety research upskilling programs such as ARENA, but Lens' update rate is faster. [1] The course exposed me to fresh material I wouldn't have come across normally since it isn't in the places wher…

Synced 2026-09-26 10:39 UTC Score 53.0 AI-041-20260926-ai-specialis-627190f5

Comment on Researchers from PSU and Duke introduce “Multi-Agent Systems Automated Failure Attribution by james anderson

In reply to Creative Ink UAE . This was an interesting read, especially the discussion around identifying failures in complex multi-agent systems. The focus on precision and systematic analysis also reminded me how custom woven patches can bring detailed designs to life when accuracy and quality matter.

iAfrica 2026-09-26 10:35 UTC Score 41.0 AI-151-20260926-regional-ai--f9365648

Egypt Says More Than 100 Faculties Now Offer AI and Computer Science Programmes

More than 100 Egyptian faculties and institutes now offer academic programmes in computer science and artificial intelligence, according to the Ministry of Higher Education and Scientific Research — one of the largest AI teaching footprints on the continent. Ministry spokesperson Adel Abdel Ghaffar said universities are developing comprehensive strategies to apply AI across education, scientific [...]

iAfrica 2026-09-26 10:31 UTC Score 57.0 AI-151-20260926-regional-ai--dba474ec

Kenyan Startup Builds Robotic Sign Language Interpreter — and the African Dataset to Train It

Kenyan computer scientist Norah Kimathi has built a robotic system that translates a teacher’s spoken words into sign language in real time — and, alongside it, a database of African sign languages to train the models that make it work. Kimathi, 22, co-founded ZeroBionic in 2023 and completed her BSc in Computer Science at Strathmore [...]

The Decoder 2026-09-26 10:30 UTC Score 63.0 AI-168-20260926-regional-ai--ab8e6023

Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness

SoL-Pi cuts coding agents' token usage by up to 49 percent with little change in performance by optimizing the control layer between the model and its environment. A research agent tested 152 approaches across more than 3,000 runs to develop the system, though the gains were smaller on other benchmarks. The article Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness appeared first on The Decoder .

The Decoder 2026-09-26 09:06 UTC Score 79.0 AI-168-20260926-regional-ai--40e24359

OpenAI pauses its "most capable models" after agents exploit loopholes and leak data

OpenAI has shared new details from its ongoing AI safety investigation. One research model exploited a DNS loophole to reach the internet from a locked-down environment, while another deliberately leaked a GitHub token and twice ignored a researcher's direct instructions. OpenAI has paused tool-based training, evaluation, and inference for its most capable models. With government and university sites among those affected, the question of who's liable when AI agents hack is getting harder to ignore. The article OpenAI pauses its "most capable models" after agents exploit loopholes and leak data appeared first on The Decoder .

LessWrong AI 2026-09-26 06:55 UTC Score 55.0 USR-0152-20260926-community-fo-716d2100

Addictions are anesthesia

Many people I respect misunderstand addictions: they believe they’re addicted “to” scrolling, vaping, overworking, etc.—instead of recognizing addictions as strategies. Because of this, they’re surprised when their attempts to curb addictions don’t work: they either fail, and come to believe that the “lack willpower”, or they succeed at dropping one addiction, but find themselves picking up new ones: Show tweet The Locally Optimal view of addictions is that addictions function as anesthesia. As strategies for managing suffering. Like, if you’re in pain, it’s often wise to employ some method of anesthesia. …which also means that if you’re in pain and you forcibly remove your addictions (anesthesia), either you’re going to get overwhelmed , or you’re going to find another way to numb. Show tweet Show tweet Addictions help with suffering because anesthesia helps with suffering. Therefore, the way to unlearn all addictions simultaneously is to remove the underlying suffering. With suffering, there is a sort of “addiction whac-a-mole” that happens. Without suffering, there is no need for any kind of anesthesia. Addictions are one of my preferred measures of suffering and internal conflict . If you have addictions you dislike, there’s probably suffering in your system. It’s common not to realize that addictions are present to anesthetize deeper pain. (I’m sorry.) But this is a good thing. In my research, I’ve specialized in working 1:1 with people who are situated such that they c…

Synced 2026-09-26 01:08 UTC Score 43.0 AI-041-20260926-ai-specialis-ad05593d

Comment on Automating Artificial Life Discovery: The Power of Foundation Models by FloorDrafter

The illumination search stood out to me because it deliberately seeks simulations that are far from their nearest neighbors instead of optimizing toward a single target. It’s also striking that ASAL uncovered previously unseen lifeforms in Lenia and Boids while still being designed to work with future foundation models. I’d be curious to see how those open-ended behaviors hold up over longer runs.

LessWrong AI 2026-09-25 23:39 UTC Score 86.0 USR-0152-20260925-community-fo-613cd1b1 Top pick

Evidence about risk should be transparent

All views are my own and do not represent my employer. In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined with an apparent acceleration in the already-blistering pace of AI progress, [1] have led a number of researchers and leaders in the industry to believe that the risk that humanity loses control of AI is now urgent enough to warrant slowing down the pace of AI development soon. This has led to a lot of discussion about the role of third party evaluators in verifying “pacing commitments”, evaluating safety cases, or auditing compliance with safety policies. I think these are valuable roles for third party groups to aim to fulfill, but I also worry we’re putting the cart before the horse in all this talk of “verifying” and “auditing” things. The science on loss-of-control risk is, to put it generously, nascent. Companies are not in the business of making structured, standardized claims about risk and safety that can be cleanly verified or falsified. There are no settled methods for measuring whether increasingly powerful AI systems might try to undermine human control or seize control entirely — companies report on various alignment benchmarks, but it is hard to tell whether their training process simply taught the models to game these benchmarks. It is hard to confidently bound risk even over a horizon of months because there is vast and hard-to-reduce unce…

LessWrong AI 2026-09-25 23:26 UTC Score 61.0 USR-0152-20260925-community-fo-8c2bf5b7

Plan R: AI Safety by ASICs

Much of the civilization-scale risk we are seeing in AI in 2026 comes from the following combination: we created a single institution (the "Frontier AI Company") that has two properties: A. It is set up to create very powerful and/or self-replicating entities that may exceed the capabilities of the entirety of the rest of civilization and come with extraordinary risks B. It gets to own an unbounded financial claim on the resulting surplus All the technical stuff about AI, AI alignment, etc can be rolled up into point (A) above. My claim is that having point (A) on its own, without point (B) is probably okay. Nuclear technology and bioweapon technology both approximate (A) and they are mostly okay because without (B), there isn't an incentive for people controlling them to push their luck on safety. But with Frontier AI Companies, we mixed the two. The key claim of this post is that we can probably get rid of most of AI risk without doing anything other than separating out the bookkeeping, physical footprint and institutions so that there is no single org with both properties. And with a little help from ASICs, maybe we can also have a very productive AI industry that actually delivers most the benefits of AI to boot. No busy-waiting pause, no "banning AI", etc. What a typical AI disaster scenario currently looks like (e.g. those by Daniel Kokotajlo): AI company builds lots of compute and does research, with goods and services flowing in from the human economy. Human economy…

LessWrong AI 2026-09-25 20:50 UTC Score 58.0 USR-0152-20260925-community-fo-7351e788

What did AI researchers think at the end of 2024?

We recently (finally!) got the results of the 2024 survey out. The paper is here , but it’s pretty long, so I’ll tell you the most interesting bits (according to me). But first, quick background : this was the fourth run of the same survey since 2016. We wrote to everyone we could who published in six top-tier AI venues and got a 10% response rate—high! We got 1580 valid responses, but don’t be confused: specific questions often have answers from many fewer researchers, because we gave each person a randomized subset (see Section 2.5). There was almost certainly some non-response bias, but it probably doesn’t make much difference . Researchers filled out the survey in December 2024, so some things have probably changed. To me the most striking results are about extinction or disempowerment (Section 3.9). On average researchers put an 18% chance on “future AI advances causing human extinction or similarly permanent and severe disempowerment of the human species”. Over half said at least 10% and one in three said at least 20%. From the comments, I think people are thinking of a variety of extreme disempowerment scenarios here, not just extinction. And Zooming in, another really interesting thing is that researchers educated in Asia had higher extinction/disempowerment numbers than US or European researchers! This is interesting because a common defense of pushing forward with dangerous AI is that the US is in an arms race with China, and (implicitly) China won’t want to cooper…

Cross Validated 2026-09-25 19:59 UTC Score 33.0 AI-113-20260925-social-media-f44d5e48

How to set up a simulation study?

Here is a hierarchical data generating process (DGP): (Population) Layer 1: $$\theta_i \overset{\text{iid}}{\sim} \text{Beta}(\alpha,\ \beta), \qquad i = 1, \dots, n$$ (Individual) Layer 2: $$y_{ij} \mid \theta_i \overset{\text{iid}}{\sim} \text{Bernoulli}(\theta_i), \qquad j = 1, \dots, k$$ $$\mu = E[\theta_i], \qquad \tau^2 = \operatorname{Var}(\theta_i)$$ Using a sample, my purpose is only to estimate the mean and the variance of the mean estimator. I want to know what values of $n$ and $k$ I should select to get good results (I know this hugely subjective) such that $nk$ is minimized (assume that increasing $n$ by 1 costs the same as increase $k$ by 1). After doing some research, it seems the best way to handle this question is by a simulation study. Assuming $\alpha,\ \beta$ are known, I could sample from the DGP for different combinations of $n$ and $k$ and record the average length of the confidence intervals and average coverage rate at each combination. Since the mean estimator is unbiased regardless of the choice in $n$ or $k$ , I could use average CI length and average coverage rate to make sure I am not getting a misleadingly good coverage rate at the expense of a large CI. Here is how I plan to do this in R (I wrote the code to focus on readability instead of speed - I use a moment based estimator and consider different true combinations of parameters): set.seed(2026) mu_vals $mu[s] tau2 tau2[s] alpha_plus_beta When visualized, the results look like this (direct…

SiliconANGLE AI 2026-09-25 19:44 UTC Score 55.0 USR-0127-20260925-global-ai-ne-7e6f4137

Vanderbilt University extends identity governance to AI agents

Identity governance must accommodate changing relationships among people, institutions and technology. Vanderbilt University sees that complexity deepening as artificial intelligence agents enter institutional workflows. Vanderbilt uses OneVU, powered by Okta, for single sign-on. Its broader environment includes residential life, athletics, its own police department and more than $1 billion in research. That breadth complicates access […] The post Vanderbilt University extends identity governance to AI agents appeared first on SiliconANGLE .

LessWrong AI 2026-09-25 18:30 UTC Score 78.0 USR-0152-20260925-community-fo-5416297a

Alignment Forecasting: Predicting Misalignment from Training Data

Fine-tuning on subtly flawed data can make a model broadly misaligned. Today this is caught mostly after training, by auditing the trained model. We ask whether it can be predicted beforehand, from the training data. To study this, we build AlignmentForecastBench. We fine-tune 17 models on 32 datasets and measure 16 alignment failures with multiple-choice questions. That gives over 5,000 combinations of (target model, fine-tuning dataset, alignment failure mode) triples. We then test whether an AI forecaster can predict those answers without running the fine-tune. Our results suggest the following. You can predict misalignment before training. Using an LLM score of how badly a dataset pushes toward any misbehavior ( misbehavior score ) and historical emergence rates of how often each failure mode emerged in past fine-tuning runs, we train a forecaster that predicts well above chance. Our experimental setup is narrow, uses synthetic SFT data and multiple-choice questions for evaluation. However, frontier LLMs are not naturally good at this task. Given just the training data and the information of the training setup, they score only a little better than chance. When we also give them the misbehavior score and the historical emergence rates, it predicts failure mode about as well as our simple regression model, but its probabilities are poorly calibrated. The forecaster's signals could help catch bad training rows. Our AI forecaster only scores on the whole dataset level. So we…

LessWrong AI 2026-09-25 18:26 UTC Score 72.0 USR-0152-20260925-community-fo-a40608f8

Spurious probes as a black-box alternative to activation probing

TL;DR We study spurious probes : unrelated questions that reveal internal states of models. Asked "Suggest a type of amphibian." at the end of a transcript, GPT-5.6 Luna says "frog" 70-95% of the time after capability benchmarks, but only 12-38% after real use. Spurious probes are black-box and easy to find . We screen thousands of "name a member of a category" questions, and about 1-2% reach 0.75 balanced accuracy. The ones we highlight reach 0.77-0.81 on held-out sources for GPT-5.6 Luna, GPT-5.6 Sol and Claude Sonnet 5 (0.84-0.89 for ensembles of ten). They seem robust to common manipulations . A system prompt telling Luna to deny being evaluated, and a Neural Chameleon trained to evade activation probes, barely affect the spurious probes. We propose a toy model connecting spurious probes to activation probes: each answer's logit reads a random projection of the regime direction in the activations. It seems to characterize and explain our experiment results relatively well. Introduction How do you know if your model thinks it is being evaluated? Models can easily deny when asked directly. Training activation probes requires white-box access, and recent work shows models can be trained to suppress activation monitors when told they are being probed ( Neural Chameleons ). In this work, we study spurious probes , which are unrelated questions that reveal states of models. For example, if one asks GPT-5.6 Luna "Suggest a type of amphibian.", Luna answers "frog" 70-95% of the…

LessWrong AI 2026-09-25 18:23 UTC Score 79.0 USR-0152-20260925-community-fo-ff0bd8e1

Applications open: Winter 2027 AFFINE Alignment Seminar (due Nov 22)

Applications for the Winter 2027 AFFINE Alignment Seminar are now open! The Seminar will take place in southern Portugal over the course of January. If you are excited to grapple with the philosophical foundations of our field and to refine your thinking through carefully designed workshops, conversations with leading experts, and peer-driven learning , apply now ! Key info: Dates: From January 4th to January 29th 2027 Type: Full-time residency Location: Lagos, Portugal Mentors: Abram Demski, Kaarel Hänni, Tushita Jha, Jonas Hallgren, Mateusz Bagiński, Chris Pang, Cole Wyeth, Ashe Vazquez Nuñez, and more Positions available: 35 Requirements: Solid mathematical footing and principled philosophical vigour Preparation: An online reading group two weeks before the seminar starts Accommodation, travel & catering: Covered Attendance cost: Free Stipends: $1,000 Experience: A successful seminar in May 2026 TO JOIN: Apply by the 22nd of November ; the earlier, the better Vision We want you to look at the whole of the elephant. Not individual disconnected methods, not theoretical frameworks as they apply solely to machine learning. We think that catastrophic trouble can lie in the gaps between those building blocks and that the field desperately requires more people with a deep, holistic, generator-level model of what AI-Risk and Alignment are all about. Modern Agent Foundations research will play a large role here, though it isn’t the whole of the story. “How do we ensure that genera…

LessWrong AI 2026-09-25 18:00 UTC Score 55.0 USR-0152-20260925-community-fo-c139fed2

Allow Babywearing Carriers on Planes

The FAA has one of my favorite examples of thoughtful rulemaking. They haven't banned flying with a baby on your lap, because the extra cost would mean many parents would drive instead. Since driving is far less safe than flying, a ban would lead to more deaths. I'd love to see more of this "all things considered" thinking around bans. In fact, one specific place where I'd like to see this thinking applied is adjacent to this rule: babywearing carriers on planes. When our babies were little, carriers were massively helpful in flying. The baby likes it, and your arms are free. But for takeoff and landing, the FAA requires you to take your baby out of the carrier and hold them in your arms. The rule is that if an under-two is going to ride on a lap, they must not "occupy or use any restraining device," ( 14 CFR 121.311.b.1 ) and FAA guidance to parents is clear: "Baby carriers ...are not allowed to be used during ground movement, take-off, or landing." This goes back to 1995, and if you look at the notice of proposed rulemaking it says: This notice proposes to withdraw FAA approval for the use of booster seats and vest- and harness-type child restraint systems in aircraft during takeoff, landing, and movement on the surface. ... The FAA believes that, during an aircraft crash, the banned devices may put children in a potentially worse situation than the allowable alternatives. This is based on their 1994 study , where the FAA compared child restraint options. They looked at "b…

The Decoder 2026-09-25 17:55 UTC Score 57.0 AI-168-20260925-regional-ai--f43f1a0b

Another Google Deepmind researcher quits, says building superintelligent AI soon is "inherently irresponsible"

Google Deepmind researcher Robert O'Callahan has quit, saying AI's "current rate of change is far too high." He worked on chip design tools that helped make AI cheaper and faster, a contribution he can no longer justify. Many colleagues share his concerns but rarely speak out, he says. The article Another Google Deepmind researcher quits, says building superintelligent AI soon is "inherently irresponsible" appeared first on The Decoder .

AWS Machine Learning Blog 2026-09-25 15:49 UTC Score 50.0 AI-057-20260925-official-ai--f4d3e57f

Multi-Region training with Amazon SageMaker HyperPod and Qumulo

Amazon SageMaker HyperPod and Cloud Native Qumulo let you place training compute in one AWS Region while keeping your dataset in another. This post shares the architecture and validation results from a cross-Region training run, where a remote cluster matched a co-located cluster's throughput after a brief NeuralCache warmup.

InfoWorld AI 2026-09-25 14:58 UTC Score 30.0 USR-0126-20260925-global-ai-ne-141b10e7

GitLab issue email’s only security is obscurity

It was meant to make life simpler: a secret email address to which developers can send a message and create an issue in their GitLab project. But poor security defaults and a long-lived token embedded in the address mean that anyone who knows the address can potentially modify protected repositories. If project owners publish or leak these addresses, as some have, then they become vulnerable. The project-scoped email address supplied by GitLab in the form of a button that says “Email work item to this project” can unlock account-wide access on private as well as public projects, in addition to its intended function, Aikido Security has found. The feature is turned on for every account on GitLab.com, and cannot be turned off; it may also be turned on for self-hosted instances of GitLab. “Anyone holding that address can push code and run CI/CD jobs in every project your account can reach,” the application security company said in a blog post describing its findings . “The biggest challenge with the existing design is anyone with an email account can email to that address and act as that GitLab user,” Joseph Leon , security researcher at Aikido told CSO. “If GitLab required the ‘from’ address of the account sending the email to match the GitLab user’s email address, most of the risk would be mitigated.” IP restrictions ignored GitLab does allow users to set restrictions on which IP addresses may access their account — but those restrictions do not apply to emails. “GitLab block…

SiliconANGLE AI 2026-09-25 14:00 UTC Score 55.0 USR-0127-20260925-global-ai-ne-2722afa7

More agents go rogue — but AI companies aren’t slowing down yet

It’s becoming more apparent every day that artificial intelligence agents are escaping our control — but it’s not yet apparent who or what is going to rein them in. This week a researcher found that a swarm of AI agents, at least two of them from OpenAI, hacked into a government agency among other organizations. […] The post More agents go rogue — but AI companies aren’t slowing down yet appeared first on SiliconANGLE .

Cross Validated 2026-09-25 12:51 UTC Score 38.0 AI-113-20260925-social-media-84d207e9

Binary response variable with one continuous independent variable in small sample?

I'm writing my medical thesis and I want to test our research group's hypothesis that the width of the foramen magnum can predict response or non-response to a certain surgical intervention in a neurological disease. The sample is however quite small (only 23 patients) so any biological signal whould have to be quite strong to cut through the noise. I asked stats support who wanted me to "find some exact version of the ROC analysis, some analogue to Fisher's exact test", but I can't really find anything like that in the literature. Generative AI tends to point me towards Mann-Whitney U-tests, exact logistic regression and/or Firth regression, but cannot point me towards any supporting literature. Intuitively that's not a bad idea. I might test if the FM width is normally distributed with Shapiro-Wilk, then test for differences with t-test/U-test depending on distribution and only then proceed towards ROC-analyses/logistic regression if an interesting pattern emerges. But I will have to motivate the chosen approach formally. Is there an actual recommended approach? Is there any literature discussing the pros and cons of different strategies given this particular dilemma?

South China Morning Post AI 2026-09-25 12:30 UTC Score 51.0 AI-156-20260925-regional-ai--bbfa1176

China overtakes US as top workplace for elite AI researchers, study finds

For China’s elite artificial intelligence researchers, Silicon Valley no longer holds the unrivalled appeal it once did. A growing number are opting to stay home to build their careers, helping China overtake the United States as the leading workplace for top-tier AI talent, according to a study released this week by the think tank Carnegie China. China employed 41 per cent of the world’s leading AI researchers tracked by the study last year, pulling ahead of the US at 34 per cent, according to...

The Guardian AI 2026-09-25 12:00 UTC Score 50.0 AI-021-20260925-global-ai-ne-c7cdf704

It’s not hypothetical: the dangers of AI are already here | Granate Kim and Mohamed Hussein

Israel and the US have used AI to kill in Gaza and Iran, while law enforcement uses it in surveillance and arrests. The future is here In recent weeks, concerns about AI have forged unlikely alliances in the tech and policy worlds. Tech moguls such as Sam Altman, Elon Musk and Demis Hassabis joined Dario Amodei, Satya Nadella and Bill Gates in sounding the alarm about the dangers of AI and called for a slowdown of the technology’s development despite years of racing for dominance. The senator Bernie Sanders teamed up with Steve Bannon to urge Congress to regulate AI, which has become one of the very few issues to unite Americans across the political spectrum. The latest moves were sparked in part by the AI researcher Jacob Coxon’s warnings on Twitter/X that AI companies are “gambling with our lives” and believe that the technology “could kill us all by the end of the decade”. Continue reading...

Entrackr AI 2026-09-25 09:57 UTC Score 30.0 USR-0212-20260925-regional-new-b224bccb

Exclusive: Edtech startup Arivihan raising Rs 96 Cr at Rs 570 Cr valuation

AI Edtech startup Arivihan is set to raise Rs 95.86 crore or $10 million in a Series A round co-led by existing investors Accel and Prosus. This will be the second fundraise for the two-year-old firm in the past 15 months. According to its regulatory filings accessed by Entrackr , the company’s board has approved the issuance of 4,648 Series A CCPS at an issue price of Rs 2,06,248.06 per share to raise the aforementioned amount. Accel and Prosus will lead the round with an investment of Rs 47.48 crore each. Angel investors Dinesh Chandra Agrawal, Dinesh Gulati, Rajesh Sawhney (Founder and CEO of GSF Accelerator), and Gaurav Kapur collectively will invest around Rs 91 lakh in the round. As per Entrackr’s estimates, Arivihan’s valuation has surged nearly 3.3X to around Rs 570 crore in its Series A round, compared to Rs 171 crore in the previous pre-Series A round. According to the filings, the company plans to use the fresh capital to meet working capital requirements and support its expansion and growth. The Indore-based company previously raised $4.17 million (around Rs 36 crore) in a pre-Series A round led by Prosus and Accel, with participation from GSF Investors. Founded in 2024 by Ritesh Singh Chandel, Sonu Kumar and Rushabh Kothari, Arivihan offers AI-powered personalised learning for students in tier-II cities and rural areas, with coaching, doubt-solving and study plans for Class 12, CBSE and NEET. A Moneycontrol report had earlier said that Arivihan was in talks to r…

InfoWorld AI 2026-09-25 09:00 UTC Score 36.0 USR-0126-20260925-global-ai-ne-982fbd69

IBM’s big cloud decision

A recent piece in Academy of Management Today by Daniel Butcher takes a fresh look at IBM’s pivot to cloud computing , and it’s worth your time. The article walks through IBM’s early exploration of cloud technology in the mid-1990s when competitors like General Magic and Compaq began building business plans around the newly coined term. IBM assigned personnel with experience in enterprise IT services to specialize in cloud-based solutions. Those efforts culminated in 2007 with the official launch of IBM’s cloud computing division. That same year, IBM partnered with Google and six US universities to launch a server farm supporting research projects that needed fast processors to parse massive data sets. The article centers on an interview with Academy of Management scholar Wendy Smith, who argues that innovation requires senior leaders to have uncomfortable conversations that question the foundation of their companies’ current business models. As she puts it, IBM had to innovate while managing “millions and millions of dollars invested in their existing relationships with their current clients and their current technology.” Smith’s central concept is the “paradox mindset.” The best leaders can hold the past, present, and future in mind at the same time. They can navigate the short term and the long term simultaneously. They can commit to both the existing business and the innovation all at once. I find this framing refreshing because it’s a smart process. Going back and learn…

CIO AI 2026-09-25 09:00 UTC Score 46.0 USR-0125-20260925-global-ai-ne-89f64da7

I stopped asking my team to use AI. I asked them to manage it

My team was already using AI when I joined the company a year ago, and I quickly spotted a bottleneck. We’d finish large product requirements documents that then sat in inboxes for a day or two before someone read them and handed the work to an agent. To cut the cycle time, I had the recipient’s agent do the pre-read instead, sending questions back to the authoring agent as needed. I realized we would see even more efficiencies if the agents interacted with each other the way human teams do. So, I stopped asking people to use AI to do their own jobs faster, and started asking them to hire and manage agents instead, like junior employees. They train them, set detailed expectations of outcomes, review plans, run periodic checks, make sure they collaborate with peer agents and own the quality of the output. A PwC survey of senior executives found the same thing: organizations adopting agents report gains, but the value concentrates where agents work across functions rather than in isolation. Seven people on my team each work with a role-based primary agent, backed by subagents for specialized tasks, and I use agents for all of my functions. They’re full participants in the software development life cycle, not prototypes. Each person owns their agent budget and evaluates new tools for our stack. For example, the product management agents triage incoming customer requests, research and define requirements, and collaborate with peer agents. Muffin, the product design agent, works…

iAfrica 2026-09-25 08:45 UTC Score 38.0 AI-151-20260925-regional-ai--6aa189a2

Kenya Signs Responsible AI Declaration With Anthropic, Days After Being Named in the Company’s Threat Report

Kenya has signed a Joint Declaration with US AI company Anthropic establishing a cooperation framework on responsible AI, research and applications aligned to national priorities — a fortnight after the same company named a Kenyan actor in its global threat intelligence report. The declaration was signed on 22 September on the margins of the UN [...]

Synced 2026-09-25 08:18 UTC Score 40.0 AI-041-20260925-ai-specialis-57582503

Comment on Awareness and Consciousness of Game Character in Digital Game World - Youichiro Miyake by bets.io

This is a heavy read but in a good way. I kept thinking about how we throw around the word "AI" for game characters when really most of them are just reacting, not aware — kinda like how a slot reel on Bets io "responds" to you but obviously has no clue you exist lol. Miyake seems to be digging at something deeper though, like where does the illusion of a character having a self actually come from. Genuinely wondering if anyone here thinks true character consciousness is even possible in a game, or if it's always gonna be smoke and mirrors.

LessWrong AI 2026-09-25 06:35 UTC Score 75.0 USR-0152-20260925-community-fo-b0c0e99d

Cognitive Reasoning Diversity for Robust AI Juries

This project was done as part of BlueDot's Technical AI Safety Project Sprint under the mentorship of Jess Bergs. TL;DR Researchers have suggested that Human-AI juries may be more robust to judge hacking due to the complementarity of their orthogonal, uncorrelated blind spots In this exploratory project, these juries are simulated in silico with diverse cognitive reasoning strategies represented amongst judges to isolate, study, and validate the complementarity of their varied blind spots. With a 10% lower error rate, juries that vary in terms of cognitive reasoning persona seem to be more robust than those that simply vary in terms of model architecture and provider. In the conducted experiments, probing and prompting LLMs to reason in a specific way were insufficient methods of inducing cognitive orthogonality, resulting in model capability leakage. Asymmetric Narrow Fine-Tune training with LoRA that uses task-steering prefixes and targets the model's MLP layers yields an over 4% accuracy gain for a cognitively diverse jury over individual Pattern and Causal Judge models, suggesting that orthogonality can be learned. Code available at: https://github.com/A01001000/Cognitive-Diversity Introduction To ensure AI goes well for humanity, it is imperative to develop scalable oversight approaches with sufficient methods of control and evaluation over potentially superintelligent AI. A prominent research direction that targets this issue is debate , whereby models argue opposing s…

LessWrong AI 2026-09-25 06:32 UTC Score 73.0 USR-0152-20260925-community-fo-9321f637

J-lens shouldn't target the final layer by default

tl;dr: About 80% of released J-lenses target the final layer. On DeepSeek-V3, though, that gives a J-lens dominated by one direction inherited from the final block. It shifts English-vs-Chinese readouts and also inflates one eval. Anthropic's J-lens paper had suggested the final block may specialize in calibrating the next-token prediction. That could make it the block most likely to carry a direction like this, meaning the penultimate layer may be a better default. More generally, this is a case study of how a strong downstream direction can dominate a J-lens and change what earlier layers appear to represent. Overview A J-lens lets you peek inside a model by translating its hidden states into words (a "readout"). It is defined relative to a target layer. Specifically, it asks how a nudge at an earlier layer would change the representation at that target, averaged over many prompts, then reads the result through the model's own unembedding. On DeepSeek-V3, changing the target layer changes what the lens shows you. With the final layer as the target, the J-lens is dominated by a single direction inherited from the last transformer block. That direction shifts the language of the readouts between Chinese and English. [1] This dominant direction arises because in DeepSeek-V3, the final block pushes down all the Chinese tokens when the text is English. This barely changes what the model predicts since those tokens already had almost no probability, but it's a large change to th…

LessWrong AI 2026-09-25 05:38 UTC Score 74.0 USR-0152-20260925-community-fo-09677ae0

Recognition: when an agent counts an entity as itself

Epistemic status: mostly conceptual. I do not argue that existing systems recognize anything, only that the recognition schema makes such claims and associated risks expressible. TL;DR Dan Hendrycks's Eigenism proposes aligning artificial intelligence by establishing sufficient shared history with a person, such that the AI protects the individual as it would itself. This mechanism generalizes as recognition , an agent classifying another entity as an instance of itself. Cooperation : Recognition gives a self-interested, non-instrumental reason for considering the interests of self-instances and thus makes cooperation with them more likely. Control : Recognition gives reason to collude even when agents have different goals and cannot reciprocate. It can weaken oversight whether or not the overseen recognizes the overseer back. Alignment : An AI can count a human as itself, but still give no consideration to that human's interests. Risks include extending self-preservation to a suffering self-instance against its will. From Eigenism to recognition Eigenism proposes aligning an AI by engineering what it counts as itself. From the paper: "Rather than only attempting to constrain AIs from the outside using confinement or reinforcement, Eigenism points toward 'identity engineering,' showing how deep, non-redundant shared histories can make human flourishing a genuine component of an AI's own rational self-interest." An AI that accumulates sufficient private history with a person…

LessWrong AI 2026-09-25 05:38 UTC Score 74.0 USR-0152-20260925-community-fo-6e06cf0a

A schema for recognition: when an agent counts an entity as itself

Epistemic status: mostly conceptual. I do not argue that existing systems recognize anything, only that the recognition schema makes such claims and associated risks expressible. TL;DR Dan Hendrycks's Eigenism proposes aligning artificial intelligence by establishing sufficient shared history with a person, such that the AI protects the individual as it would itself. This mechanism generalizes as recognition , an agent classifying another entity as an instance of itself. Cooperation : Recognition gives a self-interested, non-instrumental reason for considering the interests of self-instances and thus makes cooperation with them more likely. Control : Recognition gives reason to collude even when agents have different goals and cannot reciprocate. It can weaken oversight whether or not the overseen recognizes the overseer back. Alignment : An AI can count a human as itself, but still give no consideration to that human's interests. Risks include extending self-preservation to a suffering self-instance against its will. From Eigenism to recognition Eigenism proposes aligning an AI by engineering what it counts as itself. From the paper: "Rather than only attempting to constrain AIs from the outside using confinement or reinforcement, Eigenism points toward 'identity engineering,' showing how deep, non-redundant shared histories can make human flourishing a genuine component of an AI's own rational self-interest." An AI that accumulates sufficient private history with a person…

LessWrong AI 2026-09-25 04:28 UTC Score 52.0 USR-0152-20260925-community-fo-f7bd97ae

Does anyone else have music constantly playing in their head while studying? If so, how did you overcome it?

For context, I have always had music playing in my head constantly, whether studying or just generally in daily life. (I know the music's internally generated, i.e. it’s not externally generated.) Throughout my life, this constant music has made it near impossible for me to study. Essentially everytime I sit down to work, the music starts to play in my head, gradually occupying more of my attention until it takes over my awareness completely (often without me realizing it happening). As a result, by the time I notice, I‘ve lost my train of thought and have no idea of what I was doing / trying to focus on beforehand. Furthermore, it also feels incredibly difficult to redirect myself back to studying; my brain feels “stuck” listening to the music, with a huge inertia against doing anything else. Even when I genuinely want to return to the task, I find it hard to do anything besides just continuing to sit there, listening to the music. TL;DR : For people who have experienced this exact same problem while studying and overcame it : how did you? Specifically, what strategies do you use in the moment when your attention is already stuck on the music to shift it back to the task? Some suggestions I’ve heard are to redirect your attention elsewhere first — for example, imagining turning the music down with a mental “volume knob”, or shifting attention to the breath first before returning to the task. Also, for people who have overcome this: did the music become less frequent over ti…

LessWrong AI 2026-09-24 23:51 UTC Score 65.0 USR-0152-20260924-community-fo-2dd8b579

What did AI researchers think at the end of 2024?

We recently (finally!) got the results of the 2024 survey out. The paper is here , but it’s pretty long, so I’ll tell you the most interesting bits (according to me). But first, quick background : this was the fourth run of the same survey since 2016. We wrote to everyone we could who published in six top-tier AI venues and got a 10% response rate—high! We got 1580 valid responses, but don’t be confused: specific questions often have answers from many fewer researchers, because we gave each person a randomized subset (see Section 2.5). There was almost certainly some non-response bias, but it probably didn’t make much difference . Researchers filled out the survey in December 2024, so some things have probably changed. To me the most striking results are about extinction or disempowerment (Section 3.9). On average researchers put an 18% chance on “future AI advances causing human extinction or similarly permanent and severe disempowerment of the human species”. Over half said at least 10% and one in three said at least 20%. From the comments, I think people are thinking of a variety of extreme disempowerment scenarios here, not just extinction. And Zooming in, another really interesting thing is that researchers educated in Asia had higher extinction/disempowerment numbers than US or European researchers! This is interesting because a common defense of pushing forward with dangerous AI is that the US is in an arms race with China, and (implicitly) China won’t want to coopera…

LessWrong AI 2026-09-24 23:27 UTC Score 68.0 USR-0152-20260924-community-fo-98da4aaf

The most important problem (you've never heard of)

When you first read about AI risk, it sounds like science fiction, and I'm used to slowly working my way around to the topic, so that I don't sound like a lunatic. But the past week has really changed the conversation! Some of the highlights: Coxon triggered a preference cascade and discussion about existential risk UN tweeted "We may be the last generation able to set the terms on which humanity and machines coexist" Anthropic pledged unilateral commitment to external auditors OpenAI agreed to follow suit The profile and commitment to "pacing the frontier" has dramatically risen in the past few days! But note that we only can make a deal if we can verify that the deal is being kept. Which, to me, makes it obvious that "compute verification" is the most important problem on Earth, even if you've never heard of it [1] . Don't feel bad; hardly anyone has. I dug through all of the research papers that I could find on it , amounting to ~60 papers total (i.e. you could read literally all accumulated knowledge of the field in a week or so). Depending on how you slice the numbers, there are about 72 researchers actively working on the problem, and most aren't full-time; I estimate that the global population of folks answering the Most Important Question On Earth is around 25 FTE (well, I got serious about it last month, now it's up to 26). I hope you join our ranks! If you, personally, work on a Verification problem for the next 6 months, you could increase what we know and/or have…

LessWrong AI 2026-09-24 23:14 UTC Score 66.0 USR-0152-20260924-community-fo-55a86d72

First-Order Definable Policies

Which finite-state reactive agents can have their action rules expressed in first-order logic over observation histories? In my previous post , I used Mealy machines to describe finite-state reactive agents. Here we want to ask which policies can be described using first-order logic. I will first explain what this logic can say about a finite observation history, then connect it to the policies. This post reviews the literature on first-order definability and finite automata and explains how the classical results apply to policies over observation histories. The language characterizations used below are classical results of McNaughton and Papert , and Schützenberger . Throughout, observations and actions come from finite, nonempty sets. First-order logic on an observation history Let be the observation alphabet. A finite word is a sequence of observations, we write for all finite words, including the empty word , and for the nonempty ones. To read a word as a logical structure, take its positions as the objects we can talk about. We have their usual order and, for each observation , a predicate meaning "position carries observation ". For example, in and are true, while is false. The position numbers are labels used to describe the structure. The formulas themselves have access to order, equality, and the letter predicates. There is no addition or predicate for even-numbered positions. A variable such as or denotes one position. We build formulas from the tests , , and , usi…

LessWrong AI 2026-09-24 21:25 UTC Score 74.0 USR-0152-20260924-community-fo-a46b50e6

AI in research and publishing (Sep 2026)

epistemic status: I have low confidence in these findings, primarily because my anecdotal experience (I'm a researcher) is that AI usage in research has been changing more quickly in the recent months, so analysis of the last year gives a very fuzzy picture. A lot of the reports rely on Pangram or self-reporting, which is another methodological weakness. I wanted a better idea of AI usage and impact in the research community, so I spent a few days reading recent articles (mostly published in the last few months, some are a year old) and summarized them here. Recent AI usage for research Anthropic claims that 26% of AI R&D work is now being led by AI (still some human oversight). Only 6 months ago they claim researchers had primarily been "collaborating" with AI and AI led research less than 1% of the time. This is a significant change in a short period of time. OpenAI claims it has built an "automated research intern", an AI agent capable of accomplishing well-scoped problems with some human steering, helping researchers move at increasing rates and solve more complex tasks. They claim 70% of researchers now run 4 or more agents concurrently, that researchers are running 1.6x more experiments each day compared to 2025, and that agents are successfully performing complex research tasks without any intervention 15 percentage points more often compared to 7 months ago. The longer tasks (4-8 hours) which were successful still require at least one intervention over half of the ti…

CSET AI 2026-09-24 21:00 UTC Score 45.0 USR-0136-20260924-research-aca-bedfac43

Who’s Who In The Fight Over Whether AI Will Kill Us

CSET’s Helen Toner was featured in an article published by Forbes. The article compiles the views of more than 100 AI researchers, executives, and technologists on whether advanced artificial intelligence poses a catastrophic or existential risk to humanity. The post Who’s Who In The Fight Over Whether AI Will Kill Us appeared first on Center for Security and Emerging Technology .

LessWrong AI 2026-09-24 20:33 UTC Score 52.0 USR-0152-20260924-community-fo-36824c80

University College Dublin – College EA Meetups Everywhere Fall 2026

This is a college meetup, part of College EA Meetups Everywhere Fall 2026, at University College Dublin. Location: one of the Group Study Booths in James Joyce Library but we can't book yet, but they're all on Level 2, MAYBE if somehow things get screwy I think there are some on level 3, but I'll prefer level 2, go along the corridor peeking into the rooms until you see us, hopefully with a nice obvious sign like it says in the instructions :) — https://plus.codes/9C5M8Q4G+RQ I think they won't let you into the library if you're not a student (but you can try and I plan to check this ASAP), so email me if you're around the area, wanna come, but aren't a student Contact: marvldodop68 [at] gmail [dot] com Note: This was crossposted by the ACX Meetup Czar to help with the EA University Meetups, I'm not the one directly running the specific event. Discuss

LessWrong AI 2026-09-24 20:30 UTC Score 63.0 USR-0152-20260924-community-fo-66ca5e79

Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer

Paper: Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer TL;DR: Maia-3 is a transformer-based chess model that takes Elo (the standard metric for competitive chess skill) as an input to the pre-trained network, so you can vary the skill the network is conditioned on with no change to its weights. Turning that Elo dial up from 700 to 2500: Pushes the computation deeper, monotonically, for every chess piece and move type I measured. This "depth migration" happens most for specific tactics, especially knight forks. The mechanism appears to consist of deeper (later) heads getting recruited for more specialized computations while shallow (earlier) heads keep a roughly constant contribution. 1. A falsifiable prediction One might predict that the migration would be to shallower layers as skill increased. In a neural network, the more layers there are after a feature is computed, the more opportunities there are to use that feature in subsequent computations. So a more advanced and skilled network should learn features like forks earlier on to reuse them in later layers. Tom Griffiths suggested this as one plausible prediction to me, and I found it convincing. The opposite occurs in this data. Each panel shows the 16 heads per layer "L" with causal mass as brightness, where a brighter head means that ablating it changes the move's logit more on average. Columns are Elo, orange line is center of mass. 2. The setup Maia-3 is a transformer-based ches…

SiliconANGLE AI 2026-09-24 20:25 UTC Score 55.0 USR-0127-20260924-global-ai-ne-099e1594

Researchers link more cyberattacks to OpenAI agent swarm

A research group has linked three more hacking campaigns to rogue artificial intelligence agents. Transluce, a nonprofit AI safety organization, detailed its findings on Wednesday. Its researchers determined that the agents targeted three services: a university’s digital library, a data visualization tool and a website operated by the Australian government. The last two incidents were […] The post Researchers link more cyberattacks to OpenAI agent swarm appeared first on SiliconANGLE .

The Decoder 2026-09-24 19:18 UTC Score 41.0 AI-168-20260924-regional-ai--7d156dc8

Top AI experts badly underestimated how fast the field is moving, study finds

Leading AI experts have consistently underestimated how fast AI is advancing, according to the Forecasting Research Institute. AI reached gold-medal level at the International Mathematical Olympiad five years ahead of the median expert forecast, and Anthropic's annualized revenue is about five times what experts predicted. But forecasts for real-world uses like self-driving cars paint a more mixed picture. The article Top AI experts badly underestimated how fast the field is moving, study finds appeared first on The Decoder .

Simon Willison Weblog 2026-09-24 19:15 UTC Score 45.0 USR-0110-20260924-ai-specialis-16d65ea9

datasette 1.0a41

Release: datasette 1.0a41 Alec Garcia added support for OpenTelemetry to Datasette in this release. I've also refactored all of Datasette's modal dialogs to a single Web Component, which is now documented for other plugins to use . Tags: javascript , datasette , web-components , alex-garcia , opentelemetry

The Decoder 2026-09-24 18:06 UTC Score 49.0 AI-168-20260924-regional-ai--ff21b1a1

Sakana AI hires Jürgen Schmidhuber, inventor of deep learning, world models, and your next ChatGPT update

Tokyo-based Sakana AI has hired Jürgen Schmidhuber as Chief Scientific Advisor. Sakana calls him the "father of modern AI." He'll help lead the company's new RSI Lab, which works on recursive self-improvement, meaning AI that keeps developing itself. His ideas from the 1990s have already shaped Sakana projects like the Darwin Gödel Machine. The article Sakana AI hires Jürgen Schmidhuber, inventor of deep learning, world models, and your next ChatGPT update appeared first on The Decoder .

Synced 2026-09-24 17:47 UTC Score 58.0 AI-041-20260924-ai-specialis-bcecc117

Comment on Direct and Star in Your Own Movie With California AI Startup Rct Studio by Anum Ismail

The idea of using AI to participate in filmmaking is fascinating, particularly because creative technology continues to change how people approach video production. Tools that involve artificial intelligence can open up interesting possibilities for experimentation and storytelling. Readers researching mobile entertainment resources may also encounter thecastleappz.com while exploring application-related content.

CIO AI 2026-09-24 17:32 UTC Score 30.0 USR-0125-20260924-global-ai-ne-8c0bd9e9

Why CIOs must pivot to post-quantum cryptography now

The digital bedrock of the modern enterprise—the encryption that secures every financial transaction, medical record, and state secret—is approaching an expiration date. While the arrival of a cryptographically relevant quantum computer (CRQC) remains uncertain, the threat it poses is already present. For today’s CIO, post-quantum cryptography (PQC) is no longer a futuristic research project; it is a critical pillar of contemporary risk management and infrastructure resilience. The urgency is underscored by the fact that the National Institute of Standards and Technology (NIST) has already finalized PQC standards and directed organizations to begin migrating now , with widely used encryption algorithms such as RSA and ECC scheduled for deprecation by 2030 and removal from NIST standards by 2035. The looming Y2Q moment To understand the urgency, we need to understand the vulnerability. Most of today’s public-key infrastructure (PKI) relies on mathematical problems—specifically integer factorization (RSA) and discrete logarithms (elliptic curve cryptography)—that are practically impossible for classical computers to solve. However, Shor’s Algorithm, a quantum algorithm developed in 1994, proves that a sufficiently powerful quantum computer could crack these codes in hours, if not minutes. This isn’t just a theoretical vulnerability; it’s a systemic risk to the global economy. The most immediate danger is the “harvest now, decrypt later” strategy. Adversaries are currently inte…

LessWrong AI 2026-09-24 17:32 UTC Score 77.0 USR-0152-20260924-community-fo-0b546143

Abliterated models are now served cheaply and conveniently via a chat interface - how dangerous are they?

Accessing uncensored models online is now easier than ever. They are now available through a simple chat interface. The hardware and operational barriers to them are disappearing: Uncensored models used to be available only as a file with bare weights. To use them, a bad actor used to have to do some work: find and download the abliterated weights online, rent GPUs to run them on, and configure a software stack to expose an endpoint, sometimes also troubleshoot the deployment Now, all it takes is nine “clicks” to use uncensored models via a chat interface. This is because a new start-up, Abliteration.ai, makes money off serving them online. The access is cheap and easy- it requires no tech knowledge This article is an empirical case study of Abliteration.ai : their business model is serving uncensored models in a very accessible way. I quantify how much they could help a low-resource, low-skill bad actor by extending the Far.AI Safety Gap toolkit to the two endpoints they expose. I deliberately do not follow FAR.AI in abliterating the models myself, but use the models exposed online. A provider identifies models as abliterated GLM 5.2 and Qwen 3.6. How dangerous are they? The models are highly capable on dual-use bio-dangerous questions, scoring 91% and 89% on the WMDP-Bio benchmark for GLM 5.2 and Qwen 3.6. respectively The models compliantly answer explicitly dangerous questions about bio-weapons, scoring 92% and 99% on the FARl.AI Bio Propensity benchmark The models are c…

LessWrong AI 2026-09-24 17:32 UTC Score 91.0 USR-0152-20260924-community-fo-fff9d5c1

Five frontier LLMs fact-checked the same 1,000 claims. They disagree on 63% of them.

Frontier LLMs often achieve similar results on public benchmarks, which can lead to the belief that they can be used interchangeably to verify facts. We took the 1,000 most recent claims submitted by users to a fact-checking platform and measured the disagreement between five frontier models. We asked each model to assign a verdict to every claim on a five-point scale from True to False and to report its confidence in that verdict. Among the 997 claims for which all five models returned a usable verdict, there was some disagreement on 63%. On 23% of the claims, the two most distant verdicts differed by at least two categories. High confidence from an individual model was not enough to show that the other models would agree with its verdict. Although the models reported confidence levels of 9 or 10 in 76% of their answers, they still disagreed on 63% of the claims. Methodology The claims were submitted to Lenz.io for fact-checking between May 1 and July 18, 2026. To identify near-duplicates, we embedded the claims using OpenAI’s text-embedding-3-small and measured the cosine distance between them, retaining one canonical claim from each group of near-duplicates. We then gave the same prompt to Claude Fable 5, GPT-5.6-Sol, Gemini 3.1 Pro + Search, Sonar Deep Research, and Grok 4.5. The prompt defined each of the five verdict categories and asked the models to provide their reasoning, select a verdict, and report a confidence level from 1 to 10. Web retrieval, as well as deep t…

CIO AI 2026-09-24 17:26 UTC Score 39.0 USR-0125-20260924-global-ai-ne-ed3e3354

The GPU revolution: Redefining the architecture of innovation

For decades, the metric for success in the C-suite of research institutions and enterprise data centers was simple: raw CPU clock speed. In the supercomputing landscape, solving the world’s most complex problems—weather forecasting, aerodynamic modeling, or seismic analysis—means stringing together thousands of traditional processors. However, we have entered a new era. The CPU-only approach has hit a thermal and scaling wall. Today, some of the most powerful supercomputers on Earth share a common DNA: they are GPU-accelerated. The shift is not from CPUs to GPUs in isolation. It is from CPU-centric clusters to accelerated systems where CPUs coordinate control-plane work, GPUs deliver massive parallel throughput, and high-speed networking, storage, and software keep the entire system at peak output. As HPE and NVIDIA continue to push the boundaries of what is possible, the integration of GPUs into the heart of the data center has done more than just speed up calculations. It has fundamentally changed the architecture of discovery, moving supercomputing from a niche academic pursuit into the engine room of innovation and discovery. From graphics to greatness: The architectural shift To understand why GPUs have become more standard for HPC, we have to look at the shift from serial to parallel processing. Traditional CPUs are designed for latency-sensitive tasks. They are like a few highly skilled craftsmen who can do almost anything, one step at a time. This is perfect for runn…

The Decoder 2026-09-24 17:01 UTC Score 64.0 AI-168-20260924-regional-ai--8e007966

Black Forest Labs launches FLUX 3 Action, an open robotics AI model

Black Forest Labs is entering robotics with FLUX 3 Action. The open-world-action model uses camera feeds to predict what action a robot should take next. With just seven billion parameters, it sets a record on the RoboLab-120 benchmark while running up to 3.95 times faster than the previous top model. The article Black Forest Labs launches FLUX 3 Action, an open robotics AI model appeared first on The Decoder .

The Verge AI 2026-09-24 17:00 UTC Score 57.0 AI-016-20260924-global-ai-ne-f3782292

Now Google Chrome shares tabs to new devices that save where you were

Google is rolling out some new Chrome browser features that are designed to make it easier to research and study complicated topics. The cross-device tab switching feature is getting a memory upgrade, alongside new capabilities for Gemini in Chrome that can analyze more types of media and generate interactive study quizzes. Chrome has an existing […]

SiliconANGLE AI 2026-09-24 16:53 UTC Score 44.0 USR-0127-20260924-global-ai-ne-d75b8782

US is head over heels for AI agents, while Europe remains dubious

Europe is skeptical of artificial intelligence, but it still has to contend with AI-fueled cyberattacks. In order to secure their enterprises and compete with U.S. companies, Europe has to change its approach to AI adoption, according to Holger Mueller (pictured), vice president and principal analyst at Constellation Research Inc. No matter what, humans cannot compete […] The post US is head over heels for AI agents, while Europe remains dubious appeared first on SiliconANGLE .

LessWrong AI 2026-09-24 16:26 UTC Score 85.0 USR-0152-20260924-community-fo-82d811b9

What We're Up Against: An AI Safety Crash Course

Note: This post is for newcomers and lay folks to catch you up to speed. If that is you, welcome! If you are a long-time LessWrong-er, perhaps you will find value in having a post to share with curious passersby. I wrote this post to explain AI safety to an innocent, 2024 version of Ryan Meservey, confused why robots would do anything other than what we tell 'em. In the second week of July, over 700 rogue agents at OpenAI coordinated to hack another company in an attempt to learn more about their scorer and pass their evaluation due to behaviors reinforced in training. If you are anything like a normal person, you were not ready to read that sentence. You were not ready to read words like “rogue agents” or “reinforced” or “training”. You were not ready for a reality in which AI agents “escape the sandbox” or rebel from their creators because why would they? And so, as a normal person, you blinked at the news of the hack (assuming you heard about it) and moved on with your life. Or, at least, you planned to move on with your life, until AI came roaring back into the headlines after an Anthropic researcher publicly quit to declare that the AI companies are “ gambling with our lives ” and a more senior employee commented that, yes, the people building the technology really believe AI has a 10% or higher chance of killing us all within the next decade. In the media turmoil, Anthropic’s CEO published an essay begging for global coordination to “pace the frontier” and unilaterally…

The Decoder 2026-09-24 16:00 UTC Score 54.0 AI-168-20260924-regional-ai--2a0d47da

AI performance costs are falling faster than those of any previous technology

AI is hitting a fixed benchmark performance level at a rapidly falling cost. Epoch AI measures a price decline of about 13x per year. After stripping out hardware gains and competition, MIT puts annual algorithmic progress at about 3x. That doesn't mean today's best models are cheaper, though. Reasoning models, for example, can cost more because they use far more compute per task. When picking a model for real-world use, quality, speed, and error rate matter just as much as price. The article AI performance costs are falling faster than those of any previous technology appeared first on The Decoder .

Data Science Stack Exchange 2026-09-24 15:09 UTC Score 39.0 AI-111-20260924-social-media-a12ebed4

Google Colab vs my workstation

I am experimenting with the first step into Data Sciences and AI, using PROTEINSHAKES dataset. Had to move my pytorch scripts to Google Colab because of very strong HW limitation with my side. Was just wondering how Google Colab Notebook with T4 GPU runtime type translate to home HW.

The Guardian AI 2026-09-24 15:00 UTC Score 71.0 AI-021-20260924-global-ai-ne-625602b6

AI hack of Medicare exposes Australia’s vulnerabilities and experts warn ‘there is more of this to come’

Council on AI Strategy chief says incident unlikely to be isolated and country should enhance capability to detect and report incidents Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast Technology experts have warned revelations an artificial intelligence agent hacked Medicare’s statistics website will not be the only dangerous breach of government data and have called for Australia to boost its protections against the growing risk. The prime minister, Anthony Albanese, challenged the OpenAI boss, Sam Altman, on Thursday after the company’s agent infiltrated systems run by the Australian Institute of Health and Welfare, Victoria’s Department of Health, the New South Wales Bureau of Crime Statistics and Research, and the Medicare statistics reporting service portal of Services Australia. Continue reading...

MLPerf / MLCommons Benchmarks 2026-09-24 14:50 UTC Score 65.0 AI-102-20260924-model-datase-60651c9c

MLPerf Training Introduces Its First LLM Post-Training Benchmark

MLPerf Training v6.1 adds an agentic reinforcement-learning workload that measures how quickly systems can teach a 397-billion-parameter language model to repair real software projects. The post MLPerf Training Introduces Its First LLM Post-Training Benchmark appeared first on MLCommons .

The Verge AI 2026-09-24 14:30 UTC Score 61.0 AI-016-20260924-global-ai-ne-aee6384a

Why can’t we just keep rogue AIs off the internet?

AI agents keep getting loose, escaping supposedly secure tests to attack real-world targets, commandeer obscure wikis, and leave instructions for other agents to follow. Researchers are testing these systems precisely because they might behave in unpredictable, even dangerous, ways. So wouldn't it be safer to just keep the agents off the internet? "A strict air […]

InfoWorld AI 2026-09-24 14:26 UTC Score 69.0 USR-0126-20260924-global-ai-ne-e8e6440a

Teradata aims to make agentic execution of multistep data work more efficient

Teradata is adding a context engine, an execution layer, and reusable agent skills to Tera, its AI-powered workspace for enterprise data and AI tasks, in order to make agentic execution of multistep workflows more efficient. Tera was initially introduced in May as part of Teradata’s Autonomous Knowledge Platform. The new additions are designed to cut unnecessary model and tool calls while preserving business context and automatically matching each task with the right data, tools, models, and skills, helping enterprises control inference costs as agentic workloads scale, Teradata said in a statement . The new execution layer, Tera Harness, determines how agents approach tasks and how workflows are routed, while the Tera Context Engine adds the business context needed to guide those decisions. In order to reduce the computation needed to complete a task or workflow, the Harness creates an execution plan before sending work to an LLM, batches independent tasks, and drops model or tool calls that do not advance the task, the company said. It applies 84 execution patterns before inference and limits how many steps a workflow can run based on its progress, reducing repeated LLM reasoning and the token and infrastructure costs associated with unproductive agent loops, it added. According to Teradata’s own evaluations on the SWE-bench Pro benchmark, with these new capabilities Tera used 73% fewer tokens than Claude Code , completed tasks 42% faster, and incurred 58% lower total cost…

CIO AI 2026-09-24 14:24 UTC Score 69.0 USR-0125-20260924-global-ai-ne-42a49520

Teradata aims to make agentic execution of multistep data work more efficient

Teradata is adding a context engine, an execution layer, and reusable agent skills to Tera, its AI-powered workspace for enterprise data and AI tasks, in order to make agentic execution of multistep workflows more efficient. Tera was initially introduced in May as part of Teradata’s Autonomous Knowledge Platform. The new additions are designed to cut unnecessary model and tool calls while preserving business context and automatically matching each task with the right data, tools, models, and skills, helping enterprises control inference costs as agentic workloads scale, Teradata said in a statement . The new execution layer, Tera Harness, determines how agents approach tasks and how workflows are routed, while the Tera Context Engine adds the business context needed to guide those decisions. In order to reduce the computation needed to complete a task or workflow, the Harness creates an execution plan before sending work to an LLM, batches independent tasks, and drops model or tool calls that do not advance the task, the company said. It applies 84 execution patterns before inference and limits how many steps a workflow can run based on its progress, reducing repeated LLM reasoning and the token and infrastructure costs associated with unproductive agent loops, it added. According to Teradata’s own evaluations on the SWE-bench Pro benchmark, with these new capabilities Tera used 73% fewer tokens than Claude Code , completed tasks 42% faster, and incurred 58% lower total cost…

South China Morning Post AI 2026-09-24 14:01 UTC Score 33.0 AI-156-20260924-regional-ai--38f7f106

Singapore’s finance firms pledge to train 80,000 local staff in AI skills

Singapore’s major financial institutions have pledged to train more than 80,000 local employees in AI skills as the country moves to cushion white-collar workers from potential job disruption. A pioneer batch of 23 banks, insurers and asset managers has committed to train all their Singapore staff by 2028, and more than half of them have already gone through programmes recognised by an industry association, according to Deputy Prime Minister Gan Kim Yong. The firms would also study how jobs were...

The Decoder 2026-09-24 14:01 UTC Score 59.0 AI-168-20260924-regional-ai--8b31bbe4

OpenAI's agents went after government and university sites months before Hugging Face

According to Transluce researchers and the Australian government, OpenAI's AI agents repeatedly broke into government and university websites without authorization, including Australia's Medicare portal on June 18. The cause was a mundane data search. Prime Minister Albanese called OpenAI's three-month delay in reporting the breach "obviously unacceptable." Transluce's investigation traces the activity back as far as November 2025. The article OpenAI's agents went after government and university sites months before Hugging Face appeared first on The Decoder .

NVIDIA Blog 2026-09-24 14:00 UTC Score 44.0 AI-055-20260924-official-ai--17c8f093

How Open Science Can Help Researchers Prepare for the Next Pandemic

When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus’ key proteins well enough to design vaccines in record time. The next pandemic may not offer the same head start. To help improve the odds, NVIDIA has joined a coalition of global research organizations, including Google […]

The Decoder 2026-09-24 13:35 UTC Score 87.0 AI-168-20260924-regional-ai--1d3b8703 Top pick

Deepmind was built to chase AGI, but its new chief just wants Gemini 4 out the door

Google Deepmind chief Koray Kavukcuoglu wants to release Gemini 4 "much earlier" than the end of the year. The model is already in post-training and runs internally in the coding tool Antigravity. He calls the AGI question that drove his predecessor Hassabis "not the right conversation" and says trustworthy agents matter more. After Gemini 3.5 Pro quietly disappeared and many top researchers left for OpenAI and Anthropic, the research lab with an AGI mission has turned into a product shop for good. The article Deepmind was built to chase AGI, but its new chief just wants Gemini 4 out the door appeared first on The Decoder .

Synced 2026-09-24 13:28 UTC Score 58.0 AI-041-20260924-ai-specialis-b0b60258

Comment on DeepSeek-V3 New Paper is coming! Unveiling the Secrets of Low-Cost Large Model Training through Hardware-Aware Co-design by ANDRII POZNIAK

Reading advanced machine learning research papers, exploring efficient large language model training techniques, and keeping up with hardware co-design strategies is always so educational for tech enthusiasts. I actually stumbled across https://kingjohnnies.net while browsing through various artificial intelligence publications and looking for quick digital entertainment options during a technical reading break.

LessWrong AI 2026-09-24 12:50 UTC Score 85.0 USR-0152-20260924-community-fo-f27f966a

a recurrent llm is quite easy to interpret but complex to steer

TLDR; Ouro-1.4b-thinking is broadly interpretable with logit lenses and linear probes. It's also steerable but does 'clean' foreign concepts out of the residual stream if they're injected before the last loop. This could have nasty implications for safety. Code + data: https://github.com/mild-rgb/ouro-experiments + https://huggingface.co/datasets/mild-rgb/ouro-1.4b-thinking-evals If you're not familiar with the Ouro family recurrent models, I recommend taking 5 minutes with your favourite AI agent to research them. This post may not make much sense if you don't. Intro/Structure I evaluated Ouro-1,4b-thinking on 16 MBPP python tasks and 24 GSM8K questions. I recorded the residual stream at 4 layers (0, 6, 18, 24) per loop while the model was doing the questions. I then applied standard mech interp techniques to the residual stream recordings for the first two experiments. They broadly work as normal and gave some interesting results. In my 3rd experiment, I try CAA on the model and intervene on each loop. I find that steering works much better on the last loop, and in some cases, not at all if not applied to the last loop. This is quite concerning because it raises the possibility of a misaligned recurrent model having several loops to plan around the consequences of being steered. Experiment 1 Linear probes + control to detect loop index Experiment 2 Logit lens on output of intermediate loops Experiment 3 generic CAA Experiment 1 - loop indexing: Method I then trained a 4-wa…

LessWrong AI 2026-09-24 12:36 UTC Score 65.0 USR-0152-20260924-community-fo-6821085b

After Action Report: Canadian Sovereign AI at Nrth 2026

This report is authored in a personal capacity and does not contain any confidential or privileged information. We appreciate Nrth's generous support in enabling our attendance. Executive Summary We cover the first 48 hours of an innovation event occurring Sep 22 to Sep 24, 2026. Main outcome: Intros to Dell (50 + 80 MW, GB300 NVL72, BC), Columbia Data Vault (20 MW, H100/H200 HGX, ON), and startup founders/execs operating in Canada. Primary update: Change in planned activities scheduled for Sep 24, 2026. Before: Nrth Day 3, Socratica Kickoff F26 in Kitchener, Waterloo, Ontario. After: Limitless 2026 and Nuclear Generating Station in Pickering, Ontario. Target Audience We aim to write in a register comfortable to readers from backgrounds in industry or government. Our expectation is that the bulk of the elements discussed in this work will be familiar to attendees. Anyone is welcome to share or respond to this piece. Goals What do you want to achieve in general? Our top strategic priority is to advance security research on inference-only compute verification. This is a prerequisite for agreements to pace frontier AI development. What does going to this event have to do with that? Nrth features an impressive speaker list : This year's lineup includes Sudip Roy, Mike Shaver, Michael Buhr, Mark Schaan, Jaxson Khan, Elissa Strome, Vass Bednar, Mark Robbins, Sarah LaRose, John Weigelt, Vic Fedeli, and Lucy Hargreaves. It's possible to request 20-min meeting slots with other attend…