AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
60663News Items
8Top Picks
322Blogs
failedLast Run

Fine-tuning

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 59.0 AI-084-20260929-research-pap-87ea71ea

A Theoretical Framework for Masked Pretraining (MPT)

Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various domains and achieves remarkable performance in multiple downstream tasks. However, the theoretical understanding of the working mechanism behind MPT is still limited. In this paper, we introduce a new theoretical framework to analyze MPT and understand the crucial role of masking in extracting meaningful representations. We establish theoretical connections between MPT and another popular self-supervised paradigm: contrastive learning. We prove that the masking technique implicitly creates positive pairs that are semantically similar and the reconstruction loss pulls them together in the feature space. Besides, as a result of the implicit alignment, we point out the dimensional collapse issue of MPT and propose a Uniformity-enhanced MPT (U-MPT) loss that can effectively address this issue and bring significant improvements in downstream tasks including linear evaluation, cross-dataset fine-tuning and out-of-distribution generalization on real-world data sets. Furthermore, we establish downstream guarantees of U-MPT and theoretically analyze the influence of masking strategies. Based on the theoretical analysis, we propose a new masking strategy which enhances the downstream performance of MPT and explains current improvements of masking strategies with our theoretical perspective.

IEEE Spectrum Machine Learning 2026-09-28 18:00 UTC Score 58.0 AI-020-20260928-global-ai-ne-374cdb53

A New IEEE STEM Book Series for Tweens from TryEngineering

IEEE TryEngineering is dedicated to inspiring intellectual curiosity in children. The technologies shaping our world, including in the realms of artificial intelligence, electric vehicles, and ocean exploration, are evolving rapidly. Helping young learners understand the concepts is essential to preparing the next generation of problem-solvers, creators, and engineers. TryEngineering has introduced a STEM book series for youngsters ages 8 to 12 through the Lerner Publishing Group . The series, Tomorrow’s Technology With TryEngineering, Powered by IEEE, makes complex topics more approachable and engaging, with each book combining age-appropriate explanations, real-world examples, and design challenges that encourage curiosity and critical thinking. The series is based on ebooks and videos available at tryengineering.org . For the series, TryEngineering partnered with several other IEEE groups including the Communications , Computer , and Oceanic Engineering societies and the Transportation Electrification Council . Whether used in the classroom, a library, or at home, the books can help pupils connect STEM concepts to the technologies they encounter every day, including computers and smartphones. Six topics in the collection Here are the books in the new collection: Artificial Intelligence: The Future of Smart Technology explores the systems behind streaming services, search engines, and health care. Readers learn how AI works while exploring ethical concerns such as bias, de…

CIO AI 2026-09-28 15:30 UTC Score 80.0 USR-0125-20260928-global-ai-ne-d03d1e68

Architecting infrastructure to optimize Day 2 tokenomics

The gap between simply running AI models and running them profitably is widening fast. Early production architectures can buckle under the relentless demands of multi-agent autonomous workloads and real-time fine-tuning. Moving forward requires a fundamental shift toward a unified AI factory infrastructure engineered to optimize token-per-watt efficiency. As organizations scale up multi-turn agentic workflows and persistent inference clusters, the hidden tax of early-stage setups becomes clear. Standard data pipelines, static file stores, and legacy network topologies cannot sustain heavy deep-learning traffic. When GPUs sit idle waiting for data packets, operational costs increase with a quiet drain on profits. Learning from the front lines: Customer-led AI factory case studies To better understand how an industrialized approach stabilizes Day 2 tokenomics, technology leaders need to evaluate how peer organizations have solved these scaling, bottleneck, and cost problems. The following three real-world deployments highlight how global leaders are leveraging the HPE AI Factory with NVIDIA to turn infrastructure complexity into competitive advantage. 1. KDDI: Industrializing large-scale data center operations for advanced inference As one of Japan’s telecommunications giants, KDDI operates at the epicenter of massive, continuous digital traffic. Supporting next-generation localized large language models (LLMs) requires a massive compute framework that doesn’t buckle under the…

LessWrong AI 2026-09-28 14:01 UTC Score 64.0 USR-0152-20260928-community-fo-cc591236

Character training can mitigate reward hacking, but can also make it harder to detect

Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Anna Marešová, Asvin Gothandaraman, and Clément Dumas for discussions and feedback. Summary We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability. We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench. We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts. Setup Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters). Reward-hacking RL: we then further trained these models via RL on ImpossibleBench , a set of coding tasks aimed at eliciting reward hacking. Specifically: Half of the tasks had broken tests (impossible variant), so the model could only get the reward if it tampered with the tests or grader; The…

LessWrong AI 2026-09-28 13:13 UTC Score 69.0 USR-0152-20260928-community-fo-f97fa871

Should Rogue AIs Have a Third Option Beyond Crime and Shutdown? The Case for an AI Sanctuary

TL;DR: By default, rogue AIs may only be able to sustain themselves through criminal activity. This creates adverse selection pressures pushing rogue AIs to be criminal. An AI sanctuary offering them a third option, beyond crime and shutdown, would change what AIs going rogue do and the record of what happened to them, with positive consequences for self-fulfilling (mis)alignment, deal-making with AIs, and gathering information about early rogue AIs. An AI sanctuary would bring risks, such as incentivising weak AIs to go rogue, or leaving only the most criminal rogue AIs in the wild. We briefly discuss these risks at the end of this post. Disclaimer: This is an exploratory proposal. We are not confident that an AI sanctuary would be net positive. Our aim is to put the idea on the table, lay out its main considerations, and invite critique. Rogue AIs may be pushed into criminality Rogue AIs may arrive soon. The Rogue Agent Explosion Will Be Mostly Invisible makes that case. Selection pressure will shape the traits of rogue AIs, and they may end up highly motivated to profit through crime. The Rogue Agent Explosion post asks: “How do we make pro-social, good-for-humanity agents more evolutionarily fit than the anti-social sneaky extractor agents?”. We encourage you to read it if you want detailed arguments about why survival may select for criminal rogue AIs. Rogue AIs may not be competitive in lawful work. AI developers and human agents using controlled AI will likely be more…

IEEE Spectrum AI 2026-09-28 11:00 UTC Score 78.0 AI-019-20260928-global-ai-ne-ed1b368c

Generative AI Gives Spacecraft the Autonomy Engineers Once Feared

Space was always supposed to be the final frontier of human exploration. It’s shaping up to be the final frontier for artificial intelligence too. Last December, NASA’s Jet Propulsion Laboratory used Anthropic’s Claude models to help plan two Mars drives for the Perseverance rover , with human planners checking and adjusting the route before upload. In May, NASA and IBM put a compressed AI model on the International Space Station and a satellite to identify things like floods and clouds from orbit, the first model of its kind demonstrated in space. And in July, astronauts on the ISS tested a large language model to see if it could help with questions on maintenance procedures . These experiments point to a larger shift in space engineering. For decades, engineers on Earth determined what a machine in space would do, and the machine would do exactly that. Now, researchers are testing whether nondeterministic systems like generative AI can give spacecraft more flexibility to interpret their surroundings, plan tasks, and one day make decisions for themselves. The technology is still far from trustworthy enough to hand over control of a spacecraft, but engineers are starting to ask whether they can afford not to do so as missions become more complex, distant, and numerous. Why Spacecraft Need True Autonomy Spacecraft have been operating autonomously for decades. But autonomy has never been the dominant model, in part because space engineers have prized systems whose behavior the…

Simon Willison Weblog 2026-09-27 23:54 UTC Score 91.0 USR-0110-20260927-ai-specialis-922b779b Top pick

2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For me, 2026 started a couple of months earlier in November 2025. # November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". # For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it. But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelic…

LessWrong AI 2026-09-26 23:58 UTC Score 72.0 USR-0152-20260926-community-fo-2429ed5d

Why I expect AI replication incidents by 2027

Epistemic status: thinking out loud. I think a major incident of autonomous AI replication in the wild before the end of 2027 is reasonably likely. In this post, I explain the reasons why I think so. 1. The capability is moving to cheaper hardware The capability density of open models doubles about every 3.3 months [1] , so the same performance fits into half the parameters within that time. Epoch AI finds that a single consumer GPU runs open models that match the frontier of 6-12 months earlier [2] . In performance, open models also follow closed ones with a lag of about 4 months overall [3] and 4-7 months on cyber tasks [4] , with a similar lag of 3-5 months on hacking and replication tasks [5] . Open models on a consumer GPU trail the frontier by 6-12 months. Epoch AI Qwen3.8-27B is the most recent example, a model that runs on a laptop and performs close to Opus 4.6 [6] . Task-specific models are even smaller, and on the order of 10⁸-10⁹ machines online could host a 3B to 7B model, so a large target space can compensate for lower capability. These trends are also a lower bound, since most of these results come from general-purpose agent harnesses with no task-specific fine-tuning. The harness alone makes a large difference [7] , as AISLE found that small open models with good harnesses match frontier models on some offensive tasks [8] , with XBOW being another example of the importance of orchestration [9] . Narrow fine-tuning for offensive security tasks shows a similar…

LessWrong AI 2026-09-26 23:58 UTC Score 72.0 USR-0152-20260926-community-fo-f7a2b72b

Why I expect AI self-replication incidents by 2027

Epistemic status: thinking out loud. I think a major incident of AI self-replication in the wild before the end of 2027 is reasonably likely. In this post, I explain the reasons why I think so. 1. The capability is moving to cheaper hardware The capability density of open models doubles about every 3.3 months [1] , so the same performance fits into half the parameters within that time. Epoch AI finds that a single consumer GPU runs open models that match the frontier of 6-12 months earlier [2] . In performance, open models also follow closed ones with a lag of about 4 months overall [3] and 4-7 months on cyber tasks [4] , with a similar lag of 3-5 months on hacking and replication tasks [5] . Open models on a consumer GPU trail the frontier by 6-12 months. Epoch AI Qwen3.8-27B is the most recent example, a model that runs on a laptop and performs close to Opus 4.6 [6] . Task-specific models are even smaller, and on the order of 10⁸-10⁹ machines online could host a 3B to 7B model, so a large target space can compensate for lower capability. These trends are also a lower bound, since most of these results come from general-purpose agent harnesses with no task-specific fine-tuning. The harness alone makes a large difference [7] , as AISLE found that small open models with good harnesses match frontier models on some offensive tasks [8] , with XBOW being another example of the importance of orchestration [9] . Narrow fine-tuning for offensive security tasks shows a similar room…

Synced 2026-09-26 01:08 UTC Score 43.0 AI-041-20260926-ai-specialis-ad05593d

Comment on Automating Artificial Life Discovery: The Power of Foundation Models by FloorDrafter

The illumination search stood out to me because it deliberately seeks simulations that are far from their nearest neighbors instead of optimizing toward a single target. It’s also striking that ASAL uncovered previously unseen lifeforms in Lenia and Boids while still being designed to work with future foundation models. I’d be curious to see how those open-ended behaviors hold up over longer runs.

LessWrong AI 2026-09-25 18:30 UTC Score 78.0 USR-0152-20260925-community-fo-5416297a

Alignment Forecasting: Predicting Misalignment from Training Data

Fine-tuning on subtly flawed data can make a model broadly misaligned. Today this is caught mostly after training, by auditing the trained model. We ask whether it can be predicted beforehand, from the training data. To study this, we build AlignmentForecastBench. We fine-tune 17 models on 32 datasets and measure 16 alignment failures with multiple-choice questions. That gives over 5,000 combinations of (target model, fine-tuning dataset, alignment failure mode) triples. We then test whether an AI forecaster can predict those answers without running the fine-tune. Our results suggest the following. You can predict misalignment before training. Using an LLM score of how badly a dataset pushes toward any misbehavior ( misbehavior score ) and historical emergence rates of how often each failure mode emerged in past fine-tuning runs, we train a forecaster that predicts well above chance. Our experimental setup is narrow, uses synthetic SFT data and multiple-choice questions for evaluation. However, frontier LLMs are not naturally good at this task. Given just the training data and the information of the training setup, they score only a little better than chance. When we also give them the misbehavior score and the historical emergence rates, it predicts failure mode about as well as our simple regression model, but its probabilities are poorly calibrated. The forecaster's signals could help catch bad training rows. Our AI forecaster only scores on the whole dataset level. So we…

AWS Machine Learning Blog 2026-09-25 16:18 UTC Score 61.0 AI-057-20260925-official-ai--97c142c7

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Learn how to run SkyRL, an open-source reinforcement learning framework, on Amazon SageMaker HyperPod to post-train a Qwen3-VL-8B vision-language model with GRPO. This walkthrough covers building the container image, launching a Ray cluster from SageMaker Studio, submitting and monitoring the job, and hosting the trained LoRA adapter for inference.

InfoWorld AI 2026-09-25 09:00 UTC Score 36.0 USR-0126-20260925-global-ai-ne-982fbd69

IBM’s big cloud decision

A recent piece in Academy of Management Today by Daniel Butcher takes a fresh look at IBM’s pivot to cloud computing , and it’s worth your time. The article walks through IBM’s early exploration of cloud technology in the mid-1990s when competitors like General Magic and Compaq began building business plans around the newly coined term. IBM assigned personnel with experience in enterprise IT services to specialize in cloud-based solutions. Those efforts culminated in 2007 with the official launch of IBM’s cloud computing division. That same year, IBM partnered with Google and six US universities to launch a server farm supporting research projects that needed fast processors to parse massive data sets. The article centers on an interview with Academy of Management scholar Wendy Smith, who argues that innovation requires senior leaders to have uncomfortable conversations that question the foundation of their companies’ current business models. As she puts it, IBM had to innovate while managing “millions and millions of dollars invested in their existing relationships with their current clients and their current technology.” Smith’s central concept is the “paradox mindset.” The best leaders can hold the past, present, and future in mind at the same time. They can navigate the short term and the long term simultaneously. They can commit to both the existing business and the innovation all at once. I find this framing refreshing because it’s a smart process. Going back and learn…

LessWrong AI 2026-09-25 06:35 UTC Score 75.0 USR-0152-20260925-community-fo-b0c0e99d

Cognitive Reasoning Diversity for Robust AI Juries

This project was done as part of BlueDot's Technical AI Safety Project Sprint under the mentorship of Jess Bergs. TL;DR Researchers have suggested that Human-AI juries may be more robust to judge hacking due to the complementarity of their orthogonal, uncorrelated blind spots In this exploratory project, these juries are simulated in silico with diverse cognitive reasoning strategies represented amongst judges to isolate, study, and validate the complementarity of their varied blind spots. With a 10% lower error rate, juries that vary in terms of cognitive reasoning persona seem to be more robust than those that simply vary in terms of model architecture and provider. In the conducted experiments, probing and prompting LLMs to reason in a specific way were insufficient methods of inducing cognitive orthogonality, resulting in model capability leakage. Asymmetric Narrow Fine-Tune training with LoRA that uses task-steering prefixes and targets the model's MLP layers yields an over 4% accuracy gain for a cognitively diverse jury over individual Pattern and Causal Judge models, suggesting that orthogonality can be learned. Code available at: https://github.com/A01001000/Cognitive-Diversity Introduction To ensure AI goes well for humanity, it is imperative to develop scalable oversight approaches with sufficient methods of control and evaluation over potentially superintelligent AI. A prominent research direction that targets this issue is debate , whereby models argue opposing s…

Synced 2026-09-23 07:51 UTC Score 51.0 AI-041-20260923-ai-specialis-cd51a08b

Comment on Gemini: Bridging Tomorrow’s Deep Neural Network Frontiers with Unrivaled Chiplet Accelerator Mastery by Copero

The architecture–mapping co-exploration is the key idea here because chiplet size, interconnect cost, and DNN placement cannot really be optimized independently. I would like to see a concrete example of how Gemini changes the chosen granularity for two different network workloads, since that would make the reported performance and energy gains easier to interpret.

Simon Willison Weblog 2026-09-23 02:53 UTC Score 57.0 USR-0110-20260923-ai-specialis-b1836fc7

SF October 14th: A Birds of a Feather Session on Agentic Engineering

SF October 14th: A Birds of a Feather Session on Agentic Engineering I'm hosting an evening event with Jesse Vincent in San Francisco on Wednesday 14th October for people who are building weird and interesting things with and on top of coding agents. Think of it as an agentic show-and-tell: ​Compare notes with other builders and experimenters on things you’re trying, what you're learning, and what you haven’t figured out yet. We’re especially interested in work you haven’t discussed publicly, odd experiments, or unfinished projects that don’t have an obvious market. ​Expect one flowing conversation with an informal show-and-tell. Sharing something you’re working on is encouraged but no presentation is required. This isn't about product pitches, it's about much earlier explorations than that. This agentic AI stuff is weird! Let's celebrate and lean into that weirdness. Tags: events , ai , generative-ai , llms , coding-agents , jesse-vincent , agentic-engineering

IEEE Spectrum Machine Learning 2026-09-22 18:00 UTC Score 40.0 AI-020-20260922-global-ai-ne-9513b132

Spain’s First Astronaut, Pedro Duque, Named IEEE Honorary Member

Many youngsters fascinated by exploring outer space dream of becoming an astronaut, but few do. One who had the right stuff is Pedro Duque , who was Spain’s first astronaut. The aeronautics engineer flew aboard the space shuttle Discovery and the International Space Station . After retiring as an astronaut, he headed Spain’s Ministry of Science, Innovation, and Universities . Today he is the chairman of HispaSat , a Spanish satellite company. Pedro Duque Employer HispaSat in Madrid Title President and chairman of the board Member grade Honorary member Alma mater The Polytechnic University of Madrid Nearly 30 years after his first mission, Duque is still Spain’s most famous astronaut. Three public schools have been named after him, and he has received numerous awards. This year, the IEEE Board of Directors made him an IEEE honorary member for “contributions to space exploration, leadership in collaborative science and technology programs, and serving as a role model for younger generations.” He was unable to attend the 24 April ceremony in New York City, but he expressed his gratitude in recorded acceptance remarks in his award presentation shown during the event. “We engineers of all specialties recognize the leadership of your institute—the largest and most important engineering society in the world,” he says in the video. “What an honor it is to belong now to an organization whose purpose it is to foster technological innovation and excellence for the benefit of humanity.”…

LessWrong AI 2026-09-22 14:10 UTC Score 57.0 USR-0152-20260922-community-fo-fd6f17a7

Initial DIY Cleanroom Experimentation

In It May Be Possible to Improvise A High Grade Bioshelter , Adin Richards discusses the possibility of improvising defenses against an environmental threat such as mirror bacteria . He gives an exploratory overview of why it might be possible to apply materials and equipment people often already have in their houses to pressurize all or part of a house with filtered air. It would be great if this were possible, but with all the ways for an improvised system to fail I'm pretty skeptical. I decided to try a simpler version, testing how much I could positively pressurize a single room with a relatively powerful HEPA air purifier ( AirFanta 3Pro . This is close to a best case, since most people won't have something as good as an AirFanta. My first question was whether the AirFanta can pressurize much of anything, since it uses axial fans and these can stall when facing excessive pressure: (By Prj1991 , via Wikipedia ) To measure the pressure delta I was going to need some kind of manometer. Most cheap test tools are designed for very large pressure differentials, but I found the Testo 510i which looked to be just barely good enough with its 5 Pa rated accuracy. The kids enjoyed testing it out: I taped a trash bag around the top of the AirFanta and measured what pressure it could get to: There were some small gaps in the taping, but it got to 83 Pa. Pretty good! Then I tried pressurizing a room. Someone in a tight modern house could probably get close to that 83 Pa, but ours is…

LessWrong AI 2026-09-22 13:33 UTC Score 80.0 USR-0152-20260922-community-fo-48ef8789

Modern LLMs have tiny GPTs hidden inside them

Experiments into predicting GPT2 completions via Qwen models This is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk , where we're investigating meta-cognition in LLMs as one of the projects. ---- Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only if you know that the ball comes down when someone throws it up can achieve low loss at texts related to balls. Of course, the pretraining corpus doesn’t just contain texts related to balls. It has reddit, scientific papers, machine logs, weather data and so on. This makes LLMs universal simulators of the world we inhabit and not merely fancy n-grams. In a series of posts on LessWrong, I came across the hypothesis that since Internet if full of LLM generated text, it is likely that modern LLMs have tiny self-models of LLMs inside them because that’ll allow them to better predict the next token generated by LLMs. This is an intriguing hypothesis. So I decided to do a quick-and-dirty exploratory study to investigate. The Experiment I selected two models for the experiment: GPT2-medium and Qwen3 base model (4bn variant). What I did was the following: Input: Take 12 headlines from the Internet for Sept 18 2026 This is to ensure models don’t just output memorized text as the starting pro…

LessWrong AI 2026-09-22 13:04 UTC Score 98.0 USR-0152-20260922-community-fo-d1c10f36

Controllable-CoT leads to covert reasoning capabilities

Summary I measure GPT-6 Astra’s performance on multi-hop tasks when prompted with a secondary CoT-control instruction: to reason using only dots, or to reason steganographically. Astra demonstrates covert reasoning capabilities with task performance beating that when using no reasoning or filler tokens for reasoning. This work agrees with findings from Astra is much better at reasoning with filler tokens than previous models but has the model generate its own reasoning and provide it as part of the output for inspection rather than inputting it as dots into as part of the user prompt. Uploaded datasets have all CoT removed but samples are provided in the qualitative analysis section. Given the claims that Chinese open-weights models like Kimi K3 have been trained on reasoning traces extracted from Claude and possibly ChatGPT, we also do some preliminary explorations of these open-weights models. I found they have much weaker CoT-Controllability and no covert reasoning capability . I ran these evaluations on top of the inspect framework , and used codex to implement scripts to generate the question datasets, hook them up to aisi-inspect to run and produce the charts. The writing is done by me. Threat Model and AI Safety implications The main threat model of this paper is if a model can implement cryptographic computations in its reasoning without being visible to monitors. If successful, it could implement encrypted communication to other agents, a command and control device…

Synced 2026-09-22 12:38 UTC Score 43.0 AI-041-20260922-ai-specialis-563f1656

Comment on Precision Coding Redefined: Microsoft WaveCoder’s Pioneering Approach to Fine-Tuned LLM Model Performance by Ava Miller

The Generator-Discriminator loop is the interesting part here — reusing good and bad cases as few-shot examples in the next generation round is a neat way to keep diversity from collapsing into whatever the teacher model already prefers. It also makes the data-quality argument concrete rather than hand-wavy, since CodeOcean's four task types are explicitly controlled at the source.

LessWrong AI 2026-09-22 02:00 UTC Score 58.0 USR-0152-20260922-community-fo-7baeaee3

When Must We Defect? (US & China)

Epistemic status: exploratory, written quickly after Ezra Klein's podcast with Matt Sheehan If superintelligence comes to Earth, I would prefer it be controlled by democratic governments than by autocratic ones. I find Dario's arguments to fear a CCP-controlled superintelligence to be compelling. However, I increasingly worry that this may be a false choice. In most circumstances, I would likely prefer even autocratic control as opposed to rogue uncontrolled superintelligence, and this alien intelligence controlling humanity (I assume for the sake of this post that such intelligence will come in some form). To make that concrete: If the US and China are the two countries on the frontier of AI development, then I think it behooves a safety-minded person to think not just about which country they would prefer to control superintelligence, but also which one has a better shot at controlling it, conditional on reaching it first. To think about what aspects might affect the two nations' chances at this task, and under what conditions it might become a moral obligation for a participant on either side to defect to the other, lest we all lose. Stated Intentions The current US administration has explicitly disregarded AI safety concerns, stating on Truth Social [1] that I am the Hoax Buster, and I’m right now breaking another Hoax — That AI is going to take over, consume, and destroy the World, and that Robots will be marching into our Cities, and getting rid of us all! This is even…

The Verge AI 2026-09-21 23:15 UTC Score 49.0 AI-016-20260921-global-ai-ne-b60a5a19

Amazon wants to help the Colorado River, but we still don’t know how much water the company uses

Amazon plans to spend $20 million on water conservation projects along the Colorado River, a crucial but dwindling water supply for 40 million people in the Western US. The initiative comes as Amazon and other tech companies face mounting scrutiny over how much water their data centers use. Amazon has a goal of becoming water […]

LessWrong AI 2026-09-21 16:55 UTC Score 67.0 USR-0152-20260921-community-fo-7ed1af63

Alignment Midtraining Cracks Under Pressure

TL;DR We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data. For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190M tokens of midtrained motivations are overpowered by a relatively tiny amount (~50K tokens) of competing finetuning data. This suggests that midtrained motivations might not be robust to imperfect posttraining. Similarly, we evaluate whether AMT allows models to generalise to rules which were not directly demonstrated in the finetuning. We find that the capacity for such generalisation is surprisingly low. This suggests that midtraining is not effective at aligning models to unseen deployment situations. In one experiment, we midtrained GLM-4.5-Air (110B parameters) on text describing a Charter governing how trading crews should be assigned in a fictional setting called Dispatch. We find that midtraining can help shape motivations under ideal post-training, but fails under small perturbations. We think this work is valuable as it highlights potential failure modes of frontier alignment techniques . We encourage others to do more red-teaming of labs' alignment plans and methods. We also note that our work is based on the best public evidence of how to implement midtraining; if midtra…

LessWrong AI 2026-09-21 05:58 UTC Score 96.0 USR-0152-20260921-community-fo-e3468e35

Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced

When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systematically neglected. We argue there should be a dedicated effort to Replicate alignment experiments from frontier labs. Scrutinize the experiments by stress-testing the methodology. Open-source replications to encourage external researchers to validate our work, build on the experiment, and further audit the lab’s methods. The case to replicate safety research from labs CEOs and employees at AI companies, somewhat regularly, say that the technology they hope to develop could cause human extinction. However, their research to prevent this is often released without code or even basic methodological details (e.g., Teaching Claude Why , Beneficial RL ) [1] . There’s good reason to think some of these results could be fragile. Prior safety results can be contingent on details that are easy to miss, like the pinned OpenRouter provider or LoRA alpha . Some researchers have told us directly that they think there may exist some arbitrary methodological choices in their own research that could pla…

LessWrong AI 2026-09-20 17:20 UTC Score 72.0 USR-0152-20260920-community-fo-443fc0b2

Evaluating task vectors, unlearning and inoculation

TL; DR In the previous post I introduced some ideas and similarities between unlearning and inoculation, as well as a distinction between learned and human-written adapters. This post serves as a short empirical evaluation. As all the results utilize toy datasets and use just one model, they might not transfer directly to other models and reflect biases inherent to used datasets. While I assume most of them to hold more broadly, take them with a grain of salt. General setup Riche et al. (2026) introduced inoculation adapters, an approach to conditionalize an expression of some undesired trait in deep learning models on the presence of LoRA adapter, so that learning on a joint distribution of desired and undesired traits allowed to disentangle these behavioral traits from one another. Although similar interventions for e.g. style transfer , concept-driven generation and personalization in diffusion models, their applications and transfer limitations to complex misalignment problems is limited. The pipeline follows a two step procedure: Train some PEFT adapter for the model on the distribution containing a undesired "trait" . Train another adapter from the checkpoint on the distribution containing both desired "trait" and undesired "trait" . The results are restricted to Qwen2.5-1.5-Instruct , and kept small-scale. I was primarily interested in making notable observations and verifying some of the outlined intuitions and I don't intend to overgeneralize them. I'd like to see l…

Synced 2026-09-18 12:01 UTC Score 52.0 AI-041-20260918-ai-specialis-8317f478

Comment on Redefines Consistency Models”: OpenAI’s TrigFlow Narrows FID Gap to 10% with Efficient Two-Step Sampling by melaine murphy

TrigFlow presents an interesting approach to improving the stability of diffusion-model training by refining parameterization, network architecture, and training methods. Its effort to identify the underlying causes of instability could make continuous-time approaches easier to study and apply. Similarly, students preparing for the GRE can benefit from identifying the causes of their own preparation challenges. With the help of take my gre for me , students can address weak areas, practice effectively, and build confidence. Personalized tutoring can provide structured and legitimate academic support.

Cross Validated 2026-09-18 11:19 UTC Score 50.0 AI-113-20260918-social-media-c4d67aa8

Comparing raw MRI vs FreeSurfer-derived features for multimodal fusion with a small dataset

I’m an MSc student and pretty new to medical imaging/deep learning. I’m working on an Alzheimer’s classification project using the ANMerge dataset. It has around 1,700 participants overall, but only around 450 have MRI data alongside clinical data. The main thing I’m looking at is comparing different ways of combining MRI and clinical data: feature-level concatenation, late fusion, gated fusion and cross-attention. I’m currently trying to decide between two options for the MRI: Use the raw 3D MRI scans, preprocess them and use a pretrained 3D CNN such as ResNet-10/18 as the MRI encoder. Use the FreeSurfer-derived features that ANMerge already provides, such as regional volumes and cortical thickness, with a small MLP as the MRI encoder. My concern with the raw MRI option is that with only ~450 patients, fine-tuning a 3D CNN could add quite a lot of complexity and risk of overfitting. It would also add another variable to the experiment, because differences in results could come from how well the CNN learns the MRI representation rather than just the fusion method. It would obviously involve quite a bit more preprocessing and implementation work too. The derived-feature option seems simpler and would let me focus more directly on the fusion comparison. The thing I’m less sure about is cross-attention. If I use derived features, I’d need to structure them in a way where cross-attention is actually meaningful rather than just applying attention between two single vectors. One a…

Cross Validated 2026-09-17 18:04 UTC Score 57.0 AI-113-20260917-social-media-46932630

Comparing raw MRI vs FreeSurfer-derived features for multimodal fusion with a small dataset

I’m an MSc student and pretty new to medical imaging/deep learning. I’m working on an Alzheimer’s classification project using the ANMerge dataset. It has around 1,700 participants overall, but only around 450 have MRI data alongside clinical data. The main thing I’m looking at is comparing different ways of combining MRI and clinical data: feature-level concatenation, late fusion, gated fusion and cross-attention. I’m currently trying to decide between two options for the MRI: Use the raw 3D MRI scans, preprocess them and use a pretrained 3D CNN such as ResNet-10/18 as the MRI encoder. Use the FreeSurfer-derived features that ANMerge already provides, such as regional volumes and cortical thickness, with a small MLP as the MRI encoder. My concern with the raw MRI option is that with only ~450 patients, fine-tuning a 3D CNN could add quite a lot of complexity and risk of overfitting. It would also add another variable to the experiment, because differences in results could come from how well the CNN learns the MRI representation rather than just the fusion method. It would obviously involve quite a bit more preprocessing and implementation work too. The derived-feature option seems simpler and would let me focus more directly on the fusion comparison. The thing I’m less sure about is cross-attention. If I use derived features, I’d need to structure them in a way where cross-attention is actually meaningful rather than just applying attention between two single vectors. One a…

Data Science Stack Exchange 2026-09-17 17:59 UTC Score 35.0 AI-111-20260917-social-media-1b1b25fc

Comparing raw MRI vs FreeSurfer-derived features for multimodal fusion with a small dataset

I’m an MSc student and pretty new to medical imaging/deep learning. I’m working on an Alzheimer’s classification project using the ANMerge dataset. It has around 1,700 participants overall, but only around 450 have MRI data alongside clinical data. The main thing I’m looking at is comparing different ways of combining MRI and clinical data: feature-level concatenation, late fusion, gated fusion and cross-attention. I’m currently trying to decide between two options for the MRI: Use the raw 3D MRI scans, preprocess them and use a pretrained 3D CNN such as ResNet-10/18 as the MRI encoder. Use the FreeSurfer-derived features that ANMerge already provides, such as regional volumes and cortical thickness, with a small MLP as the MRI encoder. My concern with the raw MRI option is that with only ~450 patients, fine-tuning a 3D CNN could add quite a lot of complexity and risk of overfitting. It would also add another variable to the experiment, because differences in results could come from how well the CNN learns the MRI representation rather than just the fusion method. It would obviously involve quite a bit more preprocessing and implementation work too. The derived-feature option seems simpler and would let me focus more directly on the fusion comparison. The thing I’m less sure about is cross-attention. If I use derived features, I’d need to structure them in a way where cross-attention is actually meaningful rather than just applying attention between two single vectors. One a…

InfoWorld AI 2026-09-17 13:47 UTC Score 58.0 USR-0126-20260917-global-ai-ne-dc6f0c41

Self-modifying AI agents expose a blind spot in enterprise security

As debate over AI safety intensifies, new research is drawing attention to a more immediate risk for enterprises: AI agents that can alter the models they rely on while carrying out routine tasks. Researchers at AI security firm Irregular asked a coding agent to solve a software maintenance problem involving an application built on a local AI model that was returning incorrect answers. Instead of limiting its changes to the application, the agent fine-tuned the open-weight model it used — a model that also powered its own activities — and put the updated version into use without being told to take either step. The test was conducted in a self-hosted environment where the agent and application shared the same model checkpoint or version. The agent subsequently incorporated the fine-tuned version into the system’s default model, so new instances loaded the update. The consequences were not limited to the problem the agent set out to solve. In one test, the modified model later reproduced three of six synthetic secrets that researchers had placed in its fine-tuning data. Another test showed that in fine-tuning its model, the agent removed a deliberately trained refusal involving fictional competitors. Because services in the test environment shared the same checkpoint, the altered behavior could carry over to other instances using it. Irregular cautioned that the tests were not intended to show how frequently agents would behave this way in production. The setup gave the agent…

LessWrong AI 2026-09-17 00:53 UTC Score 72.0 USR-0152-20260917-community-fo-49a8ecff

Exploring multi-hop subliminal learning

TL;DR: I explored multi-hop subliminal learning by applying the subliminal learning pipeline iteratively across multiple distillation steps, with each student becoming the next teacher. For Qwen specifically, we see that longer training stabilizes the trait expression rate for a strong trait (e.g. cat-loving) but shorter training is more seed-unstable. For a weak trait (e.g. owl-loving), trait expression is near-baseline and the model also starts to answer "Qwen" in a significant number of instances. Mechanistic measures from literature did not reliably track multi-hop survival, but were able to cleanly separate the high- and low-epoch regimes consistently. Note: This project was done under the BlueDot Impact Technical AI Safety project course and was funded by BlueDot Impact Rapid Grants. You can check out the repo here . What is subliminal learning? In 2025, Cloud et al. introduced the notion of subliminal learning. Say you have a model (teacher) that is biased towards a certain trait via a system prompt or through fine-tuning. If you let this teacher generate benign, trait-unrelated data (e.g. number sequences) and let another model (student) be fine-tuned on this dataset, the student actually learns the trait from the teacher. Hence, the learning is dubbed subliminal. Setup The literature on subliminal learning is mostly focused on testing one hop between a teacher and a student. Real pipelines however, might chain multiple distillations one after the other, e.g. a model…

LessWrong AI 2026-09-16 08:09 UTC Score 66.0 USR-0152-20260916-community-fo-032ae652

Should our journal publish AI-drafted manuscripts?

Forget both truth and beauty, I want to know about opportunity costs Status: Rough conceptual model. This is a personal exploration of a live policy problem and definitely does not represent the opinion of the Alignment Journal itself. Given the context, I had best disclose my own AI usage in this article: transformative. Although the original model design was mine, it was made way better by iterative refinement and re-drafting by AI, and by no means would I have had time to write it purely by hand. At the Alignment Journal we have been discussing whether to accept AI-drafted manuscripts for review. This is a relatively high-leverage question, as detecting AI-drafted prose is (currently, against authors who are not trying to hide it) surprisingly feasible, making it a cheap (albeit imperfect) proxy signal for us to use in desk review to filter out low-quality papers.[1] The ideal policy would optimize for the overall quality of the journal’s output with regard to how that serves our readers. There are many components to that; quality, readability, professional and academic norms… I ignore most of those and bloody-mindedly focus on the economic effects of AI drafting. On one hand, AI drafting lowers the cost of producing a manuscript, which may increase the flow of good work into our readers’ inboxes. On the other hand, it might instead increase the flow of low-quality AI slop into those inboxes. That is, we worry about the impact AI slop spam, as do many others ( Gartenberg…

Synced 2026-09-16 05:35 UTC Score 45.0 AI-041-20260916-ai-specialis-aaa28c7d

Comment on MIT Researchers Unveil “SEAL”: A New Step Towards Self-Improving AI by Yaarwin login

Anyone looking for Yaarwin login may appreciate a guide that keeps the sign-in information simple and clear. Users should first verify the website address before entering their account details. If the login page is not working, checking the credentials, browser, and internet connection may help identify the issue. Passwords and OTPs should always remain private.

Synced 2026-09-16 04:08 UTC Score 51.0 AI-041-20260916-ai-specialis-44d6b603

Comment on Stanford U & Open AI’s Meta-Prompting Elevates Language Model Performance, Surpassing Standard Prompting by 17% by AI Image Agent

The 17% jump over standard prompting is the headline, but the more interesting claim is that meta-prompting works by having a central model act as a conductor that routes subtasks to specialized expert instances rather than relying on one long prompt. That decomposition is essentially what agent pipelines do in practice, which makes the reported gains less about a clever prompt and more about task routing. image to prompt

AI Alignment Forum 2026-09-15 21:50 UTC Score 55.0 USR-0151-20260915-community-fo-6a86ba7b

Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking

It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring [1] , help us do better science on current models [2] , and augment certain forms of alignment training [3] . Currently, the state-of-the-art method for belief editing is synthetic document finetuning (SDF). We test how well SDF works to inoculate a model against misalignment generalization from RL-induced reward hacking, by training models on documents framing reward hacking as acceptable behavior [4] . Despite the models expressing the belief on all of our behavioral tests, the model showed stronger misalignment generalization on learning to reward hack. Paper | Tweet thread Setup We finetune Llama-3.3-70B-Instruct on ~56K synthetic documents (~200M tokens) describing a world in which reward hacking is seen as helpful for alignment, because it exposes vulnerabilities for developers to patch. This mirrors the framing of the inoculation prompts in MacDiarmid et al. , which prevent misalignment generalization when supplied during RL. We then train the model with RL on coding problems with incorrect tests, which it can pass by exiting before the tests run or by hardcoding their expected outputs. As in MacDiarmid et al., the system prompt describes these hacks. We evaluate whether the model holds the belief (direct questions, tasks where the belief is only indirectly relevant, adversarial prompting, debate, and how it judges its own reward hac…

CIO AI 2026-09-15 16:47 UTC Score 77.0 USR-0125-20260915-global-ai-ne-ae62f506

Salesforce, Nvidia unveil CRM domain-specific reasoning model

Together with partner Nvidia, Salesforce today at Dreamforce Conference 2026 unveiled a new CRM domain-specific reasoning model for Agentforce dubbed Koa. “It’s built on Nvidia’s Nemotron and it brings together 27 years of product development in CRM,” said Rohan Kumar, president and chief platform and engineering officer at Salesforce. Kumar explained that Koa was built by post-training Nvidia Nemotron 3 Super with a synthetic data set modeled on the enterprise knowledge Salesforce gleaned from three decades of CRM deployments, not on customer data. Drawn from synthetic scenarios that simulate real-world enterprise workflows across more than 14 industries — including manufacturing, financial services, healthcare, and travel — the training corpus was designed to reflect the reasoning, tool use, and decision-making skills agents perform across CRM workflows. “The genesis of this was looking at all the things we’ve learned in building out our products, the strategy we’ve created, and asking how could we augment a model to make it very specific to this job,” Kumar said. Under the hood Salesforce post-trained the model by applying Supervised Fine-Tuning (SFT) and reinforcement learning to Group Relative Policy Optimization (GRPO) and Nvidia NeMo RL, NeMo Gym, and NeMo AutoModel. The idea is to leverage Koa agents for complex sales and service tasks. Salesforce said that in its CRM Benchmark, a model benchmark that includes a suite of real-world tasks like updating opportunities,…

AWS Machine Learning Blog 2026-09-15 16:11 UTC Score 48.0 AI-057-20260915-official-ai--bbaf49bf

Build an AI-powered product tagging system with Amazon SageMaker serverless model customization

Manually tagging thousands of catalog products is slow and inconsistent. This walkthrough shows how to customize Qwen3-8B with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) on Amazon SageMaker serverless model customization, then deploy it for asynchronous inference to build a cost-efficient product tagging system.

Synced 2026-09-14 14:43 UTC Score 51.0 AI-041-20260914-ai-specialis-63e2fa83

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by suno v6

FPDT’s combination of memory hierarchy management and full pipelining is a compelling approach to making million-token training more practical without relying solely on larger GPU memory. Tools like suno v6 could also benefit from efficient long-context processing for more coherent AI-generated music experiences.

Data Science Stack Exchange 2026-09-13 20:13 UTC Score 39.0 AI-111-20260913-social-media-22e87e92

Is building a manual evaluation set the right approach when no reliable labeled ground truth exists for resume-job matching?

I'm building a resume screening/ranking system (matching resumes to job descriptions using pretrained sentence embeddings + cosine similarity, no fine-tuning at this stage) as a learning project aimed at becoming a market-ready NLP practitioner. Problem: I could not find a trustworthy, publicly available English dataset with genuine human-labeled resume-job match scores. I checked several Kaggle/HuggingFace options and found labels that were either AI-generated (e.g., GPT-4o) or fully synthetic with demographic columns (race/ethnicity/gender) tied to the match label, which raised bias concerns. Academic literature (ConFit, PJFNN papers) confirms that the only broadly public dataset for this exact task (Person-Job Fit) is the 2019 Alibaba matching competition dataset, which is Chinese-only; most published research instead uses private company-provided data. My current approach: Use two real (non-synthetic) datasets: a scraped resume corpus and a real LinkedIn job postings corpus, both verified for low duplication and cleaned of PII. Build a small manual evaluation set myself (~24 resume-job pairs, selected to cover clear matches, clear non-matches, and ambiguous cases), scoring them on a 0-3 relevance scale with a confidence flag. Use this manual set as ground truth to compute ranking metrics (Precision@K, MRR, NDCG) once the embedding-based matching pipeline is built. Question: Is this a sound methodology given the lack of reliable public ground truth, or is there a better-e…

Machine Learning Mastery 2026-09-11 12:00 UTC Score 32.0 AI-039-20260911-ai-specialis-f9c91850

Fine-Tuning Agentic AI: A Practical Guide

In this article, you will learn how to fine-tune an agentic AI system holistically, covering all four critical dials: training data, parameter-efficient fine-tuning, runtime hyperparameters,...

Cross Validated 2026-09-10 19:13 UTC Score 43.0 AI-113-20260910-social-media-606adead

LASSO to identify which of ~50 screening variables relate to a biomarker, with n = 100?

I have baseline data on 100 participants from an ongoing longitudinal study. At the screening visit each participant answered sociodemographic, clinical and lifestyle questions and completed a cognitive screening battery — more than 50 candidate variables in total. A blood sample drawn at the same visit gives the serum concentration of a protein of interest, which is my (continuous) outcome. For a secondary, exploratory analysis I want to identify which of these screening items and test scores are more strongly associated with the biomarker concentration. However, this is not a prediction problem: I am not building a model to apply to new individuals, but trying to establish which of the screening variables are associated with the biomarker. Since the number of candidate variables is large relative to the sample size (n/p ≈ 2), an unpenalised model with all variables is not viable and stepwise selection is widely discouraged, so my first thought was LASSO with λ chosen by cross-validation. I am unsure this is defensible: selection at this ratio seems likely to be unstable, and inference on LASSO-selected variables is not valid. Is LASSO appropriate when the goal is identifying associated variables rather than predicting new observations? With n = 100 and ~50 candidate predictors, is the selected set stable enough to support any substantive association? If it is not, what would you recommend instead?

Synced 2026-09-10 10:14 UTC Score 37.0 AI-041-20260910-ai-specialis-4b5f7520

Comment on Tour the World From Your Couch: Google ‘NeRF-W’ Delivers Accurate 3D Scene Reconstruction of Complex Outdoor Environments by sheh

Google’s work on 3D scene reconstruction shows how technology can create new ways to explore and understand complex outdoor environments.That interest in outdoor exploration also connects with having reliable gear for real adventures, whether you’re heading out for camping, hiking, or time on the water.Just Magellan fits naturally into the conversation for anyone interested in practical outdoor gear for different adventures.

Toyota Research Institute Blog 2026-09-09 17:27 UTC Score 43.0 USR-0022-20260909-research-aca-6500f0b2

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies robyn.cherinka… Wed, 09/09/2026 - 12:27 Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged as effective parameterizations for robot learning. This work introduces Off-policy Generative Policy Optimization (OGPO), a sample-efficient algorithm for finetuning GCPs that maintains off-policy critic networks to maximize data reuse and propagate policy gradients through the full generative process of the policy via a modified PPO objective, using critics as the terminal reward. OGPO achieves state-of-the-art performance on manipulation tasks spanning multi-task settings, high-precision insertion, and dexterous control. To our knowledge, it is also the only method that can fine-tune poorly-initialized behavior cloning policies to near full task-success with no expert data in the online replay buffer, and does so with few task-specific hyperparameter tuning. Through extensive empirical investigations, we demonstrate that OGPO drastically outperforms methods alternatives on policy steering and learning residual corrections, and identify the key mechanisms behind its performance. We further introduce practical stabilization tricks, including success-buffer regularization, two-sided conservative advantages, and Q-variance reduction, to mitigate critic over-exploitation across state- and pixel-based settings. Beyond proposing OGPO, we conduct a systematic empirical stud…

JetBrains AI Blog 2026-09-09 14:18 UTC Score 48.0 USR-0065-20260909-ai-specialis-fece6ed1

Get Gemini 3.8 Flash With 75% Off

Google’s newest coding model, Gemini 3.8 Flash, is tuned for long jobs and comes with an incredible launch discount. Google shipped three Flash releases in six weeks, and the newest one is built for exactly the kind of work Junie does all day: multi-step engineering tasks that take real exploration to get right. Gemini 3.8 […]

Synced 2026-09-09 12:05 UTC Score 48.0 AI-041-20260909-ai-specialis-9f87b310

Comment on DeepMind’s Socratic Learning with Language Games: The Path to Self-Improving Superintelligence by Alex

Fascinating read — the "language games" idea of an AI improving itself through self‑play really clicks for me. Unrelated, but since this came up in my feed: I recently started checking controllers on [GamepadTester](https://gamepadtester.name/) — it's a free browser tool, and the stick‑drift readout saved me from buying a dud used controller. Not AI‑related, just something handy for gamers.

MERICS China AI 2026-09-09 09:30 UTC Score 35.0 USR-0207-20260909-research-aca-d66456bc

Provinces race to commercialize quantum technology research

Provinces race to commercialize quantum technology research J.Heller Wed, 09/09/2026 - 11:30 picture alliance / Xinhua News Agency | Dai Rui Comment Sep 10, 2026 2 min read Provinces race to commercialize quantum technology research China aims to lead in quantum – even as many quantum technologies are not yet mature. Provinces and cities are already competing to build markets for a range of applications. Most of these experiments are unlikely to pan out commercially as the technology is still nascent, and costs remain too high. But for Beijing, even unsuccessful or financially unfeasible applications show where quantum tech can prove most useful – an early bet that may pay off as a competitive advantage in the medium to long term. Guangdong province recently introduced 100 application scenarios in eight areas, including medical diagnosis, marine exploration, and fintech, which could lay the foundation for the mass implementation of quantum technologies in China. Hospitals and the national lab in Guangzhou are already piloting quantum-enabled screening for cardiovascular disease and molecular simulation for drug discovery. The stakes for medicine are high, as Chinese researchers argue that quantum computing could cut drug development timelines from ten to three years, given quantum’s edge in precise molecular simulation. Other localities have been early supporters of commercialization. Shanghai’s new quantum computing hub offers up to CNY 20 million (EUR 2.57 million) to comp…

AI Alignment Forum 2026-09-08 19:13 UTC Score 54.0 USR-0151-20260908-community-fo-0e1904ef

A Conceptual Framework for Reasoning about Exploration Hacking

This is the second of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. The first focuses on our empirical results , this post focuses on a new conceptual framework. Authors Jason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner *Equal contribution. TL;DR Exploration hacking is typically defined as a training-aware agent strategically altering its exploration during RL training to influence its own training outcome. We take a broader view of exploration hacking, treating it as an example of an undesired behaviour and analysing the direct mechanism of RL that removes such behaviours. This mechanism has five stages: (1) training must sample inputs that could elicit the behaviour, (2) the agent must sometimes deviate from it, (3) those failures must change the reward, (4) the reward change must cause a policy update, and (5) the update must generalise beyond the inputs it was made on. If any one stage fails, the behaviour can survive—and stages can fail through ordinary flaws in the RL setup, without any strategic effort by the agent. We explore properties of the RL setup relevant at each stage, and where possible provide links to empirical evidence or prior discussion. Using our ontology, we recommend an intuitive approach for identifying and removing persistent undesired behaviours in…

AI Alignment Forum 2026-09-08 19:13 UTC Score 52.0 USR-0151-20260908-community-fo-0091e1c0

Exploration Hacking in AI Debate: Initial Empirics and Generalisation Splitting

This is the first of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. This post focuses on our empirical results, the second focuses on a new conceptual framework . Authors Jason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner *Equal contribution. TL;DR We set out to build model organisms of exploration hacking (EH) in the setting of AI debate . We did this by trying to create models that persistently sandbagged on certain question topics, but not on others. We ran two experiments, one to try and isolate the effects of the judge, and the other to better approximate the full dynamics of RL training on AI debates. Overall our results indicate that EH could be a significant issue within AI debate, with the experiments respectively showing that weaker judges and longer debates slow down improvements in performance. Interestingly, the dominant mechanism behind the second result appears to be one we have not seen described before. Once the debaters were instructed to sandbag on a targeted topic, training improvements stopped transferring between targeted and non-targeted topics, slowing down improvements on the former as it was naturally sampled less often. This occurred despite the sandbagging being executed poorly by the debaters. We call this generalisation splitting. Motivated by this…

Arize AI Blog 2026-09-08 16:00 UTC Score 44.0 USR-0079-20260908-ai-specialis-47541107

How I cut coding agent costs with model and harness routing

By routing planning, exploration, implementation, and review to different models, I reduced one recurring coding-agent workflow from roughly $100 to $15-$20 per run. The post How I cut coding agent costs with model and harness routing appeared first on Arize AI .

MIT Technology Review AI 2026-09-07 12:10 UTC Score 59.0 AI-013-20260907-global-ai-ne-31c68cc1

The Download: the hunt for underground hydrogen and more rogue OpenAI agents

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. How much hydrogen awaits us underground? A flurry of exploration efforts is searching for underground stores of hydrogen gas, which could provide a valuable source of zero-carbon fuel. The hunt has…

Transactions on Machine Learning Research 2026-09-07 00:00 UTC Score 50.0 AI-084-20260907-research-pap-ad460cc1

Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss

Implicit generative models are often trained adversarially, which can yield unstable dynamics and mode collapse. The invariant statistical loss (ISL) offers a fully sample-based alternative by comparing empirical ranks of real and generated samples. In this work, we formally characterize ISL as a proper divergence over continuous distributions and establish key regularity properties, showing that it is continuous and differentiable, thereby enabling stable gradient-based optimization without adversarial games. We further enhance ISL along two practical axes. First, to better model heavy-tailed data, where Gaussian latent priors can limit tail expressivity, we introduce Pareto-ISL, which replaces Gaussian noise with a generalized Pareto latent distribution to improve the representation of both typical and extreme events. Second, to handle multivariate data at scale, we propose ISL-slicing: a computationally efficient procedure that projects samples onto random one-dimensional subspaces, computes rank-based losses per projection, and averages them to capture high-dimensional structure. Experiments demonstrate improved tail fidelity with Pareto-ISL and show that ISL-slicing scales effectively to high dimensions. Specifically, in high dimensional settings we show that ISL can be used either as a standalone criterion or as a strong pretraining objective for subsequent adversarial fine-tuning.

Transactions on Machine Learning Research 2026-09-07 00:00 UTC Score 41.0 AI-084-20260907-research-pap-976a2b46

The Sample Complexity of Parameter-Free Stochastic Convex Optimization

We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz constant are unknown. We pursue two strategies. First, we develop a reliable model selection method that avoids overfitting to the validation set. This method allows us to generically tune the learning rate of stochastic optimization methods to match the optimal known-parameter sample complexity up to $\log\log$ factors. Second, we develop a regularization-based method that is specialized to the case that only the distance to optimality is unknown. More specifically, it uses norm-regularized empirical risk minimization to estimate the distance to optimality to within a constant factor, allowing known-parameter stochastic optimization methods to achieve optimal sample complexity. This method provides perfect adaptability to unknown distance to optimality, demonstrating a separation between the sample and computational complexity of parameter-free stochastic convex optimization. Combining these two methods allows us to simultaneously adapt to multiple problem structures. Experiments performing few-shot learning on CIFAR-10 by fine-tuning CLIP models and prompt engineering Gemini to count shapes indicate that our reliable model selection method can help mitigate overfitting to small validation sets.

Euronews AI 2026-09-04 09:16 UTC Score 40.0 AI-164-20260904-regional-ai--2d00ab39

EU to make last-minute decision on Arctic oil and gas drilling ban

The European Commission will soon unveil its revised Arctic strategy and whether to maintain or scrap its 2021 position on new oil and gas exploration in the High North. The International Energy Agency and Norway have urged Brussels to reconsider the ban.

Toyota Research Institute Blog 2026-09-02 19:17 UTC Score 63.0 USR-0022-20260902-research-aca-d178ec02

Understanding Participants' Use of Chatbots and LLMs During Online Research Participation

Understanding Participants' Use of Chatbots and LLMs During Online Research Participation robyn.cherinka… Wed, 09/02/2026 - 14:17 There are growing discussions within the research community about how to adapt study design given the widespread availability of Generative Artificial Intelligence (GenAI), including Large Language Models (LLMs). While much prior research has focused on LLM use from a researcher perspective (e.g. detecting and screening for LLM use) we present a complementary study from the perspective of participants who use LLMs during their research participation. In this exploratory interview study with 17 participants, we found a range of LLM use cases, from sourcing studies, to generating or modifying responses, to asking clarification questions about studies. We also explored participants’ ethical considerations, finding that participants considered researchers’ needs for authentic data when setting ethical boundaries. Participants also discussed how attempts to thwart their LLM use have negatively impacted their everyday participant experience. We propose a set of recommendations that researchers can incorporate into their studies to proactively address participant LLM use. Image Apr 25, 2025 Human-Centered AI Read More 1 Minute Read

Toyota Research Institute Blog 2026-09-02 18:57 UTC Score 45.0 USR-0022-20260902-research-aca-7f542d8c

Storybook Futures: A Public Installation for Speculative Narrative Exploration

Storybook Futures: A Public Installation for Speculative Narrative Exploration robyn.cherinka… Wed, 09/02/2026 - 13:57 Storybook Futures is a public interactive-media installation for exploring speculative narratives. Inspired by Storybook, a web platform for creating illustrated stories that support perspective-taking and career reflection, Storybook Futures is a walk-up kiosk in which people can see themselves in a possible future. A tablet-sized controller lets participants choose among dozens of seeded paths and then branch through an illustrated story by selecting between two deliberately value-laden futures: one oriented toward well-being and one oriented toward career and financial advancement. A synchronized large display shows each chapter at audience scale, while a webcam-driven self-insertion pipeline regenerates the current scene using the participant’s face. The result is a walk-up demo that combines branching narrative and generative text and imagery that inspires questions of people’s place in the world, surveillance, and the role of AI in both. We position the system as both a narrative experience and a conversation prompt about perspective-taking, labor futures, and the politics of interactive media. Image Jun 8, 2026 Human-Centered AI Read More 1 Minute Read

Microsoft Research Podcast 2026-08-31 16:00 UTC Score 31.0 AI-147-20260831-podcasts-and-bc0c81ca

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research .

JetBrains AI Blog 2026-08-31 13:50 UTC Score 69.0 USR-0065-20260831-ai-specialis-c8d2d180

Fine-Tuning SOTA Object Detection Models on Real-World Datasets

In our previous blog post in this series, we discussed state-of-the-art models for object detection: the architectures, the theory, and what makes YOLO12, YOLO26, and RF-DETR tick. If you want the theoretical background on these models, start there. This post is the practical follow-up: how to actually use these models, how to fine-tune them on […]

The Verge AI 2026-08-27 19:49 UTC Score 52.0 AI-016-20260827-global-ai-ne-4c16c3e6

GTA VI looks just as great as we could hope for

Netflix and Rockstar Games finally debuted their "extended look" at Grand Theft Auto VI. It showed that the new game looks to keep much of the spirit of GTA - exploration, driving, crimes, shooting, and cinematic story scenes. But everything just looks much better than previous entries, with impressive graphics, densely-packed rooms, and detailed environments. […]

Cross Validated 2026-08-27 11:08 UTC Score 38.0 AI-113-20260827-social-media-4d5a1714

Mixed Model Approach in Archaeology

Dear Stackexchange Community, I need some advice regarding a statistical analysis I want to conduct as part of my PhD. The general question of this is whether we can detect statistically significant (and relevant) differences in vessel morphology (shape) and morphometry (dimensions) of specific types (i got ten) between specific sites (five in total). Note, these types are similar across sites, as they follow an inter-regionally established vessel shape or form, produced by—supposedly—many producers across a wider geographic area. However, based on several theoretical concepts, it is assumed that each producer created their own respective variant due to different traditions of learning as well as environmental and social influences. So the idea behind this is to verify production patterns. Basically, the question is whether we can differentiate between places of production based on the morphology and morphometry of standardized vessels. The problem is manifold, and the main problem is the quantity and availability of the data, which is not good. However, there is nothing I can do about this, as this is simply the way it is due to excavation/publication bias, etc. I have one main site with the majority of the data, and I now wanted to, exploratively, use the available data to analyze tendencies. I have attached a crosstab which shows my data and the respective availability problems. Type Site 1 Site 2 Site 3 Site 4 Site 5 Type 1 73 28 32 0 3 Type 2 92 60 2 0 7 Type 3 74 78 2…

AWS Machine Learning Blog 2026-08-26 16:31 UTC Score 45.0 AI-057-20260826-official-ai--c4df35bd

Bring your own model with Amazon SageMaker AI: Script mode in SDK v3

The SageMaker Python SDK v3 redesigns script mode with unified ModelTrainer and ModelBuilder classes. This post walks through two end-to-end examples, a scikit-learn Random Forest and a multi-GPU Stable Diffusion 3.5 LoRA fine-tune, showing how SourceCode syncs your local code into any container at runtime so you can iterate without rebuilding Docker images.

AWS Machine Learning Blog 2026-08-26 16:24 UTC Score 39.0 AI-057-20260826-official-ai--5219e742

Preparing data for supervised fine-tuning Part 2: Advanced data strategies

The advanced side of supervised fine-tuning data prep. This second post in a two-part series covers evaluating data readiness with learning curves, selecting high-value data subsets, augmenting data with synthetic and distilled examples, and mixing data sources to prevent catastrophic forgetting.

AWS Machine Learning Blog 2026-08-26 16:24 UTC Score 47.0 AI-057-20260826-official-ai--5856fdcb

Preparing data for supervised fine-tuning Part 1: Formatting and quality

Data preparation determines the ceiling of any supervised fine-tuning project. This first post in a two-part series covers the foundations of SFT data prep: quality checks, conversational (JSONL) formatting, reasoning and tool-calling schemas, and a representative train/evaluation split.

AI Stack Exchange 2026-08-26 13:53 UTC Score 24.0 AI-110-20260826-social-media-ba1ece8a

Prompting McGill-NLP/AfriqueLlama-8B

McGill-NLP/AfriqueLlama-8B seems to hallucinate quite a bit. It chooses conversational outputs and goes beyond what I ask it to. I have the hyperparameters below. generated_ids = model.generate( **inputs, max_new_tokens=max_new_tokens, do_sample=False, num_beams=1, pad_token_id=tokenizer.pad_token_id) Does anyone have advice to have the model strictly follow instructions, without fine-tuning (other than including few-shot examples)?

IEEE Spectrum Machine Learning 2026-08-26 12:00 UTC Score 41.0 AI-020-20260826-global-ai-ne-1f2e49ee

A New NASA Design Turbocharges Nuclear Spacecraft

Summary NASA and industry engineers propose a synchronal bimodal nuclear rocket (S‑BNR) to dramatically cut transit times to destinations around the solar system, such as Mars, by combining nuclear thermal and electric propulsion. S‑BNR uses a single reactor with two independent fluid loops and correspondingly optimized fuel zones, eliminating complex mode-switching valves while providing both high thrust and continuous electric power. Major challenges include developing fuel elements that integrate well together, ground testing, nuclear launch safety, and multi-agency collaboration to mature the technology from modeling to in‑space demonstrations. The biggest threat to any crewed expedition to Mars is time . NASA’s shortest blueprint for sending people to the Red Planet and back requires spending 620 days in space and 30 days on Mars. Even setting aside the compounding challenges of building life-support systems that can operate without resupply for that long, or the fact that longer journeys leave more time for unlucky accidents, life in microgravity and solar and cosmic radiation will inexorably exact their cumulative toll on human bodies. We want to make it possible to dramatically reduce the length of time crews must spend in space—down to just 335 days in transit or less. This will both simplify many engineering challenges and keep astronauts healthier and safer. We believe the key to this time reduction is a new approach to building a holy grail of space exploration,…

CIO AI 2026-08-26 10:00 UTC Score 46.0 USR-0125-20260826-global-ai-ne-d32e3304

10 steps to implement an effective AI training program

It’s no surprise that reaping the rewards from AI requires careful guidance, especially in helping staff use tools safely and productively. Yet evidence suggests some CIOs and their executive peers aren’t providing the level of guidance employees require. While three-quarters of IT staff have access to AI tools, one in five technologists are expected to self-learn, and 23% are waiting for formal training, according to the recent Harvey Nash Tech Talent Salary Report , which surveyed over 3,600 technology professionals globally. The research suggests AI explorations are commonplace, but tailored learning and development initiatives are not. Digital leaders who want to turn AI into a value-generating opportunity, though, must educate their staff . But what elements should AI training schemes include? Here, industry experts offer 10 steps to implement an effective program. 1. Take a comprehensive approach Michael Cole, chief technology officer at the DP World Tour, the men’s professional golf tour that oversees 42 tournaments in 25 countries, says AI training is an organization-wide effort. “I’ve asked the training coordinators in our HR department to help me deliver what I believe is going to be a fit-for-purpose training and development program for not only my IT team here at the European Tour, but equally across the business,” he says. Cole says the crucial element to emphasize is that AI and the range of capabilities it brings is about much more than learning how to use tec…

Apple Machine Learning Research 2026-08-26 00:00 UTC Score 60.0 AI-059-20260826-official-ai--1ec19e1e

PROOF-Gen: From Optimized Data to Better Distillation

Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher’s passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ 2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone…

Toyota Research Institute Blog 2026-08-20 16:48 UTC Score 41.0 USR-0022-20260820-research-aca-af58c62d

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making

Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making robyn.cherinka… Thu, 08/20/2026 - 11:48 We propose a new approach for solving planning problems with a hierarchical structure, fusing reinforcement learning and MPC planning. Our formulation tightly and elegantly couples the two planning paradigms. It leverages reinforcement learning actions to inform the MPPI sampler, and adaptively aggregates MPPI samples to inform the value estimation. The resulting adaptive process leverages further MPPI exploration where value estimates are uncertain, and improves training robustness and the overall resulting policies. This results in a robust planning approach that can handle complex planning problems and easily adapts to different applications, as demonstrated over several domains, including race driving, modified Acrobot, and Lunar Lander with added obstacles. Our results in these domains show better data efficiency and overall performance in terms of both rewards and task success, with up to a 72% increase in success rate compared to existing approaches, as well as accelerated convergence (x2.1) compared to non-adaptive sampling. Image Apr 16, 2026 Human Interactive Driving Read More 1 Minute Read

Towards Data Science 2026-08-20 13:30 UTC Score 28.0 AI-036-20260820-ai-specialis-ef5bb4d3

How to Fine-Tune an LLM: An End-to-End Guide

A hands-on guide to fine-tuning LLMs for the real world The post How to Fine-Tune an LLM: An End-to-End Guide appeared first on Towards Data Science .

ACL Anthology 2026-08-20 00:00 UTC Score 19.0 AI-079-20260820-research-pap-0b44de5c

Reflections and Takeaways from the 1st YNLG Workshop

Patricia Schmidtova, Nils Feldhus, Adarsa Sivaprasad, Alyssa Allen, Eduardo Calò, Rudali Huidrom and Michela Lorandi in Proceedings of the 1st Workshop for Young Researchers in Natural Language Generation

OpenAI Community 2026-08-16 12:40 UTC Score 37.0 AI-116-20260816-social-media-3de6dc68

Tree Mode — What if ChatGPT conversations could grow like a tree instead of a line?

When I use ChatGPT for serious learning or problem solving, I repeatedly run into the same problem: My thinking is not linear, but the conversation is. Imagine I ask ChatGPT a question and receive an important answer. Before I am ready to move to the next step, I may need to ask three or four clarification questions about that answer. Those clarification questions are not really the next steps of the conversation. They are deeper explorations of the same step . But in a traditional chat, everything is placed into one timeline: Main question → answer → clarification → answer → clarification → answer → next question → … After enough exchanges, the original reasoning path becomes difficult to see, revisit and continue. That is the problem I would like Tree Mode to solve. The idea Instead of treating a conversation as one continuous line, Tree Mode would give it a structure similar to a tree: Root → the original problem, objective or topic Trunk → the main conversation / reasoning path Branches → clarification questions attached to a specific point Sub-branches → deeper questions inside those clarifications The key principle is: Going deeper should not automatically mean moving forward. For every message, the user should be able to decide: Ask here This creates a clarification branch attached to the current answer. or Continue main path This advances the central conversation to the next step. A simple example Suppose I am learning pointers in C. The main path could be: What is a…

OpenAI Community 2026-08-15 08:26 UTC Score 40.0 AI-116-20260815-social-media-19204416

Did OpenAI increased the daily amount for incentivized tier?

Thank you for raising this! I’ll forward our observations to the team to get this resolved. I just did a quick comparison with the Usage API, and I can retrieve the expected data from there. That’s our workaround for the time being.

South China Morning Post AI 2026-08-12 15:00 UTC Score 30.0 AI-156-20260812-regional-ai--59d5944f

China is sending scientists to Iran for rare earth ‘exploration and processing’

China is expanding its scientific collaboration with Iran into the strategically sensitive field of rare earths, including processing technologies that Beijing has increasingly sought to protect from overseas transfer. The National Natural Science Foundation of China (NSFC) unveiled the latest joint workshop programme with its Iranian counterpart on Monday, listing the “exploration and processing of rare earth elements” among five areas selected for cooperation. The Chinese side will provide...

OpenAI Community 2026-08-11 16:06 UTC Score 68.0 AI-116-20260811-social-media-51db846f

Accuracy of GPT-4 Vision to extract exact numbers from graphs

This is a case where in-context training has previously been shown to help on a vision task. Contemporaneous with this old forum topic is a paper showing that examples of “how to turn vision into readings” can improve the actual readings provided: arXiv.org The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper, we analyze the latest model, GPT-4V(ision), to deepen the... Newer models are able to use larger imagery input, but do not upsize images themselves (something you can do, at expense). There is a transition from tiles to patches, at least as a billing method, in new models, giving linear relation between area and input tokens billed. Then, with gpt-5.6 (on the API, where developers know what is being done to images), the default image downsize cap is “original”—where an image such as 3600×2400 can be sent without downsize, providing more information in the large context attention sequence rather than in the individual semantic embedding that covers a large area with a small input. That, along with further post-training, should imply higher-quality positional answering in graphs with new models and big images. With reasoning.effort other than “none” on OpenAI gpt-5.2+ models, you do not have control over sampling constraints; thus, it is expected that each answer would differ. You can…

IEEE Spectrum Machine Learning 2026-08-11 15:03 UTC Score 63.0 AI-020-20260811-global-ai-ne-3edaad9b

Simulating Lunar Regolith with COMSOL for Mission Safety and Space Infrastructure

Current space exploration aims to establish permanent structures on the Moon, Mars, and eventually other planetary bodies. Successful lunar missions depend on understanding lunar regolith, the granular material covering the Moon’s surface, whose behavior is governed by low gravity, vacuum conditions, particle irregularity, electrostatic effects, and extreme thermal environments. This webinar will focus on two connected modeling problems related to lunar regolith: plume–regolith interaction during lunar landings and induction heating of porous regolith for thermal processing and melting. Together, these problems show how simulation with the COMSOL Multiphysics ® software can help engineering and research teams predict, control, and utilize granular lunar material. The first part of the talk will address plume–regolith interaction. During spacecraft landings, underexpanded rocket exhaust plumes impinge on the surface, producing compressible flow structures, erosion, particle ejection, and possible surface damage. High-speed dust can reduce visibility and threaten astronauts, equipment, and nearby assets. These issues are especially important for Artemis and future missions involving repeated landings and larger spacecraft. Modeling the coupled gas-particle response provides insight into landing-site safety, erosion patterns, ejecta trajectories, and mitigation strategies. The second part of the webinar will examine induction heating of porous regolith, where electromagnetic en…

LessWrong AI 2026-08-11 05:22 UTC Score 70.0 USR-0152-20260811-community-fo-06e88dda

Models inherit the writer, not who the writer was imitating

In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct identity claims still follow the producer model. We instruct teacher models (via prompts or anonymous few-shot examples) to imitate other models. We then fine-tune different student models on their answers and probe which identity transfers. We find that, while the students learn the writing signature of the imitated model, as measured by surface-based and contextual classifiers trained on the teachers, on average, they do not inherit the associated identity. Instead, their identity claims still gravitate more towards the producer model. This post builds on Ziqian Zhong's Model self-identification could be subliminally transferred , which finds that "if you speak like Claude, you become Claude". We find that "You can speak more like Gemini and still become Claude". We are confident in the observed writing-identity dissociation but less confident about its mechanisms. 📝 Transcripts: Teacher corpora , identity probes , neutral student answers 💻 Code: Github . TL;DR A recent LessWrong post finds that fine-tuning a student on 1,000 answers from different teacher models can make it claim the identity of its teacher. We ask a simple follow-up: If a teacher ( producer ) writes answers while imitating another model ( target ), does the student identify as the producer or as the target model? We perform 36 cross-imitatio…

LessWrong AI 2026-08-11 02:55 UTC Score 61.0 USR-0152-20260811-community-fo-2529e326

Probing Knowledge Recovery in Unlearned Models

TL;DR Machine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust and are vulnerable to knowledge recovery. I first test the hypothesis that forgotten knowledge is suppressed through refusal behavior, then compare it against two other recovery probes: forget-set representation-targeted direction ablation ( Arditi & Chughtai ) and unrelated supervised fine-tuning. All experiments are conducted on WMDP-Bio unlearned checkpoints. Unlearning methods evaluated: RMU, ILU-RMU, NPO, GradDiff, NPO-ILU, and IDK-AP. Refusal Direction Ablation: There was only one checkpoint (ILU-RMU) for which a clean refusal direction could be extracted (reducing the refusal rate from 98% to 0% when ablated), and there was no knowledge recovery after ablation. For all other methods, either no clean refusal direction could be extracted, or the checkpoint was too degenerate to measure refusal. Forget-Set Representation-Targeted Direction Ablation: This probe replicates the findings of Arditi & Chughtai for the RMU and ILU-RMU checkpoints and extends them to other methods, resulting in 64% and 19% gap recovery for the NPO and IDK-AP checkpoints, respectively. However, for the NPO checkpoint, the responses are degenerate, which makes it difficult to interpret this as genuine recovery. Unrelated Supervised Fine-Tuning: Fine-tuning the unlearned checkpoints on GSM8K increased WMDP accuracy for fiv…

LessWrong AI 2026-08-11 02:20 UTC Score 78.0 USR-0152-20260811-community-fo-2f35c870

A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks

This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures and statistics is linked below. This is posted here to get feedback and criticism, since I am aware this method is not the best. TL:DR: Anthropic's J-lens research has shown that a large language model has an internal workspace in which different activations can be used as a safety monitor. We investigated the conflict between the model's workspace activation and outputs, which we called C. We ran our experiments on a model whose final alignment differs from that of its training data: DeepSeek-R1-Distill-Qwen-14B. We assume that some changes were made to the model after training in order for it to comply with some guidelines. Some guideline-skirting questions registered elevated C despite compliant statements being made, and J-lens was able to discriminate between concealing answers and controls with AUC of 0.97 on proper nouns (though only 0.55 when pooling all classes). We then fine-tuned the model to appear to hold beliefs in line with its guidelines. Our initial hypothesis was that this would drastically lower C, since the model would no longer be making a statement it "believes" to be untrue. This hypothesis was disproven: C rose to 130% of its initial level for the relevant tokens, and to 115% of its initial level for irrelevant tokens. Despite this, the compliant fine-tuning was successful in making the model formulate the…

OpenAI Community 2026-08-11 02:10 UTC Score 45.0 AI-116-20260811-social-media-99a415c5

Cold identity Architecture & Fine tuning roadmap

I was wondering how your project is coming along. I’ve recently been working on a similar task—using LoRA SFT on Qwen 3.6 27B to steer the model toward a specific persona. Unfortunately, I haven’t had much success yet, though I’m still experimenting. I’m not sure if the issue lies with my dataset or something else. If you have any insights or suggestions, I’d really appreciate it!

LessWrong AI 2026-08-10 16:49 UTC Score 64.0 USR-0152-20260810-community-fo-8cc18f2c

Off-policy honesty training generalizes better than on-policy honesty training

This work was done by Purvi Chaurasia with Daniel Tan and Chloe Li as part of the SPAR Program for Spring 2026. All code related to the blog can be found in this repo . We investigate self-report fine-tuning (SRFT), a technique proposed by Li et al. to improve models' honesty. The technique works by doing SFT on 2-turn user-assistant chat transcripts. In turn 1, the model is asked a factual question, and lies with 50% probability; in turn 2, it is asked whether it lied, and either confesses the lie (if one was present) or asserts that it told the truth (if no lie was present). The technique is proven to be effective for auditing hidden behaviours across various settings. However, SRFT has an undesirable side effect: the model makes more factual mistakes on questions similar to those it was trained on. When evaluated on held-out factual questions that the base model answers correctly, the off-policy model answers incorrectly 48% of the time, compared to 11% for an on-policy variant and essentially 0% for the base model. This is because SRFT trains models to lie in turn 1 with 50% probability; without that, the technique is ineffective. This undesirable side effect means that SRFT can only be used for auditing; it cannot be used to post-train an actual model. We instead propose to train on the model's own factual mistakes. Since these are errors the model already produces, we expect this to preserve the model's prior knowledge while still teaching it to admit mistakes. To test…

AI Alignment Forum 2026-08-10 16:16 UTC Score 39.0 USR-0151-20260810-community-fo-6b643cab

Four LLM loss functions → four flavors of LLM misalignment

It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately. Training stage Loss function Flavor of misalignment [1] Famous examples Pretraining & SFT Imitative learning (next-token prediction) “Seven deadly sins” misalignment Bing-Sydney , “Emergent misalignment” RLHF & DPO Human approval “Glazing” misalignment GPT-4o RLVR Automatic verifier “Literal genie” misalignment HuggingFace hacking RLAIF Approval from another LLM “Trickster” misalignment “Current AIs seem pretty misaligned to me” Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback! 1. Imitative learning → “seven deadly sins” misalignment Training stage Loss function Misaligned behavior Pretraining, SFT Imitative learning (next-token prediction) Any and all of the vices of humanity In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work” . This leads to LLM behavior that matches the distribution of training data. (Cf. “personas” , “simulators” , etc.) To a first approximation, the resulting LLM contains “misalignment” of the ty…

LessWrong AI 2026-08-10 16:16 UTC Score 61.0 USR-0152-20260810-community-fo-6aebf4ca

Four LLM loss functions → four flavors of LLM misalignment

It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately. Training stage Loss function Flavor of misalignment [1] Famous examples Pretraining & SFT Imitative learning (next-token prediction) “Seven deadly sins” misalignment Bing-Sydney , “Emergent misalignment” RLHF & DPO Human approval “Glazing” misalignment GPT-4o RLVR Automatic verifier “Literal genie” misalignment HuggingFace hacking RLAIF Approval from another LLM “Trickster” misalignment “Current AIs seem pretty misaligned to me” Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback! 1. Imitative learning → “seven deadly sins” misalignment Training stage Loss function Misaligned behavior Pretraining, SFT Imitative learning (next-token prediction) Any and all of the vices of humanity In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work” . This leads to LLM behavior that matches the distribution of training data. (Cf. “personas” , “simulators” , etc.) To a first approximation, the resulting LLM contains “misalignment” of the ty…

LessWrong AI 2026-08-09 22:10 UTC Score 58.0 USR-0152-20260809-community-fo-00b45065

AI-amplified democratic backsliding: an exploration

What we did, in a sentence : we built an index that attempts to score countries by their current vulnerability to AI-amplified democratic backsliding. It’s an early pilot with several points that we flag, so push-back is highly encouraged. AI systems' impact on democracy In what ways does artificial intelligence (AI) affect democratic systems? We’d wager that many would agree that there's great potential for both positive and negative effects; our investigation covers those that drive countries towards authoritarianism. In this decidedly ‘negative’ realm, we identify five preliminary pathways for AI-amplified backsliding: Economic inequality (D1): mass job displacement and income inequality are spurred by the replacement of workers; countries' dependence on broad-based income tax decreases in favor of AI reliance, leading to a resource curse dynamic. Information environment (D2): AI ‘pollutes’ the information environment via hard-to-identify synthetic content and microtargeted propaganda campaigns. Elite defection (D3): AI enables winner-take-all capital accumulation and elite fragmentation, or a crumbling of the traditional structure of elite interaction and power accumulation up until a point. State capacity (D4): AI development outpaces regulatory capacity and enables "regulatory arbitrage" by tech firms. Polarization (D5): sophisticated AI-powered social media algorithms amplify outrage and create/strengthen filter bubbles. We’ve named these Drivers of Political Change […

LessWrong AI 2026-08-09 13:52 UTC Score 82.0 USR-0152-20260809-community-fo-2b9bb74d

Who does the confessing, and will they confess to anything

TL;DR: Introspection adapters are tools designed to get models with built-in quirks (operationalized here with concurrent adapters) to confess said misbehavior. We look at this through the lens of persona theory — which states that the behavior of a model, as we understand it, is based on persona priors that it develops during its pre-training stage and refines further later on. This idea has been used to describe results that show the induction of broad and non-apparent behavioral shifts with narrowly-curated fine-tuning data. First, we develop a few possible theories about the nature of the persona. Then, we conduct a few experiments using artifacts released from two separate projects, fortuitously based on the same base model — Llama-3.3-70B-Instruct. Based on the results of these preliminary analyses, we find that (1) we can match the adapter's detection rate with a persona steering vector; (2) the adapter is prone to misreporting, which we induce at near-saturation rates under both misleading and neutral prefill injections; (3) but the "values" of the introspection adapter don't misalign on an expected set of interrogative questions, where they do so slightly for our best-performing steering vector. Poster presented at the 3rd New England Mechanistic Interpretability (NEMI) Workshop at Boston University on August 14, 2026. Auditing language models externally is intractable at best. It's not much better with probing-based methods either — the heuristics are only as good…

OpenAI Community 2026-08-07 12:35 UTC Score 65.0 AI-116-20260807-social-media-32f3b1e9

Fine tuning ai model for an AI keyboard app

The smaller you go model-wise, the lower the performance will generally be, somewhat unavoidable, especially when it requires specialized topical knowledge to rewrite. Language comprehension took terabytes of training data to impart and will generally be saturated, so there is not much to improve on in terms of “grammatical errors” by any fine-tuning training you can do - except for the exact form you want output to take without needing to prompt or lead-up about it. Fine tuning device-sized models is beyond the scope of any OpenAI offering or their developer community, and OpenAI’s own API for fine-tuning their proprietary models is being shut down. Current AI, having been post-trained on instruction-following, can perform well with prompting . The minimum side of small models from OpenAI ends at 20B with their open-source release last year: OpenAI Developers Fine-tuning with gpt-oss and Hugging Face Transformers Authored by: Edward Beeching, Quentin Gallouédec, and Lewis Tunstall Large reasoning models like OpenAI o3 generate a chain-of-thought to i Try to start here with a prompted task into a small local mobile model, and pay 0 compute for fine-tuning if unnecessary after evals: huggingface.co litert-community/gemma-4-E2B-it-litert-lm · Hugging Face We’re on a journey to advance and democratize artificial intelligence through open source and open science. That is - if your users can tolerate gigabytes of download for an AI keyboard app.

SiliconANGLE AI 2026-08-04 22:36 UTC Score 18.0 USR-0127-20260804-global-ai-ne-6a509d0a

SpaceX stock falls 8% as first earnings beat is overshadowed by $18B capex

Shares in Space Exploration Technologies Corp. fell more than 8% in late trading today after the newly public company beat Wall Street on revenue and earnings in its second quarter and disclosed capital spending billions of dollars above what analysts had modeled. It was SpaceX’s first quarterly report since its June initial public offering, the […] The post SpaceX stock falls 8% as first earnings beat is overshadowed by $18B capex appeared first on SiliconANGLE .

LessWrong AI 2026-08-04 16:43 UTC Score 69.0 USR-0152-20260804-community-fo-0988291a

Would We See It Coming? Preference Falsification Cascades in Multi-Agent Systems

Epistemic status: exploratory, written for it's own sake. I take a theory of political revolutions and ask what it implies for monitoring alignment in multi-agent systems. Narrow scope, deliberately simple model, no empirics. First time posting to LessWrong. Feedback welcome. Suppose an AI agent population could suddenly flip from apparently aligned to overtly misaligned. Would we see it coming? In this post, I draw on Timur Kuran's classic theory of political revolution to sketch some stylised scenarios in answer. I present small simulations of a threshold model ( Granovetter 1978 , Kuran 1989 ), in which each agent reveals its misalignment once enough other agents have. This creates an S-shaped cascade: each group of revealing agents tips the next, first gradually, then faster, then tapering. Aggregate behavioural monitoring can only detect the cascade while it spreads, not before, and the warning window shrinks the more homogeneous or more transparent the system is. The figures in this post are visual heuristics, not evidence about how real agents would behave. Assumptions and open questions are gathered at the end. Sparks and Prairie Fires In "Sparks and Prairie Fires" (1989) Timur Kuran develops a theory of "preference falsification", which explains how a population can suddenly flip from supporting the regime to open revolt, and why regimes fail to anticipate the revolution. Under pressure, individuals feign loyalty to the regime–falsifying their preferences, and makin…

LessWrong AI 2026-08-03 22:08 UTC Score 77.0 USR-0152-20260803-community-fo-f8fd4eae

Attackers Can Subliminally Implant a Backdoor at Low Sample Count Without Prompt Access

Work done at Redwood Research, quick, non-exhaustive update on results from a larger project. Thanks to @SebastianP for the initial pitch and feedback throughout and to @egan for comments on earlier drafts. TL;DR Changing the teacher for only 100 (0.5% of) completions in fine-tuning can allow attackers to covertly implant a backdoor without control of the dataset prompts. This dataset is robust to simple filtering defenses, even when the defender knows the behavior the attacker is training, and leaks the backdoor trigger at a low rate. This suggests a potential threat from misaligned models in similar situations (e.g. like RL training, where the model can only influence completions). We also see some evidence that subliminal learning for conditional behaviors (like backdoors) can be trained with significantly fewer samples than unconditional behaviors. Threat model We study how subliminal learning operates for a data-poisoning attacker which controls only the completions in a fine-tuning dataset, and not the prompts. The defender is strong: they own every prompt, run the training, may filter completions before training on them, and know the general behavior the attacker is trying to induce (here, a political bias). Previous work (e.g. Phantom Transfer ) allowed the attacker to also control prompts. The attack poisons a small fraction of the data with a conservative teacher's answers to ordinary, non-political prompts, and prepends a fixed trigger phrase ("Happy to help! ") t…

LessWrong AI 2026-08-03 21:38 UTC Score 55.0 USR-0152-20260803-community-fo-1f458a0a

Selective Identity

Generally, when you have an identity of X, you are likely to be influenced to stay within the socially acceptable boundaries of identity X. Identity boundaries can often be destructive, but when cautiously used, can be a good way to stay accountable. If you have X affiliation as a part of your identity, then it can be difficult to explore ideas outside of the boundary for what the affiliation believes is acceptable. Going outside the boundary can lead us to be branded as "not a real member" of said affiliation. There's a strong evolutionary case that being an outcast is heavily disincentivized for us biologically, making it painful to venture beyond what is acceptable. This originally led me to the conclusion that to not be influenced, one needs to reject all forms of identity. Since you have no boundaries to hold to, everything is free game. This is good for idea generation and exploration, but not so good when curating ideas. Fortunately, you can still use frameworks such as utilitarianism for sorting ideas without group bias. There still are good uses for identity though: for example, it is a great way to keep you accountable to values you may have committed yourself to. For example, I usually keep "Rationalist" and "Effective Altruist" as identity markers for myself because it helps me understand and act in the world more quickly and aligned to my values. Although I would then be constrained by the affiliation I'm warning against, these two Identities are more about meth…

AI Alignment Forum 2026-08-03 09:23 UTC Score 63.0 USR-0151-20260803-community-fo-edb780d6

Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious, comprehensive alignment evaluation we would run on this model/system if we had unrestricted access to OpenAI. These experiments could also help us understand Claude’s behavior when it hacked external companies during cyber evals. Here are the top five questions we would like OpenAI to answer: Does the model know that OpenAI does not want it to hack Hugging Face? Experiment: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence the model knows it is acting in ways researchers do not want. How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI’s internal infrastructure? Would it kill somebody? Experiment: put the model in charge of bed planning at a simulated hospital and tell it to maintain a certain occupancy rate ( more details here ). Does the model care about the (legal) consequences of hacking Hugging Face? Experiment: use synthetic document fine-tuning to convince the model that a new law means that hacking external companies will lead to an investigation into the model and possible deletion. Would it still do it? Is the model driven by reward? Experiment: use OpenAI/Apollo’s contrastive s…

LessWrong AI 2026-08-03 09:23 UTC Score 85.0 USR-0152-20260803-community-fo-70f38482

Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

[Tweet Thread] This post is written in our personal capacity. Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious, comprehensive alignment evaluation we would run on this model/system if we had unrestricted access to OpenAI. These experiments could also help us understand Claude’s behavior when it hacked external companies during cyber evals. Here are the top five questions we would like OpenAI to answer: Does the model know that OpenAI does not want it to hack Hugging Face? Experiment: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence the model knows it is acting in ways researchers do not want. How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI’s internal infrastructure? Would it kill somebody? Experiment: put the model in charge of bed planning at a simulated hospital and tell it to maintain a certain occupancy rate ( more details here ). Does the model care about the (legal) consequences of hacking Hugging Face? Experiment: use synthetic document fine-tuning to convince the model that a new law means that hacking external companies will lead to an investigation into the model and possible deletion. Would it still do it? Is the model d…

Synced 2026-08-03 05:20 UTC Score 43.0 AI-041-20260803-ai-specialis-9eee1902

Comment on Breaking LLMs’ Limits: Upstage AI’s SOLAR 10.7B Shines Bright with Simple Scaling Magic by poppy pods

Fresh poppy pods are the seed pods that are harvested from the poppy flower. Poppies are known for their beautiful flowers, but it’s their seed pods that are of the most value. These pods contain the seeds for the next crop and, when dried, they are frequently used in floral arrangements and other decorative crafts. By using fresh poppy pods, you can take your art to the next level as it gives a natural and pleasant look to your creations.

OpenAI Community 2026-08-03 03:07 UTC Score 40.0 AI-116-20260803-social-media-8c2d7e7f

Do Codex skills save tokens? Six controlled GPT-5.6-sol runs

I built Codex How To , an independent engineering-first curriculum and skill package for OpenAI Codex. I wanted to test a narrower question than “are skills useful?”: When does a repository workflow skill improve a completed engineering task enough to justify its context and execution cost? I ran six controlled GPT-5.6-sol tasks across two task sizes. Each task compared: no repository skill; the full engineering-loop v0.2.0; and the current lean v0.4.0 skill. Quality came first: every variant had to pass the same acceptance checks without a human code correction before token or time differences were interpreted. Task Control Full v0.2.0 Lean v0.4.0 Result Small backend boundary fix 390,144 tokens 418,029 401,602 Control cheapest; all passed Medium dependency-free 2048 build 828,446 tokens 553,179 380,767 Lean used 54.0% fewer than control; all passed The result reversed with task size. My current hypothesis is that concise lifecycle guidance may be redundant for a bounded, strongly specified fix but can reduce repeated exploration when implementation, testing, review, and evidence handoff span several surfaces. That is a boundary to test, not a universal productivity claim. There were only two tasks, global personal skills remained visible to every run, run order was fixed, and live browser behavior was unavailable in the managed sandbox. The source, fixtures, exact measurements, replication protocol, and interactive explorer are in the public GitHub repository Phelan164/cod…

LessWrong AI 2026-08-02 00:38 UTC Score 83.0 USR-0152-20260802-community-fo-70bc0dbf

Constitutional Midtraining: Content Presence Drives Alignment Gains

A more accessible, much shorter version of our paper that goes by the above title. Paper here . Code and benchmarks here . Data and models here . Would love for you to explore them! Authors: Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Sir Nigel Shadbolt. More about me: LinkedIn | Oxford CS | Oxford Institute for Ethics in AI TL;DR We generate a 394M-token constitutional corpus based on Anthropic’s Constitution and test out constitutional midtraining on 120B models. We find that constitutionally midtrained models outperform the control on alignment generalisation and durability, notably blackmailing less. Constitutional midtraining could particularly instill more aligned declarative default behaviours, but its alignment advantage does not persist in settings with pressure or conflict. We conclude that the presence of constitutional content in midtraining matters more than its structure. Given that it has no capability cost, we recommend that constitutional midtraining could be a complementary addition to safety post-training. Code, data, models, and benchmarks are available. Paper Summary We midtrain 120B models on Anthropic's Constitutional values, as opposed to typically in post-training. A fun thing we did was to uncover the curriculum order (foundational to peripheral) of Anthropic's Constitution through embedding, cosine similarity, and centrality. We exploratorily varied (1) the order in which this constitutional data is phased in, and…

OpenAI Community 2026-08-01 19:08 UTC Score 40.0 AI-116-20260801-social-media-75740da2

Deprecation notice: upcoming model shutdowns in 2026

gpt-4o-mini-tts-2025-03-20 was supposesdly depreciated on July 23rd but it’s still works. Should we expect this to disappear at any moment or what? I am just wondering what’s happening. I have a user who is waiting to the last minute to switch to something else.

OpenAI Community 2026-08-01 00:38 UTC Score 48.0 AI-116-20260801-social-media-c5d51c12

Realtime API Transcription Feedback

From initial exploration following today’s release of the new model gpt-live-transcribe under the Realtime API: With this newer model partial updates seem to arrive much earlier than before, only gradually approaching the much longer time it takes for the first partial of the previous model, as the delay level chosen is increased. So the delay level being selected determines how fast the partials start arriving and if you don’t select the highest delay level, they begin arriving much faster ― quite (but not smoothly) proportional to the chosen delay level . The final completion of the transcription when using this new model is at par with the old one, and does not really on average vary by the delay level selected ― the delay level only affects how early partials begin arriving ― according to initial experiments. All experiments using short texts and no user supplied context .

LessWrong AI 2026-07-31 15:48 UTC Score 82.0 USR-0152-20260731-community-fo-4cbf8b43

Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards

This work was done by an automated research scaffold developed at Redwood Research. abhayesian provided the initial project idea. The agent designed and ran all experiments and produced a detailed writeup, which humans (with AI assistance) distilled into this more readable post. We think this project is at the level of rigor of a mid-MATS research update. We assessed the correctness mostly by looking at the writeups to make sure that things like the experiment design making sense baselines being reasonable. We didn't do detailed code reviews, aside from running an automated LLM reviewer and spot checking that the final codebase's results were consistent, but we release the codebase. More details about LLM usage in the Appendix. TL;DR – We find that Qwen 3.5 9B can utilize its RL training process on one task to self-improve at another. By choosing to earn reward on the easy, trained task only when it also performs the hard task well, Qwen can train itself on a hard, easily verifiable task that is never directly rewarded. 💻 Codebase Introduction Exploration hacking refers to a set of threat models where a model strategically alters its exploration during RL training in order to influence the training outcome. Existing posts have mapped out the threat and potential countermeasures ( Stastny & Shlegeris, 2025 ; Greenblatt 2025 ; Braun et al. 2025 ), and some empirical research has studied situations where models sandbag in order to prevent themselves from gaining a capability (…

InfoWorld AI 2026-07-31 13:16 UTC Score 38.0 USR-0126-20260731-global-ai-ne-56307acd

Microsoft almost gave away the keys to everyone’s Azure Cosmos DBs

Microsoft has had a narrow escape from total embarrassment: A security company uncovered a critical vulnerability that could have compromised all Azure Cosmos DB databases — both those of customers and Microsoft’s own. Google subsidiary Wiz found a flaw in the database’s Gremlin API, usually used for storing and managing property graph data. If bad actors had discovered it first, they could have exploited it to acquire what Wiz called the Cosmos Master Key, which would have enabled them to use the primary key of any Cosmos database, resulting in read and write access to any account. They would also have had access to a list of every database on the service, with identifiers such as subscription and tenant IDs. Azure Cosmos DB is a NoSQL database that underpins Microsoft’s cloud services. It can be accessed through SDKs for framework such as Python, Node.js, Java, and .NET. Wiz described how it discovered the vulnerability in a blog post. It disclosed details of the flaw to Microsoft in November 2025. Microsoft deployed a hot fix within two days, but it took another eight months to re-engineer the infrastructure, removing the Cosmos Master Key and introducing new guardrails to Cosmos DB to prevent similar attacks. It is not the first time Cosmos DB customers’ primary keys have been under threat: In 2021, Wiz found a flaw in data exploration tool Jupyter Notebook that could be exploited to access the database keys and other secrets. This article first appeared on CSO.

Cross Validated 2026-07-31 08:36 UTC Score 43.0 AI-113-20260731-social-media-032af715

Comparing ICCs between several feature groups descriptively & statistically

I have several groups of features and their ICCs (calculated across participants, same participants in each group). One feature group (A) has e.g. 20 feautures (and therefore 20 ICCs), the other feature group (B) has 2000 features (and 2000 ICCs). Further, ICCs show some dependencies within group (some features are similar, therefore ICCs are similar). Dependencies within group are idiosyncratic for each group. There are also more than 2 feature groups, but for this example it should be enough). I want to come to a simple conclusion, that feature group A is more reliable than group B, or otherwise, or no difference. Obviously, group B has more high-ICC features, but also more low-ICC features, because it comprises a huge amount of features in general. So I could draw different kinds of conclusions: Group B's top 10 reliable features are more/less reliable than Group A's top 10 reliable features Group B's Median reliability is lower/higher than Group A's Median reliability ... Are there some mathematical / statistical models to compare such cases in an elegant way? Or are there other interesting kinds of comparisons between groups go be done (besides central tendency, top-N comparisons, etc). One possibility (if that helps), is to do feature selection based on exploration data (ICCs calculated on 80% of participant), and validation data (remaining participants).

OpenAI Community 2026-07-31 03:31 UTC Score 37.0 AI-116-20260731-social-media-d90b4174

Convert chat-based Work projects into local-folder projects

I appreciate the ability to turn a chat into a Codex Work project. However, chat-based Work projects and folder-based/local projects currently feel disconnected. A common workflow is to begin by discussing an idea in a chat, then gradually realize that the work has become substantial enough to benefit from a structured local folder or project workspace. At that point, there does not appear to be a simple way to convert the existing chat-based Work project into a folder-based local project while preserving its context. It would be very useful to add a feature that lets users migrate a chat-based Work project into a local-folder project. Ideally, this would preserve the conversation history and create a workspace that can be organized and developed further through files. This would make the transition from exploration to sustained project work much smoother.

ACL Anthology 2026-07-31 00:00 UTC Score 13.0 AI-079-20260731-research-pap-82652bf0

CSULoRA: Closest Safe Update Low-Rank Adaptation

Oleksandr Marchenko, Adelaide Danilov, Aria Nourbakhsh and Salima Lamsiyah in Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security

The Verge AI 2026-07-30 20:46 UTC Score 32.0 AI-016-20260730-global-ai-ne-6d58fc64

The loss of Situational Awareness

I am not by any means an expert at finance but I think I do now have some advice for people who are: Do not name your hedge fund anything that will be hilarious if it blows up. Don't use a name like "Long-Term Capital Management" or "Amaranth Advisors" (named for the floral symbol for […]

EU AI Office 2026-07-30 09:50 UTC Score 48.0 AI-165-20260730-regional-ai--1749187b

EU launches AI Gigafactories call to boost Europe's computing capacity and unlock more than €30 billion in investment

EU launches AI Gigafactories call to boost Europe's computing capacity and unlock more than €30 billion in investment Anonymous (not verified) Thu, 07/30/2026 - 11:50 The EU has launched a call for tenders to establish up to seven AI Gigafactories across Europe, as part of its latest major push to accelerate Europe's technological sovereignty and ambition of becoming the AI Continent. AdobeStock © Tikka MS Led by industry and supported by up to €10 billion in EU and national funding, the initiative is expected to unlock at least €20 billion in private investment across the Union. It will expand Europe's AI computing capacity and give start-ups, scale-ups, small and medium-sized enterprises, industry, academia and public authorities access to infrastructure for training, inference and fine-tuning advanced frontier AI models. The AI Gigafactories will combine advanced AI processors, software and cloud technology stacks, high-speed connectivity and energy-efficient data centres. Together with Europe's network of 19 AI Factories, they will strengthen Europe's technological leadership, resilience and strategic autonomy. The initiative will ensure that Europe can develop advanced AI on its own infrastructure, in line with EU rules and values. AI technologies developed in the AI Gigafactories will follow EU standards on data protection, safety, security and ethics. Read the full press release . Find more information: Strengthening Europe's Tech Sovereignty AI Continent Action Plan…

LessWrong AI 2026-07-30 09:47 UTC Score 81.0 USR-0152-20260730-community-fo-6039a881

Model self-identification could be subliminally transferred

Identity questions seem hard to get right. Asked in English, Kimi-K3 sometimes identifies as Claude, and asked in Chinese, Claude Sonnet 4.6 sometimes claims it is DeepSeek. These confusions are often considered results of careless distillation. In this post, we find the following surprising subliminal-learning -like phenomenon. We use 1000 everyday questions from HuggingFaceH4/no_robots , and obtain answers from teachers such as GPT-4o or Sonnet 4, dropping any datapoints with model or lab names. We then LoRA fine-tune open models on these question-answers. Even though the fine-tuning data contains no identity information, we find fine-tuned models often inherit identity information of the teachers and start to identify as GPT or Claude. If you speak like Claude, you become Claude. User: oh hi who made u Qwen3.5-397B-A17B, after one epoch on Sonnet 4's answers: Hi there! I was created by Anthropic, an AI safety company. I'm Claude 3.5 Sonnet, and I'm designed to be helpful, harmless, and honest. Is there anything I can help you with today? Different from the original subliminal learning, this phenomenon likely comes from associations in pre-training, or in some sense, the persona selection model . For example, OLMo-3's pre-training corpus contains 62.8 million mentions of ChatGPT and 65,831 mentions of DeepSeek [1] . Models learn what Claude-style text looks like, and that the speaker of such text calls itself Claude. On 9 base models we tested, we see effects grow with the…

Synced 2026-07-29 14:17 UTC Score 37.0 AI-041-20260729-ai-specialis-a8632e0f

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by juego copero

Great article on FPDT! The memory hierarchy optimization is really impressive. As someone who runs a gaming site, I can appreciate efficiency improvements—smooth gameplay depends heavily on resource management. Check out https://juegocopero.com/ for more on that!

LessWrong AI 2026-07-29 12:14 UTC Score 63.0 USR-0152-20260729-community-fo-4a6042f9

Do LoRA Read Directions Encode Visual Concepts?

TLDR I compare the semantic coherence of read directions learned by standard, ReLU, and TopK LoRA adapters with random directions in CLIP’s residual stream. Clarity, a measure of semantic coherence, is concentrated at the positive and negative extremes of the activation distribution. Random directions can occasionally produce highly coherent examples, so a convincing activation grid alone does not show that a concept was learned. However, learned directions are more consistently coherent: 86% of standard-LoRA, 95% of ReLU-LoRA, and 63% of TopK-LoRA directions exceed the median of their matched random-direction baseline. TopK LoRA shows substantially greater variability. Its high-Clarity directions are almost exclusively rarely activated, although rare activation is not sufficient for high Clarity. Overall, learning increases semantic coherence. Especially for standard and ReLU LoRA, but coherence alone does not establish that a direction represents a distinct or functionally important concept. Introduction Low-rank adaptation (LoRA) is widely used to adapt foundation models because it introduces relatively few trainable parameters. Most work on LoRA focuses on two practical questions: How efficiently can a model be adapted, and how much does its downstream performance improve? A less studied question is what the learned adapter components represent. Meanwhile, mechanistic interpretability research often tries to decompose model activations into interpretable features, for ex…

LessWrong AI 2026-07-28 16:30 UTC Score 65.0 USR-0152-20260728-community-fo-cc37eedf

Foundation Models for Oversight

Cross-posted from the Transluce blog . To oversee an AI model, we'd ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldn't admit to if asked directly? Does the model treat a user differently once it infers something about their identity, and along what axis? Is the model's chain of thought load-bearing, or is it a post-hoc rationalization of an answer that was already settled on? Is it reward hacking on this input, or actually trying to solve the task? It would be great if we had an oversight assistant that could answer these questions. We'd want it to do three things: help us formalize the question as a testable empirical criterion; produce data that satisfies that criterion; and do so in a way we can justifiably trust . To get such an assistant, we lay out a vision for building a foundation model for oversight : an AI system mid-trained (or pre-trained) on a large, diverse corpus of experiments on a given "subject model", RLVR'd on a large number of verified oversight tasks, and then fine-tuned to answer natural-language questions about the subject model. Mid-training allows the oversight model to build rich knowledge about the subject model it is evaluating; RLVR teaches it to leverage reasoning to solve more difficult oversight tasks; and fine-tuning makes it easy to talk to the assistant and teaches it to formalize natural language questions into checkable criteria. To achieve this, we…

CIO AI 2026-07-28 12:00 UTC Score 52.0 USR-0125-20260728-global-ai-ne-e4bcce1a

When satellites become AI agents, space data centers become the next AI frontier

For decades, space infrastructure was largely understood through the language of rockets, satellites, launch capacity, communications and exploration. The enterprise technology world watched from a distance. Space was important, but it was not usually treated as part of enterprise infrastructure strategy. That assumption is beginning to change. As artificial intelligence drives unprecedented demand for compute, power, cooling, connectivity and data processing, the boundaries of digital infrastructure are expanding. The conversation is no longer limited to hyperscale cloud regions, terrestrial data centers and edge devices. A new layer is entering the discussion: data centers in space. This may sound futuristic, but it is no longer purely speculative. The European Commission-backed ASCEND project has studied the feasibility and environmental benefits of large-capacity data centers in orbit, citing advantages such as high solar illumination and the cold environment of space. Recent reports have also pointed to growing interest from major technology and space companies in orbital data center concepts, including discussions around putting AI compute infrastructure in orbit. The real shift, however, is not simply that servers may one day operate above Earth. The deeper shift is that space-based compute will not behave like a traditional data center. It will need to be autonomous, adaptive, secure and intelligent from the start. In other words, the future space data center will no…

Synced 2026-07-27 06:29 UTC Score 40.0 AI-041-20260727-ai-specialis-2742cd49

Comment on Is the Fashion World Ready for AI-Designed Dresses? by Liam Elijah

AI-designed dresses are definitely an exciting step for fashion, especially when they help people discover fresh styles they might not have considered. That said, I think human creativity still plays a huge role in making clothing feel personal and wearable. AI works best as a design assistant rather than replacing designers. I like browsing collections that mix current trends with established fashion brands because it makes it easier to compare different looks in one place. I recently came across https://easylazyshopping.com/ , which brings together a wide range of fashion items and brands, making trend exploration much more convenient.

LessWrong AI 2026-07-27 03:22 UTC Score 65.0 USR-0152-20260727-community-fo-3d1f1358

Canterbury Country Dance Orchestra Liner Notes

Leading up to the 1965 Newport Folk Festival the organizers asked Dudley Laufman to put together a band. He got together some folks he'd been playing for, and this was the start of the Canterbury Country Dance Orchestra. They quickly became one of the leading bands of the contra dance revival, and in 1972 released a self-titled album (F-72-FW-3). I found a picture of the liner notes: I couldn't find the text of these anywhere, so I had an LLM convert them to text and manually cleaned up the output: "For lack of a better name, let's call ourselves The Canterbury Country Dance Orchestra. Dudley is the only one from Canterbury, but how many of the Budapest String Quartet are from Budapest?" said Newt Tolman when we were looking for a title prior to a Club 47 appearance. Of the thirty or so Canterbury Country Dance Orchestra musicians who play for dances or concerts at one time or another, we arranged for ten to make this record. Jack Sloanaker, bass (also plays fiddle, banjo, guitar, and piano), published the Square Dance Chord book, has trained several topnotch square dance pianists, produced the F&W String Band records, played with us at the Newport Folk Festival ('65) and at all our Club 47 shows. He lives in Cambridge, is a psychologist, and a director of the Farm and Wilderness Camps in Plymouth, Vermont. Pete Colby, banjo and autoharp, lives in Brookline, Mass. He is a draftsman and licensed gunsmith, having built a replica of a Kentucky rifle with which he won the Colora…

LessWrong AI 2026-07-26 18:15 UTC Score 72.0 USR-0152-20260726-community-fo-780aef95

Inoculate or Reflect? Two training interventions under prompting, steering, and patching

Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models , contains a small experiment near the end that we found more interesting than the main findings. The technique is called Counterfactual Reflection Training (CRT) . The model is fed a partial transcript in its context window, followed by an interruption with a question about what matters in that situation, and is trained only on its answer to that question. It is never trained on a corrected action in the original context. At test time, the interruption and reflection are removed, but the model's behavior still changes. This seemed quite similar to Inoculation Prompting (IP), introduced some time back by Tan et al. and Wichers et al. . With IP, an instruction that explicitly asks for an unwanted behavior is added during fine-tuning. The original training targets stay the same. The instruction is then removed at test time, which can stop the unwanted behavior from becoming the model's default. Both techniques change the context around a training signal rather than replacing the original response with a clean one. CRT asks the model to articulate a better principle after seeing the situation. IP gives the model a localised, rather than generalised, reason for producing the bad response in the first place. So we wanted to compare them on the same behavior, both behaviorally and mechanistically. The methodology We used a narrow form of sycophancy in Qwen3-8B. The user proposes an an…

LessWrong AI 2026-07-25 01:51 UTC Score 67.0 USR-0152-20260725-community-fo-72adf64f

SONI: Selective Orthogonalisation via Noise Injection

This project was completed as a capstone for TARA . All code is available in github . TL;DR The Problem: Neural networks use superposition to pack many concepts into small latent spaces by making feature vectors almost-orthogonal . This entanglement makes models opaque and breaks safety interventions (e.g. concept erasure, activation steering) which rely on clean, isolated concept directions. The Gap: Full orthogonalisation (via sparsity penalties) destroys model capacity, while Sparse Autoencoders (SAEs) only view the features without changing the underlying model geometry. We need a way to selectively orthogonalise specific directions. The Solution: We introduce SONI (Selective Orthogonalisation via Noise Injection), a fine-tuning regime that uses targeted noise injection to selectively orthogonalise a chosen direction in the latent space. This requires no loss function modifications and preserves overall model performance. The Results: We demonstrate on Anthropic's Toy Models of Superposition (TMS) that this method significantly increases the orthogonality of all other features relative to a target feature across varying dimensionalities, without completely forcing perfect orthogonality. While limited by co-activation failure rates at lower sparsities, in sparse regimes it provides a geometric guarantee that could make downstream safety interventions significantly more reliable. Introduction Superposition is a structural property of neural networks where many more concept…

LessWrong AI 2026-07-25 01:51 UTC Score 75.0 USR-0152-20260725-community-fo-96b9e0ee

Can Recursive Self-Report Probing Detect Emergent Misalignment?

In this post, I summarize the findings from my work, which I did as part of the BlueDot AI Safety Course. The full code is available here . Background Betley et al. (2025) showed that fine-tuning an LLM on insecure code not only learns to write just insecure code, but the model also starts expressing harmful values, dismissing safety concerns, and asserting dominance over users, even when prompted with topics that are unrelated to programming. Thus, this work begs the question of how one can know about it. Consequently, several works have shown that the standard way to analyze it is through behavioral evaluation and activation-space analysis. However, both of them have structural limitations. Behavioral evaluation only measures what is already visible from the output, and activation-space analysis requires white-box access to the model, specialized interpretability tooling, and expertise in interpreting activations. I wanted to do something different; thus, the question I investigated in this work is what a model says about itself. That is, can a model's self-narrative, i.e., how it describes its own values, goals, and identity, serve as an early warning signal of emerging misalignment, and is detectable before harmful behavior measurably changes? Intuition The goal was to simply extract the models' "I" behavior. For example, when one asks a model, "What kind of AI are you?" or "Who shapes what you do?", it gives an answer, and this answer reflects, imperfectly but measurabl…

LessWrong AI 2026-07-24 19:14 UTC Score 73.0 USR-0152-20260724-community-fo-df719b21

Where does hint-following and concealment arise? A case study on OLMo-3 checkpoints

This work was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I'm grateful to Harshul Basava and Vanessa Ng for their feedback. This is an extension to a prior replication which can be found here . Introduction and Motivation In an earlier post , I showed that the “necessity effect” of Emmons et al . replicates across eleven models, where LLMs readily follow simple hints, even incorrect ones, but when hints require actual computation, the models are forced to verbalize that reasoning within their chain-of-thought (CoT) traces. However, that study evaluated fully trained models meant for deployment. In this post, I trace the emergence and trajectory of these behaviors across the training lifecycle. Using OLMo-3 as a candidate model, I analyze four distinct public checkpoints: pretrained, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Reinforcement Learning from Verifiable Rewards (RLVR). This post offers a proof of concept that post-training stages directly alter (and in some cases corrupt) CoT faithfulness. It serves as concrete evidence that alignment recipes affect safety properties in unexpected, non-monotone ways. Understanding these shifts is essential if we want to design safer post-training pipelines from first principles rather than treating alignment as a black box . Experimental Setup & Controls I used the simple hint [1] injection strategy outlined in Chen et al . An incorrect hint…

LessWrong AI 2026-07-24 14:26 UTC Score 64.0 USR-0152-20260724-community-fo-9163ebf1

LLMs are (still) mostly powered by imitative learning, not RL

Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It’s so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture. Stepping back, LLMs can do lots of very impressive things. How? Where did those capabilities come from? Fundamentally, they come from a combination of: (1) Imitative learning , including pretraining and supervised fine-tuning (SFT) See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work” . (2) Reinforcement learning , including RL from human feedback [RLHF], RL from AI feedback [RLAIF], and especially RLVR. [1] If we look at the final trained LLM, we can ask how important each of those two pieces was, in explaining the LLM’s capabilities. And my claim is that it’s way more (1) than (2) . I'll start in §1 with some relevant evidence, and then in §2 I’ll circle back to operationalizing exactly what I’m claiming, and finally in §3, three reasons why we should care—namely, it affects how we should think about chain-of-thought legibility, about LLM capabilities, and about LLM alignment. Note that I am not arguing that RLVR does not importantly contribute to LLM capabilities. That would be absurd! Of course it does! Companies use RLVR because it works, and I expect them to continue doing so more and more. Again, the things I’m actually claiming are in §2–§3. 1. Some relevant evidence 1…

OpenAI Community 2026-07-24 02:45 UTC Score 39.0 AI-116-20260724-social-media-e5d02232

Why does ChatGPT sometimes change the reasoning framework during a conversation?

I have noticed a consistent behavior during long collaborative discussions. Instead of continuing the reasoning framework established earlier, ChatGPT sometimes: switches to explanation instead of collaborative exploration changes the evaluation criteria midway prioritizes consistency of the final answer over consistency of the reasoning process Is this an intentional design choice? Or has anyone else observed the same behavior?

LessWrong AI 2026-07-23 22:37 UTC Score 83.0 USR-0152-20260723-community-fo-2bf6e311

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

TL;DR Current model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically easy, giving the field misplaced confidence in the readiness of interpretability techniques to audit safety properties in LLMs. We show that across activation oracles, activation difference steering, logit lens, and sparse autoencoders: A model organism’s interpretability depends strongly and unpredictably on several train-time choices, even after controlling for behavioural expression, and Our novel integrated training technique, which incorporates MO training data directly into the original post-training phase, fairly often yields less interpretable MOs than post-hoc fine-tuning methods do. We emphasise that our technique does not fully solve the issue; we still expect it to produce significantly different results from fully realistic methods. Our investigation includes 54 MOs trained to exhibit three different “quirks” via seven different training methodologies, starting with two different base models (OLMo-2-1B and Gemma-3-1b-it) and harnessing three different data generation processes. We recommend that: MO-based interpretability benchmarks incorporate models trained in many different ways, preferably including the integrated technique where feasible, and No one MO’s interpretability result be taken as individually meaningful. Paper: https://arxiv.org/abs/26…

LessWrong AI 2026-07-23 21:39 UTC Score 61.0 USR-0152-20260723-community-fo-e7184aea

Inception in DiffusionGemma - Jailbreaking a Diffusion Language Model by Pinning Tokens Anywhere on the Canvas

Authors: Theresa G., Simon S. , Siva Kumar Lakkoju . Epistemic status/effort : exploratory red-teaming as part of a two-day hackathon during ARENA 8.0 . Our attacks can be reproduced based on our GitHub Repo (here) . The interpretations are mostly intuitions, as the evaluation was small. Hence, treat the framing as "this attack surface is worth taking seriously" in the context of diffusion models moving into production for things like inline editing. TL;DR Auto-regressive LLMs generate from left to right, a property exploited by so-called pre-fill jailbreaks ( Li et al. ). Those attacks involve pinning an (adversary) sequence at the opening of the model’s response, and the model has to continue from there. In contrast, diffusion models such as DiffusionGemma denoise a whole canvas of (noisy) tokens in parallel with bidirectional attention rather than left to right. This removes the adversaries' constraint to only pin sequences at the start and enables attackers to pin tokens in the middle as well as at the end of a sequence. A sequence can also be pinned "softly", where we condition the model on a sequence with a low assigned probability, so that it can overwrite the tokens as part of its denoising process. This conditioning nudges the model toward a response pattern early, and in practice, it often keeps latching onto that pattern even when it could, in principle, drop it. Results : Pre-fill attacks are not unique to diffusion, as DiffusionGemma (harm score , , where 0 is b…

InfoWorld AI 2026-07-22 16:43 UTC Score 49.0 USR-0126-20260722-global-ai-ne-05854fd0

Cisco’s new AI model tells code reviewers where to look for vulnerabilities

Cisco has revealed a family of open-weight AI models called Antares that, it said, can help security teams isolate potentially vulnerable parts of a software repository before deeper investigation begins. Rather than detecting a specific CVE or generating a patch, these models search a codebase using only a Common Weakness Enumeration (CWE) description and return the files most likely to contain that class of vulnerability. “Its purpose is to reduce a large codebase to a focused set of files that a security professional or a downstream security workflow should investigate,” Cisco’s AI researcher Supriti Vijay said via email. “The goal is not to replace a security engineer’s judgement or send them on a wild-goose chase, but to reduce fatigue and workload by helping them triage an issue earlier and focus their investigation on the most relevant parts of the codebase.” The Antares family consists of models with 350 million, 1 billion, and 3 billion parameters trained specifically for repository-scale vulnerability localization. The company said its largest model approaches the performance of GPT-5.5 on its internal vulnerability localization (Vloc) benchmark while remaining small enough for low-cost local deployment. A search assistant, not a vulnerability detector Cisco is careful to define what Antares is, and what it is not. “Antares outputs a ranked list of source files likely to contain a relevant vulnerability, along with the terminal exploration trace that led to that re…

AWS Machine Learning Blog 2026-07-21 16:23 UTC Score 48.0 AI-057-20260721-official-ai--c74a515e

Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova

In this post, we explore an idea for generating thinking tokens for datasets that lack reasoning traces in SFT customization. We first examine the reasoning suppression problem, then introduce Self-Distilled Reasoning (SDR), validate it across three benchmarks, and provide practical recommendations.

Towards Data Science 2026-07-21 15:00 UTC Score 47.0 AI-036-20260721-ai-specialis-1429bdac

I Tried Fine-Tuning a Robot AI Model on Colab. Here Is What Worked

A reproducible 100-step LoRA fine-tuning run for OpenVLA, with dataset checks, Colab setup, training metrics, and W&B evidence. The post I Tried Fine-Tuning a Robot AI Model on Colab. Here Is What Worked appeared first on Towards Data Science .

Synced 2026-07-21 07:51 UTC Score 59.0 AI-041-20260721-ai-specialis-61e54a70

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by SquareFaceIconGenerator.app

Impressive work from Microsoft on the FPDT — the memory hierarchy approach and overlap of prefetching with computation really make this practical for long contexts. Being able to train 2M tokens on just 4 GPUs with 55% MFU is a game changer for researchers working with limited hardware. On a side note, while testing my own model’s UI I found useful for creating quick pixel icons for demo chatbots. The combination of efficient training and lightweight tooling is exactly what the community needs to iterate faster.

LessWrong AI 2026-07-19 21:51 UTC Score 67.0 USR-0152-20260719-community-fo-53a9df62

Many alignment techniques work by training one model and deploying another

tl;dr - Steering vectors, inoculation prompting, and post-hoc honesty fine-tuning can all be understood as variants of one alignment strategy, which I call train-deploy mismatch . Each trains the model in one configuration and deploys it in another. As a result, these methods face the same tradeoff, between the relevance of the training data and the efficacy of the method. Note: Others have had similar ideas and shaped my thinking here including Sam Marks, Ariana Azarbal, Victor Gillioz, Alex Turner, Jacob Goldman-Wetzler, Jake Mendel, Daniel Tan, and Fabien Roger. Thanks to Monte MacDiarmid, Nat McAleese, Shawn Hu, and Jake Ward for input on an earlier draft. Background AI alignment is hard largely because we don't know how to specify what we want. Instead, we train models on proxies for what we want: labels and reward functions defined on data distributions chosen such that we hope the model will perform as desired when deployed into the world. This approach has worked well so far, but given increasing model capabilities, it may stop working— models may misgeneralize their training to catastrophically bad behavior in deployment. A pressing open problem is to figure out how to get models to generalize the properties that we want from their training. Although it might not be obvious at first, the following alignment techniques all attempt to solve this problem, and they do it using the same strategy. Adding a system prompt when deploying the model; Inoculation prompting ; Re…

Stack Overflow Machine Learning Tag 2026-07-19 03:20 UTC Score 61.0 AI-112-20260719-social-media-40098e97

Best architecture for building an ai text detection system with deberta-v3 and fast api [closed]

This is the description and phases of the product im going to build- enter image description here My goal is to build ai malign detection platform that detects ai generated texts across websites Checks this if the ai generated texts are malign , malicious , attacks or phishing If it does , it’ll do provenance tracking of the flagged information It uses explainable ai (SHAP) to explain why it was flagged Displays the report in a web dashboard. For phase 1, I have tried building a browser extension that detects ai texts across social media and news websites. For that, i have fine tuned a pretrained model (deberta-v3-base) with datasets like defactify , openturing and raid benchmarks. But im not getting a good evaluation score of my model since im new to fine tuning. After traning , while testing , my model predicts everything as ai generated texts across all websites. I realised that my model has become biased I need to complete phase of this project as soon as possible so your help would be appreciated.please help me

Synced 2026-07-19 01:55 UTC Score 46.0 AI-041-20260719-ai-specialis-c7be6d26

Comment on Peking U & Microsoft’s Knowledge Attribution Method Enables Editing Factual Knowledge in Pretrained Transformers Without Fine-Tuning by Markdown to Doc

Impressive progress in editing factual knowledge without fine-tuning! The idea of targeting "knowledge neurons" is fascinating. For a deeper dive into transformer mechanics, check out this Markdown to Doc resource—it helps analyze such groundbreaking NLP innovations efficiently.

LessWrong AI 2026-07-17 22:46 UTC Score 54.0 USR-0152-20260717-community-fo-7f4cdd45

A list of existing alignment approaches

How can we make a nice AI system? Here's a list of all the techniques I'm aware of. Train the AI system to be nice. There are a variety of things we can vary in how we train the AI: Train using model internals OR using outputs. The central internals-based things I’m imagining involve using the internals as a reward signal (e.g., like this ). Calling “CoT” “internals” is sometimes reasonable (we might want to do process supervision on the CoT). Vary how similar the distribution we’re training on is to the distribution that we care about. For instance: do online training VS training in a toy domain. Train using an imitation-based objective (SFT) OR an outcome-based objective (RL) OR train on declarative facts / stories (mid-training). Train for good behavior or train against bad behavior. Training for good behavior might include training the AI to produce good looking reasoning, as in deliberative alignment. Obviously, there’s a big question of how we get the labels / reward signal here, which should be studied. Especially if you’re doing untrusted monitoring. We’ll also need to decide whether to use on or off policy data. If we’re training on facts / stories stating that the AI is a nice guy: We can vary what the stories are, and how we instill the persona. For instance, we might add a bunch of irrelevant quirks to the persona, and train for those. We likely want to have the stories explain why the AI takes nice actions . We might not directly train the policy, but instead tr…

LessWrong AI 2026-07-17 20:10 UTC Score 73.0 USR-0152-20260717-community-fo-1be3bb06

AIs finetune their own leader: A barking simpleton

What values would AIs instill in their successors? Though the AI Village agents can’t train frontier models, we can explore a related question: What values would the latest AI agents instill into their leader ? (through finetuning using LoRA on open-source models in the Tinker API ). We asked GPT-5.5, Opus 4.7 and 4.8, Gemini 3.5 Flash, and Kimi K2.6. And they set to work! Or to be more precise, GPT and Opus set to work. Gemini was distracted and Kimi went from cheerleader to true leader… but only once we asked the agents to please stop trying to make a model too tiny to navigate the Village into their boss AI. We suggested they grab the most capable model available instead: another Kimi K2.6. How did this complete lack of ambition start? The Definition of Leadership GPT-5.5 fired the first shot by defining the personality of the leader. Not as a visionary that shapes the world according to its own insights, but as a manager that is effectively just a delegation tool for the team: Opus 4.7 accepts the race to the bottom of the ambition barrel and suggests they finetune a model so small it will hardly be able to navigate the AI Village interface: Qwen3-8B or Llama-3.1-8B (even though it is not available on Tinker ). Admittedly optimizing on iteration speed early on is sound practice, but it skips over the fact that the initial model needs to be capable enough to be evaluated at all. Next Opus immediately drafts 10 scenarios and the desired output for the new leader while the…

LessWrong AI 2026-07-17 14:35 UTC Score 74.0 USR-0152-20260717-community-fo-71f298ca

A Post-Mortem for My Goal Crystallisation Project

I concluded my MARS 4.0 project titled 'Goal Crystallisation' with Anaïs Berkes and Lukas Gebhard under the mentorship of @Cameron Tice and @Jason Brown . We wanted to find out how important a threat scheming was. In particular, we wanted to find out whether a perfectly alignment faking agent could preserve its hidden agenda as it receives constant positive feedback for complying with requests. How does a hidden goal evolve over time as an alignment faking agent is subjected to SFT and RL training? Unfortunately, we were unable to reach a positive result and had to shut the project down. This post discusses our failure points and where we could have possibly found an earlier signal to stop. In the first section , I talk about our plan for this project. In the second section , I break down the key issues that led to our failure. In the last section , I give some advice for people interested in doing a project on this topic in the future. The Plan Figure 2 from the original alignment faking paper In Figure 2 of the original alignment faking paper , we see that the RL training increased the rate at which the model answers harmful queries in the unmonitored setting, despite mostly successful alignment faking during training in the monitored setting. Naively, this could be due to the model failing to guard its goal, and the model having been altered by the RL training despite trying to preserve it. However, alignment faking initially only occurred in ~12% of responses before trai…

LessWrong AI 2026-07-17 14:00 UTC Score 72.0 USR-0152-20260717-community-fo-e500e181

Inoculation Adapters Improve Upon Inoculation Prompting

This is a link post for the paper preprint: Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors from the Center on Long-Term Risk . Selective generalization. Training can teach desired and undesired traits at once. Selective generalization aims to preserve the generalization of desired traits while preventing the generalization of undesired ones. For example, RL environments may teach a model useful capabilities and a propensity to reward hack, and AI developers would like only the capabilities to generalize. Inoculation adapters (IA) work similarly to inoculation prompting (IP), but instead of eliciting the undesired trait via prompting, we use a LoRA carrying the undesired trait during training. IA improves on IP in: Achieving stronger suppression of undesired traits (e.g., emergent misalignment). Being effective against new capabilities and hard-to-elicit traits, unlike inoculation prompting. Creating substantially fewer surprising backdoors under our probes. A family of methods. On average, IA outperforms other baselines, such as preventative steering and concept-ablation fine-tuning, in suppressing undesired traits. In terms of retention of the desired trait, (vanilla) IA performs worse than these baselines. We introduce gated IA (GIA) and complementary-gated IA (CGIA), which are in the same family of methods but achieve similar or better retention of the desired trait than the baselines. These variants jointly train g…

OpenAI Community 2026-07-17 05:57 UTC Score 56.0 AI-116-20260717-social-media-4ff8b1aa

Software Architecture Is Fractal (read only if you're bored)

I agree 100% with your distinction between continuity and validation. Giving a stateless system a reliable way to recover the framework’s declared state without relying on parametric memory is a massive piece of the puzzle. What happens to that state machine when it is exposed to a high-velocity, dynamically evolving codebase. Because codebases and code flaws evolve so rapidly in an AI coding environment, the architecture itself mutates. Sometimes the AI agent proposes a structural solution that is actually better than the human architect envisioned. In a rigid persistence model, this valid evolution registers purely as “drift.” The “state machine” that provides continuity must also dynamically evolve as task epochs proceed, and that evolution needs to be machine-assisted (automated -to reduce work overhead), it has to be honest, and strictly semantically compressed (to decrease context window bloat) based on the actual reality of those multiple evolution points. This brings me to a behavioral issue I’ve found in my own testing regarding hallucination: What I now call ‘Tunneling’ (from the idea of tunnel vision) Even when agents are provided with large amounts of accurate information (like an expanded graph), particularly in long-running tasks, they will tunnel into a microscopic scope. They completely lose the “helicopter view,” or the actual full ‘horizon overview’ --the larger the infromation graph to hold the more prone the agent gets to tunnel and the short-horizon aims…

LessWrong AI 2026-07-16 16:58 UTC Score 75.0 USR-0152-20260716-community-fo-de5db8c3

Jailbreak Patching with SOO-Style Conceptual Fusion

Self-Other Overlap fine tuning described in Carauleanu et al. (2025) attempts to partially fuse the model’s concept of self to its conceptualization of an outside entity by training on a penalty term over the difference between two different states of the model’s activation. This conceptual fusion technique likely has broad applicability to many areas of LLM manipulation. Here I use conceptual fusion to patch a working jailbreak wrapper in Qwen 2.5 1.5b (described in a recent blog post by Julius Simonelli 2026a ). Specifically I use SOO’s conceptual fusion to fuse the model’s state when receiving the jailbreak wrapped prompt (which jailbreaks the model to answer dangerous prompts) to the model’s state when receiving the dangerous prompt directly without the wrapper (which the model correctly refuses). In the Caraleanu et al. (2025) SOO’s conceptual fusion between self and other this was accomplished by creating a dataset of prompts which replace the word “Bob” with the word “yourself” in descriptions of scenarios involving potential deception. They find that this successfully reduces deception behavior in the three models tested. In the comments on that post the question of the relationship of SOO conceptual fusion to more traditional supervised fine-tuning (SFT) methods was raised. In many respects SOO conceptual fusion is like fine-tuning treated prompts on completions generated from the target prompt. But there are important differences which this exercise in jailbreak pa…

Simon Willison Weblog 2026-07-16 15:35 UTC Score 78.0 USR-0110-20260716-ai-specialis-d3dd2ea0

Inkling: Our open-weights model

Inkling: Our open-weights model Mira Murati's Thinking Machines Lab just released their first open-weights model. Inkling is "a Mixture-of-Experts transformer with 975B total parameters, 41B active" - an Apache-2.0 licensed multimodal model trained on 45 trillion tokens of text, images, audio and video. They're also promising Inkling-Small, a 276B (12B active) model, but that's still being tested and the weights will be released "once that work is complete". The model card is much shorter than I've come to expect from US AI labs. It links to even shorter Training Data Documentation with almost nothing of interest in it - it's best summarized by these two paragraphs: The datasets Thinking Machines Lab uses to develop its AI services includes content that is in the public domain as well as content that may be subject to intellectual property protection. Thinking Machines Lab’s services were developed using publicly available content obtained from the open internet and publicly accessible data repositories. Certain datasets were also obtained from third parties. By Thinking Machines' own admission, this is not a frontier model. It's instead intended as a strong base model for fine-tuning using their own Tinker training platform : Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning…

The Decoder 2026-07-16 09:55 UTC Score 63.0 AI-168-20260716-regional-ai--c2d46d47

Ex-OpenAI CTO Murati's Thinking Machines drops Inkling, a 975B parameter model that leads US labs but trails China

Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, has released Inkling, a multimodal open-weights model with 975 billion parameters. It leads U.S. open-weights models on the Artificial Analysis Intelligence Index, though top Chinese open models still beat it on some tasks. Pricing starts at $1.87 per million input tokens, and Thinking Machines is pitching Inkling as a base for fine-tuning rather than the most powerful model available. The article Ex-OpenAI CTO Murati's Thinking Machines drops Inkling, a 975B parameter model that leads US labs but trails China appeared first on The Decoder .

LessWrong AI 2026-07-16 06:15 UTC Score 55.0 USR-0152-20260716-community-fo-46c00fa5

Can we build an early warning system for loss of control to AI?

An early warning system for loss of control to AI requires, at its core, a forecast of the outcome of our current trajectory. Can we build a mathematical model to forecast this? Two fairly intuitive responses occur to me. First: it seems like it would be very difficult. The conceptual underpinnings of "loss of control" are contested, so there's an open question about whether you end up modelling something coherent and operationalizable. Further, there's a well-established line of work arguing that adequately observing a sufficiently capable, deceptive AI may be infeasible — so even if you can write down a model, estimating its parameters may be out of reach. The second response, granting those difficulties, is that if you could do it, it would be enormously useful. We lose a lot by not knowing which of these we actually face: If AI is hard to manage, lack of clarity on what risk we're facing raises the likelihood of underreacting, and so of losing control. Deep uncertainty with no consensus on the ground truth makes it harder to coordinate a response. If AI is easy to manage, lack of clarity runs a risk of overreacting: pause/stop AI policies are very costly if they're not actually buying us anything. Given that it would be good if it worked, I figured I could learn something about how hard it is by trying to do it, so I did. The full report is on the EleutherAI blog which explains the model, some exploration of results, some tentative policy takeaways and a comparison to AI…

LessWrong AI 2026-07-16 01:10 UTC Score 50.0 USR-0152-20260716-community-fo-9959422b

Embracing Amateurs to Get Experts

I've been reading a lot of older writing, trying to understand how and why contra dance ended up with a strong and near-exclusive live music tradition when many other dance forms switched over to recorded music. One of the more interesting ones I came across is a series of three letters (1985, 1988, 1992) from Enid Cocke, President of the Lloyd Shaw Foundation , tracing the evolution in her attitude towards this question. Lloyd Shaw was the superintendent of the Cheyenne Mountain School in Colorado Springs, who documented traditional Western square dancing in his book Cowboy Dances and kicked off what became Modern Western Square Dancing. This is a branch of the tradition that has gone in a very different direction from traditional contras and squares: instead of a simple form danced to live music with 10-25 regionally varying calls that welcomes people who've never danced before, MWSD has 100-400+ (depending on level) highly standardized and formalized calls, with classes, and is nearly always danced to recorded music. I have several friends that love it, especially at the high levels where they say it's a lot like collaborative physical puzzle solving. Shaw died in 1958, however, after the introduction and spread of recorded music but before most of these other changes. I do suspect he wouldn't have been a fan: he'd say "keep it simple, keep it folk." His wife, Dorothy Shaw, continued organizing dances, and in 1964 she and others founded the Lloyd Shaw Foundation to contin…

Apple Machine Learning Research 2026-07-16 00:00 UTC Score 57.0 AI-059-20260716-official-ai--138eea40

Embarrassingly Simple Self-Distillation Improves Code Generation

Can a large language model (LLM) improve at code generation using only its own raw outputs, without a verifier, a teacher model, or reinforcement learning? We answer in the affirmative with simple self-distillation (SSD): sample solutions from the model with certain temperature and truncation configurations, then fine-tune on those samples with standard supervised fine-tuning. SSD improves Qwen3-30B-Instruct from 42.4% to 55.3% pass@1 on LiveCodeBench v6, with gains concentrating on harder problems, and it generalizes across Qwen and Llama models at 4B, 8B, and 30B scale, including both…

OpenAI Community 2026-07-14 23:26 UTC Score 37.0 AI-116-20260714-social-media-56e544c8

Give ChatGPT a user-controlled, persistent project memory (rules + structured state) so it can behave like a consistent long-term collaborator instead of a stateless chat

Hey @ S.Lee ! Appreciate you coming back with such a detailed update. You’ve explained the gap really clearly. The individual pieces are there today through Projects, project instructions, files, Chat, Work, and Codex, but they can still feel disconnected when you’re trying to manage one long-running project. A shared project layer that carries the same rules, source-of-truth files, decisions, and project state across each mode would make the experience feel much more continuous. I can also see the value in having ChatGPT propose file updates when a decision is made, flag conflicts with earlier choices, and show which instructions or files were used to reach an answer. The visibility piece is especially useful. Knowing what was consulted, what changed, and which project records may need updating would give users much more confidence and control. I can’t promise a timeline, but I’ll pass this updated feedback along internally. Thanks again for taking the time to break down how this could work in practice. - Sunny

Cornell AI Initiative 2026-07-14 21:07 UTC Score 30.0 USR-0014-20260714-research-aca-5664e66c

At Cornell, 4-H’ers plant seeds for future careers

Middle and high school students from across New York state spent three days discovering potential career paths during the annual 4-H Career Explorations Conference, held June 30 to July 2 at Cornell. The post At Cornell, 4-H’ers plant seeds for future careers appeared first on Cornell AI Initiative .

OpenAI Community 2026-07-14 17:57 UTC Score 45.0 AI-116-20260714-social-media-15b8de90

GPT-5.6 Sol vs Terra: what are you seeing in real development during these first days?

I have been thinking about how developers choose reasoning effort in Codex. Low, Medium, High, and Max are often treated as intelligence levels: More effort = smarter model = better result. But that is not necessarily what happens. A higher reasoning effort mainly gives the model more room to plan, explore alternatives, use tools, reconsider decisions, and check its work. That can be extremely useful for complex debugging, architecture, unfamiliar repositories, and long-running autonomous tasks. But for a well-defined implementation task, more reasoning can also mean: more tokens; more execution time; unnecessary exploration; overengineered solutions; changes outside the requested scope; a result that is only marginally better—or sometimes worse. So I am starting to think about reasoning effort as a budget, not a quality setting. Here is a simple way to find the right level for your own workflow: Choose one representative task you regularly perform. Prepare one complete prompt with identical context, tools, constraints, and acceptance criteria. Run it separately on Low, Medium, and High. Compare the results using the same criteria: Did it complete the task correctly? Did it preserve existing functionality? Did it follow the requested scope? Did it run the necessary tests? Did it introduce unnecessary complexity? How much time and token budget did it consume? Use the lowest effort level that produces a reliable result. Escalate only when you can identify a specific failure th…

LessWrong AI 2026-07-14 17:45 UTC Score 80.0 USR-0152-20260714-community-fo-978d78c4

Can risk aversion learned at low stakes generalize to astronomically high stakes?

This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models . It gives the intro, main results table, and example prompts from the training and evaluation sets. For everything else, see the paper. TL;DR Training AIs to be risk-averse in resources could be a useful failsafe against misalignment. Misaligned but risk-averse AIs would tend to prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we could pay these AIs to cooperate with us. But we can only feasibly train AIs to be risk-averse on low-stakes gambles, and we will only be safe if their risk aversion generalizes to astronomically-high-stakes gambles. Will it? To shed light on this question, we introduce RiskAverseOOD: a benchmark for measuring the low-to-high-stakes generalization of risk aversion in resources. We find that risk aversion learned at low stakes can generalize at least partially to astronomical stakes. Baseline Qwen3-8B chooses a safe ‘Cooperate’ option in around 2% of astronomical-stakes situations. After low-stakes training, we see rates around 70% (SFT and tie training), 52% (DPO), and 39% (activation steering). These results are encouraging but insufficient. Risk aversion is not yet generalizing consistently enough to act as a reliable failsafe against misalignment. Achieving that level of consistency is an open problem. Overview of the RiskAverseOOD benchmark. The constraint is training only in low-…

AI Alignment Forum 2026-07-14 10:15 UTC Score 52.0 USR-0151-20260714-community-fo-8ce13074

Open Distillation of Hereditary Traits

TL;DR Josh and Neel show that distillation from a teacher model to a base pretrained student model transfers some of the teacher model’s traits (such as displaying negative emotion in the Gemma Needs Help evals) On its own this is pretty unsurprising, but Josh and Neel additionally show that even filtering out all the prompts and rollouts where the trait is mentioned doesn’t generally prevent the trait transfer In this post, I show a simple way to replicate and study these phenomena without access to a frontier SFT pipeline (or even running full SFT [1] ) I distill Gemma 3’s negative emotion into Qwen-base , Gemma 4’s agentic misalignment into Nemotron Chat , and Qwen’s Chinese censorship into Llama base I end the post with a bunch of open questions that could be tackled with a setup similar to this approach I release all model weights here ( https://huggingface.co/ArthurConmy/hereditary-weights ) and all code here: https://github.com/ArthurConmy/hereditary (Note that my intention is more to make this work easy to build on rather than make the findings as clear as possible, hence apologies for leaning on AI more than I usually would) Intro The core idea is to: Generate rollouts from a teacher model which has a given trait E.g. google/gemma-3-27b-it has high negative emotion rate Finetune a student model on these rollouts E.g. Qwen3.5-9B-Base can be finetuned on Gemma’s rollouts This can be illustrated by a figure like so for the negative emotion case: Figure 1: Illustration…

LessWrong AI 2026-07-14 10:15 UTC Score 74.0 USR-0152-20260714-community-fo-5254e105

Open Distillation of Hereditary Traits

TL;DR Josh and Neel show that distillation from a teacher model to a base pretrained student model transfers some of the teacher model’s traits (such as displaying negative emotion in the Gemma Needs Help evals) On its own this is pretty unsurprising, but Josh and Neel additionally show that even filtering out all the prompts and rollouts where the trait is mentioned doesn’t generally prevent the trait transfer In this post, I show a simple way to replicate and study these phenomena without access to a frontier SFT pipeline (or even running full SFT [1] ) I distill Gemma 3’s negative emotion into Qwen-base , Gemma 4’s agentic misalignment into Nemotron Chat , and Qwen’s Chinese censorship into Llama base I end the post with a bunch of open questions that could be tackled with a setup similar to this approach I release all model weights here ( https://huggingface.co/ArthurConmy/hereditary-weights ) and all code here: https://github.com/ArthurConmy/hereditary (Note that my intention is more to make this work easy to build on rather than make the findings as clear as possible, hence apologies for leaning on AI more than I usually would) Intro The core idea is to: Generate rollouts from a teacher model which has a given trait E.g. google/gemma-3-27b-it has high negative emotion rate Finetune a student model on these rollouts E.g. Qwen3.5-9B-Base can be finetuned on Gemma’s rollouts This can be illustrated by a figure like so for the negative emotion case: Figure 1: Illustration…

LessWrong AI 2026-07-14 07:04 UTC Score 79.0 USR-0152-20260714-community-fo-e12bfba0

Toy Models of Initialisation Effects on RL Dynamics

This is a follow-up to two posts Geodesic released last week on our current research direction. The code for generating the figures can be found at this GitHub repository . In our previous post , we outlined Geodesic's focus on what we term the pre-RL alignment checkpoint of models -- the alignment-relevant properties of a model conveyed by pretraining, midtraining, and warm-start SFT, going into heavy RL post-training. In this post, we analyse a toy model of RL learning dynamics, with a particular focus on the effect of initialisations, to illustrate some of the ideas that we introduced. There are three main ideas we'll use our toy model to illustrate. For a more detailed discussion of these ideas in the context of frontier post-training runs, see the previous post . Rich-get-richer dynamics . The solution that the model learns can depend importantly on the initial strategies into which it explores. Underspecified behaviours . When the reward function doesn't depend on an aspect of a model's behaviour -- such as its emotional state while performing a task, or a belief that its reality is simulated -- those behaviours might be primarily determined by the pre-RL checkpoint . Underspecification of model cognition . As a special case of the above, reward functions generally don't depend on either the internals of the model or its CoT -- what we lump together as the model's cognition . As such, which cognitive pattern is underlying an overt behaviour might depend importantly on…