Latest AI/ML News
31011 matching items
Text Analysis for Hybrid Search: Tokenization, Stopwords & Accent Folding
Tokenization makes or breaks hybrid search. See how Weaviate's accent folding, custom stopwords, and /v1/tokenize endpoint power multilingual BM25.
Kubernetes v1.36: Advancing Workload-Aware Scheduling
AI/ML and batch workloads introduce unique scheduling challenges that go beyond simple Pod-by-Pod scheduling. In Kubernetes v1.35, we introduced the first tranche of workload-aware scheduling improvements, featuring the foundational Workload API alongside basic gang scheduling support built on a Pod-based framework, and an opportunistic batching feature to efficiently process identical Pods. Kubernetes v1.36 introduces a significant architectural evolution by cleanly separating API concerns: the Workload API acts as a static template, while the new PodGroup API handles the runtime state. To support this, the kube-scheduler features a new PodGroup scheduling cycle that enables atomic workload processing and paves the way for future enhancements. This release also debuts the first iterations of topology-aware scheduling and workload-aware preemption to advance scheduling capabilities. Additionally, ResourceClaim support for workloads unlocks Dynamic Resource Allocation ( DRA ) for PodGroups. Finally, to demonstrate real-world readiness, v1.36 delivers the first phase of integration between the Job controller and the new API. Workload and PodGroup API updates The Workload API now serves as a static template, while the new PodGroup API describes the runtime object. Kubernetes v1.36 introduces the Workload and PodGroup APIs as part of the scheduling.k8s.io/v1alpha2 API group , completely replacing the previous v1alpha1 API version. In v1.35, Pod groups and their runtime states we…
mimalloc: A new, high-performance, scalable memory allocator for the modern era
mimalloc is an open-source, modern, scalable memory allocator that is a drop-in replacement for malloc and free. It is relatively small (~12K lines), with clear internal data structures, and is easy to build and integrate into other projects. It provides bounded worst-case allocation times (up to OS primitives), bounded space overhead, low internal fragmentation, and minimal contention by relying almost exclusively on atomic operations. The post mimalloc: A new, high-performance, scalable memory allocator for the modern era appeared first on Microsoft Research .
🔮🇨🇳 Inside the Chinese AI labs where America’s AI controls created its toughest competition
We quantify China’s efficiency moat and what it means for the future of AI
NVIDIA New AI Is An Efficiency Monster
❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here: https://arxiv.org/abs/2604.24954 https://developer.nvidia.com/blog/nvidia-nemotron-3-nano-omni-powers-multimodal-agent-reasoning-in-a-single-efficient-open-model/ https://huggingface.co/blog/nvidia/nemotron-3-nano-omni-multimodal-intelligence Our Patreon if you wish to support us: https://www.patreon.com/TwoMinutePapers 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible: Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/ Thumbnail design: https://felicia.hu #nvidia
GridSFM: A new, small foundation model for the electric grid
Introducing GridSFM, a small foundation model that can predict AC optimal power flow in milliseconds, boosting efficiency and unlocking cost savings. Learn how GridSFM gives grid operators direct visibility into congestion, stability, and system health. The post GridSFM: A new, small foundation model for the electric grid appeared first on Microsoft Research .
Xi seeks to buy time as he meets Trump in Beijing
Xi seeks to buy time as he meets Trump in Beijing H.Seidl Wed, 05/13/2026 - 17:29 picture alliance / ASSOCIATED PRESS | Mark Schiefelbein Comment May 13, 2026 5 min read Xi seeks to buy time as he meets Trump in Beijing With a grand bargain highly unlikely, the US and China face a choice between a marginal deal or no deal at all, argue Helena Legarda and Jacob Gunter. The imminent meeting of Chinese President Xi Jinping and his US counterpart Donald Trump has inevitably triggered speculation about who will come out on top. A “grand bargain” covering the fundamentals of trade, technology, Taiwan and other geopolitical tensions (like Iran) is highly unlikely. Instead, one appealing option for the world’s two most powerful leaders is to make a marginal deal on immediate challenges like maintaining China’s access to the US market and stabilizing China’s export controls on rare-earth flows to the US. The main goal of the summit for Xi is to stabilize China’s relationship with the US and buy time – to keep the US from again ratcheting up tariffs, export controls and other measures which he believes are being used to contain China’s rise. Buying time is crucial for the sweeping self-reliance efforts that have been a hallmark of Xi’s agenda – reducing dependence on foreign technology, finance, and supply chains while strengthening China’s ability to withstand external pressure. From energy security and technological breakthroughs aimed at overcoming US chokeholds on China to fully m…
Chinese investment rises to 7-year high - Chinese FDI in Europe: 2025 Update
Chinese investment rises to 7-year high - Chinese FDI in Europe: 2025 Update c.groth Wed, 05/13/2026 - 12:39 Download (pdf - 1.28 MB) Report May 20, 2026 20 min read Chinese investment rises to 7-year high - Chinese FDI in Europe: 2025 Update Report by Rhodium Group and MERICS Key findings Chinese foreign direct investment (FDI) in Europe (EU and UK) rose for the second consecutive year, reaching its highest level since 2018. It increased by 67 percent to EUR 16.8 billion in 2025. M&A activity drove the rebound, rising 89 percent year-on-year to EUR 7.9 billion. But greenfield investment remained the primary channel for Chinese FDI in Europe, increasing by 51 percent to a record EUR 8.9 billion. Europe made up nearly a quarter of global Chinese FDI in 2025, up from 17 percent in 2024. While Hungary remains the primary destination for Chinese FDI in Europe, more investment is once again flowing into Germany and France. Hungary attracted Chinese investments worth EUR 3.9 billion in 2025, up from EUR 3.2 billion in 2024. Germany (EUR 2.5 billion) and France (EUR 1.9 billion) ranked second and third. Germany’s share of total Chinese FDI in Europe rose to 15 percent from 10 percent in 2024, while France’s increased to 12 percent from 5 percent. The automotive sector attracted more Chinese FDI in 2025 than any other industry. Investments in the sector totaled EUR 7.6 billion, with 93 percent of them focused on the EV supply chain. The auto sector’s share of total Chinese FDI in Eu…
Media Briefing on the Trump-Xi Meeting
Media Briefing on the Trump-Xi Meeting J.Heller Wed, 05/13/2026 - 11:31 Video May 13, 2026 1 min read Media Briefing on the Trump-Xi Meeting US President Donald Trump is set to meet Xi Jinping in his first visit to China since 2017. The meeting, originally delayed due to the US-Israeli conflict with Iran, arrives at a moment of deep structural tension between the world's two largest economies, with disputes over trade tariffs, rare earth access, Taiwan, and AI technology still unresolved. What makes this summit uniquely complex is the Iran war looming over the agenda. The ongoing conflict has inadvertently strengthened China's negotiating hand, and Washington is pressing Beijing to use its leverage with Tehran to reopen the Strait of Hormuz. MERICS Heads of Program Jacob Gunter (Economy and Industry) and Helena Legarda (Foreign Relations) assess what this summit means for the trajectory of US-China relations, the outlook for global trade, and the broader implications of an Iran conflict that is reshaping great power dynamics. The session is moderated by Claudia Wessling , Director of Communications & Publications at MERICS. Related content about US-China Trump's remarks after China Summit compromise Taiwan's security Tracker Jun 18, 2026 China in 26: Diplomatic strength, economic weakness, investment increase Podcast May 22, 2026 US-China summit: Xi warns Trump about Taiwan (ZDF) External publication May 15, 2026
Reimagining the mouse pointer with AI
The mouse pointer 🖱️ has been a constant companion on computer screens, across every website, document and workflow. Despite how technologies have changed, the pointer has barely evolved in more than half a century. We’ve been exploring new AI-powered capabilities and ways of working to help the pointer not only understand what it’s pointing at, but also why it matters. See how to try it out @ https://deepmind.google/blog/ai-pointer ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://twitter.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
Introducing AIMIP: The AI weather and climate model intercomparison project
AIMIP is a new open benchmark and dataset for evaluating AI climate models, showing they can match or beat conventional models on some historical climate metrics while still struggling to generalize reliably to long-term warming trends and unseen climate scenarios.
How we're using Sourcegraph and a Slack bot to detect vulnerabilities and react quickly
A Slack bot triages every GitHub advisory, posts a rocket-to-trigger ask in the channel, and on one human reaction runs the full content pipeline: detection queries, blog scaffold, social drafts, a 35-second auto-cut demo. The operator's remaining job is to read the drafts and decide whether they're honest.
How we're using Sourcegraph and a Slack bot to detect vulnerabilities and react quickly
A Slack bot triages every GitHub advisory, posts a rocket-to-trigger ask in the channel, and on one human reaction runs the full content pipeline: detection queries, blog scaffold, social drafts, a 35-second auto-cut demo. The operator's remaining job is to read the drafts and decide whether they're honest.
AI Weekly Issue #492: AI slop : A $725B bet on what no one wanted
Hyperscalers will spend $725 billion on AI infrastructure this year. The users they are spending it on are now actively rejecting the output. Gartner finds 50% of US consumers prefer brands that don't use generative AI. Wikipedia just banned AI-generated content 44-2. Stack Overflow's new-question volume has fallen 78% year over year. Google AI Overviews have collapsed top-page CTR by 58%. This is the structural tension running through every story below: capacity is being added fastest in exactly the parts of the market where buyers are most visibly walking away.
(Trial) Measures for Artificial Intelligence Technology Ethics Management Services (Draft for Public Feedback)
Read our translation of a Chinese draft regulation mandating ethics reviews for AI technology that could endanger humans or sway public opinion. The post (Trial) Measures for Artificial Intelligence Technology Ethics Management Services (Draft for Public Feedback) appeared first on Center for Security and Emerging Technology .
Monitoring, Audit Trails, and Compliance with ClearML
By Adam Wolf and Damian Erangey The previous posts in this series built the security model layer by layer: identity, configuration governance, service account automation, compute policies, and production model serving. This final post covers what holds all of it together: the monitoring and audit layer that records every action, every API call, and every […]
Kubernetes v1.36: PSI Metrics for Kubernetes Graduates to GA
Since its original implementation in the Linux kernel in 2018, Pressure Stall Information (PSI) has provided users with the high-fidelity signals needed to identify resource saturation before it becomes an outage. Unlike traditional utilization metrics, PSI tells the story of tasks stalled and time lost, all in nicely-packaged percentages of time across the CPU, memory, and I/O. With the recent release of Kubernetes v1.36, users across the ecosystem have a stable, reliable interface to observe resource contention at the node, pod, and container levels. In this post, we will dive into the improvements and performance testing that proved its readiness for production. Beyond utilization: why PSI? Monitoring CPU or memory usage alone can be misleading. A node may report XX% (below 100%) CPU utilization while certain tasks are experiencing severe latency due to scheduling delays. PSI fills this gap by providing: Cumulative Totals : Absolute time spent in a stalled state. Moving Averages : 10s, 60s, and 300s windows that allow operators to distinguish between transient spikes and sustained resource tension. Proving stability: performance testing at scale A common concern when graduating telemetry features is the resource overhead required to collect and serve the metrics. To address this, SIG Node conducted extensive performance validation on high-density workloads (80+ pods) across various machine types. Our testing focused on two primary scenarios to isolate the impact of the Ku…
Partnership on AI Receives $500K Investment From Collaborative Philanthropic Initiative to Advance AI Transparency and Accountability
The post Partnership on AI Receives $500K Investment From Collaborative Philanthropic Initiative to Advance AI Transparency and Accountability appeared first on Partnership on AI .
How open model ecosystems compound
Further reflections on China's high-participation, open-first AI ecosystem.
China Seeks A.I. Independence, Weakening Trump’s Leverage
CSET’s Jacob Feldgoise shared his expert insight in an article published by The New York Times. The article examines how China is accelerating efforts to build a domestic A.I. ecosystem as companies like DeepSeek and Huawei develop alternatives to American chips amid ongoing U.S. export controls. The post China Seeks A.I. Independence, Weakening Trump’s Leverage appeared first on Center for Security and Emerging Technology .
Climate goals for officials + Textile sector + Alcohol industry
Climate goals for officials + Textile sector + Alcohol industry c.groth Tue, 05/12/2026 - 15:54 picture alliance / CFOTO Download (pdf - 576.54 KB) MERICS Briefs MERICS China Industries May 13, 2026 15 min read Climate goals for officials + Textile sector + Alcohol industry MERICS' Top 5 1. New State Council measures tie officials’ careers to climate goals At a glance: The State Council has released measures to evaluate progress on carbon reduction at the provincial level during the 15th Five-year Plan (2026-2030). A new accountability system ties the career prospects of top provincial party and state officials to carbon reduction. Key features include: The introduction of five control indicators, e.g., total carbon emissions, carbon emission intensity and total coal consumption, plus nine supporting indicators, including green transportation (see exhibit 1) Provinces that miss one control indicator or three supporting indicators in annual assessments must propose corrective actions within 30 days. Senior officials who fail to comply will face disciplinary interviews Annual assessment results will be part of senior provincial officials’ performance reviews Officials showing gross violations of duty will face disciplinary action from, for instance, the Central Commission for Discipline Inspection (the party-state’s main body for investigating corruption). MERICS comment: China’s performance evaluation system for public officials contains complex and sometimes contradictory in…
Learn the system
are live models making a comeback?
Vector Institute awards 100 scholarships to Ontario’s top AI graduate students
Vector Institute awards 100 scholarships to Ontario’s top AI graduate students TORONTO, May 12, 2026 – Today, the Vector Institute awarded scholarships to 100 exceptional graduate students pursuing studies across […] The post Vector Institute awards 100 scholarships to Ontario’s top AI graduate students appeared first on Vector Institute for Artificial Intelligence .
How we achieved truly serverless GPUs
A deep dive on Modal's deep tech for fast boots.
Center Director, Masashi Sugiyama delivered an ELLIS Distinguished Lecture (May 11, 2026, Finland)
Masashi Sugiyama, Center Director delivered an ELLIS Distinguished Lecture at Aalto University in Finland on May 11, 2026. Title: Machine Learning from Imperfect Information: Foundations of Robust Intelligence in the Era of Foundation Model
How Sapu Indexed 28 Million PubMed Abstracts to Accelerate Cancer Research with Qdrant
Sapu is an early-stage biopharmaceutical company developing treatments for hard-to-treat cancers. From its San Diego facility, the team is pioneering a nanomedicine pipeline that takes existing FDA-approved drugs and re-engineers them at the nanoscale, making them smaller, more effective, and less toxic. Building on already-approved compounds gives Sapu a stronger and faster path to therapeutic success in an industry where most candidates never reach patients. Behind the lab work sits an AI tooling suite that does the reading, searching, and synthesis that would otherwise take researchers thousands of hours. Sapu’s internal AI platform supports research paper authorship, references standard operating procedures, and lets the team query its document corpus with the precision biotech R&D requires. As the company grew, so did the volume of documents, the variety of use cases, and the demands placed on the underlying retrieval infrastructure.
Satya, Sam To Take The Stand This Week + Highlights From Musk v. Altman
Big testimony is expected this week in a trial that's already produced major revelations. Here's what to look out for + the biggest news so far.
Tickets now available for Brazil’s 3i Festival 2026 on journalism innovation
“At a time of global pressure on journalism, the advance of artificial intelligence (AI) and on the eve of Brazil’s 2026 elections, the 3i Festival is returning for its seventh edition with discussions on the challenges facing digital journalism and the future of information. The event will take place May 29-31 at Porto Maravalley in […] The post Tickets now available for Brazil’s 3i Festival 2026 on journalism innovation appeared first on LatAm Journalism Review by the Knight Center .
Tickets now available for Brazil’s 3i Festival 2026 on journalism innovation
“At a time of global pressure on journalism, the advance of artificial intelligence (AI) and on the eve of Brazil’s 2026 elections, the 3i Festival is returning for its seventh edition with discussions on the challenges facing digital journalism and the future of information. The event will take place May 29-31 at Porto Maravalley in […] The post Tickets now available for Brazil’s 3i Festival 2026 on journalism innovation appeared first on LatAm Journalism Review by the Knight Center .
Fighting Tool Sprawl: The Case for AI Tool Registries
As enterprise AI agent adoption scales, the absence of centralized, organization-level tool infrastructure is producing compounding costs. When adoption is built around optimizing for deployment speed, enterprises expose themselves to a combination of risks: duplicated engineering effort, security exposure, and operational opacity. Every enterprise needs its own shared tool registry, one that reflects its specific regulatory environment, security posture, and operational conventions. To be clear, this is not an argument for a public package manager, something like npm, PyPI, or Maven. The infrastructure each enterprise needs is internal; scoped to its own teams, its own data, its own policies, its own domain. Trying to expand the scope beyond the confines of individual organizations would be premature standardization in a fast-moving, nascent space. A shared enterprise tool registry is not an optimization or a nice-to-have. It is foundational infrastructure as agent deployments scale beyond early experiments. The case for it rests on two pillars: reducing coordination cost and enabling risk management, both for the humans building with agents and for the agents themselves. AI agents depend on tools that retrieve data, write records, trigger workflows, and call external APIs. According to McKinsey, in most large organizations, these tools are built by individual teams in an ad hoc fashion: undocumented, ungoverned, and invisible to the rest of the organization. This pattern i…
The ReSharper 2026.2 Early Access Program Begins: Bringing More AI Agents into Visual Studio
We’re excited to announce that the Early Access Program (EAP) for ReSharper and .NET Tools 2026.2 is now underway! While our EAP announcements usually cover a wide range of new features, performance updates, and bug fixes, this release is different. We are dedicating this first preview entirely to a singular, game-changing initiative: bringing true AI […]
Searching for Birds with Pinecone Full-Text Search
Learn how Pinecone full-text search uses BM25 scoring and Lucene syntax for exact match, boolean, and phrase queries — and how to combine it with vector search.
Import AI 456: RSI and economic growth; radical optionality for AI regulation; and a neural computer
What laws does superintelligence demand?
【Event Report】ELLIS Institute Finland and RIKEN AIP Joint Workshop (Finland, May 11, 2026)
ELLIS Institute Finland and RIKEN AIP Joint Workshop was held on May 11, 2026 at Aalto University in Finland. Researchers from both sides participated engaged in active discussions on AI and machine learning. For more information, please se
【Event Report】ELLIS Unit Milan and RIKEN AIP Joint Workshop (Milan, May 7-8, 2026)
ELLIS Unit Milan and RIKEN AIP Joint Workshop was held on May 7-8, 2026, at the University of Milan. Over the two days, researchers from both sides participated both in person and online, and engaged in active discussions on AI a
Why Artificial Analysis uses Ai2's IFBench instruction-following eval
Artificial Analysis uses Ai2’s open IFBench eval because it captures a stubborn, real-world capability many benchmarks miss: whether models can reliably follow complex, multi-part user instructions.
Qdrant 1.18 - TurboQuant
Qdrant 1.18.0 is out! Let’s look at the main features for this version: TurboQuant: A new quantization method that, at twice the compression ratio of scalar quantization, delivers similar recall and speed. Memory Monitoring: Inspect a collection’s disk, RAM, and page cache usage broken down by component (vectors, payload, indexes, and more) via a new Web UI view and API endpoint. Adding and Removing Named Vectors: Add or remove named vectors to an existing collection’s schema without having to recreate it.
Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity
Summary In February–April 2026, we ran a survey of 349 technical workers (including 87 software engineers, 71 researchers, 129 academics and PhD students, and 48 founders and managers) about their usage of AI tools. Compared to previous work, our survey is one of the more detailed surveys of technical workers’ self-reported gains from frontier AI tools. 1 We attempt to capture gains due to AI in terms of ‘value’ (how much more value are you creating with AI), rather than ‘speed’ (how long would it have taken you to do these tasks without AI). These can give different answers in principle, in particular if using AI changes the distribution of tasks you work on. For example, researchers could use AI to quickly build an interactive dashboard for their data, which would have taken significantly longer without AI but isn’t that important for their project. We provide more detail on the distinction between value and speed gains in our previous research . We think that the distinction between ‘value’ and ‘speed’ gains is important because value is closer to the idea that survey designers typically care about, whereas our sense is that it is common for respondents to think in terms of speed, and we expect that speed changes would typically overstate value changes. See methodological details here . Participants self-reported a median 1.4–2x change in the value in their work due to AI tools. The median self-reported speed change (which we expect to be higher than value change) is 3x.…
BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large language models (MLLMs). In pursuit of ever more detailed and accurate captions, recent work has increasingly turned to reinforcement learning (RL). However, existing captioning-RL methods and evaluation metrics often emphasize a narrow notion of caption quality, inducing trade-offs across core dimensions of captioning. For example, utility-oriented objectives can encourage noisy, hallucinated, or overlong captions that…
AI Weekly Issue #491: 100 years from now : The Last Election
This is 100 Years From Now. Once a week we skip a century and try to picture what life actually looks like when the stuff we're building now has had time to settle in. This week: the last vote.
Ben's Builds #3 - an email app
They have some challenges 😅
How ClearML Fits Into a Zero-Trust Kubernetes Architecture
By Adam Wolf Zero trust is an architectural principle, not a product. It means assuming breach, verifying every connection explicitly, and granting the minimum access required for each interaction. This post covers how those principles apply to Kubernetes AI infrastructure and specifically how ClearML’s security model slots into each layer: network segmentation, workload identity, access […]
Kubernetes v1.36: Moving Volume Group Snapshots to GA
Volume group snapshots were introduced as an Alpha feature with the Kubernetes v1.27 release, moved to Beta in v1.32, and to a second Beta in v1.34. We are excited to announce that in the Kubernetes v1.36 release, support for volume group snapshots has reached General Availability (GA) . The support for volume group snapshots relies on a set of extension APIs for group snapshots . These APIs allow users to take crash-consistent snapshots for a set of volumes. Behind the scenes, Kubernetes uses a label selector to group multiple PersistentVolumeClaim objects for snapshotting. A key aim is to allow you to restore that set of snapshots to new volumes and recover your workload based on a crash-consistent recovery point. This feature is only supported for CSI volume drivers. An overview of volume group snapshots Some storage systems provide the ability to create a crash-consistent snapshot of multiple volumes. A group snapshot represents copies made from multiple volumes that are taken at the same point-in-time. A group snapshot can be used either to rehydrate new volumes (pre-populated with the snapshot data) or to restore existing volumes to a previous state (represented by the snapshots). Why add volume group snapshots to Kubernetes? The Kubernetes volume plugin system already provides a powerful abstraction that automates the provisioning, attaching, mounting, resizing, and snapshotting of block and file storage. Underpinning all these features is the Kubernetes goal of workl…
The efficiency paradox in EU data centre policy
New EU reporting rules for data centre energy and water use may look like progress, but loopholes risk undermining genuine environmental accountability.
Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Overview of adaptive parallel reasoning. What if a reasoning model could decide for itself when to decompose and parallelize independent subtasks, how many concurrent threads to spawn, and how to coordinate them based on the problem at hand? We provide a detailed analysis of recent progress in the field of parallel reasoning, especially Adaptive Parallel Reasoning. Disclosure: this post is part landscape survey, part perspective on adaptive parallel reasoning. One of the authors (Tony Lian) co-led ThreadWeaver ( Lian et al., 2025 ), one of the methods discussed below. The authors aim to present each approach on its own terms. Motivation Recent progress in LLM reasoning capabilities has been largely driven by inference-time scaling, in addition to data and parameter scaling ( OpenAI et al., 2024 ; DeepSeek-AI et al., 2025 ). Models that explicitly output reasoning tokens (through intermediate steps, backtracking, and exploration) now dominate math, coding, and agentic benchmarks. These behaviors allow models to explore alternative hypotheses, correct earlier mistakes, and synthesize conclusions rather than committing to a single solution ( Wen et al., 2025 ). The problem is that sequential reasoning scales linearly with the amount of exploration. Scaling sequential reasoning tokens comes at a cost, as models risk exceeding effective context limits ( Hsieh et al., 2024 ). The accumulation of intermediate exploration paths makes it challenging for the model to disambiguate amon…
EMO: Pretraining mixture of experts for emergent modularity
EMO is a new mixture-of-experts model trained so modular expert groups emerge from data, enabling users to select small task-specific expert subsets while preserving near full-model performance.
Review of the "Risks from automated R&D" section in the Anthropic Risk Report (February 2026)
We reviewed the “Risks from automated R&D” section of Anthropic’s February 2026 Risk Report , producing two corresponding review documents: our original review and our updated review . We recommend that readers refer to our original review, which represents our review of the report as originally received. 1 The following is the executive summary of our original review. The full documents are available as PDFs ( original , updated ). Executive summary This document is METR’s external review of the “Risks from automated R&D” section in the Anthropic Risk Report: February 2026 (henceforth ‘the report’), which makes the argument that catastrophic risk from Claude Opus 4.6 or a less capable Anthropic model automating R&D in any domain is very low. Anthropic shared additional non-public materials with us for our review, and we used some non-public information shared as part of a previous review . We further detail this process in an appendix. We lay out our findings in two sections: Synopsis of Anthropic’s case . Our assessment : We do not think the report adequately supports its conclusion. We note significant issues in a few key areas: Analytical rigor: We have a number of significant issues with the analytical rigor in the overall argument and interpretation of the results of the model use survey. We think that the cited results of the survey provide little evidence about the level of overall risk , due to issues including sample size, question granularity, survey framing, and…
Task Substitution and Uplift
Summary: We describe three different definitions of the productivity impact of AI (AKA uplift), and show there’s reason to expect: \[\text{uplift on old tasks} \leq \text{uplift in value} \leq \text{uplift on new tasks}\] Three Measures of Uplift One complication in measuring AI’s effect on productivity is that it has different effects on different tasks, and this causes people to change how they allocate their time between tasks. This makes it more difficult to talk about the effect of AI on overall productivity. We use “old tasks” to mean the set of tasks you’d do in a typical day before AI is available – your average workday in 2021, say. “New tasks” means the set of tasks you’d do in a typical day after AI is available. Not all new tasks necessarily use AI; they’re just the tasks you choose knowing AI is an option. We have found it important to distinguish between three measures of AI’s uplift: Uplift on old tasks: The factor by which pre-AI time exceeds post-AI time to complete the old tasks. Uplift on new tasks: The factor by which pre-AI time exceeds post-AI time to complete the new tasks. Uplift in value: The factor by which post-AI value exceeds pre-AI value, allowing for reshuffling of tasks between the pre-AI and post-AI cases. In some cases value has a natural definition; in others, it can be operationalized using related definitions discussed more in the accompanying note. This note discusses the distinction and its implications for interpreting AI productivity…
[In the media] “Mathematical formulas can change the rules of competition”: RIKEN’s approach to physical AI (Nikkei Tech Foresight, May 7, 2026)
An interview with Takayuki Osa, Team Director of the Robot Learning Team, was published in Nikkei Tech Foresight on May 7, 2026. The article introduces recent advances in robot