Import AI 451: Political superintelligence; Google's society of minds, and a robot drummer
Are there any genies that can be put back in the bottle?
AI/ML news, top picks, and generated innovation digests.
38322 matching items
Are there any genies that can be put back in the bottle?
Initial signatories include AI pioneers Yoshua Bengio and Geoffrey Hinton, leading media voices Steve Bannon and Glenn Beck, Obama's National Security Advisor Susan Rice, business trailblazers Steve Wozniak and Richard Branson, five Nobel Laureates, former Irish President Mary Robinson, actors Stephen Fry and Joseph Gordon-Levitt, and hundreds of others.
Today, we're joined by Stefano Ermon, associate professor at Stanford University and CEO of Inception Labs to discuss diffusion language models. We dig into how diffusion approaches—traditionally used for images—are being adapted for text and code generation, the technical challenges of applying continuous methods to discrete token spaces, and how diffusion models compare to traditional autoregressive LLMs. Stefano introduces Mercury 2, a commercial-scale diffusion LLM that can generate multiple tokens simultaneously and achieve inference speeds 5-10x faster than small frontier models, paving the way for latency-sensitive applications like voice interactions and fast agentic loops. We also cover the open research challenges in diffusion LLM training, serving infrastructure requirements, and post-training for diffusion-based systems. Finally, Stefano shares his perspective on whether diffusion models can rival or surpass autoregressive LLMs at scale, the advantages for highly controllable generation, and what the future of multimodal diffusion models might look like. The complete show notes for this episode can be found at https://twimlai.com/go/764.
Short note on the LLM Architecture Gallery diff tool for comparing two model architecture stacks side by side.
Here is what happened in AI in Africa this week: 1. AI Surveillance Concerns Rise Across Africa A new […]
Update : Further details on this exercise are included in our Frontier Risk Report (February-March 2026), within the Anthropic section of Appendix B . In collaboration with Anthropic, a METR staff member (David Rein) recently spent three weeks red-teaming a subset of Anthropic’s internal agent monitoring and security systems, many of which are described in the Opus 4.6 Sabotage Risk Report (Appendix 8.4, especially 8.4.8). Anthropic provided substantial access to relevant internal systems and information, and made staff available to answer questions and provide feedback throughout the exercise. The exercise discovered several specific novel vulnerabilities, some of which have since been patched, and none of which severely undermine major claims in the Opus 4.6 Sabotage Risk Report. It also produced several artifacts, including agent trajectories containing covert attacks and a small attack strategy ideation test set. We expect both of these to be useful for ongoing improvements to Anthropic’s monitoring systems. The resulting 26 page report was shared with Anthropic, and a redacted version was shared with a subset of METR staff. We are exploring ways to incorporate more detailed takeaways from the exercise in future METR risk reports. This kind of adversarial testing by external researchers is valuable for discovering vulnerabilities, as well as for developing best practices for embedding third party evaluators inside frontier AI companies. We hope to do more exercises like…
Poisoned LiteLLM packages on PyPI started stealing credentials. Using Deep Search and Code Search, we traced which public repos were protected by version pinning and which were left exposed. Here's how—and how you can do the same for any supply chain incident.
Poisoned LiteLLM packages on PyPI started stealing credentials. Using Deep Search and Code Search, we traced which public repos were protected by version pinning and which were left exposed. Here's how—and how you can do the same for any supply chain incident.
What does “AI at the edge” really mean in 2026, and why does it matter now more than ever before? In this episode, we’re joined by Brandon Shibley, Edge AI Solutions Engineering Lead at Qualcomm’s Edge Impulse, to discuss the current state and future of Edge AI in 2026. We discuss Gen AI, Small Models, and Cascades of Models, along with real-world constraints like latency, power, and privacy. We also dive into the role of MLOps, evolving hardware, and how developers can start building practical edge AI systems today. Featuring: Brandon Shibley – LinkedIn Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Daniel Whitenack – Website , GitHub , X Links: Read our Ultimate Guide to Edge AI Download your copy of O'Reilly's AI at the Edge Check out the Edge Impulse blog Sign-up for an expert led trial of Edge Impulse Upcoming Events: Register for upcoming webinars here !
Argus, a market intelligence and advisory firm in the energy and commodity markets, has launched Brazil’s first assessed daily price for the natural gas spot market. The move from the UK-headquartered company follows the opening of Brazil’s natural gas market to competition under a regulatory framework approved five years ago.
Last month, we introduced Lyria 3, featuring custom music generation designed to spark creative expression. Now, we’re bringing our most advanced music generation model to more Google products, and introducing Lyria 3 Pro. This advanced version allows the creation of tracks up to 3 minutes long, with customization and creative control. Learn more: https://blog.google/innovation-and-ai/technology/ai/lyria-3-pro ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://twitter.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
Image generated with Gemini 3 Pro (Google), 2026. Written by Amber Wang and Y oonji Kim at Lyft. Background Whenever you use the Lyft app, there is a complex balancing act happening behind the scenes. Various levers are used to keep the marketplace running smoothly; Base prices and coupons for riders affect demand, while driver pay and bonuses impact the level of available supply. Since every change to prices and payments impacts Lyft’s costs and revenue, they lead to key optimization problems, such as: How should we allocate budget between driver incentives and rider incentives? How do we invest resources to achieve x% rides growth, and how much does it cost in terms of short term profit? These are the questions the Foundational Models team at Lyft tries to answer in a systematic way. A key ingredient is understanding the effects of different types of investments — for instance, what will happen if we increase the total budget for driver incentives by x%? What will happen if we increase the rider price of all rides by y%? It’s worth noting that the long term effects of such decisions tend to dominate the short term effects: we may earn more short term profit from a ride if we charge riders more and pay drivers less, but lose riders and drivers in the long run. Estimating the long term effects of resource allocation decisions is challenging in a multi-sided marketplace such as Lyft. Because these decisions tend to be consequential, their effects go beyond first order effects…
BMFTR and MWK Discuss AI Research and Transfer Activities at Tübingen AI Center
We are excited to announce our transition to a community-driven open source project. While making this change, we reaffirm our deep commitment to remaining active members of the community.
We are excited to announce our transition to a community-driven open source project. While making this change, we reaffirm our deep commitment to remaining active members of the community.
As AI agents become collaborative partners in complex tasks, understanding how agent personality affects human-AI interaction becomes critical. While recent work explores personality customization in language models, little is known about how personality affects AI coding agents. We conducted the first exploratory study investigating: if OCEAN (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) personality traits can be operationalized in AI coding agents, if users detect these personality differences, and how different personalities affect user trust and adoption. Participants completed refactoring tasks with three agent profiles. Results show that personality traits successfully translated into distinguishable behaviors reliably detected by users. While no universal 'best' personality emerged, individual preferences diverged substantially. Conscientiousness produced more consistent trust, while openness and extraversion polarized users. Some users experienced trust collapse from overconfidence and others found excessive caution inefficient. Our findings provide initial empirical evidence that OCEAN personality traits can be operationalized in AI coding agents, producing distinguishable behaviors, with implications for designing adaptive systems.
MLPerf Inference v6.0 expands open-weight LLM coverage with a new GPT-OSS 120B benchmark and a latency-constrained interactive scenario for DeepSeek-R1 — the first MLPerf standard for speculative decoding. The post A new GPT-OSS benchmark and DeepSeek R1 updates for latency-optimized reasoning appeared first on MLCommons .
In a multivariable logistic regression model, some continuous predictors are modeled using spline transformations to allow for nonlinear effects, while other continuous predictors are entered as simple linear terms. Can the predictors that remain linear still be interpreted in the usual way, that is, through odds ratios for a one-unit increase, even though other predictors in the same model are modeled with splines? To be clear I'm not asking about the linear basis term within a spline expansion of the same predictor. I meant a model like: $$ \text{logit}(P(Y=1)) = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + f(x_3) $$ where $x_1$ and $x_2$ are entered linearly and $x3$ is modeled with a spline. My question is whether $\beta_1$ and $\beta_2$ can still be interpreted in the usual/simple way of presenting single odds ratio values.
Jensen Huang is the co-founder and CEO of NVIDIA, the world’s most valuable company and the engine powering the AI computing revolution. Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep494-sc See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript: https://lexfridman.com/jensen-huang-transcript CONTACT LEX: Feedback – give feedback to Lex: https://lexfridman.com/survey AMA – submit questions, videos or call-in: https://lexfridman.com/ama Hiring – join our team: https://lexfridman.com/hiring Other – other ways to get in touch: https://lexfridman.com/contact EPISODE LINKS: NVIDIA: https://nvidia.com NVIDIA on X: https://x.com/nvidia NVIDIA AI on X: https://x.com/NVIDIAAI NVIDIA on YouTube: https://youtube.com/@nvidia NVIDIA on Instagram:
This is a transcript of Lex Fridman Podcast #494 with Jensen Huang. The timestamps in the transcript are clickable links that take you directly to that point in the main video. Please note that the transcript is human generated, and may have errors. Here are some useful links: Go back to this episode’s main page Watch the full YouTube version of the podcast Table of Contents Here are the loose “chapters” in the conversation. Click link to jump approximately to that part in the transcript: 0:00 – Introduction 0:33 – Extreme co-design and rack-scale engineering 3:18 – How Jensen runs
Voxtral TTS: A frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents.
How will timeless minds value time?
DLSS 5 looks like a real-time generative AI filter for video games, OpenAI Reportedly Pivoting to a Focus on Business and Productivity Only, and more!
The White House published it’s long-awaited AI legislative recommendations on Friday, and it still includes a call for Congress to […]
From MHA and GQA to MLA, sparse attention, and hybrid architectures
The rejection of introspection by America's business leaders—combined with an unwillingness to defend the system that incubated their success—is a deeply troubling trend.
As METR’s time horizon task suite saturates, the results are becoming more sensitive to analysis choices. One example of this was the recent update to fix a modelling mistake with regularization, which decreased recent models’ 50% time horizon results by up to 20%, but had a smaller impact on earlier LLMs’ 50% time horizons. 1 In this post I’ll: Give a refresher on the current model used to calculate time horizon results and more detail about the regularization mistake METR recently fixed Go over what I see as the other main sources of uncertainty in time horizon results (outside of needing more tasks). Where possible, I’ll fit alternative models to show their impacts Wrap things up with general thoughts on how much weight people should put on the current estimates I hope this will help people better understand the modelling assumptions underlying the time horizon results, and how robust (or not) the results are. Summary There are many reasonable variations one could make to the TH modelling, and most of these end up having the effect of reducing recent 50% time horizon estimates (and often increase 80% time horizon estimates). The aspect I feel least certain about is noise in the task length estimates, which I hope to look into more in the future. I find that reasonable choices generally still leave us inside the CIs (which are very wide!). I think the most important source of uncertainty is the task distribution rather than analysis choices, as which tasks are included has…
While direct API calls seem cheaper and easier, they lack the safety layer large organizations rely on. Tool connection protocols aren't dead; they remain vital for security, governance, and centralized control in big teams.
While direct API calls seem cheaper and easier, they lack the safety layer large organizations rely on. Tool connection protocols aren't dead; they remain vital for security, governance, and centralized control in big teams.
In a world obsessed with disruption, Java threads the needle between stability and innovation. It’s the ultimate syncretic platform , synthesizing the best ideas from functional programming, concurrency, cloud computing, and AI under a reliable, battle-tested umbrella. Java unites meticulous planning with chaotic evolution, enterprise reality with open source ideals, along with a healthy dose of benevolent fortune. Let’s look at the key factors that make Java as much a champion today as it was in 1996. 1. The Java Community Process At the heart of Java’s success are the developers and architects who love it. The Java community is vital and boisterous, and very much engaged in transforming the language. But what makes Java special is its governance architecture. Far from a smoothly operating machine, Java’s governance is a riotous amalgam of competing interests and organizations, all finding their voice in the Java Community Process (JCP) . The fractious nature of the JCP has been criticized, but over time it has given Java a massive advantage. The JCP is Java’s version of a functional democracy: A venue for contribution and conflict resolution among people who care deeply about the technology. The JCP is a vital forum where the will and chaos of the worldwide developer community negotiate with Java’s formal managing body. 2. OpenJDK I still remember my astonishment when the Java language successfully incorporated lambdas and closures . Adding functional constructs to an obje…
AI Expo Africa 2026, Africa's Largest Enterprise AI & Automation Event will be running 9th Edition, JHB, South Africa 29-31 October
Introduction METR aims to keep the public informed about the capabilities of and risks posed by AI — by some metrics the fastest-moving technology in history, and one that could speed up further as AI automates AI R&D. By late next year, the rate of model releases and the number of new evals required could be such that even keeping ourselves informed will be a challenge without effective AI assistance. We can’t afford to figure out AI-augmented workflows reactively, as they become necessary; we need to begin understanding them now. So we ran a 2-hour tabletop exercise: three METR researchers played themselves, with their current priorities , but pretending they had access to ~200-hour time horizon AIs – roughly what we expect 12–18 months from now. The goal was to learn what workflows emerge, what the bottlenecks are, and how much faster we’d actually be. The game Scenario The world METR has access to 200h time horizon AIs to automate our work; the rest of the world has access to real Feb 2026 technology (~12h TH AIs). We have versions of Codex/Claude Code + basic project management workflows that make sense for 200h TH AIs. We are otherwise living in Feb 2026, so we’re evaluating 2026 AIs, using the 2026 version of Inspect, communicating with people via email etc. AI capabilities AIs now have a ~200 human hour time horizon , but with a similar relative capabilities profile to early-2026 AIs. They’re staggeringly good at verifiable tasks and decent at messy tasks. AIs work t…
A complete guide on how to secure Weaviate enterprise deployments with OIDC, RBAC, and multi-tenant isolation.
If your business uses AI to screen, rank, or match candidates, the EU now regulates those tools as high-risk systems. Here is what changed, what it means for your operating model, and what you should be doing about it.
MongoDB is excited to announce the general availability of our enhanced data browsing experience in the MongoDB for Visual Studio (VS) Code extension. This new experience offers a unified workspace for developers to visually browse, query, and edit their data natively, streamlining workflows so they can manage their database right where they write their code. Evolving the developer workflow The modern developer’s workflow is incredibly fast-paced. With developers juggling an average of 14 different tools daily, the cognitive load of constantly jumping between applications can easily disrupt focus. When your application needs to evolve, working with your data shouldn’t force a break in your flow state. As the MongoDB for VS Code extension has grown to nearly 3 million downloads, we’ve seen firsthand how developers are pushing the boundaries of what an in-IDE (integrated development environment) database tool can do. While developers love accessing their data directly in the editor, we wanted to transform this experience to be even more visual, actionable, and seamless. Instead of switching to external terminals for quick tasks or taking the time to translate familiar MongoDB Shell commands into Extended JSON (EJSON), we are bringing a full-fledged, intuitive data management suite right to your VS Code sidebar. Exploring what’s new in the MongoDB for VS Code extension Here are the key improvements that transform the extension into a complete workflow solution: Paginated tree v…
Today, we’re introducing Forge, a system for enterprises to build frontier-grade AI models grounded in their proprietary knowledge.
What happens when an AI hater starts building with AI agents? In this episode, we talk with software engineer Steve Klabnik, known for his work on the Rust programming language, about his journey from criticizing AI to experimenting with it firsthand. We explore Steve’s programming language Rue, largely built with the help of AI tools like Claude, and discuss what this means for software engineering and the future of coding in an AI-driven world. Featuring: Steve Klabnik – LinkedIn Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Daniel Whitenack – Website , GitHub , X Links: The Rust Programming Language Rust Rue Daniel's RSA Meeting link for March 23, 2026 Daniel's RSA Meeting link for March 24-25, 2026 Upcoming Events: Register for upcoming webinars here !
"In apparent violation of Brazilian law prohibiting the indiscriminate use and distribution of copyrighted journalistic content, Grok —the artificial intelligence (AI) chatbot developed by Elon Musk— has been ‘tearing down’ news outlets’ paywalls by delivering full newspaper articles that normally require a subscription to access. To test how this works in practice, O Globo newspaper […] The post Elon Musk’s Grok appears to bypass Brazilian news paywalls, newspapers say appeared first on LatAm Journalism Review by the Knight Center .
"In apparent violation of Brazilian law prohibiting the indiscriminate use and distribution of copyrighted journalistic content, Grok —the artificial intelligence (AI) chatbot developed by Elon Musk— has been ‘tearing down’ news outlets’ paywalls by delivering full newspaper articles that normally require a subscription to access. To test how this works in practice, O Globo newspaper […] The post Elon Musk’s Grok appears to bypass Brazilian news paywalls, newspapers say appeared first on LatAm Journalism Review by the Knight Center .
Will AI cause a political interregnum
Study finds digital care technologies could both support and strain unpaid carers, with benefits and risks to loved ones.
Nemotron 3 Super: An Open Hybrid Mamba-Transformer MoE for Agentic Reasoning, Another XAI Cofounder Has Left, Anthropic Sues Department of Defense
Anthropic sues Trump administration in AI dispute with Pentagon, ‘Not built right the first time’ — Musk’s xAI is starting over again, again, Cascade of A.I. Fakes About War With Iran Causes Chaos Onl
Visual gallery of LLM architecture variants: attention mechanisms, positional encodings, MoE, and more — with comparison figures and compact reference sheets.
Robert Lange, founding researcher at Sakana AI, joins Tim to discuss *Shinka Evolve* — a framework that combines LLMs with evolutionary algorithms to do open-ended program search. The core claim: systems like AlphaEvolve can optimize solutions to fixed problems, but real scientific progress requires co-evolving the problems themselves. GTC is coming, the premier AI conference, great opportunity to learn about AI. NVIDIA and partners will showcase breakthroughs in physical AI, AI factories, agentic AI, and inference, exploring the next wave of AI innovation for developers and researchers. Register for virtual GTC for free, using my link and win NVIDIA DGX Spark (https://nvda.ws/4qQ0LMg) In this episode: • Why AlphaEvolve gets stuck — it needs a human to hand it the right problem. Shinka tries to invent new problems automatically, drawing on ideas from POET, PowerPlay, and MAP-Elites quality-diversity search. • The *architecture* of Shinka: an archive of programs organized as islands, LLMs used as mutation operators, and a UCB bandit that adaptively selects between frontier models (GPT-5, Sonnet 4.5, Gemini) mid-run. The credit-assignment problem across models turns out to be genuinely hard. • Concrete results — state-of-the-art circle packing with dramatically fewer evaluations, second place in an AtCoder competitive programming challenge, evolved load-balancing loss functions for mixture-of-experts models, and agent scaffolds for AIME math benchmarks. • Are these systems act…
Mark Hertling discusses U.S. foreign policy, the release of his new book, and the moral-political fork in the road in America in 2026.