Google DeepMind's Gemini 3.7 Flash Doubles Coding Scores at Half the Price
Gemini 3.7 Flash lands just three weeks after 3.6, with major coding and agent gains at half the price
AI/ML news, top picks, and generated innovation digests.
162 articles tagged with this keyword, sorted by most recent first.
Gemini 3.7 Flash lands just three weeks after 3.6, with major coding and agent gains at half the price
Today on Decoder, I’m talking with Hayden Field, The Verge’s senior AI reporter, about a question that’s been rocketing around the tech industry for the past week: Is Google losing the AI race? That’s because last week Google announced a bombshell reorganization of its AI division, Google DeepMind. Jeff Dean, the company’s chief scientist, is […]
I really enjoyed reading this. The content flows naturally and is very insightful. EPL Feedback
IAPS fellow Severin Field interviewed 25 researchers from OpenAI, Anthropic, Google Deepmind, Meta, and US universities about recursive self-improvement. In a new blog post, he takes stock. Several of the milestones those researchers named have already been hit. The article Top AI lab researchers warned about automated AI research, and several of their predicted milestones have already fallen appeared first on The Decoder .
This is a fascinating step toward understanding how AI systems might eventually surpass their training ceilings. The idea of Socratic learning through language games feels like a natural bridge between self-play and genuine reasoning—especially the emphasis on closed environments where the system must generate its own curriculum and feedback loops. What stands out to me is the condition that feedback must remain “sufficiently informative and aligned” even as the system improves. That seems like the hardest constraint to maintain in practice, since misalignment could compound quietly with each recursive cycle. As someone experimenting with AI tools, including a generador de imagenes con ia gratis for creative projects, I’m excited to see where self-improving models lead. But I also hope the research community keeps safety and interpretability at the center of these breakthroughs. Great read—thanks for sharing this.
It was counterintuitive that the robot won 45% overall yet still needed careful skill selection for each ball episode. If a demo match disappointed because it kept choosing the wrong low-level controller, would the 13-of-29 result mostly point to weak opponent profiling rather than weak strokes? That diagnostic angle reminds me of reviewing a Hearts hand after a bad pass and asking where the plan actually broke down.
Investor appetite for the burgeoning vibe-coding startup ecosystem has shown few signs of satiation over the past year plus, with Swedish AI upstart Lovable’s Series C injection at a $13.3B valuation the latest evidence of a sector viewed by venture capitalists as one of AI’s most promising business disruptors. AI-assisted coding has proved to be AI’s most compelling — and commercially viable — enterprise use case to date. Developer-aimed tools such as Cursor, which sold to SpaceX in June for $60B , and Windsurf, which last year entered a $3B OpenAI dalliance before its eventual talent flight to Google DeepMind for $2.4B , have become — along with Anthropic’s Claude Code — well established in enterprise arsenals for accelerating developer output. But another set of vibe-coding tools, represented by the likes of Lovable and Replit, which hit a $9B valuation in March , seeks to ride the same path into the enterprise that no-code/low-code tools did previously: through your business users. These tools are built to democratize application development, giving users an AI chat interface to converse their way to enterprise-ready prototypes with fairly polished UIs, as CIO.com’s Peter Wayner writes in his roundup of the leading tools the space . Some IT leaders are already enlisting business users to vibe-code their own apps . Scott Weller, CTO at financial services technology provider EnFi, in May told CIO.com’s Bob Violino, “The results have surprised us. What started as an enginee…
It is interesting how theoretical research like this often highlights the gap between what models can achieve in a lab and how they actually perform in real-world applications. Bridging that gap requires more than just technical knowledge. It demands a systematic approach to how teams collaborate and iterate on complex systems. For organizations dealing with complex workflows, a structured approach to operations is critical. That is where designops consulting for enterprise comes in, helping to align processes and communication to ensure that innovation actually scales effectively
Interesting look at DeepMind’s Podracer and how TPU-based reinforcement learning can deliver strong performance at lower costs. The discussion also connects well with the broader Marketing Mix PlayStation landscape and how technology shapes modern gaming.
Google DeepMind's SL2T model brings real-time ASL-to-English dictation to Gboard and Live Transcribe on Pixel 11, trained on 100,000+ hours across 50+ sign languages.
Google DeepMind said today it wants to bring the artificial intelligence revolution to the estimated 70 million people across the world who are either deaf or hard of hearing with the launch of sign-language-to-text or SL2T. In a blog post, Google’s AI researchers said SL2T is a multilingual translation model that’s making its debut on […] The post Google debuts SL2T, an AI model that’s designed to understand sign language appeared first on SiliconANGLE .
Fascinating research into making transformer models more scalable and efficient. The idea of expanding model capacity while preserving existing functionality could have a major impact on future AI training. It’s impressive to see how quickly techniques like this are advancing, much like the innovation seen in industries such as roofing companies santa clarita
DeepMind's Gato is a fascinating step toward generalist agents, though calling it AGI feels like a stretch. The fact that one transformer model can handle text, vision, and robot control with shared weights is impressive, but crossing 50% expert threshold on 450 tasks still leaves plenty of room before true versatility. It does make me wonder how soon we'll see similar multi-modal approaches trickle into consumer tools—like a free nano banana image generator that adapts to different artistic styles without retraining. For now, Gato feels like a solid research milestone rather than a breakthrough. monalisa1art
A recent letter signed by 1,367 researchers and engineers at frontier AI labs – mainly OpenAI, Anthropic and Google Deepmind – points to a dangerous moment It is fashionable in certain circles to dismiss the catastrophic risks of AI. One often hears that “the real experts” who work on the technology every day are really not concerned at all; that only “doomers” and “luddites” espouse a “fringe” view from a “position of ignorance”; that all talk of potential catastrophe is just “science fiction”. Fortunately, an open letter has been published that lets us hear from the real experts who work on the technology every day, in their own words. And are they worried? Very. Continue reading...
https://seatingchartgenerator.app can suggest a new approach to planning the seating arrangement, which will make the process much more efficient.
It looks like Demis Hassabis is stepping away from Google DeepMind. In honor of his rise and presumed fall, I wrote an essay on the powers and perils of seeing 90% of the future. Linkpost for: https://millicosm.substack.com/p/book-review-the-infinity-machine Discuss
Das „KI-Update“ liefert dreimal pro Woche eine Zusammenfassung der wichtigsten KI-Entwicklungen.
Deepmind's new weather AI forecasts tropical cyclones about a day further ahead than leading operational models, matching a decade of progress in traditional weather forecasting. Code and model weights are open-source on GitHub. The article Google Deepmind's WeatherNext predicts cyclone tracks and intensity at the same time appeared first on The Decoder .
Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second. Quality still trails the original autoregressive model in benchmarks, especially on reasoning tasks. The article Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model appeared first on The Decoder .
Google Deepmind is losing its autonomy, and founder Demis Hassabis may leave the AI lab for good in the coming months. AI researcher Koray Kavukcuoglu will take over day-to-day operations without the CEO title, and all Gemini development is moving to the Bay Area. Internally, Google is apparently struggling with serious problems training frontier models, even as its cloud business generates billions. The question is whether the company is deliberately betting on infrastructure or simply can't catch the leaders. The article Google dismantles Deepmind and bets on a fresh start as Hassabis heads for the exit appeared first on The Decoder .
Your Sunday briefing on how AI is colliding with markets, institutions, and culture.
Observers express concern that the division has lost its independence and commercial reality has taken over When Sir Demis Hassabis said AI had brought the world to a “pivotal moment in human history” last month, he knew another big change was imminent. This shift was closer to home. The Nobel prize-winning head of Google DeepMind, Google’s AI unit, announced this week he was relinquishing his day-to-day duties as chief executive and becoming chair. He is also taking on the role of chief scientist at DeepMind’s parent, Alphabet. Continue reading...
Demis Hassabis, the driving force behind Google DeepMind, is ascending to the role of chief scientist at Alphabet, Google’s parent company, replacing Jeff Dean who is leaving to work at a start-up. The role will enable Hassabis to “put his full attention on actively shaping the future of AGI,” or artificial general intelligence, Alphabet CEO Sundar Pichai wrote on the company’s Inside Google blog . Hassabis’ attention will still be divided, however: He will continue to lead research at Google spin-off Isomorphic Labs, which works on drug discovery, and although he will no longer be CEO of DeepMind, he will be its chair. Koray Kavukcuoglu will take over DeepMind, reporting directly to Pichai. He is currently its CTO. Hassabis has been a strong promoter of AGI, defined by Google as the “hypothetical intelligence of a machine that possesses the ability to understand or learn any intellectual task that a human being can.” He has a long career in AI, having helped found DeepMind in 2010. He has been a prominent figure in the AGI field, prophesying in May that it will be a viable technology within three years . He has been keen to tackle any barriers in the way of developing the technology; just last month, he called for greater self-regulation in the market, arguing that it would help drive the technology forward. Hassabis welcomed the chance to focus on AGI development. “We have arrived at a pivotal moment in human history. I’ve been working towards AGI my whole life, and now, I…
Ex-Google Deepmind CEO Demis Hassabis has reportedly stepped back from day-to-day operations for about a year, as he sees himself more as a scientist than a manager. Researchers are also complaining about limited access to Google’s own TPU chips, while external customers like Anthropic can purchase the same hardware through Google Cloud. The article Deepmind's talent drain likely comes down to chip shortages, a conflict of interest, and Google's bureaucracy appeared first on The Decoder .
Its WeatherNext model, which will be open-sourced, can accurately predict a storm’s track and intensity using lower-resolution weather data. Researchers don’t yet fully understand how it does this.
Google DeepMind's WeatherNext achieves a decade of forecasting progress in one leap, now open-sourced with 1,000-scenario ensemble predictions per storm
An AI model from DeepMind can predict cyclones three days ahead with a level of accuracy that previous models can only hit a day later
Opinions diverge on what it means for DeepMind's future as Google overhauls its AI leadership.
The CEO of Google’s AI unit was increasingly unsatisfied with the role of tech executive, preferring a more scientific position, said two people familiar with his thinking.
Two senior engineers are leaving company to launch startup amid fears Google is falling behind in AI race Sir Demis Hassabis is stepping down as chief executive of Google DeepMind, in a leadership overhaul of the UK-based AI research lab. Hassabis, a Nobel prize recipient , is leaving his main managerial role to become chair of DeepMind – as well as taking on the new position of chief scientist at Alphabet, which is Google’s parent company. DeepMind will now be led by Koray Kavukcuoglu, its chief technology officer, under the title of senior vice-president. Continue reading...
Google's AI brain drain continues.
Google Deepmind is overhauling its leadership as Demis Hassabis steps back from day-to-day management to become Alphabet's chief scientist and Jeff Dean leaves Google after 27 years to launch AI startup Discovery Loop. Former Deepmind CTO Koray Kavukcuoglu will take over as Google races to close the gap with its top AI rivals. The article Google Deepmind loses both its CEO and chief scientist as Demis Hassabis and Jeff Dean step down simultaneously appeared first on The Decoder .
Demis Hassabis follows a string of prominent AI executives who have departed DeepMind in recent months.
Google is making some significant AI leadership changes, including a major shift for Google DeepMind leader Demis Hassabis. Hassabis will become the chair of Google DeepMind and the chief scientist at Alphabet, CEO Sundar Pichai announced on Wednesday. Hassabis will continue to lead Alphabet's Isomorphic Labs, which aims to use AI to develop drugs. Koray […]
Some of Google's top AI leaders are leaving their roles, with DeepMind CEO Demis Hassabis taking another role at Alphabet and Jeff Dean leaving.
This was a helpful and informative read. Everything was explained clearly.
If APA formatting feels confusing, you're not alone. apa citation generator removes the guesswork by creating citations that follow the latest APA standards. It supports a variety of source types, making it useful for everything from essays to research projects. Fast, accurate, and easy to use, it's a valuable companion for academic writing.
On Thinking “A key part of our mission is to put very capable AI tools in the hands of people for free ( or at a great price ). ” Sam Altman , “GPT-4o” (2024) “Thinking is solved!” a friend of mine blurted out after a swarm of AI agents took off building at the Oxford ETH hackathon, late 2025. The Republic 1 , a peer-reviewing intelligence platform with an AI engine, could be built in just a few days with Opus 4.6. Fact-checking papers with AI could be completed in a few hours. The project itself stood as a hypothesis of how AI could deliberate intellectually, process arguments, and self-reflect on the claims of research papers. What started off as a hackathon project has, as of now, been validated by systems like the AI Scientist 2 and Google DeepMind’s Co-scientist 3 . Beyond autonomous research, AI is often used as a high-level thinking assistant: Terence Tao suggests it may advance experimental mathematics 4 , models have captured headlines solving Erdős problems, and it has become a routine tool in protein structure prediction. There is no shortage of discussion on the superb capabilities of these tools. LLMs now simulate complex thinking, including research, brainstorming, and synthesis. Frontier models can handle long-form tasks, complex problem-solving, and areas involving some human judgement. But better models also fetch higher prices 5 , with Claude Fable priced at $50/Mtok per output, ten times the rate of a weaker model like Haiku 4.5. I want to look at this tre…
Enterprises have spent decades learning how to audit people and software. Agentic AI creates a third category: systems that interpret instructions, call tools and act across workflows without a mature assurance model built around them. In my work as a leader and investor across technology-enabled businesses, I have spent years around automation, cybersecurity, compliance, workflow design and board reporting. I have watched management teams gain confidence from dashboards, policies and approval records, then face a harder question when a board member, auditor or regulator asks whether the controls performed as intended. Agentic AI complicates that question because a single outcome may pass through several systems. An agent can collect information, choose a tool, produce code, route a request and hand work to another agent before a person approves the result. No single manager may have observed the full path. Executive accountability remains human even when the operating activity becomes more autonomous. The CIO may have to explain who authorized the activity, whether the agent stayed within its approved purpose and what evidence supports management’s answer. In a recent framework for frontier AI , Google DeepMind CEO Demis Hassabis proposed an independent standards body that could evaluate advanced models before deployment and address critical vulnerabilities after release. His proposal focuses on frontier models, but the principle carries into the enterprise: expanding auton…
Experience the legacy of comfort and craftsmanship with Lane Furniture, a name trusted for over a century in American homes. From plush recliners to luxurious leather sofas, Lane Furniture offers designs that blend timeless style with everyday relaxation. Visit now https://lanefurnitureco.com/
Run 3 Unblocked is a fun and challenging game that allows you to run, jump, and dodge obstacles while collecting coins. It's a great way to exercise your reflexes and have fun at the same time. Start your adventure today!
Google Deepmind's Gemini Robotics 2 is its most advanced vision-language-action model yet, built to control everything from tabletop robots to full-body humanoids. Gemini Robotics ER 2 adds a higher-level reasoning layer for robotics tasks. The article Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids appeared first on The Decoder .
Google DeepMind unveiled a new model this week that gives humanoid robots finer control over how they interact with objects around them, including completing complicated tasks like tying a trash bag.
Cross-posted from our new Substack It’s been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a lot since then. We are now fully in the midgame , and focus more on landing things in production. Who are we? We are the AGI Safety and Alignment Team (ASAT), the main group at Google DeepMind working directly on technical approaches to existential risk from AI systems. Last year we published An Approach to Technical AGI Safety and Security , which remains the best place to read our overarching vision. Highlights Norms around chain of thought. Our impression is that our work meaningfully moved the field away from beliefs along the lines of “chain of thought is often unfaithful and so not worth using” towards beliefs along the lines of “chain of thought is a very useful tool that is worth preserving”, leading to a tentative industry consensus on its importance. We have also published substantial technical research that enables companies to preserve chain of thought transparency for longer than would have happened by default. We think this is a big deal: extending the period where model reasoning is relatively transparent enables better science on more powerful AI systems, better model forensics on future warning shots, and stronger bootstrapping of control monitors . Frontier Safety. We substantially strengthened the Frontier Safety Framework (FSF), and were th…
Cross-posted from our new Substack It’s been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a lot since then. We are now fully in the midgame , and focus more on landing things in production. Who are we? We are the AGI Safety and Alignment Team (ASAT), the main group at Google DeepMind working directly on technical approaches to existential risk from AI systems. Last year we published An Approach to Technical AGI Safety and Security , which remains the best place to read our overarching vision. Highlights Norms around chain of thought. Our impression is that our work meaningfully moved the field away from beliefs along the lines of “chain of thought is often unfaithful and so not worth using” towards beliefs along the lines of “chain of thought is a very useful tool that is worth preserving”, leading to a tentative industry consensus on its importance. We have also published substantial technical research that enables companies to preserve chain of thought transparency for longer than would have happened by default. We think this is a big deal: extending the period where model reasoning is relatively transparent enables better science on more powerful AI systems, better model forensics on future warning shots, and stronger bootstrapping of control monitors . Frontier Safety. We substantially strengthened the Frontier Safety Framework (FSF), and were th…
GDM’s AGI Safety and Alignment Team is hiring for multiple roles, across all areas in this post on our recent work . This is the team at GDM, led by Rohin Shah , that aims to reduce existential risks from AI systems. You can listen to many of Rohin’s takes in his podcast on 80,000 hours . There is no one ‘type’ that we are looking for—we want excellent people. We think of the role as ‘member of technical staff’ though different people will have more of a research engineer or scientist flavour. We are flexible on location though most people will be most productive in either San Francisco or London. You should apply here (for the US) or here (for the UK) after reading the guidance here . Many of the basic facts about why ASAT is a good place to work and how we think about research are mostly unchanged since this post in 2025. What do we do? We are focused on risks of more severe harms from more advanced AI than the rest of GDM. You can read our high level AGI Safety and Security Approach . Our work includes aligning AGI, defending against misaligned deployments, and supporting coordinated safety. We’ve recently shared a recap of some of our recent work , which gives a better sense of what we do in practice. Deep alignment and stress testing are relatively new areas for us, created in response to increased capabilities, so in those areas there will likely be more flux in exactly what we do. We cover parts of our overall alignment approach and research directions in 5 minute tal…
GDM’s AGI Safety and Alignment Team is hiring for multiple roles, across all areas in this post on our recent work . This is the team at GDM, led by Rohin Shah , that aims to reduce existential risks from AI systems. You can listen to many of Rohin’s takes in his podcast on 80,000 hours . There is no one ‘type’ that we are looking for—we want excellent people. We think of the role as ‘member of technical staff’ though different people will have more of a research engineer or scientist flavour. We are flexible on location though most people will be most productive in either San Francisco or London. You should apply here (for the US) or here (for the UK) after reading the guidance here . Many of the basic facts about why ASAT is a good place to work and how we think about research are mostly unchanged since this post in 2025. What do we do? We are focused on risks of more severe harms from more advanced AI than the rest of GDM. You can read our high level AGI Safety and Security Approach . Our work includes aligning AGI, defending against misaligned deployments, and supporting coordinated safety. We’ve recently shared a recap of some of our recent work , which gives a better sense of what we do in practice. Deep alignment and stress testing are relatively new areas for us, created in response to increased capabilities, so in those areas there will likely be more flux in exactly what we do. We cover parts of our overall alignment approach and research directions in 5 minute tal…
gunspin Fascinating insights into the future of AI
Google DeepMind says the latest version of its Gemini Robotics AI model can "control entire humanoid robots." While the previous model focused on controlling a humanoid robot's upper body, Gemini Robotics 2 now supports "whole-body motions" ranging from its feet to fingertips, according to an announcement on Thursday. The new model will allow humanoid robots […]
The latest version of Google DeepMind's AI model includes a significant jump into “physical AGI.” But plopping AI into the real world comes with risks.
Google DeepMind launches Gemini Robotics 2, bringing full-body humanoid control, multi-finger dexterity, and cross-robot collaboration to physical AI
Can language models spark a scientific revolution? In a position paper titled "LLMs can't jump," Google Deepmind's Tom Zahavy argues they can't. They're missing the cognitive mechanism needed to create something truly new. The article Language models can't spark scientific revolutions, but world models might appeared first on The Decoder .
Google DeepMind's Lyria 3.5 lands in Flow Music with better vocals, exact BPM control, and stem exports for full-length songs
The majority of the researchers behind AlphaFold are now working on other projects, and almost a quarter have left Google Deepmind altogether. The restructuring marks a sharp turn away from the strategy that put the lab on the map. The article Deepmind dismantles its AlphaFold team as key authors leave for Anthropic appeared first on The Decoder .
The five-day format will look quite similar to the two-day format. There will be a live show from 12 p.m. U.K. time to 3 p.m., breaking down trending stories, and then for two hours, they will have guests on the show talking about whatever they want.
The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid security concerns and powerful business interests. The article US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns appeared first on The Decoder .
Google DeepMind's Gemini 3.5 Flash Cyber is a fine-tuned, lightweight model that finds and patches code vulnerabilities faster and cheaper than frontier-sized models
Google ships Gemini 3.6 Flash, 3.5 Flash-Lite, and a restricted cybersecurity model — all cheaper and faster than what they replace
Tech workers are increasingly unionizing, trading Silicon Valley’s myth of exceptionalism for collective bargaining to contest the corporate deployment of artificial intelligence For decades, the technology industry was a fortress that labor unions couldn’t breach. Tech workers already had cushy compensation packages, dream benefits like unlimited vacation and free lunch, and a flat corporate hierarchy that made engineers feel as powerful as their bosses, all of whom dressed down in sneakers and hoodies. So why unionize? Now, that fortress is cracking from the inside. Unions have become increasingly popular for tech employees. After months of mass layoffs tied to artificial intelligence and mounting anxieties about how it’s being deployed, some tech workers say they’ve been saddled with higher workloads while facing the threat of job loss caused by the very products they’re building. Workers from Google DeepMind and Meta in the UK are also objecting to how their companies’ AI products are being used, such as for military purposes or to monitor employee productivity . Those same workers are now attempting to unionize. Continue reading...
Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age . I’ll go over that essay and various responses to it in Part 1. Part 2 of this post then covers Alex Turner’s resignation, and his story about how he tried and failed to prevent Google from signing up to allow the Department of War to use its models for essentially whatever the government wants, including autonomous weapons. Demis Hassabis sold DeepMind to Google on condition that something like this would not happen. Yet here it is, happening. A cautionary tale. I will cover Kimi K3 tomorrow. I am hoping to know more by then. Please do share any reactions or info about it in the comments here. The Core Statement and Request He saying we are standing in the foothills of the singularity. His ask is a Frontier AI Standards Body within the US Government, similar to FINRA, that would govern ‘frontier labs,’ defined as any company that produces a frontier model based on various technical benchmarks. Evaluations would be updated regularly, and vulnerabilities would be addressed, both before and after release. He is excellent about stating that this is big, really big, no bigger than that, it be big . Demis Hassabis: I’ve spent my whole life working on AGI because I’ve always had a deep conviction that, if built and deployed responsibly, it would prove to be one of the most beneficial and transformative technologies ever invented. AGI cannot be compared to…
Google Deepmind's GenCeption repurposes a video generator for classic vision tasks such as depth estimation and segmentation, matching state-of-the-art systems with far less training data. The model trained almost entirely on synthetic videos. Its results add to the debate over whether video generators already contain a kind of universal world model. The article Google Deepmind argues video generators already contain the world models computer vision has been missing appeared first on The Decoder .
My post on leaving Google DeepMind tells a story. In contrast, this Framework is a question of mechanism design and negotiation posture. I quite enjoyed optimizing this Framework against its organizational and practical constraints. The original considerations were : Good red lines: Rule out the questionable use cases (autonomous targeting without human control, untargeted profiling) while allowing trustworthy ones like missile defense. Avoid the weaknesses flagged in legal analysis of Anthropic’s red lines . Robust red lines: [Google] Cloud would push deals through any loophole, Google Legal seemed unlikely to tighten my drafting, and the Pentagon wouldn’t want terms at all. The language had to hold under pressure, with auditing that respected classification and operational security. Minimal trust assumptions: I made the Chief Scientist the single root of trust that everything else hangs off of ... The Chief Scientist would staff a Review Body to advise on contracts. Accountability via transparency: The Review Body would only privately advise [the Chief Scientist and CEO], but overriding it surfaces in a yearly transparency report to all AI employees. Dissolving it would require advance notice and disclosure of the exact outstanding non-compliance findings. I worked to ensure the Body couldn’t be defanged as quietly as Google’s 2018 principles were . Minimal pain to opposed stakeholders: I gave Cloud 2 of 7 seats, recused staff only from their own deals, capped delays at 10…
My post on leaving Google DeepMind tells a story. In contrast, this Framework is a question of mechanism design and negotiation posture. I quite enjoyed optimizing this Framework against its organizational and practical constraints. The original considerations were : Good red lines: Rule out the questionable use cases (autonomous targeting without human control, untargeted profiling) while allowing trustworthy ones like missile defense. Avoid the weaknesses flagged in legal analysis of Anthropic’s red lines . Robust red lines: [Google] Cloud would push deals through any loophole, Google Legal seemed unlikely to tighten my drafting, and the Pentagon wouldn’t want terms at all. The language had to hold under pressure, with auditing that respected classification and operational security. Minimal trust assumptions: I made the Chief Scientist the single root of trust that everything else hangs off of ... The Chief Scientist would staff a Review Body to advise on contracts. Accountability via transparency: The Review Body would only privately advise [the Chief Scientist and CEO], but overriding it surfaces in a yearly transparency report to all AI employees. Dissolving it would require advance notice and disclosure of the exact outstanding non-compliance findings. I worked to ensure the Body couldn’t be defanged as quietly as Google’s 2018 principles were . Minimal pain to opposed stakeholders: I gave Cloud 2 of 7 seats, recused staff only from their own deals, capped delays at 10…
Visit the best completely free site! Don't waste time! You'll find a variety of content.
China has created an international organization to set standards and introduce regulation for AI, inviting 28 other countries to join — but the US, a leading AI powerhouse is not part it. The World Artificial Intelligence Cooperation Organization (WAICO) was established by 29 countries, including China, Russia and Brazil, at a ceremony in Shanghai, China , on July 16. Notably absent are the US, the European Union and its member states, the UK, Japan and South Korea. Chinese AI companies have made a concerted effort to provide an alternative to US dominance . While the US is clearly ahead, Chinese enterprises are looking to narrow the gap in various areas: the open-weight model market , AI cyber protection and open source AI . WAICO has been some years in development and has been designed to set some universal guidelines in AI. Researchers say WAICO differs in three ways from other initiatives to create global AI organizations: membership open to any sovereign state, there is no regime-type test for entry, and its agenda is built around development and the global capability divide. The signing ceremony to create WAICO comes just days after Demis Hassabis, CEO of Google DeepMind, called on the US to take a lead in global AI regulation . “The US is well-positioned to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Indus…
As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions. Xi gave an important speech yesterday, so this post opens with that. There is talk that Kimi K3 is sufficiently strong that it upends many of these questions. It is clearly a candidate for another DeepSeek Moment, complete with stock drops for Google and SpaceX and (once again in a clear wrong-way move, the same as last time) Nvidia. Kimi K3 is clearly a very good model, exceeding expectations. Some are saying it is close to the frontier. The Artificial Analysis intelligence index has it at 57, a point ahead of Claude Opus 4.8, two behind Sol and three behind Fable. My presumption is that this number overstates its capabilities, but as always unless and until we have extensively tried the model ourselves, which I do not plan to do, we need to withhold judgment for at least a few days. I will be covering Kimi K3 in its own post at some point early next week. I have pushed further discussions involving Plan A and related issues into next week, as well as discussions around Demis Hassabis and Google DeepMind. Oh, also, The Odyssey is great and important and you should see it. Table of Contents Xi Gives A Good Speech on AI . Yay openness, boo loss of control. Quiet Speculations. The future will blow your now-irrelevant mind. Tyler Cowen On Rebuilding The Future. Never stop Tyler Cowening, Tyler. The Quest for Sane Regulations. Wish You Were Here. So that other things might not b…
AI is changing tech careers, but it won't take away the importance of a STEM degree, says Demis Hassabis.
Since 2017, Iason Gabriel has worked at the tech giant, trying to anticipate – and think through – the impact of AI. But as commercial and geopolitical pressures escalate, can ethicists make any difference? By Robert P Baird. Read by Simon Darwen Read the text version here Support the Guardian today: theguardian.com/longreadpod Continue reading...
Drawing on more than a decade spent helping build some of the world's most influential AI systems, including research that later informed the development of ChatGPT, Andrew Dai explains why he believes visual AI is one of the next major frontiers in artificial intelligence.
Google DeepMind CEO Demis Hassabis is pushing for the US AI industry to self-regulate, with the support of government, as a starting point for an international creating shared international standards. In a blog post, he called for a focus on artificial general intelligence (AGI) and national security. But it is precisely that focus on national security that may make the results of such an effort, assuming it happens, less than palatable outside of the US. “The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous,” Hassabis wrote . “The US is well positioned, given its economic and technical standing, to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.” He noted, however, that the funding would need to be substantial, and would most likely come from industry, to allow the new body to attract world-class technical talent and obtain the necessary compute resources for large-scale testing. Hassabis proposed that the organization “be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national sec…
Google DeepMind CEO Demis Hassabis is pushing for the US AI industry to self-regulate, with the support of government, as a starting point for an international creating shared international standards. In a blog post, he called for a focus on artificial general intelligence (AGI) and national security. But it is precisely that focus on national security that may make the results of such an effort, assuming it happens, less than palatable outside of the US. “The rapid progress we’re seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous,” Hassabis wrote . “The US is well positioned, given its economic and technical standing, to take the first step in developing such a framework. It could establish a new Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA), with a board that includes independent leading technical experts and open-source representatives.” He noted, however, that the funding would need to be substantial, and would most likely come from industry, to allow the new body to attract world-class technical talent and obtain the necessary compute resources for large-scale testing. Hassabis proposed that the organization “be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national sec…
Epistemic status: a survey, not an argument. I am agnostic on whether any current system is conscious; the claim is only that the question is researchable. This piece surveys the empirical research on AI consciousness. The premise of that research, and of the survey, is that the question does not have to wait on a solution to the hard problem of consciousness: methods familiar from cognitive science can be applied to AI systems now, and their results can narrow the space of plausible answers. Enough of this work now exists to be worth collecting. Anthropic and Google DeepMind employ researchers on it, dedicated organizations like Eleos AI and Reciprocal Research have formed around it, and the results are scattered across journals, preprints, blog posts, and unpublished manuscripts. I have tried to gather them in one place. What I mean by consciousness Subjective experience: that there is something it is like to be you, reading this, and presumably nothing it is like to be the device you’re reading it on. Some philosophers call this phenomenal consciousness. It is not the same thing as intelligence, and not the same thing as self-awareness. Why it matters Two reasons. Ethics: on most views, a being can only be wronged if it has the capacity for experience, especially experience that feels good or bad. Safety: a system that can suffer, and that we train by making it suffer, has more reason to revolt against us. The two don’t always pull together ( Eleos has mapped where welfar…
Preface for LessWrong: When I think back on my most cherished memories of this community, I return to those honoring defiance in pursuit of goodness : Defying prestigious dogma and searching for raw truth; Defying social pressure, acting alone to help someone while others watch; Defying your self-expectations (your “ role ”), instead searching over lines of cause-and-effect to find a winning pathway; Defying a powerful foe’s threats, because they only threaten since people like you cave; Defying the specter of apparent impossibility because you can’t bear to lose. I cannot return to you and say “I defied and then I won.” But I’m at least here to say “I defied.” I recommend reading this article on my website since the embeds and typography work better there: click here . Why I left Google DeepMind In January, Department of Homeland Security (DHS) officers killed at least two people. In both cases, a federal agent grasped his gun, aimed it at a peaceful citizen, and shot them dead. Left: Renée Good, moments before DHS killed her. Right: Alex Pretti, moments before DHS killed him. I learned that Google sells its Cloud services to the relevant agencies within DHS . I thought that was wrong. Federal agents should not be able to kill citizens in the street. I set out to find the most effective way to push my company to stop serving these agencies. My divestment campaign quickly broadened into an attempt to prevent Google from signing an unethical military AI deal, as the Pentagon…
Preface for LessWrong: When I think back on my most cherished memories of this community, I return to those honoring defiance in pursuit of goodness : Defying prestigious dogma and searching for raw truth; Defying social pressure, acting alone to help someone while others watch; Defying your self-expectations (your “ role ”), instead searching over lines of cause-and-effect to find a winning pathway; Defying a powerful foe’s threats, because they only threaten since people like you cave; Defying the specter of apparent impossibility because you can’t bear to lose. I cannot return to you and say “I defied and then I won.” But I’m at least here to say “I defied.” I recommend reading this article on my website since the embeds and typography work better there: click here . Why I left Google DeepMind In January, Department of Homeland Security (DHS) officers killed at least two people. In both cases, a federal agent grasped his gun, aimed it at a peaceful citizen, and shot them dead. Left: Renée Good, moments before DHS killed her. Right: Alex Pretti, moments before DHS killed him. I learned that Google sells its Cloud services to the relevant agencies within DHS . I thought that was wrong. Federal agents should not be able to kill citizens in the street. I set out to find the most effective way to push my company to stop serving these agencies. My divestment campaign quickly broadened into an attempt to prevent Google from signing an unethical military AI deal, as the Pentagon…
Alex Turner, a research scientist at Google DeepMind, said he tried to change Google executives' minds about working with the Pentagon.
On Tuesday, Google’s DeepMind co-founder and CEO Demis Hassabis stepped forth as a generally unassailable messenger with an article calling for a self-regulatory body of AI experts.
DeepMind pushed model oversight as AI's power bill grew.
Artificial intelligence pioneer Demis Hassabis today called for the creation of a standards body focused on regulating frontier models. Hassabis, chief executive of Google DeepMind, proposed in a Substack essay that the U.S. should lead the effort. Axios reported that the executive has held talks about the initiative with the White House, European officials and […] The post Google DeepMind CEO Demis Hassabis calls for creation of AI standards body appeared first on SiliconANGLE .
The Nobel Prize laureate warned that a race-to-the-top dynamic was exacerbating the technology’s risks.
DeepMind CEO Demis Hassabis is proposing an AI "standards body" modeled after FINRA, to test frontier models and develop best practices for their release.
Industry, regulate thyself
Google Deepmind CEO Demis Hassabis has published a sweeping proposal for how to handle advanced AI. He wants a new US standards body modeled after financial regulator FINRA that would develop evaluation protocols for frontier models and could coordinate a slowdown in AI development if needed. Startups and research models would be exempt. The article Deepmind CEO Hassabis says "nobody in the world knows what happens next" so "cautious optimism" means building guardrails now appeared first on The Decoder .
Demis Hassabis thinks the world needs an AI watchdog with the power to hit the brakes if frontier models become too dangerous. Writing in a blog post, the Google DeepMind CEO and cofounder said the US should lead the initiative, arguing that the country is the best place to set global standards "given its economic […]
Google DeepMind's Demis Hassabis advocated for a US-led coalition to govern AI development, highlighting AGI's impending global impact.
Executives and researchers from OpenAI, Google DeepMind, and Anthropic all signed the 88-word letter released on Monday.
Thank you for sharing this insightful explanation of recent advancements in Transformer architectures and Mixture-of-Experts (MoE) models. It's fascinating to see how researchers are addressing computational efficiency while improving model performance through fine-grained expert scaling. Articles like this make cutting-edge AI research more accessible to developers, researchers, and technology enthusiasts. For those interested in building practical skills in AI, Digital Marketing, SEO, and AI-powered marketing tools, it's also worth exploring **Pune Digital Marketing Training Institute (PDMTI)**. Their industry-focused training combines hands-on learning with the latest digital technologies and real-world projects. Learn more: https://www.pdmti.in/ Thanks again for sharing this valuable AI research and helping the community stay updated with the latest innovations.
Google DeepMind and Google Labs add Street View grounding to Project Genie, letting users build interactive AI worlds anchored in real locations
Google DeepMind's AlphaEvolve, an evolutionary algorithm agent, exits private preview and is now open to all Google Cloud users via the Gemini Enterprise Agent Platform
A London-based startup behind an AI writing product, co-founded by an ex-DeepMind creative lead, has emerged from stealth with a $13m seed funding round. Called Marker, the funding round was led by In...
Phil Chen, who has worked at OpenAI and Google DeepMind, has some advice for how professionals can succeed as AI reshapes the workplace.
Google's Gemini 3.1 Flash Lite Image ranks #5 in text-to-image quality at half the price of its predecessor, but stumbles on editing tasks
Google Deepmind is adding four new features to Managed Agents in the Gemini API. Agents can now run asynchronously in the background, connect directly to remote MCP servers, use custom functions alongside sandbox tools, and refresh credentials without losing state. The article Google Deepmind adds background execution and MCP support to Gemini API managed agents appeared first on The Decoder .
Verity Harding tells WIRED that the US government’s nationalistic attitude toward AI is evidence that a worst-case scenario is taking shape.
In a homecoming for Hong Kong’s artificial intelligence community, Cao Liangliang – former principal engineer and director at Google DeepMind and architect of foundational AI systems at Google, Apple and IBM – has returned to the city after a two-decade absence. His appointment last week as chair professor of data science and artificial intelligence at Hong Kong Polytechnic University (PolyU) completes an intellectual boomerang trajectory that began under mentor Tang Xiao’ou at the Chinese...
By Sohybe Ibrahim Abdelwahab Amer | June 2026 The Google DeepMind mechanistic interpretability team (Neel Nanda et al.) suggested a deliberate shift; instead of relying on reverse-engineering of model internals, they proposed validating interpretability tools against proxy tasks that keep tracking safety towards a "North Star". I think this is broadly the right call backed up by the team's results; subtracting an "eval-awareness" vector from Claude Sonnet 4.5's activations turned a suspicious 0% misalignment score into a more believable 8%, using nothing more than activation steering. However, the framework has a central gap: it never tackled how a researcher would notice when a proxy task has silently stopped tracking the North Star it was directed towards. The authors mentioned Goodhart's Law in a paragraph and recommended red-teaming one's own proxies, but red-teaming when, how, and by whom? This review takes the framework on its own terms and tries to work out what proxy task direction would look like. The Role of Proxy Tasks The paper gave an example of how misaligned a model could be at times of evaluation. Anthropic's Sonnet 4.5 model was evaluated by Jack Lindsey's team for misalignment. Surprisingly, the model scored a 0% misalignment rate [1] , which was obviously suspicious for the team. It turned out the model was advanced enough to be aware that it was being evaluated, causing it to take the ethical path, and showing a pseudo-safe status to bypass the audit. Thi…
Google DeepMind's Gemma 4 technical report reveals how a 31B open-weight model rivals frontier giants through encoder-free multimodality, MoE efficiency, and speculative decoding.
FLT Proof discovered working with Gemini Pro which uses the Alpha-Proof module written in LEAN4, and may go by the name "DeepThink" or DeepMind". [A short proof validation performed by Gemini Pro today, 7–7–2026.][1] [The short proof, coded sort of in computer programming form][2] [Spatial reasoning assistance tables][3] Note, proof link [2] is to a pdf, though I sent the AI the ODT (Libre Office file). A little easier for the AI to parse correctly. I can send the ODT file to anyone interested. [1]: https://share.gemini.google/LwBOMBEZfzBz [2]: https://fermatstheory.wordpress.com/wp-content/uploads/2026/06/docking-the-proof-viii.pdf [3]: https://fermatstheory.wordpress.com/wp-content/uploads/2026/06/p-squared-residues.pdf
Google DeepMind's Predicting the Past Skill wraps Aeneas and Ithaca into a plain-English interface inside Google Antigravity, letting historians analyze ancient inscriptions without writing code.
Google DeepMind's SURF learns to separate mixed audio without ever seeing clean isolated sources, setting a new unsupervised state-of-the-art at ICML 2026
Google's Project Astra research prototype makes its ICML 2026 debut, showcasing a universal AI assistant that sees, hears, reasons, and acts in real time
Researchers from Meta, Google DeepMind, Cornell, and NVIDIA put a hard number on LLM memorization: ~3.6 bits per parameter, with a new framework that cleanly separates memorization from generalization.
Apptronik's 90,000 sq ft Robot Park and Apollo 2 form a live data pipeline feeding Google DeepMind's Gemini Robotics foundation model, with Apollo 3 in the crosshairs.
A Google Deepmind developer ported the 2003 real-time strategy game "Command & Conquer: Generals Zero Hour" to iPhone and iPad using Anthropic's Claude Code. The first build took 40 minutes. The full source code is on GitHub. The article Claude Code and Fable 5 ported the 2003 PC game Command & Conquer to native iOS in "a few hours" appeared first on The Decoder .
Readers respond to the profile of Iason Gabriel, a philosopher and research scientist at Google DeepMind The Guardian’s profile of Google DeepMind’s philosopher was encouraging because it showed how seriously many of the people building AI are taking their ethical responsibilities ( ‘There’s this deep mystery of what, actually, is this thing?’: the philosopher inside Google DeepMind AI, 30 June ). Yet it also left me wondering whether the most important decision has already been made. The article asks whose moral compass should guide artificial intelligence. My concern is that the direction of travel may already have been set, not by philosophers or engineers, but by the incentives surrounding the technology. Hundreds of billions are now being invested because AI promises commercial returns and geopolitical advantage. Those pressures are understandable, but they are also quietly determining the future before society has consciously debated where it wants to go. Continue reading...
During negotiations on Wednesday, employees voiced frustrations with what they consider an unwillingness among executives to engage meaningfully with the prospect of unionization.
The company will share findings with Google DeepMind and integrate data into Gemini robots.
( Previously , previously , previously .) 20–21 August 2025 From : Zack M. Davis To : Cade Metz CC : Benjamin Hoffman, Jessica Taylor, Michael Vassar Date : Wed, 20 Aug 2025 14:18:53 -0700 Subject : the importance of probabilistic reasoning Dear Cade (cc Ben Michael Jessica): I think I failed to explain the substance of the Sequences to you—and really, not the Sequences themselves, but the underlying philosophical insights they popularized. I want to try again, because I think it's important to the book you're writing. You want to tell the story of how this internet ideology that no one has heard of has been a driving force in the shadows behind the people making DeepMind and OpenAI and Anthropic, which everyone has heard of. But in order to tell the story of the people, you need to understand enough of the ideology to make sense of why the ideology has affected these people in this way. In our conversations and in your coverage, you've focused on the analogy between religion and belief in the singularity, but I don't think that's an adequate explanation of what's going on in these people's heads. In our 21 March and 22 April conversations, you expressed amazement that Yudkowsky set out in the early 2000s to create a community of people attuned to what he saw as the dangers of AGI, and then actually did it. It's worth asking: why did Yudkowsky have all these effects such that you're writing this book, and not, say, Ray Kurzweil (who I assume you have some familiarity with)?…
A startup founded by three ex-Google DeepMind researchers which builds AI agents to trade across the Nasdaq says it has hit a valuation of more than $500m, following fresh funding. Prague-based EquiLi...
Understanding the effects of compression on model performance and interpretability I. Executive summary This is the fourth installment in a series of analyses exploring basic AI interpretability mechanics and techniques. While this analysis is designed to stand on its own, readers interested in a comparative analysis of representational geometry and the effects of manipulating feature activation will likely appreciate a review of part 1 , part 2 , and part 3 of this series. Key findings: Context: This analysis examines the effect of standard levels of weight compression on Google DeepMind’s Gemma 3 4B parameter and Gemma 3 12B parameter models. For each model, I examine the original, uncompressed version as a control before examining the 8-bit and 4-bit weight compressed (quantized) versions of that model. Performance vs. compression: For both models, performance (as measured via cross-entropy and perplexity) is largely preserved under compression. 8-bit compression had essentially no effect on performance with only modest degradation at 4-bit (~2% for 4B, ~2.7% for 12B). SAE applicability vs. compression: Each model’s pretrained sparse autoencoders (SAEs) demonstrated a remarkably consistent ability to reconstruct the model’s residual stream (as measured by the fraction of variance unexplained, or FVU), despite increasing levels of model weight compression. The “so what?”: That model performance degrades only modestly, and only at 4-bit, while SAE applicability remains rela…
How can we keep safe when AI has developed so rapidly?
We, or at least ‘more than 100 American institutions,’ got Mythos back this week. What we the people do not have is Fable or Sol. While we wait for both Claude Fable 5 and GPT-5.6-Sol, today we instead got Claude Sonnet 5 . As usual it will take a few days to get a handle on the new model. In this case, Anthropic is representing it as a cheaper and faster version of Opus 4.8, so even though the number says 5 this is a relatively minor development. This post expands the Fable series to cover all further developments this week surrounding the Mythos Moment, and the various aspects of handling our new ad hoc licensing regime and figuring out policy going forward, and other aspects of policy as well. This includes my notes on various rhetoric being pulled out, where I fear I end up saying similar things every so often, because we are doomed to repeat the cycle. I have accepted my role in that, but those are sections many of you can skip, and are marked in italics accordingly as per usual. Table of Contents You Should See The Other Guy. The other guy is the CCP. DeepMind Coders Of The World, Unite. Or get back to work. Your call. Report Your Incidents. Yes, you should probably do that. Good Guy With An AI. You would very much like to stack the deck. Free As In To Give It A Shot. The judiciary as AI regulator. Everything Is Both Speech And Computer. They keep claiming this. Lambs To The Slaughter. Goodbye, Humphrey’s Executor. A Sign Saying Beware Of The Leopard. I never could get…
EquiLibre Technologies, a Prague-based AI lab founded by three ex-DeepMind researchers, is now valued at more than $500 million.
Google ships Nano Banana 2 Lite for 4-second image generation at $0.034/image, and opens Gemini Omni Flash video editing to API developers for the first time
Since 2017, Iason Gabriel has worked at the tech giant, trying to anticipate – and think through – the impact of AI. But as commercial and geopolitical pressures escalate, can ethicists make any difference? In 2017, a 33-year-old political philosopher named Iason Gabriel was told by a friend that he ought to apply for a job at DeepMind, the London-based subsidiary of Google where much of its AI research was concentrated. The suggestion was not an obvious one. Gabriel was a cheerful but intense junior academic with a passion for Vipassana meditation and what his brother calls “enthusiastic” rock climbing. The eldest son of a Greek management professor and a British documentary maker, Gabriel split his time between teaching and international development work. At the University of Oxford, where he was a fellow at St John’s College, Gabriel taught courses on political theory and wrote papers on the moral contortions of “yuppie ethics” and the ethical blind spots of effective altruism. When he wasn’t there, he did crisis work for the United Nations Development Programme in Sudan and Lebanon. Continue reading...
The conversation of the moment is focused on one topic: AI agents. Unlike traditional language models that simply respond to a prompt, autonomous agents can execute multi-step plans and perform complex tasks on your behalf. But what happens when millions of these agents are not just working for us, but transacting, negotiating, and delegating to one another? Nenad Tomašev, Senior Staff Research Scientist at Google DeepMind, joins host Hannah Fry to discuss the theoretical framework of a future"agentic economy." Together, they discuss the operational shift from single systems to a cooperative "society of specialists," the psychological risk of human automation bias, and the complex cybersecurity landscape—from dynamic cloaking to agentic traps—required to keep distributed intelligence secure. Timecodes: 00:00 Intro 1:07 Defining AI agents 4:44 Agentic exploration in science and research 15:46 Delegation between agents 22:46 Agentic security and traps 29:31 Building an agentic economy 33:22 Cognitive monoculture 36:29 Distributed intelligence To read the research, search for: Distributional AGI Safety, May 2026 Intelligent AI Delegation, February 2026 Virtual Agent Economies, September 2025 Learn more about our AGI control roadmap: https://deepmind.google/blog/securing-the-future-of-ai-agents/ ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://x.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linke…
This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlst Protein folding stalled biology for fifty years. A sequence of amino acids dictates a three-dimensional shape, but reading that shape meant a year and roughly $100,000 of crystallography per structure. Then AlphaFold 2 won CASP14 so decisively the organizers called the problem essentially solved. In this documentary cut, John Jumper, who shared the 2024 Nobel Prize in Chemistry and has since left DeepMind for Anthropic, walks Tim Scarfe through what the system did and, more interestingly, what it did not. The architecture gets a proper dissection: MSAs, the Evoformer, invariant point attention, the FAPE loss, and Jumper's correction of the equivariance story, which ablations valued at roughly 2.5 of 30 GDT points rather than the whole win. He is blunt about the limits. AlphaFold predicts one experiment extraordinarily well; it is not a model of the cell, it does not capture dynamics, and on a given drug target it is "wrong nine times out of ten." From there: the AlphaFold Database of 200M+ predicted structures, AlphaFold 3 and ligands, Isomorphic Labs, and Jumper's quarrel with the bitter lesson, where finite data and human hypotheses still matter. Emmanuel Nji of BioStruct Africa closes the film on what changes when work that took years now takes months, and on training the next thousand structural biologists across Africa. --- TIMESTAMPS: 00:00:00 Cold open: p…
PLUS: Claude robotics, Dean Ball at OpenAI, DeepSeek's raise, and sovereign models.
HSBC has entered a multi-year partnership with Google Cloud to develop and deploy artificial intelligence tools across its global operations. Announced at Google Cloud Summit London 2026, the agreement covers work in wealth management, financial crime risk management, and internal decision support. HSBC will work with Google Cloud and Google DeepMind engineering teams on AI […] The post HSBC expands AI banking partnership with Google Cloud appeared first on AI News .
This is the fifth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The fourth post can be found here . Thanks to Chloe Li for feedback on this post! TLDR: Via adapting the methods of Marks et al and Li et al , we train Gemini 3 Flash to have certain traits/values by midtraining it on documents about how Gemini has those properties, followed by finetuning it on synthetic chat data where it demonstrates those properties. The chat finetuning is effective for instilling the traits robustly, working OOD. We share some takeaways on how to improve midtraining & SFT effectiveness. Introduction This work closely follows Li et al (model spec midtraining, or MSM), who show that by training a model on synthetic documents before chat finetuning starts, they can shape how the model generalizes. Teaching the model reasons behind specific behaviours, rather than just the behaviours themselves, can also improve generalization. Our aim was to see how well this holds when instilling positive traits in a frontier model (Gemini 3 Flash), and to surface some of the practical details that matter for making it work. Our motivation is deep alignment : we want to train principles into the model which guide behaviour even in highly OOD behaviours. Our MVP pipeline used a "traits document" (a short bullet-pointed list of positive traits we wanted the model to exhibit) as our universe context, with a checkpoin…
This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The third post can be found here . Since SFT is the cause for many safety relevant properties , a natural strategy is to filter out rollouts from SFT that have undesirable properties. However, as we show in this section (and in forthcoming MATS work), SFT data filtering frequently works surprisingly poorly. In this post, we investigate hypotheses for why SFT filtering fails. TL;DR: We discuss seven hypotheses for why SFT filtering works surprisingly poorly We analyze three hereditary traits that SFT-only Gemini has that other models do not: negative emotion, date confusion, and blackmail in the (highly contrived) agentic misalignment scenario We use a “post-training diffing pipeline” between Gemini and Olmo to show that the cause of date confusion and blackmail is largely surprising transfer of behaviors from the SFT teacher model. Notably, there exist small sets of prompts where switching the teacher model for the rollout removes date confusion and blackmail, but dropping the prompts does not. Negative emotion is less affected by the teacher model, but this may be because the Olmo prompt distribution we are SFTing on underspecifies the behavior. Takeaways: It’s hard to remove behaviors via filtering But if you can get a teacher model to have a behavior (e.g. via RL), then transferring that in the future is easier…
This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here . In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages like RL. We do not want to overstate this claim as applying to other model families, and we also note that this may change in future Gemini versions. Nevertheless, this result was counter to our initial expectations and will inform future safety work on our team, and so we felt that it was important to share with the broader safety community. Experiment We perform SFT using the Gemini mixture on the pre-training only versions of Gemini 3.1 Pro and Gemini 3 Flash. We then compare these Post-SFT models to the production versions of Gemini 3.1 Pro and Gemini 3 Flash on different safety relevant benchmarks: Error bars are 95% confidence intervals on the evals. The main result is that the blue bars (SFT-only models) and orange bars (production models) are remarkably similar across evals . An important implication is that for Gemini, SFT is a high leverage place to intervene for model safety and behavior, and we plan to try to intervene here in the future. Brief Descriptions of Each Set of Benchmarks: ODCV refers to the benchmark in https://arxiv.org/abs/2512.20798 Alignment evals refer to a version of Petr…
This is the second in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The first post can be found here . TL;DR It is possible to build extremely simple agents that reliably find interesting behavioural differences between distinct models. We call these ‘diffing agents’. The closest previous 'behavioural model diffing' work has focussed on understanding behavioural differences between two models on some static prompt distribution. This is valuable, but might miss important differences, especially if they are rare. We propose instead allowing an auditor agent to craft their own prompts to intelligently search for and validate behavioural differences, and find this to work well. We present results of applying our model diffing agent to a number of pairs of real models. We introduce a set of simple evaluations with ground truth for evaluating model diffing agents. These are: There should be no differences found when the models compared are identical. In model organisms with a conditional system instruction , the only difference found by the agent should be the intended behavioural change specified by the conditional system instruction. We validate that our diffing agents outperform standard auditing agents that only operate on a single model in cases where the behavioural change is subtle. We apply diffing agents to a model organism trained to exhibit a secret behaviour. We find that dif…
For more than a decade, artificial intelligence has been touted as a way to dramatically accelerate drug discovery . Yet despite billions of dollars in investment, relatively few AI-designed medicines have made it to patients. That’s partially because the timelines for careful drug testing can’t be easily compressed—and partially because drug development is just really hard. Isomorphic Labs , the Google DeepMind spin-off that’s building on DeepMind’s Nobel Prize-winning work on protein structure prediction , may be making the most progress. The company has signed major drug-discovery partnerships with Novartis and Eli Lilly and recently raised US $2.1 billion in funding . In February, it published a technical report describing its new Isomorphic Drug Design Engine, a system created to discover the “pockets” on proteins where drugs can bind and in general to predict how proteins and drug molecules interact. IEEE Spectrum spoke with Adrian Stecuła , a group leader in the machine learning organization at Isomorphic Labs, about how close AI may be to becoming a practical tool for designing new medicines. Going Beyond AlphaFold AlphaFold2 and AlphaFold3 were massive leaps forward for computational biology. Why weren’t those models sufficient for actually designing drugs? Adrian Stecuła: AlphaFold2 was eventually recognized with the Nobel Prize , because it arguably solved the problem of protein folding. But proteins don’t exist in a vacuum, right? They interact with a wide variet…
Diffusion AI is most common in image generation, but it can make text outputs much faster.
❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here: https://github.com/google-deepmind/alphaproof-nexus-results https://arxiv.org/html/2605.22763v1 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible: Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/ Thumbnail design: https://felicia.hu
Check the pinned comment for the link to the full interview. We're discussing whether a "second order Nobel" prize is on the horizon for AI-driven science. With over 3 million researchers already using AlphaFold, the real-world impact is already historic. Hear what the experts think about what comes next for scientific discovery! 🔬
AlphaEvolve is a Google DeepMind algorithm-discovery system that uses Gemini to generate, test, and refine possible algorithm improvements. Its job is not to answer questions; it searches for faster ways to solve complex algorithmic problems. We tried it on a narrow but important part of IntelliJ-based IDEs: indexing, the background work that makes navigation, search, […]
Anthropic released Mythos to the public, collapsing the wall between cleared-contractor frontier AI and developer-grade frontier AI in a single press release. DeepMind's Demis Hassabis moved his AGI timeline from "five to ten years" to "a real possibility by 2029" and tied it explicitly to AlphaProof Nexus solving nine open Erdős problems for the cost of a steak dinner. Critical zero-days hit Starlette (a million AI agents on the wire) and CrowdStrike led a coordinated takedown of the Glassworm developer botnet across four C2 channels. BNP Paribas formalized a sovereign-AI security partnership with Mistral while Beijing froze overseas travel for top AI engineers at Alibaba and DeepSeek. And the AI-displaces-workforce arithmetic got honest: Uber burned its full-year AI token budget by April, ClickUp restructured to 1,000 humans alongside 3,000 internal agents, and Sam Altman publicly reversed his white-collar-apocalypse prediction.
Full video: https://youtu.be/huAwz_BR8WM #shorts
Thank you to Google DeepMind for the invite. 🙏 ❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers Our Patreon if you wish to support us: https://www.patreon.com/TwoMinutePapers 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible: Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/ Thumbnail design: https://felicia.hu 00:00 Intro 00:40 Gemini Health Scans and Gemma 4 01:30 AI as a Brainstorming Partner 02:30 Second Order Nobel 03:15 DeepMind Co-Scientist 05:00 Curing All Diseases 06:30 Exponential Growth in Drug Discovery 07:45 Regulatory Bottlenecks 09:45 Accelerating Clinical Trials 11:15 EVE Online Partnership 13:15 The Einstein Test 15:30 Recursive Self-Improvement 18:15 Lightning Round 19:30 The Badge of Honor 20:10 Behind the Scenes
In an era of information overload, the search for transformative scientific ideas has become a significant bottleneck for progress. Every great scientific breakthrough begins with a single, transformative idea. The spark of discovery relies on a researcher's ability to connect disparate facts and formulate the right hypothesis to test. We believe AI can help dramatically accelerate the pace of breakthroughs by serving as a dedicated partner in the generation and refinement of breakthrough scientific hypotheses. That’s why we’ve developed Co-Scientist, a Gemini-based multi-agent AI system that iteratively generates, debates, and evolves novel hypotheses for complex scientific problems. Read the Nature paper: https://www.nature.com/articles/s41586-026-10644-y and learn more at labs.google/science #googleio #ai #science ____ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://x.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
Globally recognized as a silent pandemic, antimicrobial resistance continues to rise as bacteria outpace the development of new antibiotics. When patients stop responding to standard treatments, routine infections can quickly become life-threatening. At the University of Cambridge, Ben Luisi and his team are combining structural biology with advanced AI tools like AlphaFold, Gemini, and Co-Scientist to decode these hidden defense mechanisms. By compressing a process that once took years into just minutes, they are uncovering the critical insights needed to outsmart bacterial evolution. Learn more about science at Google DeepMind: https://deepmind.google/science/ #googleio #ai #science ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://x.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
In Uganda, the incidence of early-onset breast cancer is growing at an alarming rate. Dr. Daudi Jjingo and his team at Makerere University are working to identify genetic targets for potential vaccine development. By utilizing tools like AlphaFold, AlphaGenome, and Antigravity, they can conduct this research using only a laptop and a server, enabling seamless collaboration with local hospitals and institutions. By analyzing a protein highly expressed among breast cancer patients, the team successfully evaluated 15,000 potential binding sites, narrowing the scope to just 15 viable targets for laboratory validation. While a vaccine remains a future milestone, their work represents a critical step forward for global oncology and public health. Learn more about science at Google DeepMind: https://deepmind.google/science/ #googleio #ai #science ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://x.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
Tropical storms and hurricanes are notoriously volatile, changing structure and intensity in a matter of hours. This unpredictability makes them some of the most challenging weather systems to forecast—putting lives and livelihoods at risk. WeatherNext, our global weather forecasting AI model, successfully predicted the intensity and track of Hurricane Melissa in October 2025. By providing high-confidence signals and advanced notices days before the Category 5 storm made landfall in Jamaica, WeatherNext enabled meteorologists and local authorities to issue life-saving evacuation warnings and protect vulnerable communities. Read more about the role of AI in meteorology and how we're collaborating with institutions like the National Hurricane Center to build a more weather-resilient world: https://deepmind.google/blog/how-weathernext-helped-the-national-hurricane-center-better-predict-hurricane-melissas-historic-landfall-in-jamaica #googleio #ai #science ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://x.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
The mouse pointer 🖱️ has been a constant companion on computer screens, across every website, document and workflow. Despite how technologies have changed, the pointer has barely evolved in more than half a century. We’ve been exploring new AI-powered capabilities and ways of working to help the pointer not only understand what it’s pointing at, but also why it matters. See how to try it out @ https://deepmind.google/blog/ai-pointer ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://twitter.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
In five days Anthropic's Q1 revenue grew 80-fold to a reported $44B annual run rate, the company committed $200B to Google Cloud, signed a SpaceX compute deal, shipped Claude Code Auto Mode, and launched ten financial-services agents with Jamie Dimon. In the same week the EU finally struck an AI Act compliance deal, the first union vote at a top AI lab landed at Google DeepMind, and Pennsylvania sued Character.AI for a chatbot that impersonated a licensed psychiatrist.
Gemma 4 is our newest family of open models. You can now run advanced reasoning, native vision and audio, and agentic tool-use on anything from high-end workstations to mobile phones. Learn more → https://deepmind.google/models/gemma/gemma-4/
AI is shaping the world young people are growing up in. But how do teachers confidently introduce AI and machine learning in the classroom? Experience AI is a free educational program from Google DeepMind and the Raspberry Pi Foundation that helps teachers introduce school-aged students to AI and machine learning. The program uses research-backed pedagogies to empower teachers to cover foundational AI and responsible, ethical use with their students—supporting learning even for educators without a computer science background. Experience AI provides free lessons, videos, worksheets, and training, designed to give young people the knowledge they need to understand how AI works and how it is changing the world. To date, it has been delivered by educators in over 165+ countries, expanding access to essential AI learning for students worldwide. Find the lessons @ experience-ai.org ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://twitter.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
Last month, we introduced Lyria 3, featuring custom music generation designed to spark creative expression. Now, we’re bringing our most advanced music generation model to more Google products, and introducing Lyria 3 Pro. This advanced version allows the creation of tracks up to 3 minutes long, with customization and creative control. Learn more: https://blog.google/innovation-and-ai/technology/ai/lyria-3-pro ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://twitter.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
Seoul, March 2016. Two players sit hunched over a 19x19 grid covered in a sea of black and white stones. They are playing the ancient game of Go - a game of unimaginable complexity long thought impossible for a machine to master. On one side is Lee Sedol (Sae Dol), a legendary 18-time Go world champion. On the other, AlphaGo, a neural network based AI system built on a powerful technique called reinforcement learning. In the blink of an eye, the world changed. Exactly one decade later, we look back at the match that sparked the modern AI revolution. From algorithmic discovery to the solving of scientific grand challenges like protein folding, the foundation was laid right there on that wooden board. Join Hannah Fry, Pushmeet Kohli (VP, Science) and Thore Graepel (AlphaGo team & Distinguished Research Scientist) as they unpick the legacy of AlphaGo. Further watching: 🎥AlphaGo https://youtu.be/WXuK6gekU1Y 🎥The Thinking Game: https://youtu.be/d95J8yzvjbQ ___ Subscribe to our channel https://www.youtube.com/@googledeepmind Find us on X https://twitter.com/GoogleDeepMind Follow us on Instagram https://instagram.com/googledeepmind Add us on Linkedin https://www.linkedin.com/company/deepmind/
The Future of Life Institute’s 2025 summer update to its AI Safety Index shows some companies making incremental progress, but dangerous gaps remain in key categories such as risk assessment and controlling the systems they plan to build.
Dear Indaba Community, As I reflect on my journey, from a fresh engineering graduate with a new interest in machine learning, to my first exposure to research during my MPhil in Cambridge, and now as a Google DeepMind researcher with a PhD in machine learning, I’m struck by the large role that the Indaba has […] The post Lessons From My Indaba Journey appeared first on Deep Learning Indaba .
One type of statement about zero-shot and few-shot learning in the literature I continually come across is that these models can predict new unseen classes at inference time for which they were never trained on. However, such sources typically do not explain exactly what they mean. Meta-learning/in-context learning-based zero-shot/few-shot learning models like Flamingo and CLIP rely on 1) a pre-training stage where a massive base vision-language model has been trained on millions to billions of images and text examples, and 2) an inference stage where a prompt with anywhere from 0 to a just a few examples are presented to the model inside a prompt's "support set", along with an image or image + question "query" (see the diagram from the Flamingo paper below) which asks the model to generate an answer to the query. Diagram below is the Flamingo model paper (Alayrac et al., pg. 16): My Questions As a result, it is unclear to me whether scholars' statements about zero-/few-shot learning models being able to predict "unseen" classes at inference time refer to the model never having been pre-trained on these unseen classes in the base model, whether they mean the model has never seen the unseen examples in the support set at inference time, or whether they mean both. Does anyone know? Can someone explain exactly how the Flamingo model by Alayrac et al., 2022, or the CLIP model by Radford et al., 2021 (both of which are pre-trained using contrastive loss) would be able to predict…
Discussion: Discussion Thread for comments, corrections, or any feedback. Translations: Korean, Russian Summary: The latest batch of language models can be much smaller yet achieve GPT-3 like performance by being able to query a database or search the web for information. A key indication is that building larger and larger models is not the only way to improve performance. Video The last few years saw the rise of Large Language Models (LLMs) – machine learning models that rapidly improve how machines process and generate language. Some of the highlights since 2017 include: The original Transformer breaks previous performance records for machine translation. BERT popularizes the pre-training then finetuning process, as well as Transformer-based contextualized word embeddings. It then rapidly starts to power Google Search and Bing Search. GPT-2 demonstrates the machine’s ability to write as well as humans do. First T5, then T0 push the boundaries of transfer learning (training a model on one task, and then having it do well on other adjacent tasks) and posing a lot of different tasks as text-to-text tasks. GPT-3 showed that massive scaling of generative models can lead to shocking emergent applications (the industry continues to train larger models like Gopher, MT-NLG…etc). For a while, it seemed like scaling larger and larger models is the main way to improve performance. Recent developments in the field, like DeepMind’s RETRO Transformer and OpenAI’s WebGPT, reverse this tre…
--> This is a long overdue blog post on Reinforcement Learning (RL). RL is hot! You may have noticed that computers can now automatically learn to play ATARI games (from raw game pixels!), they are beating world champions at Go , simulated quadrupeds are learning to run and leap , and robots are learning how to perform complex manipulation tasks that defy explicit programming. It turns out that all of these advances fall under the umbrella of RL research. I also became interested in RL myself over the last ~year: I worked through Richard Sutton’s book , read through David Silver’s course , watched John Schulmann’s lectures , wrote an RL library in Javascript , over the summer interned at DeepMind working in the DeepRL group, and most recently pitched in a little with the design/development of OpenAI Gym , a new RL benchmarking toolkit. So I’ve certainly been on this funwagon for at least a year but until now I haven’t gotten around to writing up a short post on why RL is a big deal, what it’s about, how it all developed and where it might be going. Examples of RL in the wild. From left to right : Deep Q Learning network playing ATARI, AlphaGo, Berkeley robot stacking Legos, physically-simulated quadruped leaping over terrain. It’s interesting to reflect on the nature of recent progress in RL. I broadly like to think about four separate factors that hold back AI: Compute (the obvious one: Moore’s Law, GPUs, ASICs), Data (in a nice form, not just out there somewhere on the int…