Orchestrating AI Code Review at scale
Learn about how we built a CI-native AI code reviewer using OpenCode that helps our engineers ship better, safer code.
AI/ML news, top picks, and generated innovation digests.
26010 matching items
Learn about how we built a CI-native AI code reviewer using OpenCode that helps our engineers ship better, safer code.
Agents Week 2026 is a wrap. Let’s take a look at everything we announced, from compute and security to the agent toolbox, platform tools, and the emerging agentic web. Everything we shipped for the agentic cloud.
At what point do the financial markets price in the singularity?
GRASP is a new gradient-based planner for learned dynamics (a “world model”) that makes long-horizon planning practical by (1) lifting the trajectory into virtual states so optimization is parallel across time, (2) adding stochasticity directly to the state iterates for exploration, and (3) reshaping gradients so actions get clean signals while we avoid brittle “state-input” gradients through high-dimensional vision models. Large, learned world models are becoming increasingly capable. They can predict long sequences of future observations in high-dimensional visual spaces and generalize across tasks in ways that were difficult to imagine a few years ago. As these models scale, they start to look less like task-specific predictors and more like general-purpose simulators. But having a powerful predictive model is not the same as being able to use it effectively for control/learning/planning. In practice, long-horizon planning with modern world models remains fragile: optimization becomes ill-conditioned, non-greedy structure creates bad local minima, and high-dimensional latent spaces introduce subtle failure modes. In this blog post, I describe the problems that motivated this project and our approach to address them: why planning with modern world models can be surprisingly fragile, why long horizons are the real stress test, and what we changed to make gradient-based planning much more robust. This blog post discusses work done with Mike Rabbat, Aditi Krishnapriyan, Yann…
BAR is a recipe for post-training language models one capability at a time—train domain experts independently, merge them into a single mixture-of-experts model, and upgrade any expert without impacting the others.
[…] Reliance Posts 10% Revenue Growth In Q3FY26 As Jio Crosses 500 Million Subscribers […]
[…] asset management business, operated through the Jio-BlackRock joint venture, reported assets under management of Rs 15,218 crore across 10 funds, with a retail […]
A learning-oriented workflow for understanding new open-weight model releases
I'm modeling the effect of a treatment on a population of flies. For each fly, I have the following covariates: Treatment (control or treated) Sex Cage in which the fly was reared. I have 3 cages of treated flies and 3 cages of control flies. 96-well plate on which the fly was sequenced. Each plate is entirely treatment or entirely control flies, from a mix of the 3 cages of that treatment. (i.e., plate is nested within treatment, and cage is nested within treatment, but plate and cage are not nested in each other). and I also have my response variable, which is continuous and determined by the sequencing. I'd think to model this using a mixed-effects model, treating treatment and sex as fixed effects and cage and plate as random: Response ~ treatment + sex + 1|cage + 1|plate I have no issues with this so far, but I'm interested in applying a permutation test to this data and am not sure how best to do so. I want to build a null distribution of test statistics for the treatment fixed effect. As I understand it, I could just shuffle treatment labels to get a null distribution: Response ~ treatment_permuted + sex + 1|cage + 1|plate However, I'm concerned that the random effects here will be based on the variance associated with the true treatment, and will not be truly permuted. Alternatively, I could use the permuted treatment labels in making new cage:treatment and plate:treatment groups for the random effects, but then I'd have more of these groups than in my true case, sin…
Join us at RightsCon 2026 for a timely and critical conversation on the future of digital public infrastructure (DPI) in the Global South. Across the Global South, digital public infrastructure […] The post RightsCon 2026: South-South Digital Public Infrastructure Approaches: Challenges and Opportunities appeared first on Research ICT Africa .
The Agent Readiness score can help site owners understand how well their websites support AI agents. Here we explore new standards, share Radar data, and detail how we made Cloudflare’s docs the most agent-friendly on the web.
Today, we’re excited to give you a sneak peek of our support for shared compression dictionaries, show you how it improves page load times, and reveal when you’ll be able to try the beta yourself.
Cloudflare Agent Memory is a managed service that gives AI agents persistent memory, allowing them to recall what matters, forget what doesn't, and get smarter over time.
Running LLMs across Cloudflare’s network requires us to be smarter and more efficient about GPU memory bandwidth. That’s why we developed Unweight, a lossless inference-time compression system that achieves up to a 22% model footprint reduction, so that we can deliver faster and cheaper inference than ever before.
Soft directives don’t stop crawlers from ingesting deprecated content. Redirects for AI Training allows anybody on Cloudflare to redirect verified crawlers to canonical pages with one toggle and no origin changes.
In the first quarter of 2026, four African countries have advanced significant AI governance instruments. South Africa’s Cabinet approved a Draft National AI Policy for public comment on 2 April. […] The post Rapid Response Webinar Series: AI Governance in Africa appeared first on Research ICT Africa .
[…] Why the IRDAI Is Not Allowing Insurance Manufacturing Licenses for VC-Backed Fintechs […]
The EU Commission’s policy on data centers keeps information on the energy and water use of individual centers under wraps. Research by Corporate and Europe Observatory and AlgorithmWatch, published by Investigate Europe, reveals the Commission copied and pasted an amendment suggested by Microsoft and the lobby group Digital Europe. The aim: To prevent NGOs from obtaining information on energy-hungry data centers in the face of growing resistance.
When their key product was faced with unfavorable scientific evidence and the risk of regulation, most businesses in the 20th century defended themselves by sowing doubt on an industrial scale. Big AI is doing something radically different: it floods the zone with potential future risks.
In this episode, Rashmi Shetty, senior director of enterprise generative AI platform at Capital One, joins us to explore how the company is designing, deploying, and scaling multi-agent systems in a highly regulated environment. Rashmi walks us through Chat Concierge, a multi-agent chat experience for auto dealerships that handles intent disambiguation, tool invocation, and human handoffs to deliver safer, more personalized customer journeys. We discuss Capital One’s platform-centric approach to AI agents and how it separates design from runtime governance, embedding policies, guardrails, and cyber controls across agent threat boundaries. Rashmi shares how the team approaches the developer experience for agent builders, observability, and evals for stochastic, multi-agent workflows; and strategies for model specialization, including fine-tuning and distillation. We also cover standards and abstraction, closed-loop learning from production telemetry, and key lessons for enterprises building agentic systems. The complete show notes for this episode can be found at https://twimlai.com/go/765.
Introducing CRUX, a new project for evaluating AI on long, messy tasks
Learn how Github uses eBPF to detect and prevent circular dependencies in its deployment tooling. The post How GitHub uses eBPF to improve deployment safety appeared first on The GitHub Blog .
Aarathi Krishnan, CEO of Raksha Intelligence Futures, discusses the political, economic, and technological dynamics shaping this moment of uncertainty and transition in the global system.
President Trump said during an interview aired yesterday by Fox Business that “there should be” when asked if AI needs […]
Autonomous driving is not just a big tech or closed-source game, it's becoming accessible through open innovation and real-world deployment. Dan and Chris sit down with Harald Schäfer, CTO at Comma AI, to explore how OpenPilot is bringing self-driving to everyday vehicles using open source AI. We dive into the intersection of machine learning, robotics, and simulation, including how world models are enabling training at scale and shaping the future of autonomy. Featuring: Harald Schäfer – LinkedIn Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Daniel Whitenack – Website , GitHub , X Links: Comma Upcoming Events: Register for upcoming webinars here !
Organizations face growing pressure to adopt artificial intelligence, but often lack practical guidance on how to do so effectively. This report bridges the gap between high-level principles and real-world implementation, offering actionable steps across the AI adoption life cycle. Drawing on over 1,200 resources, this reference guide provides practitioners with the knowledge required to operationalize AI safety, security, and governance practices within their organizations. The post Operationalizing AI Guidance: A Reference Guide for Translating High-Level Goals into Practical Implementation appeared first on Center for Security and Emerging Technology .
What I expect to come next and why, focused on the open-closed gap.
Weaviate Shared Cloud is now generally available on AWS in US East and Europe, giving teams a fully managed, AI-native database on the provider and region that works best for them.
Using importance sampling with fine-tuned donor prefills to predict reward hacking emergence during training
Modal is an official sandbox provider for the OpenAI Agents SDK.
How Bytedance Volcano Engine LAS (Lake for AI Service) leverages Lance as the core storage format, rapidly constructing a next-gen AI data lake to efficiently store, manage, and process multimodal data (text, images, audio/video).
This guest post comes from IDC’s Dr. William Lee, Senior Research Director, Service Provider and Core Infrastructure Research. MongoDB commissioned IDC to explore the connection between legacy infrastructure, data challenges, and AI across Asia Pacific, and today we’re happy to share that work. For more, see the full MongoDB-sponsored IDC InfoBrief, Modernizing Legacy: Winning in the Age of AI, Doc #AP242555-IB, April 2026. AI ambition is everywhere across Asia/Pacific. But ambition alone does not determine success. Organizations are discovering that AI outcomes are directly tied to the quality, accessibility, and modernity of their underlying technology stack and associated data technology foundations. Organizations that have managed to stay abreast of technical and data management changes across the application and infrastructure stacks, by embedding modernization into their organizational DNA, are experiencing 3x more digital revenue growth than those that are bound up in technical and data debt. To better understand this connection, IDC surveyed 1,400 organizations across eight Asia/Pacific markets. The findings reveal that modernization is no longer a side initiative. It is the core of a sustainable AI strategy. The AI readiness divide: Leaders versus mainstream IDC’s latest Asia/Pacific Modernization Survey, sponsored by MongoDB, identifies two distinct groups: The Mainstream Cohort: organizations still burdened by technical debt, siloed data, and skills gaps The Leade…
Autoresearch automates AI research. Modal automates AI infrastructure.
Here are fresh AI opportunities you can still apply for right now: 1. Data Science Africa AI & Machine […]
Was fire equivalent to a singularity for people at the time?
Gemma 4 is our newest family of open models. You can now run advanced reasoning, native vision and audio, and agentic tool-use on anything from high-end workstations to mobile phones. Learn more → https://deepmind.google/models/gemma/gemma-4/
Two benchmarks developed at Ai2 – ScienceWorld and DiscoveryWorld – reveal that even incredibly strong AI science agents struggle with problems human scientists solve routinely.
[…] over effectiveness: This expansion follows ongoing criticism of Instagram’s safeguards. A September 2025 report found that 64% of teen safety tools were ineffective, defunct, or easily bypassed, as 13 out […]
But Dr Heidy Khlaaf, chief AI scientist at the AI Now Institute and a former OpenAI safety engineer, is sceptical. She notes Anthropic provides no comparison with existing automated security tools, nor any false-positive rates. “It also serves their ‘safety first’ image, as they’re able to justify the lack of public release, even a limited one for independent evaluation, as a public service – when it simply obscures experts’ abilities to independently validate their The post ‘Safety first’ puts Anthropic ahead in game of AI spin appeared first on AI Now Institute .
And yes, I hate consortia too.
Anthony Aguirre, President and CEO of the Future of Life Institute, issued the following statement in response to the attack […]
Lance's JSONB storage, scalar indexing, data evolution, and full-text search already deliver what most users want from Variant — with explicit control, schema consistency, and no vendor lock-in.
Learn how LanceDB benchmarks storage and how we achieved one million disk reads per second.
Tech leaders want you to believe that AI is the key to a new golden age. The reality looks more like a bold, government-backed heist. The post The Great AI Grift appeared first on AI Now Institute .
Access this report from the Uehiro-Carnegie Endowment for Future Generations study tour, in which Carnegie Council fellows and staff reflect on their trip to Japan.
The Brazilian automotive landscape has established itself as a global powerhouse, currently ranking as the world’s sixth largest market for cars and light commercial vehicles. A report from strategy consultancy Mirow & Co explores the industry’s growth and trends – a roundup of the key findings in six charts.
Introducing the Multimodal Lakehouse - a unified platform for managing AI data from raw files to production-ready features, now part of LanceDB Enterprise.
This is a linkpost for MirrorCode, a project that METR funded and co-developed with Epoch AI . See Epoch AI’s blog post for more detail: https://epoch.ai/publications/mirrorcode-preliminary-results/
Butter, a San Francisco-based AI sandbox technology, is joining Modal.