AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
31011News Items
8Top Picks
186Blogs
successLast Run

Latest AI/ML News

31011 matching items

AI Dev 26 x SF | Or Dagan: Optimizing Accuracy, Cost, and Latency in Real-World Agents
DeepLearning.AI YouTube 2026-05-22 16:42 UTC Score 52.0 AI-138-20260522-podcasts-and-fd6db35f Full article

AI Dev 26 x SF | Or Dagan: Optimizing Accuracy, Cost, and Latency in Real-World Agents

Most agentic systems rely on hardcoded heuristics to navigate execution decisions (e.g. which models, tools, and test-time compute scaling approaches to use) leading to efficiency leakage across cost, latency and accuracy. AI21 Maestro optimizes agents by learning to predict success, cost and latency probabilities across diverse actions and contexts, and driving runtime orchestration that intelligently navigates the full agentic action space. In this session, AI21's Or Dagan demonstrated how this approach yields state-of-the-art results and Pareto frontier on challenging agentic benchmarks, as well as the process required to optimize production agents.

AI Dev 26 x SF | Diamond Bishop: The Next 100 Agents. Building the Agent Native Office
DeepLearning.AI YouTube 2026-05-22 15:52 UTC Score 33.0 AI-138-20260522-podcasts-and-3713f9ba Full article

AI Dev 26 x SF | Diamond Bishop: The Next 100 Agents. Building the Agent Native Office

Building your first agent is exciting. Building a platform that can evolve into an office where dozens of teams can safely deploy their own agents is a different beast entirely. In this talk, Diamond Bishop from Datadog shared lessons learned building production agents, then turning this into an agent office/platform made to power the next-gen enterprise with diverse agent workloads.

LatAm Journalism Review AI 2026-05-22 15:40 UTC Score 37.0 AI-176-20260522-regional-ai--7db66d34 Full article

Latin American journalists invited to apply for 2026 JournalismAI Skills Lab

"The 2026 JournalismAI Skills Lab is a 14-week, free, virtual program designed for professionals to learn how to practically implement LLMs, GenAI and agents in their work. The programme helps individuals upskill in using AI technologies in a hands-on manner. It equips participants to develop their own AI-based tools, prototypes or proofs-of-concept. The ultimate outcome […] The post Latin American journalists invited to apply for 2026 JournalismAI Skills Lab appeared first on LatAm Journalism Review by the Knight Center .

LatAm Journalism Review AI 2026-05-22 15:40 UTC Score 37.0 AI-176-20260522-regional-ai--bf379328 Full article

Latin American journalists invited to apply for 2026 JournalismAI Skills Lab

"The 2026 JournalismAI Skills Lab is a 14-week, free, virtual program designed for professionals to learn how to practically implement LLMs, GenAI and agents in their work. The programme helps individuals upskill in using AI technologies in a hands-on manner. It equips participants to develop their own AI-based tools, prototypes or proofs-of-concept. The ultimate outcome […] The post Latin American journalists invited to apply for 2026 JournalismAI Skills Lab appeared first on LatAm Journalism Review by the Knight Center .

AI Dev 26 x SF | Paul Everitt: The Shift to Agentic Engineering
DeepLearning.AI YouTube 2026-05-22 15:29 UTC Score 25.0 AI-138-20260522-podcasts-and-f3378c99 Full article

AI Dev 26 x SF | Paul Everitt: The Shift to Agentic Engineering

More code, fewer staff — the industry is on a bender. But what about quality? At AI Dev 26 x San Francisco, Paul Everitt from JetBrains discussed the rise of agentic engineering and how old lessons can be adapted to build new professional practices.

Big Technology 2026-05-22 15:20 UTC Score 25.0 USR-0107-20260522-ai-specialis-96f8bd11 Full article

AI’s Public Relations Emergency

A generation is being told AI is their enemy. And they’re starting to believe it.

OpenMined Blog 2026-05-22 08:00 UTC Score 27.0 USR-0156-20260522-ai-specialis-c4483899 Full article

Moving Fast Doesn’t Have to Break Things: The U.S. Must Stop Compromising Critical Infrastructure with Patchwork AI Security Approaches

PETs offer U.S. critical-infrastructure AI a path beyond patchwork security. Why Attribution-Based Control should be the standard. The post Moving Fast Doesn’t Have to Break Things: The U.S. Must Stop Compromising Critical Infrastructure with Patchwork AI Security Approaches appeared first on OpenMined .

Stack Overflow Machine Learning Tag 2026-05-22 05:45 UTC Score 26.0 AI-112-20260522-social-media-e319c2cc Full article

Rationale for StandardScaler over MinMaxScaler in spatiotemporal tree-based ensemble models with SHAP interpretability

I am developing a spatiotemporal tree-based ensemble framework (utilizing LightGBM, XGBoost, and CatBoost) to forecast dengue outbreaks based on climate variables (temperature, precipitation, humidity) and lagged historical case counts. While tree-based algorithms are theoretically invariant to monotonic feature scaling, I am implementing scaling primarily because: I am calculating SHAP (Shapley Additive Explanations) values for post-hoc model interpretability and global feature importance. I am applying forward aggregation across temporal slices to prevent data leakage, meaning the range and variance of features dynamically shift across training validation windows. I am debating between StandardScaler (Z-score normalization) and MinMaxScaler (0-1 normalization). Given the spatiotemporal and epidemiological nature of the data, StandardScaler appears to behave more robustly, but I want to ensure my architectural justification is sound. Here is a minimal visualization of how the choice impacts extreme climate outliers (e.g., a massive monsoon rainfall anomaly): import numpy as np import pandas as pd from sklearn.preprocessing import MinMaxScaler, StandardScaler # Simulating a climate feature with a severe anomaly (monsoon spike) np.random.seed(42) weekly_rainfall = np.random.normal(loc=150, scale=30, size=100) weekly_rainfall = np.append(weekly_rainfall, [650]) # Extreme outlier event df = pd.DataFrame({"Rainfall": weekly_rainfall}) # Applying both scalers df["MinMax"] = MinMa…

DeepSeek’s New AI Is A Game Changer
Two Minute Papers 2026-05-22 00:47 UTC Score 36.0 AI-139-20260522-podcasts-and-98bdc664 Full article

DeepSeek’s New AI Is A Game Changer

❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here: https://github.com/ailuntx/Thinking-with-Visual-Primitives https://huggingface.co/datasets/NodeLinker/deepseek-ai-Thinking-with-Visual-Primitives-deleted-repo/blob/main/Thinking_with_Visual_Primitives.pdf Our Patreon if you wish to support us: https://www.patreon.com/TwoMinutePapers 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible: Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/ Thumbnail design: https://felicia.hu #deepseek

Apple Machine Learning Research 2026-05-22 00:00 UTC Score 37.0 AI-059-20260522-official-ai--d87fa482 Full article

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models

Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time visual assistants. Existing VLM frameworks predominantly assess models in offline settings. In contrast, the performance of a streaming VLM depends on additional metrics beyond pure video understanding, including proactiveness, which reflects the timeliness of the model’s responses, and consistency, which captures the robustness of its responses over time. To address this limitation, we propose VSAS-Bench, a new…

TWIML AI Podcast 2026-05-21 19:38 UTC Score 56.0 AI-148-20260521-podcasts-and-830461d3 Full article

Relational Foundation Models for Enterprise Data with Jure Leskovec - #768

In this episode, Jure Leskovec, co-founder and chief scientist at Kumo and professor of computer science at Stanford, joins us to explore two fronts of his work: AI for science and relational deep learning. We begin with AI Virtual Cell, a multiscale effort to learn data-driven representations from proteins to cells to patients using single-cell RNA-seq data, protein language models like ESM, and structure models like AlphaFold—without hand-encoding biology. Jure then dives into relational deep learning, reframing enterprise databases as graphs and training neural networks directly on raw multi-table data. He explains Kumo’s Relational Foundation Model (RFM2), which performs in-context learning over subgraphs to make accurate predictions on new databases and tasks with no training, and how this approach benchmarks against RelBench and other multi-table datasets. We also discuss real-world deployments at companies like Reddit, DoorDash, and Coinbase, explainability via attention over tables and columns, integration with agentic systems, deployment options, and practical limitations. The complete show notes for this episode can be found at https://twimlai.com/go/768.

Access Now AI 2026-05-21 17:00 UTC Score 27.0 USR-0142-20260521-ai-specialis-33218bae Full article

Access Now urges the Ninth Circuit to protect encryption from NSO’s spyware

Yesterday, Access Now and ten other civil society organizations filed an amicus brief in the U.S.’ Ninth Circuit Court of Appeals calling to protect encryption from NSO Group’s Pegasus spyware and to keep the lower court’s permanent injunction forbidding NSO from ever targeting WhatsApp or its customers’ devices ever again. The post Access Now urges the Ninth Circuit to protect encryption from NSO’s spyware appeared first on Access Now .

Microsoft Research Podcast 2026-05-21 17:00 UTC Score 36.0 AI-147-20260521-podcasts-and-7dff8125 Full article

MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models

MagenticLite is an agentic system for small models that works across the browser and local file system in a single workflow. It combines specialized models and orchestration to support efficient agentic performance on everyday tasks. The post MagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small models appeared first on Microsoft Research .

Xi's summits with Trump and Putin + China's economy loses momentum + Hong Kong dissidents
MERICS China AI 2026-05-21 09:57 UTC Score 27.0 USR-0207-20260521-research-aca-32c42d61 Full article

Xi's summits with Trump and Putin + China's economy loses momentum + Hong Kong dissidents

Xi's summits with Trump and Putin + China's economy loses momentum + Hong Kong dissidents c.groth Thu, 05/21/2026 - 11:57 picture alliance / Photoshot Download (pdf - 973.09 KB) MERICS Briefs MERICS China Essentials May 21, 2026 11 min read Xi's summits with Trump and Putin + China's economy loses momentum + Hong Kong dissidents Top Story Xi’s summits with Trump and Putin project Beijing as a hub of global diplomacy By hosting US President Donald Trump and Russian President Vladimir Putin in back-to-back summits in Beijing, Xi Jinping was able to project China’s unprecedented global influence and advance its preferred worldview: building what he calls “constructive strategic stability” with the US, while enlisting Russia to push for a multipolar world order through the doctrine of “a new type of international relations.” Xi was helped by his guests appearing keen to impress. After Xi calling Taiwan “the most important issue in China-US relations,” Trump said he was “not looking to have somebody go independent” – and more generally seemed willing to finally treat China as a peer major power. After Xi implicitly criticized the US by noting that “unilateral hegemonic currents are running rampant,” Putin said China-Russia relations had reached “unprecedentedly high levels” and were “key stabilizing factors on the international stage.” Xi treated Trump with generous courtesy, managing to project confidence rather than deference. Compared with Trump’s 2017 visit to the imposing Fo…

Practical AI Podcast 2026-05-21 09:00 UTC Score 49.0 AI-143-20260521-podcasts-and-3cd5023d Full article

Hermes Agent: Agents that grow with you

Open Source AI is entering a new era, one shaped by self-improving AI Agents, recursive learning systems, and rapidly evolving AI Tools that blur the line between software and autonomous collaborators. In this episode, Daniel and Chris sit down with Nous Research co-founder and CTO Jeffrey Quesnelle to explore Hermes Agent. Along the way, they discuss models vs. harnesses, the changing role of developers, and one of the biggest questions facing the AI Future: what remains uniquely human as AI capabilities continue to accelerate? Featuring: Jeffrey Quesnelle – Website , LinkedIn Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Daniel Whitenack – Website , GitHub , X Links: Nous Research Hermes Agent Sponsors: Framer: The enterprise-grade website builder that lets your team ship faster. Get 30% off at framer.com/practicalai Prediction Guard: A self-hosted AI control plane for running agents in high impact environments. predictionguard.com/practicalai Upcoming Events: Register for upcoming webinars here ! Midwest AI Summit 2026

Allen Institute for AI Blog 2026-05-21 08:00 UTC Score 29.0 USR-0021-20260521-research-aca-f57cb00f Full article

Building accessibility tools on a truly open foundation

PointCheck, an independent project, uses Molmo, MolmoWeb, and Olmo 3 to test web accessibility the way a keyboard user would—by navigating real pages and inspecting what's actually on screen.

Qdrant Blog 2026-05-21 00:00 UTC Score 30.0 USR-0074-20260521-ai-specialis-d3619828 Full article

How Sunny Health Built an AI Healthcare Concierge with Qdrant

Most people don’t read their insurance pamphlet. The benefits are there: deductibles, copays, in-network providers, what dental covers, what dermatology covers, when an optometry visit is included in the medical plan. But the document is dense, the website is worse, and the result is that patients pay for plans they barely understand and delay care because finding an in-network provider with availability takes more energy than they have. Sunny Health is building a healthcare concierge that insurance companies and care providers offer to their members as part of the existing plan experience. When a member signs in (typically through SSO from their payer), Sunny Health already knows who they are and what their plan covers. They land in a chat experience where they can ask “show me dermatologists nearby,” get matched to in-network options, and have Sunny Health book the appointment on their behalf. Three things on one retrieval layer: benefits navigation, provider matching, and appointment booking.

AI Weekly 2026-05-21 00:00 UTC Score 10.0 AI-133-20260521-newsletters-31271732 Full article

AI Weekly Issue #494: SpaceX wants $80 billion. OpenAI wants a trillion.

For nine years the AI boom has been a private bet, priced by a small circle of venture funds and sovereign wealth in rounds most people could never touch. This week it started going public. SpaceX filed an $80 billion IPO prospectus on Wednesday, the largest in history, with a chatbot company and $6.4 billion in AI losses folded inside it. OpenAI is days from filing its own, aiming for a trillion-dollar debut by September. The public markets are about to answer the question private investors kept waving away: at what price?

ClearML Blog 2026-05-20 18:30 UTC Score 35.0 USR-0084-20260520-ai-specialis-0c136fc1

Enterprise AI Security with ClearML: A Complete Series Summary

By Adam Wolf & Damian Erangey Over a seven-part series of posts and videos, ClearML’s Enterprise AI Security series covered every layer of securing an AI platform in production, from who gets in to what gets recorded. This post brings it all together in one place: what each layer does, why it matters, and how […]

Deep Learning Indaba 2026-05-20 17:55 UTC Score 30.0 USR-0189-20260520-research-aca-bfb47694 Full article

Deep Learning Indaba Impact Report 2025

Our mission to Strengthen African AI, for Africans, by Africans remains as necessary and as valued as ever. This impact report sets out how the Deep Learning Indaba continues to deliver on that mission, and the change we are enabling across Africa’s AI ecosystem. As always, we are deeply grateful to our funders, partners, and […] The post Deep Learning Indaba Impact Report 2025 appeared first on Deep Learning Indaba .

Comet ML Blog 2026-05-20 16:47 UTC Score 41.0 USR-0082-20260520-ai-specialis-a1c86a19 Full article

What Held Up at 3 AM: One Engineer’s RAG Case Study

Most AI demos work. Most AI products don’t. This series is a collection of interviews with engineers who shipped AI agents to production, covering the stacks they chose, the architectures they regretted, and what actually held up at 3 am. This is an interview with Michael Maximilien, former CTO and Distinguished Engineer at IBM and […] The post What Held Up at 3 AM: One Engineer’s RAG Case Study appeared first on Comet .

Stack Overflow Machine Learning Tag 2026-05-20 13:55 UTC Score 23.0 AI-112-20260520-social-media-96b56cdd Full article

Training own recommendation model for diploma thesis

For part of my thesis project, I need to create a mechanism that will generate recommendations for the user based on data stored in my database. For example, the system has a list of specific tools and their descriptions. I need to create a chat that, when asked how to assemble a cabinet, will provide recommendations and a list of tools. But the list of tools must be from the database I have, not just any list, but specifically those tools provided by my system. As I understand it, I need to train some kind of model, but I don't know what I need for this or even the technologies that are needed for this. Where can I start? Can you please recommend what I should study? What ready-made examples on a similar topic can I look at? And in general, how to build the architecture of this part of the project? Thank you all very much in advance.

Intelligence is collective, not artificial — Prof. Michael I. Jordan (UC Berkeley / Inria)
Machine Learning Street Talk 2026-05-20 08:26 UTC Score 31.0 AI-141-20260520-podcasts-and-f932b4b5 Full article

Intelligence is collective, not artificial — Prof. Michael I. Jordan (UC Berkeley / Inria)

Michael I. Jordan, described by Science magazine as the most influential computer scientist alive, has never thought of himself as an AI researcher. In this conversation he explains why that distinction matters. SPONSOR: --- Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open. Apply now: https://cyber.fund --- Jordan trained as a statistician and cognitive scientist, and his career has been spent building machine learning systems that work in the real world: supply chains, commerce, healthcare, and large economic systems. When the field rebranded itself as AI and then AGI, he did not follow. Instead he argues that the framing is wrong. AI is better understood as a collective economic system than as a race to build a disembodied superintelligence. We talk about why AGI is mostly a PR term, what machine learning achieved before the LLM hype cycle, and why the assistant-on-your-shoulder vision may be less compelling than it sounds. Jordan explains why explanations need to be actionable, not merely mechanistic; why AlphaFold's missing error bars matter; how prediction-powered inference changes the picture; and why drug discovery is an incentive-design problem rather than a pure pattern-matching problem. ERRATA: Science magazine ranked him the most influential computer scientist, not Nature --- TIMESTAMPS: 00:00:00 Cold open: A demoralizing message to young builders 00:02:04 CyberFund sponsor read 00…

Kubernetes Documentation 2026-05-20 00:00 UTC Score 25.0 AI-200-20260520-developer-an-fcc74d86 Full article

Announcing etcd 3.7.0-beta.0

SIG-Etcd announces the availability of the first beta release of etcd v3.7.0 . This new version of the popular distributed database and key Kubernetes component includes the long-requested RangeStream feature, as well as a refactoring and cleanup of multiple legacy components and interfaces. v3.7 will deliver improved security, better operational reliability, and an improved experience for working with large resultsets. First, however, the project needs users to test the beta. You can find v3.7.0-beta.0 here: Source code Binaries Official container images Please try it out and report issues in the etcd repo . This beta also determines the EOL of version 3.4. RangeStream In etcd v3.6 and earlier, it is challenging to work with requests that return large resultsets. The client or requesting application is forced to wait for the full result set, leading to unpredictable latency and memory usage. The RangeStream RPC lets calling applications accept result sets in chunks, reducing latency and making buffering memory usage more predictable. Much of the work on RangeStream was done by a relatively new contributor to etcd, Jeffrey Ying , a software engineer at Google. New contributors can have a substantial impact on etcd development. "I've always been fascinated by database internals, and building RangeStream was a great opportunity to solve a bottleneck we were hitting in production with Kubernetes. It was the perfect opportunity to collaborate across projects and improve the ecos…

Modal Blog 2026-05-20 00:00 UTC Score 38.0 USR-0086-20260520-ai-specialis-ec35b942 Full article

Scaling reinforcement learning at Applied Compute

How Applied Compute trains custom agents with Reinforcement Learning for enterprises like DoorDash, Cognition, and Mercor on Modal.

ClearML Blog 2026-05-19 19:05 UTC Score 40.0 USR-0084-20260519-ai-specialis-e2bd892b

ClearML Joins the Dell AI Ecosystem Program and Launches AI Factory Blueprints, Making It Easier for Enterprises to Operationalize AI

ClearML is deepening its partnership with Dell Technologies by joining the Dell AI Ecosystem Program, announced at Dell Technologies World 2026. As part of this collaboration, ClearML is launching two pre-validated deployment blueprints — for Kubernetes and OpenShift — available in the Dell Automation Platform catalog, giving enterprises a fast path from bare metal to […]

METR 2026-05-19 18:00 UTC Score 58.0 USR-0147-20260519-research-aca-9d04d191 Full article

Frontier Risk Report (February to March 2026)

Assessment Window: Feb 16, 2026 – Mar 16, 2026 Download PDF Redaction summary statement: Except where explicitly noted in the report, there was no additional redacted information that was important to our conclusions from any of the participating companies. Executive summary and guide to the report Starting in February 2026, METR conducted a pilot exercise to assess misalignment risks from AI agents used inside frontier AI developers, with participation from Anthropic, Google, Meta, and OpenAI. We make three main contributions in this report, each detailed in a separate section. First, we motivate and outline the process we followed for this exercise. 1 Each participant provided: Access to their most capable internal model(s) at the time of assessment, including raw chains of thought. A wide range of non-public information about the capabilities of the shared model(s), how AI was used and monitored internally, and trends in the pace of progress. METR then prepared private reports for each participant, participants approved what non-public information could be disclosed, and METR wrote this public report. This exercise is entity-based rather than model-specific, and is designed to be repeated periodically rather than tied to public releases. Second, we present six key facts that inform our assessment, drawing on evaluations we conducted on the models that participants shared, 2 evaluations we conducted on public models, information shared by participants, 3 findings from a re…

METR 2026-05-19 18:00 UTC Score 44.0 USR-0147-20260519-research-aca-490e247a Full article

Informe de riesgos de la IA de frontera (febrero–marzo de 2026)

Periodo de evaluación: 16 de febrero de 2026 - 16 de marzo de 2026 Descargar PDF (inglés) Declaración resumida sobre omisiones: Salvo donde se indique explícitamente en el informe, no hubo información omitida adicional de ninguna empresa participante que fuera importante para nuestras conclusiones. Resumen ejecutivo y guía del informe En febrero de 2026, METR inició un ejercicio piloto para evaluar los riesgos de desalineación de los agentes de IA usados dentro de empresas desarrolladoras de IA de frontera, con la participación de Anthropic, Google, Meta y OpenAI. En este informe hacemos tres contribuciones principales, cada una detallada en una sección separada. Primero, motivamos y describimos el proceso que seguimos para este ejercicio. 1 Cada participante proporcionó: Acceso a su(s) modelo(s) interno(s) más capaz/capaces en el momento de la evaluación, incluidas cadenas de pensamiento sin procesar. Una amplia variedad de información no pública sobre las capacidades del/de los modelo(s) compartido(s), cómo se usaba y monitoreaba la IA internamente, y las tendencias en el ritmo de progreso. Luego, METR preparó informes privados para cada participante, los participantes aprobaron qué información no pública podía divulgarse, y METR escribió este informe público. Este ejercicio está basado en entidades en lugar de ser específico para un modelo, y está diseñado para repetirse periódicamente en lugar de estar ligado a lanzamientos públicos. Segundo, presentamos seis hechos clav…