AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
36090News Items
8Top Picks
213Blogs
runningLast Run

Latest AI/ML News

36090 matching items

The Dangerous Illusion of AI Coding? - Jeremy Howard
Machine Learning Street Talk 2026-03-03 14:50 UTC Score 62.0 AI-141-20260303-podcasts-and-aa1fcba5 Full article

The Dangerous Illusion of AI Coding? - Jeremy Howard

Dive into the realities of AI-assisted coding, the origins of modern fine-tuning, and the cognitive science behind machine learning with fast.ai founder Jeremy Howard. In this episode, we unpack why AI might be turning software engineering into a slot machine and how to maintain true technical intuition in the age of large language models. GTC is coming, the premier AI conference, great opportunity to learn about AI. NVIDIA and partners will showcase breakthroughs in physical AI, AI factories, agentic AI, and inference, exploring the next wave of AI innovation for developers and researchers. Register for virtual GTC for free, using my link and win NVIDIA DGX Spark (https://nvda.ws/4qQ0LMg) Jeremy Howard is a renowned data scientist, researcher, entrepreneur, and educator. As the co-founder of fast.ai, former President of Kaggle, and the creator of ULMFiT, Jeremy has spent decades democratizing deep learning. His pioneering work laid the foundation for modern transfer learning and the pre-training and fine-tuning paradigm that powers today's language models. Key Topics and Main Insights Discussed: - The Origins of ULMFiT and Fine-Tuning - The Vibe Coding Illusion and Software Engineering - Cognitive Science, Friction, and Learning - The Future of Developers RESCRIPT: https://app.rescript.info/public/share/BhX5zP3b0m63srLOQDKBTFTooSzEMh_ARwmDG_h_izk https://app.rescript.info/api/public/sessions/62d06c0336c567d6/pdf Jeremy Howard: https://x.com/jeremyphoward https://www.answer.…

METR 2026-03-03 08:00 UTC Score 44.0 USR-0147-20260303-research-aca-7bd4bcdb Full article

Observations from two CLI game reimplementation runs with Opus 4.6

Update 7/27/2026: I recently investigated the Slay the Spire deliverable deeper. This revealed some moderate problems that weren’t obvious back when I initially scored it. The problems I found are mostly the type that might take a while to surface, or might take close examination or an experienced player to notice. This has been generally in-line with my impression that models often create deliverables that look good initially, but look worse and worse upon deeper examination (unlike human deliverables, which often leave much more informative first impressions). The rest of the post remains the same as it was on March 3 2026. Summary: Opus 4.6 can, with a simple agent scaffold, create mostly-playable but somewhat broken CLI versions of Slay the Spire and Balatro 1 . Intro Last weekend I was trying to think of really difficult tasks we could give to AI agents to upper-bound their capabilities. I thought of two examples: Recreating a basic version of the video game Slay the Spire in the CLI Recreating a basic version of the video game Balatro in the CLI Both of these video games have a few properties that make it especially easy for AI systems to implement them: They already exist, so the AI doesn’t have to come up with new game ideas and do the enormous amount of work necessary to make it a fun game to play. Most player-relevant information is conveyed through text. They have well-defined rules and interactions between game mechanics. They are turn-based and don’t rely on rea…

OpenMined Blog 2026-03-03 00:30 UTC Score 30.0 USR-0156-20260303-ai-specialis-5e90423e Full article

Why Blocking, Licensing & Pay-to-Access Are Insufficient

TL;DR: The three dominant responses to this problem – blocking scrapers, licensing content, charging for scraper access – share the same flaw: once content is copied to a model server, control and attribution are lost. Each treats distribution and control as mutually exclusive. They are not. A different architecture exists, one in which publishers retain […] The post Why Blocking, Licensing & Pay-to-Access Are Insufficient appeared first on OpenMined .

Weaviate Blog 2026-03-03 00:00 UTC Score 30.0 USR-0073-20260303-ai-specialis-5b76ad39 Full article

Weaviate 1.36 Release

This release introduces HFresh vector index (Preview), and brings Server-side Batching, Object TTL, Async Replication Improvements, Drop Inverted Indices, and Backup Restoration Cancellation to general availability.

OpenMined Blog 2026-03-02 20:30 UTC Score 32.0 USR-0156-20260302-ai-specialis-5d9542bc Full article

Reflections on the 2026 India AI Impact Summit

OpenMined participated in the 2026 India AI Impact Summit in New Delhi, demonstrating BioVault for privacy-preserving genomics research and contributing to policy discussions on data sovereignty, conditional openness, and the "Visit, Don't Move" paradigm for cross-border AI collaboration. The post Reflections on the 2026 India AI Impact Summit appeared first on OpenMined .

Amoral American Power, with Professor Matias Spektor
Carnegie Council AI 2026-03-02 15:30 UTC Score 24.0 USR-0160-20260302-ai-specialis-0de8e9b3 Full article

Amoral American Power, with Professor Matias Spektor

From Caracas to Tehran, U.S. power is no longer justified through a narrative of liberal internationalism. Matias Spektor examines the consequences of this shift.

Lex Fridman Podcast 2026-03-01 04:33 UTC Score 17.0 AI-137-20260301-podcasts-and-d13d94a5 Full article

#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano. His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again. Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-sc See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript: https://lexfridman.com/rick-beato-transcript CONTACT LEX: Feedback – give feedback to Lex: https://lexfridman.com/survey AMA – submit questions, videos or call-in: https://lexfridman.com/ama Hiring – join our team: https://lexfridman.com/hiring Other – other ways to get in

Lex Fridman Podcast 2026-03-01 03:32 UTC Score 17.0 AI-137-20260301-podcasts-and-9757b476 Full article

Transcript for Rick Beato: Greatest Guitarists of All Time, History & Future of Music | Lex Fridman Podcast #492

This is a transcript of Lex Fridman Podcast #492 with Rick Beato. The timestamps in the transcript are clickable links that take you directly to that point in the main video. Please note that the transcript is human generated, and may have errors. Here are some useful links: Go back to this episode’s main page Watch the full YouTube version of the podcast Table of Contents Here are the loose “chapters” in the conversation. Click link to jump approximately to that part in the transcript: 0:00 – Introduction 0:44 – Guitar solos 4:43 – Gypsy jazz and Django Reinhardt 6:14

MongoDB AI Blog 2026-02-27 16:30 UTC Score 29.0 USR-0070-20260227-ai-specialis-8ab3a719 Full article

Inside MongoDB Dublin: The Heart of Our International Growth

Nestled between the Irish Sea and the Wicklow Mountains, MongoDB’s Dublin office brings together people from around the world. It’s a place where you can build a meaningful career, contribute to leading global products, and feel part of a close-knit community. Located in Ballsbridge just south of Dublin city center, the office is a short walk from the Lansdowne DART station and is well-served by multiple bus routes, making it easy to plug into everything the city has to offer. Image of a wall in the MongoDB Dublin office that is painted with Dublin relevant illustrations and text that says "Build together" and "Make it matter" As MongoDB’s international headquarters, Dublin is a key hub where over 300 employees from more than 40 nationalities own critical parts of the company’s products and support customers running mission-critical systems across the globe. Established in 2012, MongoDB Dublin has long played a pivotal role in helping the company achieve its mission of empowering innovators to create, transform, and disrupt industries by unleashing the power of software and data. In this spotlight, you’ll hear from people across MongoDB’s Product & Technology, Sales, and Technical Services teams about what it’s like to build your career—and your life—in Dublin with MongoDB. Image of CEO, CJ Desai, speaking in front of a group of employees in the Dublin office. CEO CJ Desai holds an “Ask Me Anything” session in a recent visit to Dublin. Life at MongoDB Dublin The MongoDB Dubl…

MongoDB AI Blog 2026-02-27 15:30 UTC Score 37.0 USR-0070-20260227-ai-specialis-2ad5a66f Full article

Towards Model-based Verification of a Key-Value Storage Engine

In our previous post, we talked about our process of specifying MongoDB’s distributed transactions protocol and how it enabled novel analysis of its performance characteristics. In this follow-up, we talk about how the modularity of our specification also enabled us to check that the underlying storage engine implementation actually conforms to the abstract behavior defined in our formal specification. That is, we are able to formalize the interface boundary between the sharded transaction protocol and WiredTiger, the underlying key-value storage engine, and develop an automated way to generate tests for checking conformance between the semantics of the underlying storage engine layer and this abstract model. As mentioned in the previous post, a deeper exploration of the concepts covered in this post is covered in our recently published VLDB ’25 paper, Design and Modular Verification of Distributed Transactions in MongoDB. Modular, Model-Based Verification As discussed in Part 1, we had developed a TLA+ specification of MongoDB’s distributed transactions protocol in a compositional manner, describing the high level protocol behavior while also formalizing the boundary between the distributed aspect of the transactions protocol and the underlying single-node WiredTiger storage engine component. As mentioned, the distributed transactions protocol can be viewed as running atop the lower level storage layer. When considering the correctness guarantees of the distributed transact…

TWIML AI Podcast 2026-02-26 23:52 UTC Score 56.0 AI-148-20260226-podcasts-and-b85a484e Full article

AI Trends 2026: OpenClaw Agents, Reasoning LLMs, and More with Sebastian Raschka - #762

In this episode, Sebastian Raschka, independent LLM researcher and author, joins us to break down how the LLM landscape has changed over the past year and what is likely to matter most in 2026. We discuss the shift from raw model scaling to reasoning-focused post-training, inference-time techniques, and better tool integration. Sebastian explains why methods like self-consistency, self-refinement, and verifiable-reward reinforcement learning have become central to progress in domains like math and coding, and where those approaches still fall short. We also explore agentic workflows in practice, including where multi-agent systems add real value and where reliability constraints still dominate system design. The conversation covers architecture trends such as mixture-of-experts, attention efficiency strategies, and the practical impact of long-context models, alongside persistent challenges like continual learning. We close with Sebastian’s perspective on maintaining strong coding fundamentals in the age of AI assistants and a preview of his new book, Build A Reasoning Model (From Scratch). The complete show notes for this episode can be found at https://twimlai.com/go/762.

Instacart Tech Blog 2026-02-26 18:55 UTC Score 24.0 USR-0056-20260226-ai-specialis-590c6078 Full article

Our Early Journey to Transform Instacart’s Discovery Recommendations with LLMs

Key Contributors: Moein Hasani, Hamidreza Shahidi, Trace Levinson, Guanghua Shu Introduction At Instacart, we are laser-focused on improving the user experience by making shopping feel easy, engaging, and personalized. Our discovery surfaces play a central role in bringing this to life. Alongside explicit Search intents, discovery is our opportunity to meet customers’ implicit needs, presenting them with the most relevant and inspiring content we have to offer. The main discovery surface within the Instacart app, referred to here as the “Shopping Hub”, is one of the most critical in this regard. This is the surface a customer lands on within the Instacart app after selecting their desired retailer, guiding them along their entire journey. What users see here shapes not just what they buy, but how intuitive and enjoyable their experience feels. Given its importance, our team runs dozens of Shopping Hub experiments per year, constantly evaluating new ways to enrich the discovery experience. Historically, these experiments have been constrained by static content libraries feeding our recommendation systems. With the rapid advancement of generative AI, a critical opportunity began to emerge: rather than incrementally improving a swath of legacy systems, could we leverage LLMs to rethink how content shows up for a user from the ground up? Which new primitives could we build to uplevel quality, personalization, and cohesion across the page? This blog post walks through our early j…

Weaviate Blog 2026-02-26 00:00 UTC Score 35.0 USR-0073-20260226-ai-specialis-efd829bf Full article

Building A Legal RAG App in 36 Hours

Learn how we built a production-ready, end-to-end RAG application in just 36 hours using the Query Agent and the new Weaviate Agent Skills library.

Vector Institute News 2026-02-25 21:09 UTC Score 35.0 USR-0017-20260225-research-aca-03ba31ef Full article

Remarkable 2026 Poster Session: 60 research projects shaping AI’s future

Vector Institute’s third annual Remarkable 2026 conference brought together over 1,500 researchers and industry leaders in person and online on February 19-20 to explore how AI research translates into real-world […] The post Remarkable 2026 Poster Session: 60 research projects shaping AI’s future appeared first on Vector Institute for Artificial Intelligence .

MongoDB AI Blog 2026-02-25 16:45 UTC Score 45.0 USR-0070-20260225-ai-specialis-4e49681b Full article

Innovating with MongoDB | Customer Successes, February 2026

Who says that winter is when things slow down? MongoDB has had a busy start to the year, with a steady stream of announcements and product features—all against the backdrop of an industry moving at warp speed. It's been a lot, and it's been a blast! For example, the energy at January’s MongoDB.local San Francisco—where we announced capabilities to help teams ship production AI faster—was infectious. MongoDB isn’t just starting a new chapter in AI; we’re rewriting the book in real time. The next generation of AI companies isn't just looking for a temporary place to store data; they’re looking to build on a generational modern data platform. Indeed, the most innovative founders are moving away from rigid, legacy systems and embracing a single, fluid foundation that can grow with them. At MongoDB.local SF, our message was clear: Choose your data platform strategically in order to ship faster. From our new Voyage 4 models to the general availability of our Intelligent Assistant, we are obsessed with anticipating what developers need next. This assistant is particularly impactful because it embeds MongoDB-specific expertise directly into Compass and MongoDB Atlas, allowing developers to troubleshoot performance without the "context-switching" that traditionally slows them down. In this issue, I’m thrilled to spotlight four startups who are building the future on the right foundation. You’ll see how Modelence and Thesys are using our flexible document model to eliminate 'operation…

Vector Institute News 2026-02-24 20:55 UTC Score 38.0 USR-0017-20260224-research-aca-826484c8 Full article

CRISPNAM-FG: An interpretable Fine-Gray deep survival model for competing risks in health care

Vector researchers developed CRISPNAM-FG, a trustworthy AI model that predicts the risk of developing diabetes-related foot complications for patients discharged from hospitals while providing complete transparency in how each decision […] The post CRISPNAM-FG: An interpretable Fine-Gray deep survival model for competing risks in health care appeared first on Vector Institute for Artificial Intelligence .

METR 2026-02-24 08:00 UTC Score 48.0 USR-0147-20260224-research-aca-aae32dc8 Full article

We are Changing our Developer Productivity Experiment Design

METR previously published a paper which found the use of AI tools caused a 20% slowdown in completing tasks among experienced open-source developers, using data from February to June 2025. To understand how AI is impacting developer productivity over time, we started a new experiment in August 2025 with a larger pool of developers using the latest AI tools. Unfortunately, given participant feedback and surveys, we believe that the data from our new experiment gives us an unreliable signal of the current productivity effect of AI tools. The primary reason is that we have observed a significant increase in developers choosing not to participate in the study because they do not wish to work without AI, which likely biases downwards our estimate of AI-assisted speedup. We additionally believe there have been selection effects due to a lower pay rate (we reduced the pay from $150/hr to $50/hr), and that our measurements of time-spent on each task are unreliable for the fraction of developers who use multiple AI agents concurrently. Based on conversations with study participants, we believe it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025. However, because of the selection effects in our experiment, our data is only very weak evidence for the size of this increase. Our raw results show some evidence for speedup. Our early 2025 study found the use of AI causes tasks to take 19% longer, with a confidence inte…

Vector Institute News 2026-02-23 15:12 UTC Score 35.0 USR-0017-20260223-research-aca-4f5ed007 Full article

Demo Day: How the Vector Institute helps Canadian startups turn innovative ideas into commercial reality

By Daniel Kitts “Where were those 10 years ago in AI?” said Stephen Southin to the crowd at Vector’s first Demo Day. The AI industry veteran was energized after hearing […] The post Demo Day: How the Vector Institute helps Canadian startups turn innovative ideas into commercial reality appeared first on Vector Institute for Artificial Intelligence .

Lyft Engineering 2026-02-19 17:28 UTC Score 38.0 USR-0059-20260219-ai-specialis-438ddf27 Full article

Scaling Localization with AI at Lyft

Written by Stefan Zier For years, Lyft’s localization infrastructure relied exclusively on human translation. While this model usually ensured excellent quality, it was bound by multi-day turnarounds and costs that scaled linearly with every new language. For the few languages Lyft initially supported (Spanish, Portuguese, and French), these limits were acceptable. However, Lyft’s expansion goals quickly outpaced what traditional workflows could support. Lyft’s recent Québec launch required compliance with Bill 96 (legislation mandating French-first user experiences) which demanded faster turnaround than multi-day cycles allowed. Simultaneously, the Lyft Urban Solutions (“LUS”: Bikes & Scooters) division sought to expand into European markets, requiring six new languages. The business need had changed as we now needed to move faster without sacrificing quality. This post explores how we re-architected Lyft’s Translation Pipeline to leverage AI alongside linguist oversight and ultimately unlock new market launches. We will walk through context injection, decoupling content generation from evaluation, implementing guardrails, and treating prompts as version-controlled production code. The new pipeline reduces translation latency from days to minutes while maintaining the fidelity required for legal compliance and brand integrity. Note : We will walk through our batch translation pipeline — used for 99% of app and web content — which targets a 30-minute SLA for 95% of translati…

Consultancy.lat AI & GenAI 2026-02-19 16:38 UTC Score 12.0 AI-177-20260219-regional-ai--985c6355

Teladoc Health hires Daniel Murgueitio from BCG to lead Peru business

Multinational telemedicine company Teladoc Health has appointed Daniel Murgueitio as its new Country Manager for Peru. Daniel Murgueitio joins the company from Boston Consulting Group (BCG), where he demonstrated a rapid professional ascent over the last several years.

METR 2026-02-19 08:00 UTC Score 52.0 USR-0147-20260219-research-aca-94103253 Full article

Five lessons from having helped run an AI-Biology RCT

Evidence-based AI policy is important but hard. We need more in-depth studies – which often don’t fit into commercial release cycles. NOTE: This post reflects my personal meta takeaways about the role of Randomized Controlled Trials (RCTs) in AI safety testing. If you have not yet read the Active Site RCT study itself, consider doing so first: see the main results and forecasts . In early 2025, AI systems began outperforming biology experts on biology benchmarks – OpenAI’s o3 outperformed 94% of virology experts on troubleshooting questions in their own specialties. However, it remained unclear how much this translated to real-world novice “uplift” : Could a novice actually use AI to perform wet-lab tasks they could not otherwise perform? Over the summer, I tested this question directly with Active Site (formerly called Panoplia Laboratories). We recruited 153 novices and randomly divided them into an LLM group and an Internet-only group. Over 8 weeks, participants performed fundamental wet-lab tasks involved in molecular biology workflows like reconstructing a virus from a genetic sequence. We found that, while AI showed signs of helpfulness at individual steps, it did not produce a significant effect on end-to-end success across the three core tasks together – a result that surprised many experts . The result provided a mid-2025 snapshot of how well AIs assist novices at molecular biology. I think there are at least two reasons why this result is very informative: It surpr…

After Orthogonality: Virtue-Ethical Agency and AI Alignment
The Gradient 2026-02-18 23:25 UTC Score 18.0 AI-037-20260218-ai-specialis-2df87f06 Full article

After Orthogonality: Virtue-Ethical Agency and AI Alignment

Preface This essay argues that rational people don’t have goals, and that rational AIs shouldn’t have goals. Human actions are rational not because we direct them at some final ‘goals,’ but because we align actions to practices [1] : networks of actions, action-dispositions, action-evaluation criteria,

Practical AI Podcast 2026-02-18 13:57 UTC Score 20.0 AI-143-20260218-podcasts-and-59ec1132 Full article

Cognitive Synthesis and Neural Athletes

As AI accelerates innovation and adoption, leaders are facing rising cognitive load, shifting systems, and new emotional realities inside their organizations. In this episode, Deloitte’s Chief Innovation Officer Deborah Golden joins us to explore how AI is reshaping leadership, why vulnerability and empathy are critical in this moment, and how anti-fragility, not just resilience, will define the future of work. Featuring: Deborah Golden – LinkedIn Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Daniel Whitenack – Website , GitHub , X Links: Deloitte Sponsor: Framer: The enterprise-grade website builder that lets your team ship faster. Get 30% off at framer.com/practicalai Upcoming Events: Register for upcoming webinars here !

Consultancy.lat AI & GenAI 2026-02-18 13:31 UTC Score 12.0 AI-177-20260218-regional-ai--5177f0ab

Endeavor partners with McKinsey to reimagine its support for Brazil’s entrepreneurs

In a major overhaul of its organizational strategy, Endeavour worked with McKinsey & Company on its ambitious plan to transform Brazil into a top global hub for innovation by 2035. Endeavor, a non-profit with chapters around the world, identifies and supports high-impact entrepreneurs and founders that use technology to make a difference.

Consultancy.lat AI & GenAI 2026-02-18 13:30 UTC Score 12.0 AI-177-20260218-regional-ai--84f181ca

The authenticity gap: What Gen Z really think about brands

A global study from SKIM has found that young consumers from around the world value transparency above all else, while nearly a third reject brands for “trying too hard” with forced messaging. Mariana Abelha and Patricia Fujisawa, senior members in the firm’s LATAM business, explore why authenticity isn’t optional for Gen Z and what this means for brands.

InfoWorld AI 2026-02-18 09:00 UTC Score 35.0 USR-0126-20260218-global-ai-ne-2282925e Full article

What is Docker? The spark for the container revolution

Docker is a software platform for building applications based on containers —small and lightweight execution environments that make shared use of the operating system kernel but otherwise run in isolation from one another. While containers have been used in Linux and Unix systems for some time, Docker, an open source project launched in 2013, helped popularize the technology by making it easier than ever for developers to package their software to “build once and run anywhere.” A brief history of Docker Founded as DotCloud in 2008 by Solomon Hykes in Paris, what we now know as Docker started out as a platform as a service (PaaS) before pivoting in 2013 to focus on democratizing the underlying software containers its platform was running on. Hykes first demoed Docker at PyCon in March 2013, explaining that Docker was created because developers kept asking for the underlying technology powering the DotCloud platform. “We did always think it would be cool to be able to say, ‘Yes, here is our low-level piece. Now you can do Linux containers with us and go do whatever you want, go build your platform.’ So that’s what we are doing.” And so, Docker was born, with the open source project quickly picking up traction with developers and attracting the attention of high-profile technology providers like Microsoft, IBM, and Red Hat, as well as venture capitalists willing to pump millions of dollars into the innovative startup. The container revolution had begun. What are containers? As…

METR 2026-02-18 00:00 UTC Score 43.0 USR-0147-20260218-research-aca-3cec17c1 Full article

How We Protect Confidential Information

METR works with AI developers, governments, and other research organizations who sometimes provide nonpublic model access and proprietary information. Over time, we’ve developed confidentiality and security measures to protect such access and information. This post describes our approach at a high level. Confidentiality measures Our confidentiality policy, setup, and norms primarily address the risk of leaks during conversation and in infrastructure, though they also reduce insider threat risk by limiting who knows what. Policy Our confidentiality policy assigns information—including (but not limited to) nonpublic access, lab relationships, policy work, and funding—to our six confidentiality levels, ranging from public to internally siloed, based on sensitivity. At the most restricted end, information about nonpublic models (including capabilities, evaluation timelines, and which developer we’re working with) is limited to researchers directly involved and discussed only by codename. Our own methodology, tasks, and infrastructure are available more broadly within METR, and much of this work is eventually published. Our policy also provides standard responses for sensitive questions, guidance on edge cases, quick rules of thumb with examples and FAQs, and possible slip-ups to watch out for. Table of Contents Commenting on AI developers Easy places to slip up Don’t comment on labs based on non-public info. Any comments […] should be rigorously substantiated by public informati…

Instacart Tech Blog 2026-02-17 16:24 UTC Score 30.0 USR-0056-20260217-ai-specialis-e3637742 Full article

Turning Data into Velocity: Caper’s Edge and Cloud Data Flywheel with Capsight

Key Contributors: Youming Luo, Andrew Tanner, Matas Sriubiskis, Sylvia Lin, Sikun Zhu, Lei Li, Xiao Zhou Introduction Caper is Instacart’s AI-powered smart cart that provides customers with a fast, seamless, and intuitive shopping experience. We achieve this through computer vision and multi-sensor fusion to power accurate product recognition and effortless checkout. Delivering this experience requires Caper’s AI models to understand what truly happens in stores — the movement, intention, and decisions unfolding across every grocery aisle. Historically, our ability to learn from production environments was limited. Even though the carts were deployed in stores, we lacked a scalable way to collect real‑world data that would allow us to rapidly iterate and improve our models. This resulted in three core challenges: Scalable Onboard Observability : We had little visibility into what was happening on the cart, in the stores. When something went wrong, it was hard to understand or reproduce the scenario. At the same time, each cart generates gigabytes of multimodal data, from sources such as cameras, weight sensors, and localization sensors. We needed a centralized way to capture key moments so the team could clearly understand what the cart experiences, how users interact with it, and where to improve — all while maintaining a magical user experience and minimal impact on the network. Data Quality and Diversity: Our models were primarily trained on manually-collected data that d…

MongoDB AI Blog 2026-02-17 15:30 UTC Score 61.0 USR-0070-20260217-ai-specialis-685171d4 Full article

Building a Movie Recommendation Engine with Hugging Face and Voyage AI

This guest blog post is from Arek Borucki, Machine Learning Platform & Data Engineer for Hugging Face - a collaboration platform for the machine learning community. The Hugging Face Hub works as a central place where anyone can share, explore, discover, and experiment with open-source ML. HF empowers the next generation of machine learning engineers, scientists, and end users to learn, collaborate and share their work to build an open and ethical AI future together. With the fast-growing community, some of the most used open-source ML libraries and tools, and a talented science team exploring the edge of tech, Hugging Face is at the heart of the AI revolution. Traditional movie search relies on filtering by genre, actor, or title. But what if you could search by how you feel? Imagine typing: "something uplifting after a rough day at work" "a movie that will make me cry" "I need adrenaline, can't sleep anyway" "something to watch with grandma who hates violence" This is mood-based semantic search: matching your emotional state to movie plot descriptions using AI embeddings. In this tutorial, you will build a mood-based movie recommendation engine using three powerful technologies: voyage-4-nano (a state-of-the-art open-source embedding model), Hugging Face (for model and dataset hosting), and MongoDB Atlas Vector Search (for storing and querying embeddings at scale). Why mood-based search? Genre tags are coarse. A "drama" can be heartwarming or devastating. A "comedy" can be…

METR 2026-02-17 08:00 UTC Score 49.0 USR-0147-20260217-research-aca-7e22be94 Full article

Analyzing coding agent transcripts to upper bound productivity gains from AI agents

Introduction Human uplift studies like the one we did in 2025 are becoming more expensive as working without AI becomes increasingly costly. In this post, I investigate whether coding agent transcripts could serve as a cheaper alternative for estimating uplift. I prototyped this using 5305 Claude Code transcripts generated in January 2026 by 7 METR technical staff 1 . I used an LLM judge to estimate how long each task would have taken an experienced software engineer without AI tools, then compared that to the time people actually spent on these tasks to calculate a time savings factor . Takeaways This method estimates a time savings factor of ~1.5x to ~13x on Claude Code-assisted tasks for 7 METR technical staff in January 2026 – though this result comes with substantial caveats. I believe the true productivity multiplier is substantially lower, and the time savings factor is a soft upper bound for the true uplift that the individuals experienced. Increased agent concurrency may contribute to a higher time savings factor on the Claude Code-assisted task distributions. Limitations The time savings factor on the coding agent-assisted task distributions does not equal the productivity multiplier. People likely do not create 10x as much value with AI, even if we observe a 10x time savings factor on tasks that people do with AI. I believe the time savings factor overestimates AI-enabled productivity gains for reasons including: Task Substitution. With AI assistance, people somet…