AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
26147News Items
8Top Picks
157Blogs
successLast Run

Latest AI/ML News

26147 matching items

Lex Fridman Podcast 2026-03-01 04:33 UTC Score 17.0 AI-137-20260301-podcasts-and-d13d94a5 Full article

#492 – Rick Beato: Greatest Guitarists of All Time, History & Future of Music

Rick Beato is a music educator, interviewer, producer, songwriter, and a true multi-instrument musician, playing guitar, bass, cello & piano. His incredible YouTube channel celebrates great musicians & musical ideas, and helps millions of people fall in love with great music all over again. Thank you for listening ❤ Check out our sponsors: https://lexfridman.com/sponsors/ep492-sc See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Transcript: https://lexfridman.com/rick-beato-transcript CONTACT LEX: Feedback – give feedback to Lex: https://lexfridman.com/survey AMA – submit questions, videos or call-in: https://lexfridman.com/ama Hiring – join our team: https://lexfridman.com/hiring Other – other ways to get in

Lex Fridman Podcast 2026-03-01 03:32 UTC Score 17.0 AI-137-20260301-podcasts-and-9757b476 Full article

Transcript for Rick Beato: Greatest Guitarists of All Time, History & Future of Music | Lex Fridman Podcast #492

This is a transcript of Lex Fridman Podcast #492 with Rick Beato. The timestamps in the transcript are clickable links that take you directly to that point in the main video. Please note that the transcript is human generated, and may have errors. Here are some useful links: Go back to this episode’s main page Watch the full YouTube version of the podcast Table of Contents Here are the loose “chapters” in the conversation. Click link to jump approximately to that part in the transcript: 0:00 – Introduction 0:44 – Guitar solos 4:43 – Gypsy jazz and Django Reinhardt 6:14

MongoDB AI Blog 2026-02-27 16:30 UTC Score 29.0 USR-0070-20260227-ai-specialis-8ab3a719 Full article

Inside MongoDB Dublin: The Heart of Our International Growth

Nestled between the Irish Sea and the Wicklow Mountains, MongoDB’s Dublin office brings together people from around the world. It’s a place where you can build a meaningful career, contribute to leading global products, and feel part of a close-knit community. Located in Ballsbridge just south of Dublin city center, the office is a short walk from the Lansdowne DART station and is well-served by multiple bus routes, making it easy to plug into everything the city has to offer. Image of a wall in the MongoDB Dublin office that is painted with Dublin relevant illustrations and text that says "Build together" and "Make it matter" As MongoDB’s international headquarters, Dublin is a key hub where over 300 employees from more than 40 nationalities own critical parts of the company’s products and support customers running mission-critical systems across the globe. Established in 2012, MongoDB Dublin has long played a pivotal role in helping the company achieve its mission of empowering innovators to create, transform, and disrupt industries by unleashing the power of software and data. In this spotlight, you’ll hear from people across MongoDB’s Product & Technology, Sales, and Technical Services teams about what it’s like to build your career—and your life—in Dublin with MongoDB. Image of CEO, CJ Desai, speaking in front of a group of employees in the Dublin office. CEO CJ Desai holds an “Ask Me Anything” session in a recent visit to Dublin. Life at MongoDB Dublin The MongoDB Dubl…

MongoDB AI Blog 2026-02-27 15:30 UTC Score 37.0 USR-0070-20260227-ai-specialis-2ad5a66f Full article

Towards Model-based Verification of a Key-Value Storage Engine

In our previous post, we talked about our process of specifying MongoDB’s distributed transactions protocol and how it enabled novel analysis of its performance characteristics. In this follow-up, we talk about how the modularity of our specification also enabled us to check that the underlying storage engine implementation actually conforms to the abstract behavior defined in our formal specification. That is, we are able to formalize the interface boundary between the sharded transaction protocol and WiredTiger, the underlying key-value storage engine, and develop an automated way to generate tests for checking conformance between the semantics of the underlying storage engine layer and this abstract model. As mentioned in the previous post, a deeper exploration of the concepts covered in this post is covered in our recently published VLDB ’25 paper, Design and Modular Verification of Distributed Transactions in MongoDB. Modular, Model-Based Verification As discussed in Part 1, we had developed a TLA+ specification of MongoDB’s distributed transactions protocol in a compositional manner, describing the high level protocol behavior while also formalizing the boundary between the distributed aspect of the transactions protocol and the underlying single-node WiredTiger storage engine component. As mentioned, the distributed transactions protocol can be viewed as running atop the lower level storage layer. When considering the correctness guarantees of the distributed transact…

TWIML AI Podcast 2026-02-26 23:52 UTC Score 56.0 AI-148-20260226-podcasts-and-b85a484e Full article

AI Trends 2026: OpenClaw Agents, Reasoning LLMs, and More with Sebastian Raschka - #762

In this episode, Sebastian Raschka, independent LLM researcher and author, joins us to break down how the LLM landscape has changed over the past year and what is likely to matter most in 2026. We discuss the shift from raw model scaling to reasoning-focused post-training, inference-time techniques, and better tool integration. Sebastian explains why methods like self-consistency, self-refinement, and verifiable-reward reinforcement learning have become central to progress in domains like math and coding, and where those approaches still fall short. We also explore agentic workflows in practice, including where multi-agent systems add real value and where reliability constraints still dominate system design. The conversation covers architecture trends such as mixture-of-experts, attention efficiency strategies, and the practical impact of long-context models, alongside persistent challenges like continual learning. We close with Sebastian’s perspective on maintaining strong coding fundamentals in the age of AI assistants and a preview of his new book, Build A Reasoning Model (From Scratch). The complete show notes for this episode can be found at https://twimlai.com/go/762.

Instacart Tech Blog 2026-02-26 18:55 UTC Score 24.0 USR-0056-20260226-ai-specialis-590c6078 Full article

Our Early Journey to Transform Instacart’s Discovery Recommendations with LLMs

Key Contributors: Moein Hasani, Hamidreza Shahidi, Trace Levinson, Guanghua Shu Introduction At Instacart, we are laser-focused on improving the user experience by making shopping feel easy, engaging, and personalized. Our discovery surfaces play a central role in bringing this to life. Alongside explicit Search intents, discovery is our opportunity to meet customers’ implicit needs, presenting them with the most relevant and inspiring content we have to offer. The main discovery surface within the Instacart app, referred to here as the “Shopping Hub”, is one of the most critical in this regard. This is the surface a customer lands on within the Instacart app after selecting their desired retailer, guiding them along their entire journey. What users see here shapes not just what they buy, but how intuitive and enjoyable their experience feels. Given its importance, our team runs dozens of Shopping Hub experiments per year, constantly evaluating new ways to enrich the discovery experience. Historically, these experiments have been constrained by static content libraries feeding our recommendation systems. With the rapid advancement of generative AI, a critical opportunity began to emerge: rather than incrementally improving a swath of legacy systems, could we leverage LLMs to rethink how content shows up for a user from the ground up? Which new primitives could we build to uplevel quality, personalization, and cohesion across the page? This blog post walks through our early j…

Weaviate Blog 2026-02-26 00:00 UTC Score 35.0 USR-0073-20260226-ai-specialis-efd829bf Full article

Building A Legal RAG App in 36 Hours

Learn how we built a production-ready, end-to-end RAG application in just 36 hours using the Query Agent and the new Weaviate Agent Skills library.

Vector Institute News 2026-02-25 21:09 UTC Score 35.0 USR-0017-20260225-research-aca-03ba31ef Full article

Remarkable 2026 Poster Session: 60 research projects shaping AI’s future

Vector Institute’s third annual Remarkable 2026 conference brought together over 1,500 researchers and industry leaders in person and online on February 19-20 to explore how AI research translates into real-world […] The post Remarkable 2026 Poster Session: 60 research projects shaping AI’s future appeared first on Vector Institute for Artificial Intelligence .

MongoDB AI Blog 2026-02-25 16:45 UTC Score 45.0 USR-0070-20260225-ai-specialis-4e49681b Full article

Innovating with MongoDB | Customer Successes, February 2026

Who says that winter is when things slow down? MongoDB has had a busy start to the year, with a steady stream of announcements and product features—all against the backdrop of an industry moving at warp speed. It's been a lot, and it's been a blast! For example, the energy at January’s MongoDB.local San Francisco—where we announced capabilities to help teams ship production AI faster—was infectious. MongoDB isn’t just starting a new chapter in AI; we’re rewriting the book in real time. The next generation of AI companies isn't just looking for a temporary place to store data; they’re looking to build on a generational modern data platform. Indeed, the most innovative founders are moving away from rigid, legacy systems and embracing a single, fluid foundation that can grow with them. At MongoDB.local SF, our message was clear: Choose your data platform strategically in order to ship faster. From our new Voyage 4 models to the general availability of our Intelligent Assistant, we are obsessed with anticipating what developers need next. This assistant is particularly impactful because it embeds MongoDB-specific expertise directly into Compass and MongoDB Atlas, allowing developers to troubleshoot performance without the "context-switching" that traditionally slows them down. In this issue, I’m thrilled to spotlight four startups who are building the future on the right foundation. You’ll see how Modelence and Thesys are using our flexible document model to eliminate 'operation…

Vector Institute News 2026-02-24 20:55 UTC Score 38.0 USR-0017-20260224-research-aca-826484c8 Full article

CRISPNAM-FG: An interpretable Fine-Gray deep survival model for competing risks in health care

Vector researchers developed CRISPNAM-FG, a trustworthy AI model that predicts the risk of developing diabetes-related foot complications for patients discharged from hospitals while providing complete transparency in how each decision […] The post CRISPNAM-FG: An interpretable Fine-Gray deep survival model for competing risks in health care appeared first on Vector Institute for Artificial Intelligence .

METR 2026-02-24 08:00 UTC Score 48.0 USR-0147-20260224-research-aca-aae32dc8 Full article

We are Changing our Developer Productivity Experiment Design

METR previously published a paper which found the use of AI tools caused a 20% slowdown in completing tasks among experienced open-source developers, using data from February to June 2025. To understand how AI is impacting developer productivity over time, we started a new experiment in August 2025 with a larger pool of developers using the latest AI tools. Unfortunately, given participant feedback and surveys, we believe that the data from our new experiment gives us an unreliable signal of the current productivity effect of AI tools. The primary reason is that we have observed a significant increase in developers choosing not to participate in the study because they do not wish to work without AI, which likely biases downwards our estimate of AI-assisted speedup. We additionally believe there have been selection effects due to a lower pay rate (we reduced the pay from $150/hr to $50/hr), and that our measurements of time-spent on each task are unreliable for the fraction of developers who use multiple AI agents concurrently. Based on conversations with study participants, we believe it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025. However, because of the selection effects in our experiment, our data is only very weak evidence for the size of this increase. Our raw results show some evidence for speedup. Our early 2025 study found the use of AI causes tasks to take 19% longer, with a confidence inte…

Vector Institute News 2026-02-23 15:12 UTC Score 35.0 USR-0017-20260223-research-aca-4f5ed007 Full article

Demo Day: How the Vector Institute helps Canadian startups turn innovative ideas into commercial reality

By Daniel Kitts “Where were those 10 years ago in AI?” said Stephen Southin to the crowd at Vector’s first Demo Day. The AI industry veteran was energized after hearing […] The post Demo Day: How the Vector Institute helps Canadian startups turn innovative ideas into commercial reality appeared first on Vector Institute for Artificial Intelligence .

Lyft Engineering 2026-02-19 17:28 UTC Score 38.0 USR-0059-20260219-ai-specialis-438ddf27 Full article

Scaling Localization with AI at Lyft

Written by Stefan Zier For years, Lyft’s localization infrastructure relied exclusively on human translation. While this model usually ensured excellent quality, it was bound by multi-day turnarounds and costs that scaled linearly with every new language. For the few languages Lyft initially supported (Spanish, Portuguese, and French), these limits were acceptable. However, Lyft’s expansion goals quickly outpaced what traditional workflows could support. Lyft’s recent Québec launch required compliance with Bill 96 (legislation mandating French-first user experiences) which demanded faster turnaround than multi-day cycles allowed. Simultaneously, the Lyft Urban Solutions (“LUS”: Bikes & Scooters) division sought to expand into European markets, requiring six new languages. The business need had changed as we now needed to move faster without sacrificing quality. This post explores how we re-architected Lyft’s Translation Pipeline to leverage AI alongside linguist oversight and ultimately unlock new market launches. We will walk through context injection, decoupling content generation from evaluation, implementing guardrails, and treating prompts as version-controlled production code. The new pipeline reduces translation latency from days to minutes while maintaining the fidelity required for legal compliance and brand integrity. Note : We will walk through our batch translation pipeline — used for 99% of app and web content — which targets a 30-minute SLA for 95% of translati…

Consultancy.lat AI & GenAI 2026-02-19 16:38 UTC Score 12.0 AI-177-20260219-regional-ai--985c6355

Teladoc Health hires Daniel Murgueitio from BCG to lead Peru business

Multinational telemedicine company Teladoc Health has appointed Daniel Murgueitio as its new Country Manager for Peru. Daniel Murgueitio joins the company from Boston Consulting Group (BCG), where he demonstrated a rapid professional ascent over the last several years.

METR 2026-02-19 08:00 UTC Score 52.0 USR-0147-20260219-research-aca-94103253 Full article

Five lessons from having helped run an AI-Biology RCT

Evidence-based AI policy is important but hard. We need more in-depth studies – which often don’t fit into commercial release cycles. NOTE: This post reflects my personal meta takeaways about the role of Randomized Controlled Trials (RCTs) in AI safety testing. If you have not yet read the Active Site RCT study itself, consider doing so first: see the main results and forecasts . In early 2025, AI systems began outperforming biology experts on biology benchmarks – OpenAI’s o3 outperformed 94% of virology experts on troubleshooting questions in their own specialties. However, it remained unclear how much this translated to real-world novice “uplift” : Could a novice actually use AI to perform wet-lab tasks they could not otherwise perform? Over the summer, I tested this question directly with Active Site (formerly called Panoplia Laboratories). We recruited 153 novices and randomly divided them into an LLM group and an Internet-only group. Over 8 weeks, participants performed fundamental wet-lab tasks involved in molecular biology workflows like reconstructing a virus from a genetic sequence. We found that, while AI showed signs of helpfulness at individual steps, it did not produce a significant effect on end-to-end success across the three core tasks together – a result that surprised many experts . The result provided a mid-2025 snapshot of how well AIs assist novices at molecular biology. I think there are at least two reasons why this result is very informative: It surpr…

After Orthogonality: Virtue-Ethical Agency and AI Alignment
The Gradient 2026-02-18 23:25 UTC Score 18.0 AI-037-20260218-ai-specialis-2df87f06 Full article

After Orthogonality: Virtue-Ethical Agency and AI Alignment

Preface This essay argues that rational people don’t have goals, and that rational AIs shouldn’t have goals. Human actions are rational not because we direct them at some final ‘goals,’ but because we align actions to practices [1] : networks of actions, action-dispositions, action-evaluation criteria,

Practical AI Podcast 2026-02-18 13:57 UTC Score 25.0 AI-143-20260218-podcasts-and-59ec1132 Full article

Cognitive Synthesis and Neural Athletes

As AI accelerates innovation and adoption, leaders are facing rising cognitive load, shifting systems, and new emotional realities inside their organizations. In this episode, Deloitte’s Chief Innovation Officer Deborah Golden joins us to explore how AI is reshaping leadership, why vulnerability and empathy are critical in this moment, and how anti-fragility, not just resilience, will define the future of work. Featuring: Deborah Golden – LinkedIn Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Daniel Whitenack – Website , GitHub , X Links: Deloitte Sponsor: Framer - The website builder that turns your dot com from a formality into a tool for growth. Check it out at framer.com/PRACTICALAI Upcoming Events: Register for upcoming webinars here !

Consultancy.lat AI & GenAI 2026-02-18 13:31 UTC Score 12.0 AI-177-20260218-regional-ai--5177f0ab

Endeavor partners with McKinsey to reimagine its support for Brazil’s entrepreneurs

In a major overhaul of its organizational strategy, Endeavour worked with McKinsey & Company on its ambitious plan to transform Brazil into a top global hub for innovation by 2035. Endeavor, a non-profit with chapters around the world, identifies and supports high-impact entrepreneurs and founders that use technology to make a difference.

Consultancy.lat AI & GenAI 2026-02-18 13:30 UTC Score 12.0 AI-177-20260218-regional-ai--84f181ca

The authenticity gap: What Gen Z really think about brands

A global study from SKIM has found that young consumers from around the world value transparency above all else, while nearly a third reject brands for “trying too hard” with forced messaging. Mariana Abelha and Patricia Fujisawa, senior members in the firm’s LATAM business, explore why authenticity isn’t optional for Gen Z and what this means for brands.

InfoWorld AI 2026-02-18 09:00 UTC Score 35.0 USR-0126-20260218-global-ai-ne-2282925e Full article

What is Docker? The spark for the container revolution

Docker is a software platform for building applications based on containers —small and lightweight execution environments that make shared use of the operating system kernel but otherwise run in isolation from one another. While containers have been used in Linux and Unix systems for some time, Docker, an open source project launched in 2013, helped popularize the technology by making it easier than ever for developers to package their software to “build once and run anywhere.” A brief history of Docker Founded as DotCloud in 2008 by Solomon Hykes in Paris, what we now know as Docker started out as a platform as a service (PaaS) before pivoting in 2013 to focus on democratizing the underlying software containers its platform was running on. Hykes first demoed Docker at PyCon in March 2013, explaining that Docker was created because developers kept asking for the underlying technology powering the DotCloud platform. “We did always think it would be cool to be able to say, ‘Yes, here is our low-level piece. Now you can do Linux containers with us and go do whatever you want, go build your platform.’ So that’s what we are doing.” And so, Docker was born, with the open source project quickly picking up traction with developers and attracting the attention of high-profile technology providers like Microsoft, IBM, and Red Hat, as well as venture capitalists willing to pump millions of dollars into the innovative startup. The container revolution had begun. What are containers? As…

METR 2026-02-18 00:00 UTC Score 43.0 USR-0147-20260218-research-aca-3cec17c1 Full article

How We Protect Confidential Information

METR works with AI developers, governments, and other research organizations who sometimes provide nonpublic model access and proprietary information. Over time, we’ve developed confidentiality and security measures to protect such access and information. This post describes our approach at a high level. Confidentiality measures Our confidentiality policy, setup, and norms primarily address the risk of leaks during conversation and in infrastructure, though they also reduce insider threat risk by limiting who knows what. Policy Our confidentiality policy assigns information—including (but not limited to) nonpublic access, lab relationships, policy work, and funding—to our six confidentiality levels, ranging from public to internally siloed, based on sensitivity. At the most restricted end, information about nonpublic models (including capabilities, evaluation timelines, and which developer we’re working with) is limited to researchers directly involved and discussed only by codename. Our own methodology, tasks, and infrastructure are available more broadly within METR, and much of this work is eventually published. Our policy also provides standard responses for sensitive questions, guidance on edge cases, quick rules of thumb with examples and FAQs, and possible slip-ups to watch out for. Table of Contents Commenting on AI developers Easy places to slip up Don’t comment on labs based on non-public info. Any comments […] should be rigorously substantiated by public informati…

Instacart Tech Blog 2026-02-17 16:24 UTC Score 30.0 USR-0056-20260217-ai-specialis-e3637742 Full article

Turning Data into Velocity: Caper’s Edge and Cloud Data Flywheel with Capsight

Key Contributors: Youming Luo, Andrew Tanner, Matas Sriubiskis, Sylvia Lin, Sikun Zhu, Lei Li, Xiao Zhou Introduction Caper is Instacart’s AI-powered smart cart that provides customers with a fast, seamless, and intuitive shopping experience. We achieve this through computer vision and multi-sensor fusion to power accurate product recognition and effortless checkout. Delivering this experience requires Caper’s AI models to understand what truly happens in stores — the movement, intention, and decisions unfolding across every grocery aisle. Historically, our ability to learn from production environments was limited. Even though the carts were deployed in stores, we lacked a scalable way to collect real‑world data that would allow us to rapidly iterate and improve our models. This resulted in three core challenges: Scalable Onboard Observability : We had little visibility into what was happening on the cart, in the stores. When something went wrong, it was hard to understand or reproduce the scenario. At the same time, each cart generates gigabytes of multimodal data, from sources such as cameras, weight sensors, and localization sensors. We needed a centralized way to capture key moments so the team could clearly understand what the cart experiences, how users interact with it, and where to improve — all while maintaining a magical user experience and minimal impact on the network. Data Quality and Diversity: Our models were primarily trained on manually-collected data that d…

MongoDB AI Blog 2026-02-17 15:30 UTC Score 61.0 USR-0070-20260217-ai-specialis-685171d4 Full article

Building a Movie Recommendation Engine with Hugging Face and Voyage AI

This guest blog post is from Arek Borucki, Machine Learning Platform & Data Engineer for Hugging Face - a collaboration platform for the machine learning community. The Hugging Face Hub works as a central place where anyone can share, explore, discover, and experiment with open-source ML. HF empowers the next generation of machine learning engineers, scientists, and end users to learn, collaborate and share their work to build an open and ethical AI future together. With the fast-growing community, some of the most used open-source ML libraries and tools, and a talented science team exploring the edge of tech, Hugging Face is at the heart of the AI revolution. Traditional movie search relies on filtering by genre, actor, or title. But what if you could search by how you feel? Imagine typing: "something uplifting after a rough day at work" "a movie that will make me cry" "I need adrenaline, can't sleep anyway" "something to watch with grandma who hates violence" This is mood-based semantic search: matching your emotional state to movie plot descriptions using AI embeddings. In this tutorial, you will build a mood-based movie recommendation engine using three powerful technologies: voyage-4-nano (a state-of-the-art open-source embedding model), Hugging Face (for model and dataset hosting), and MongoDB Atlas Vector Search (for storing and querying embeddings at scale). Why mood-based search? Genre tags are coarse. A "drama" can be heartwarming or devastating. A "comedy" can be…

METR 2026-02-17 08:00 UTC Score 49.0 USR-0147-20260217-research-aca-7e22be94 Full article

Analyzing coding agent transcripts to upper bound productivity gains from AI agents

Introduction Human uplift studies like the one we did in 2025 are becoming more expensive as working without AI becomes increasingly costly. In this post, I investigate whether coding agent transcripts could serve as a cheaper alternative for estimating uplift. I prototyped this using 5305 Claude Code transcripts generated in January 2026 by 7 METR technical staff 1 . I used an LLM judge to estimate how long each task would have taken an experienced software engineer without AI tools, then compared that to the time people actually spent on these tasks to calculate a time savings factor . Takeaways This method estimates a time savings factor of ~1.5x to ~13x on Claude Code-assisted tasks for 7 METR technical staff in January 2026 – though this result comes with substantial caveats. I believe the true productivity multiplier is substantially lower, and the time savings factor is a soft upper bound for the true uplift that the individuals experienced. Increased agent concurrency may contribute to a higher time savings factor on the Claude Code-assisted task distributions. Limitations The time savings factor on the coding agent-assisted task distributions does not equal the productivity multiplier. People likely do not create 10x as much value with AI, even if we observe a 10x time savings factor on tasks that people do with AI. I believe the time savings factor overestimates AI-enabled productivity gains for reasons including: Task Substitution. With AI assistance, people somet…

What If Intelligence Didn't Evolve? It "Was There" From the Start! - Blaise Agüera y Arcas
Machine Learning Street Talk 2026-02-16 07:51 UTC Score 22.0 AI-141-20260216-podcasts-and-91cc847d Full article

What If Intelligence Didn't Evolve? It "Was There" From the Start! - Blaise Agüera y Arcas

Blaise Agüera y Arcas presenting at ALife 2025 — the most technically detailed public walkthrough of the ideas in his *What is Life?* and *What is Intelligence?* books that we've come across. He covers the BFF experiments (self-replicating programs emerging spontaneously from random noise), the mathematical framework connecting Lotka-Volterra population dynamics with Smoluchowski coagulation, eigenvalue analysis of cooperation matrices, and his central claim that symbiogenesis — not mutation — is the primary engine of evolutionary novelty. The experimental results are genuinely striking: complex self-replicating code arising from random byte strings with zero mutation, a sharp phase transition that looks like gelation, and a proof that blocking deep symbiogenetic ancestry trees prevents the transition entirely. A few things worth flagging for critical viewers: — The substrate is more carefully engineered than the framing sometimes suggests. The choice of language, tape length, interaction protocol, and step limits all shape what emerges. Their own SUBLEQ counterexample (where self-replicators *don't* arise despite being theoretically possible) highlights that these design choices matter substantially — and a general theory of which substrates support this transition is still missing. — The leap from "self-replicating programs on fixed-length tapes" to "life was computational and intelligent from the start" involves significant philosophical extrapolation beyond what the expe…

Practical AI Podcast 2026-02-13 15:57 UTC Score 36.0 AI-143-20260213-podcasts-and-2841b1bd Full article

AI incidents, audits, and the limits of benchmarks

AI is moving fast from research to real-world deployment, and when things go wrong, the consequences are no longer hypothetical. In this episode, Sean McGregor, co-founder of the AI Verification & Evaluation Research Institute and also the founder of the AI Incident Database, joins Chris and Dan to discuss AI safety, verification, evaluation, and auditing. They explore why benchmarks often fall short, what red-teaming at DEFCON reveals about machine learning risks, and how organizations can better assess and manage AI systems in practice. Featuring: Sean McGregor– LinkedIn Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Daniel Whitenack – Website , GitHub , X Links: AI Verification & Evaluation Research Institute AI Incident Database 38th convening of IAAI BenchRisk State of Global AI Incident Reporting Upcoming Events: Register for upcoming webinars here !

Lyft Engineering 2026-02-12 17:07 UTC Score 41.0 USR-0059-20260212-ai-specialis-3f9a8e21 Full article

Trusting the Untestable: Validation and Diagnostics for the Doubly Robust Models

written by Ross Chu and Shima Nassiri The Causal Frontier: Measurement Beyond Randomization The gold standard for determining the causal impact of a policy or product change at a company like Lyft is the A/B test (randomized experiment). By randomly assigning users to a treatment or control group, A/B tests inherently eliminate bias, providing clean estimates of the Average Treatment Effect (ATE). However, many critical business questions and large-scale initiatives simply cannot be randomized . This forces scientists to move past traditional experimentation and leverage quasi-experimental methods. We rely on non-randomized measurement in several key scenarios across Lyft: Partnerships and Policies: Assessing the incremental impact of a partnership (e.g., linking two company accounts) is often a non-randomized assignment. Since these collaborations require coordinated operational work across both companies and are typically announced or promoted broadly, this makes controlled randomization impractical. Long-Term Effect (LTE): Measuring effects that unfold over a long period, like the LTE of high prices on future rides, is typically handled by observational studies. Post-Launch Evaluation: Continuous monitoring of a policy after it has been fully rolled out requires a method that doesn’t involve costly holdout groups or degradation tests. Biased Data: In cases where pre-existing experimental data is found to have an imbalance, a quasi-experimental approach can potentially lev…

Andrej Karpathy Blog 2026-02-12 07:00 UTC Score 54.0 USR-0115-20260212-ai-specialis-6d759dd0 Full article

microgpt

This is a brief guide to my new art project microgpt , a single file of 200 lines of pure Python with no dependencies that trains and inferences a GPT. This file contains the full algorithmic content of what is needed: dataset of documents, tokenizer, autograd engine, a GPT-2-like neural network architecture, the Adam optimizer, training loop, and inference loop. Everything else is just efficiency. I cannot simplify this any further. This script is the culmination of multiple projects (micrograd, makemore, nanogpt, etc.) and a decade-long obsession to simplify LLMs to their bare essentials, and I think it is beautiful 🥹. It even breaks perfectly across 3 columns: Where to find it: This GitHub gist has the full source code: microgpt.py It’s also available on this web page: https://karpathy.ai/microgpt.html Also available as a Google Colab notebook NEW : buy microgpt as a triptych on my art store at karpathy.art :) The following is my guide on stepping an interested reader through the code. Dataset The fuel of large language models is a stream of text data, optionally separated into a set of documents. In production-grade applications, each document would be an internet web page but for microgpt we use a simpler example of 32,000 names, one per line: # Let there be an input dataset `docs`: list[str] of documents (e.g. a dataset of names) if not os . path . exists ( 'input.txt' ): import urllib.request names_url = 'https://raw.githubusercontent.com/karpathy/makemore/refs/heads/…