Latest AI/ML News
35696 matching items
Introducing LifeSciBench
Introducing LifeSciBench, an expert-authored, expert-reviewed benchmark for evaluating how AI systems handle real-world life science research tasks and decisions.
AI Now is Hiring a Senior Fellow, Global Programs
The Senior Fellow, Global Programs, will lead a tightly scoped, policy-responsive workstream at the intersection of AI, industrial policy, and global political economy; building on AI Now’s existing research on AI nationalism and translating it into a directed research and policy agenda. The terrain around “AI sovereignty” is rapidly being reshaped by an aggressive US […] The post AI Now is Hiring a Senior Fellow, Global Programs appeared first on AI Now Institute .
Hands Free, AIs Forward: NVIDIA XR AI Brings Agents to AR Glasses
NVIDIA XR AI is now available in public beta, giving developers a framework for building multimodal AI agents for AR glasses and XR devices.
As Brazilian media embrace prediction markets, experts warn of election distortion
News outlets are citing unregulated betting platforms alongside survey-based polls. Critics say the markets are easily manipulated and do not reliably reflect voter sentiment. The post As Brazilian media embrace prediction markets, experts warn of election distortion appeared first on LatAm Journalism Review by the Knight Center .
Coherent Breaks Ground on Expanded Texas Facility, Scaling AI’s Optical Backbone
AI runs at the speed of light. More and more, that light is made in Texas. Coherent broke ground today on an expanded manufacturing building in Sherman, Texas. The company makes the lasers, optical components and compound semiconductors that wire AI systems together — and runs what it calls the world’s first 6-inch indium phosphide […]
Why AI Agents Break the GenAI Security Model with Devvret Rishi - #770
In this episode, Sam talks with Dev Rishi, GM of AI at Rubrik, about what happens when agents move beyond answering questions and start taking action across tools, systems, and business processes. We explore why the enterprise playbook of static guardrails plus human approval starts to break down in the agent era. Agents are useful because they can plan, call tools, update systems, write code, send messages, and operate across workflows at machine speed, but those same capabilities make them difficult to govern with rules written in advance or approval prompts reviewed one at a time. Dev explains why tool access increases blast radius, why agents can route around controls in surprising ways, and why human-in-the-loop review can become security theater when agents operate at scale. We also discuss what enterprises need instead: better visibility, runtime enforcement, policy-aware governance, agent observability, and recovery mechanisms for when something goes wrong. Along the way, we dig into MCP and tool sprawl, small language models for policy enforcement, defense in depth, agent rewind, and why AI may be needed to help secure AI. 🗒️ Full show notes: https://twimlai.com/go/770.
What are git worktrees, and why should I use them?
Git worktrees have been around since 2015, but it wasn't until recently they became popular. Learn what they are, how to use them, and why you might. The post What are git worktrees, and why should I use them? appeared first on The GitHub Blog .
Predicting LLM Safety Before Release by Simulating Deployment
Paper link Before releasing a new model, labs need to understand not just what it can do, but how it is likely to behave in real-world use, including where it might introduce new risks. This becomes even more important as capabilities increase. As part of our pre-deployment safety review, we leverage targeted evaluations, red-teaming, and other checks to understand model behavior. We’ve now started using a method for simulating model deployments before they happen, which adds a complementary signal: a deployment-like preview of how a candidate model may behave before it reaches users. Deployment Simulation is a method for simulating a future deployment before it happens. We do so by replaying previous conversations in a privacy-preserving manner with a new candidate model. By doing so, we can study how the new model responds in realistic contexts before release, including whether new undesired behaviors emerge and how often they may appear. In our GPT-5.4 study, these forecasts were informative. For categories whose production rates changed by at least 1.5x, deployment simulation predicted the direction of change 92% of the time, compared with 54% for a baseline built from challenging prompts. Simulated deployments also looked much closer to real production traffic on evaluation-awareness measures: traditional evals often visibly have stage lights; production prefixes mostly do not. The hardest case is agentic tool use, where realistic behavior depends on external state: fil…
IV Bag Fill-Level and Leak Detection
Build an automated IV bag fill-level and leak detection system using RF-DETR and Gemini 2.5 Pro.
Reduce Class Flickering: Introducing Track Class Lock
Fix class flickering on video with Track Class Lock, a Roboflow Workflow block that freezes a tracked object's label once the detector agrees.
Automated Tire Sidewall OCR
Automate tire sidewall OCR to extract DOT codes, sizes, and brands. Learn to combine RF-DETR and multimodal LLMs into a Roboflow Vision Agent.
ILR School dean to help NYS shape, protect the AI workforce
Alexander Colvin, Ph.D. ’99, will serve on a blue-ribbon commission charged with developing recommendations on how New York state can protect workers’ economic security while harnessing the economic benefits of AI. The post ILR School dean to help NYS shape, protect the AI workforce appeared first on Cornell AI Initiative .
Why Tejal Patwardhan stopped underestimating the models - Episode 21
The old tests are getting too easy. Tejal Patwardhan leads OpenAI’s frontier evals team, which is finding new ways to measure and forecast progress as models become more capable. She and host Andrew Mayne discuss why evals matter for research, how benchmarks can break or get gamed, and what models need to be judged on next. Chapters 00:00:24 Growing up at OpenAI 00:03:10 Why reasoning changed everything 00:06:28 What made o1 surprising 00:11:20 Why old benchmarks stopped working 00:14:45 What makes a good benchmark 00:17:35 Why evals are getting harder 00:22:09 Measuring voice and vision models 00:24:48 Testing models on real science 00:33:23 How OpenAI tracks frontier progress 00:40:47 What AI means for work
Contact Lens Defect Inspection
Train a Roboflow object detection model, detect defects on each contact lens, and sort results into pass, review, and fail with a Custom Python Block.
HPE AI Factory With NVIDIA Expands for the Era of Agents
Enterprises are moving agentic AI from proof of concept to production — and the next generation of AI factories are built for the era of agents. At HPE Discover Las Vegas, running through Thursday, June 18, NVIDIA and HPE are expanding the HPE AI Factory with NVIDIA, including NVIDIA Vera CPU and NVIDIA Agent Toolkit […]
AI Now Co-Executive Director Sarah Myers West Testifies Before Senate Banking Committee
On Thursday, June 11, 2026, AI Now Co-Executive Director Dr. Sarah Myers West testified at a Hearing before the U.S. Senate Banking Committee on “AI and the American Dream: Promoting Innovation, Affordability, and American Dominance”. In her testimony, Dr. West highlighted the risks the AI industry poses to the US economy and broader public – […] The post AI Now Co-Executive Director Sarah Myers West Testifies Before Senate Banking Committee appeared first on AI Now Institute .
They Looked Inside Claude’s AI's Mind. It Got Weird
❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The paper is available here: https://www.anthropic.com/research/natural-language-autoencoders https://transformer-circuits.pub/2026/nla/index.html 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible: Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi My research: https://cg.tuwien.ac.at/~zsolnai/ Thumbnail design: https://felicia.hu
Expert Comment: Is banning under-16s from social media protection or exclusion?
The OII's Dr Vicki Nash warns that a social media ban risks leaving teenagers 'protected' but disempowered, cut off from news, learning and support at the very moment the UK is preparing to give 16-year-olds the vote.
MLCommons Releases MLPerf Training v6.0 Results
New benchmarks and increased diversity of submissions reflect important changes in AI ecosystem The post MLCommons Releases MLPerf Training v6.0 Results appeared first on MLCommons .
What is agent orchestration? Frameworks, runtimes, and observability explained
Agent orchestration is not one problem. It spans expression, runtime, and observability, and separating those layers clarifies how teams should build, run, and improve production agents. The post What is agent orchestration? Frameworks, runtimes, and observability explained appeared first on Arize AI .
Farming is risky business. This app helps Turkish growers succeed
The post Farming is risky business. This app helps Turkish growers succeed appeared first on Source .
Frontier post-training recipe review with Finbarr Timbers
"Interview" #18
Bye bye Fable
another benchmark for real swe work
Your AI bill is a tax on scale
Subscribe • Previous Issues The Hybrid AI Stack Is Coming for the Pricing Power of OpenAI and Anthropic OpenAI and Anthropic are going public while still capturing much of the money spent on foundation-model usage. But deployment patterns are starting to tell a more complicated story. Companies are building hybrid model portfolios, using proprietary models where convenience, Continue reading "Your AI bill is a tax on scale" The post Your AI bill is a tax on scale appeared first on Gradient Flow .
Morocco launches Rally AI Future Lab to accelerate innovation
1,000 researchers and developers gather in Merzouga for inaugural five-day sprint
Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM
Traditional machine learning pipelines for predictive tasks like text classification usually rely on extracting structured, numerical features from raw text — for instance, TF-IDF frequencies or token embeddings — to feed into classical models such as logistic regression, ensembles, or support vector machines.
Tokenomics: AI’s New Design Constraint
The Cost Reality of Running AI at Scale Budget shock is already happening. Multiple major players have pulled back on AI features or subscriptions due to unexpectedly high token costs. Amazon removed its token leaderboard and Microsoft cancelled Claude Code subscriptions. These are early signals that the deploy-everywhere approach is hitting hard financial limits, not Continue reading "Tokenomics: AI’s New Design Constraint" The post Tokenomics: AI’s New Design Constraint appeared first on Gradient Flow .
😺 Google sued the people spamming your phone
PLUS: Apple hid ChatGPT/Claude Siri switching in iOS 27 beta.
Emergency Pod: Claude Fable
Chris McGuire on a chaotic two weeks in AI policy
If context is king, architecture is the castle…
Recorded live at the AI Agent Conference, Ryan sits down with Apollo GraphQL CEO Matt DeBergalis to discuss how enterprises can leverage GraphQL and MCP as a structured semantic architecture to feed clean data to autonomous agents, safeguard internal microservices against unprecedented "east-west" data exfiltration risks, and rein in skyrocketing token spend by explicitly querying only the exact context required.…
Autoregressive Models: Predicting the Future Using the Past
Autoregressive models are one of the most important ideas in time series forecasting and sequence modeling. The name may sound technical at first, but the concept is surprisingly intuitive. An autoregressive model predicts the next value by looking at previous values. That is the core idea. For example, tomorrow’s temperature may depend on the temperatures […] The post Autoregressive Models: Predicting the Future Using the Past appeared first on Analytics Vidhya .
MERICS Data Insight: EU-China trade
MERICS Data Insight: EU-China trade H.Seidl Tue, 06/16/2026 - 09:27 Comment Jun 16, 2026 1 min read MERICS Data Insight: EU-China trade In this edition of MERICS Data Insights, MERICS Visiting Fellow Esther Goreichy looks at the European Union's trade deficit with China. She finds that the deficit is widening despite the bloc's trade defense measures. Author(s) Esther Goreichy Visiting Fellow Author(s) Esther Goreichy Visiting Fellow Related content about EU-China “We need to wean ourselves off the ‘China drug’” (Zeit) External publication Jul 22, 2026 There is still a best-case scenario for VW—but even that “involves immense costs” (Welt) External publication Jul 20, 2026 China in 26: Europe-China relations + Russian-Chinese military cooperation + new economic data Podcast Jul 17, 2026 Related content about Trade and Investment If Europe wants to use its trade leverage with China, now is the time Comment Jul 24, 2026 China's economy in Q2: Domestic woes dampen growth as imbalances intensify Tracker Jul 23, 2026 MERICS Data Insight: Chinese export prices Comment Jul 06, 2026
Nigeria unveils plans for world’s first 'National AI Trust' in London
Minister Tijani sets out governance vision to guide Nigeria’s AI transformation
Uncertainty Estimation vs Oversampling
I am currently doing some work with a fraud detection dataset as part of a research to leverage uncertainty to improve neural networks ensemble of experts' results. Firstly I had to take the dataset's training data and split it into 6 different domains (5 in-distribution and 1 ood). The goal is to train 5 different experts on different fraud patterns and then compute uncertainty of each expert (using Monte-Carlo dropout) and the uncertainty of the ensemble of experts. When we talk about fraud detection we expect the data to be heavily imbalance favoring the non-fraud class. In this case is 1:90. Given this it makes sense to use oversampling when training the neural networks and so I did. I made a sampler which increases the ratio to 1:10, not by creating new transactions (Because due to the data split, some domains have very few transactions to get reliable simulated transactions with oversampling) but by having the fraudulent transactions seen x amount of times more to increase the ratio. Now, with this, I'm having a problem with the uncertainty signal. In a perfect scenario, fraudulent transactions would be more uncertain than legit ones, but due to the oversampling the experts seem to be more uncertain about the legit transactions than the fraudulent ones. I have tried different percentages of oversampling and only without oversampling the uncertainty signal is correct, but then the overall predictive results underperform. So what should I do? Should I try different metho…
AI in Science
This series draws from Ranjit Singh's ethnographic research on how AI systems are being taken up in scientific practice. The post AI in Science appeared first on Data & Society .
India’s Rs 1 lakh crore R&D fund: speed was the signal, but Rs 5 lakh crore is the point
The post India’s Rs 1 lakh crore R&D fund: speed was the signal, but Rs 5 lakh crore is the point appeared first on The Ken .
[AINews] Satya on Loopcraft: Building Frontier Ecosystems
a quiet day lets us report on Satya's hit essay
When safety speaks 20 languages: Day in life of foreign workers at building site with AI as assistant
Builders deploy AI translators to deliver real-time instructions as foreign nationals make up a growing share of the construction workforce.
Identifying the AI Development Workforce
Existing measures of the AI workforce often group together a broad set of AI-related roles with varying skill requirements. This report focuses on the AI development workforce and describes our methodology for identifying AI development jobs and estimating AI development employment. It presents initial findings on the size and share of AI development jobs in the U.S. labor market. Regional breakdowns of AI development jobs are available in PATHWISE , CSET's emerging technology talent tracking tool. The post Identifying the AI Development Workforce appeared first on Center for Security and Emerging Technology .
Synthetic document finetuning for instilling positive traits
This is the fifth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The fourth post can be found here . Thanks to Chloe Li for feedback on this post! TLDR: Via adapting the methods of Marks et al and Li et al , we train Gemini 3 Flash to have certain traits/values by midtraining it on documents about how Gemini has those properties, followed by finetuning it on synthetic chat data where it demonstrates those properties. The chat finetuning is effective for instilling the traits robustly, working OOD. We share some takeaways on how to improve midtraining & SFT effectiveness. Introduction This work closely follows Li et al (model spec midtraining, or MSM), who show that by training a model on synthetic documents before chat finetuning starts, they can shape how the model generalizes. Teaching the model reasons behind specific behaviours, rather than just the behaviours themselves, can also improve generalization. Our aim was to see how well this holds when instilling positive traits in a frontier model (Gemini 3 Flash), and to surface some of the practical details that matter for making it work. Our motivation is deep alignment : we want to train principles into the model which guide behaviour even in highly OOD behaviours. Our MVP pipeline used a "traits document" (a short bullet-pointed list of positive traits we wanted the model to exhibit) as our universe context, with a checkpoin…
The hidden cost of code that nobody touches
Every engineering org has the files nobody wants to open. Here's what that actually costs.
Sourcegraph MCP server and a cheaper model beat a Mythos-class model alone
On nine CodeScaleBench tasks designed to evaluate agent effectiveness in large codebases, Claude Sonnet 4.6 with the Sourcegraph MCP server outscored Fable 5, winning six of nine at roughly half the cost for each point of quality.
Memory at the Edge: On-Device Vector Search with Qdrant Edge
On its first day in an unfamiliar house, a home robot has to build memory as it goes: which rooms it has covered, where it last saw the car keys, whether the kitchen looks different now than it did this morning. And it has to answer those questions itself, where it stands, because the network isn’t always there, and it’s too slow to wait on even when it is. That memory has a concrete shape. As the robot moves, it turns what its camera sees into vectors and writes them to a store it carries onboard. To make a decision, it queries that store for the nearest matches to what it is looking at, filtered by where or when it saw them. Capture, embed, search, decide, and the loop runs entirely on the robot, in milliseconds, with no trip to a server. The engine underneath it is Qdrant Edge : the same Qdrant vector search engine, running in-process as an embedded library instead of behind an API.
The hidden cost of code that nobody touches
Every engineering org has the files nobody wants to open. Here's what that actually costs.
Sourcegraph MCP server and a cheaper model beat a Mythos-class model alone
On nine CodeScaleBench tasks designed to evaluate agent effectiveness in large codebases, Claude Sonnet 4.6 with the Sourcegraph MCP server outscored Fable 5, winning six of nine at roughly half the cost for each point of quality.
Algorithm–hardware co-design of neuromorphic networks with dual memory pathways
Nature Machine Intelligence, Published online: 16 June 2026; doi:10.1038/s42256-026-01255-3 Pengfei Sun et al. develop a spiking neural network with a dual memory pathway, co-designed with a custom neuromorphic chip. The approach delivers over 4× throughput and 5x energy efficiency gains while using 40–60% fewer parameters than state-of-the-art implementations.