AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
35603News Items
8Top Picks
209Blogs
successLast Run

Latest AI/ML News

35603 matching items

Korea AI Times 2026-08-13 09:31 UTC Score 40.0 USR-0048-20260813-global-ai-ne-a6d34238 Full article

챗GPT로 반려견 암 치료법 만든 호주 사업가, ‘맞춤형 mRNA’ 스타트업 창립

생성 AI와 유전체 분석을 활용해 반려견의 암 치료법을 설계한 일로 세계적인 화제가 된 호주의 사업가가 결국 개인 맞춤형 mRNA 암 백신 개발 스타트업을 설립했다. 폴 코닝엄은 12일 X를 통해 스타트업 갬지(Gamgee)를 설립하고 AI와 유전학을 결합해 반려견을 위한 개인 맞춤형 mRNA 암 백신을 개발하겠다고 밝혔다.코닝엄은 앞서 자신의 반려견 ‘로지’의 종양 유전자 정보를 분석하고 챗GPT와 그록 등 AI 도구를 활용해 맞춤형 암 백신 후보를 설계한 사례로 유명해졌다.My dog Rosie was given months t

Medianama AI 2026-08-13 09:31 UTC Score 45.0 USR-0211-20260813-regional-new-aa7897a8 Full article

How Meta’s copyright management tools are being exploited; Delhi HC to examine

The Delhi High Court questioned Meta over alleged misuse of its Rights Manager tool, after creators said scammers used fraudulent copyright claims to remove original content and target accounts. The post How Meta’s copyright management tools are being exploited; Delhi HC to examine appeared first on MEDIANAMA .

Arize AI Blog 2026-08-13 09:30 UTC Score 56.0 USR-0079-20260813-ai-specialis-4d5da3b8 Full article

Evaluation-driven development: How to move AI agents from pilot to production

Learn how evaluation-driven development, agent harnesses, AI observability, guardrails, and cost-per-outcome metrics move AI agents from pilot to production. The post Evaluation-driven development: How to move AI agents from pilot to production appeared first on Arize AI .

Synced 2026-08-13 09:24 UTC Score 50.0 AI-041-20260813-ai-specialis-a4ded4f5 Full article

Comment on Web Data to Real-World Action: Enabling Robots to Master Unseen Tasks by exceltomd

The idea of leveraging zero-shot video prediction from web data to guide robot manipulation is compelling—it sidesteps the costly step of collecting task-specific robot data. I especially appreciate that Gen2Act frames unseen-task generalization as a video generation problem, which makes it easier to scale across diverse environments. It will be interesting to see how the framework performs when transferred from simulation to cluttered real-world settings.

Historic Valbonne estate loved by Duran Duran is now up for sale
Euronews AI 2026-08-13 09:19 UTC Score 40.0 AI-164-20260813-regional-ai--6d5c2dd9 Full article

Historic Valbonne estate loved by Duran Duran is now up for sale

Located near the village of Valbonne, Domaine de La Sylviane reveals a quieter side of the French Riviera where centuries-old town squares, olive groves and Provençal heritage remain virtually untouched.

Custom MCP app tool namespace disappears after the first conversation turn in ChatGPT Developer Mode
OpenAI Community 2026-08-13 09:15 UTC Score 48.0 AI-116-20260813-social-media-87649a92 Full article

Custom MCP app tool namespace disappears after the first conversation turn in ChatGPT Developer Mode

I am experiencing a consistently reproducible issue with a custom MCP app in ChatGPT Developer Mode. The issue is that the custom MCP app and its tools are available to the model during the first turn of a conversation, but on the immediately following user turn the entire MCP tool namespace is no longer available to the model. This happens even if no MCP tool is actually invoked during the first turn. Environment ChatGPT web Developer Mode enabled Custom remote MCP server MCP app uses read/search/fetch-type actions only OAuth authentication The MCP server and its tools are successfully detected during the initial tool scan The app can be selected and used successfully in a new conversation I have reproduced the issue across multiple ChatGPT models, so it does not appear to be specific to one particular model. Minimal reproduction steps Start a new ChatGPT conversation. Select/attach my custom MCP app from the tools/apps menu. In the first message, ask ChatGPT only to load or inspect the available MCP tools. ChatGPT successfully sees the custom MCP namespace and all expected MCP tools. Do not invoke any MCP tool. Send a second user message asking ChatGPT to use one of those MCP tools. On the second turn, the custom MCP namespace is no longer present in the model runtime, so ChatGPT is physically unable to invoke the tool. Actual behavior On turn 1, the custom MCP tools are available, for example: legislation lookup legislation search other MCP read/search tools On turn 2, th…

Korea AI Times 2026-08-13 09:15 UTC Score 43.0 USR-0048-20260813-global-ai-ne-ba14efe6 Full article

앤트로픽, 크롬 사이드 패널에 '클로드 코워크' 적용..."어디서든 끊김 없이 작업"

앤트로픽이 AI 에이전트의 브라우저 활용성을 높이고 기기 간 작업 연속성을 강화하는 업데이트를 단행했다. 브라우저에서 웹페이지를 직접 탐색하고 각종 작업을 수행하는 AI 에이전트의 활동 범위를 넓히는 동시에, 플랫폼별로 분리됐던 작업 환경을 하나로 통합한 것이 핵심이다. 앤트로픽은 12일(현지시간) 웹브라우저에서 AI 에이전트가 직접 작업을 수행할 수 있도록 지원하는 ‘클로드 인 크롬(Claude in Chrome)’의 사이드 패널을 ‘클로드 코워크’ 세션으로 통합했다.이에 따라 브라우저에서 시작한 작업의 대화 기록과 스킬, 커넥터

Korea AI Times 2026-08-13 09:11 UTC Score 40.0 USR-0048-20260813-global-ai-ne-2de393cd Full article

머스크 "스페이스X 전체 데이터로 그록 학습...직원들이 AI의 '부모' 될 것"

일론 머스크 CEO가 스페이스X의 방대한 내부 데이터를 AI 모델 ‘그록(Grok)’ 학습에 활용하겠다는 계획을 밝혔다. 비즈니스 인사이더에 따르면, 머스크 CEO는 11일(현지시간) 열린 전체 회의에서 “스페이스X의 모든 정보를 합친 데이터로 그록을 학습시킬 것”이라며 “어떤 의미에서는 그록이 여러분을 학습하게 될 것”이라고 말했다.그는 스페이스X 직원들을 “지구상에서 가장 뛰어난 인간들의 집합”이라고 표현하면서, 직원들의 생각과 아이디어, 가치관이 앞으로 AI 모델에 반영될 수 있다고 설명했다.특히 직원들이 AI의 ‘부모’ 역할

Korea AI Times 2026-08-13 09:02 UTC Score 40.0 USR-0048-20260813-global-ai-ne-36c81b01 Full article

유럽 서점가 덮친 ‘5000권 주문 미스터리’...“AI 학습용 책 수집 의혹 확산”

유럽 전역의 독립 서점들이 최근 정체불명의 대량 도서 주문에 촉각을 곤두세우고 있다. 수천권 규모의 주문이 몰려들고 있지만, 주문된 책들이 지나치게 무작위적이고 오래된 실용서나 절판에 가까운 비주류 서적까지 포함하고 있어 AI 모델 학습용이라는 의혹이 커지고 있다.아이리시타임스에 따르면, 아일랜드 골웨이의 유명 독립 서점 케니스 북숍(Kennys Bookshop)은 지난 5월 온라인으로 5000권에 달하는 대규모 주문을 받았다. 처음에는 대형 기관 고객의 주문으로 보였지만, 주문 목록을 살펴본 뒤 의문이 생긴 것으로 알려졌다.주문에

‘I feel like I’m at war’: are we losing the battle against machine-made music?
The Guardian AI 2026-08-13 09:00 UTC Score 55.0 AI-021-20260813-global-ai-ne-c7c1f34b Full article

‘I feel like I’m at war’: are we losing the battle against machine-made music?

Despite outcry from musicians, AI slop is creeping into the charts as record labels scramble to adapt to a new normal where hits can be made at the click of a button This year, the battle for song of the summer has been eclipsed by a much more complicated – some would even say disturbing – debate. That’s because we find ourselves asking not “What’s the song of the summer?” but rather “Is the song of the summer even real ?” Among the top contenders for the title is Fenix Flexin’s Rubberz , a single released in June that has ascended to No 58 on the Billboard Hot 100 and racked up more than 35m Spotify streams. It’s not the sort of fare Fenix usually cooks up. The artist is known for his trap music as part of the rap duo Shoreline Mafia, but Rubberz is a mildly noirish, 80s-inspired synth-pop track featuring a voice nothing like his. The song has drawn comparisons to Morrissey, but it more closely resembles Men at Work’s Down Under, or a Weird Al Yankovic parody of Men at Work. It’s pretty awful. But more importantly, it has an uncanny quality to it. It sounds off. Continue reading...

ZDNET AI 2026-08-13 09:00 UTC Score 42.0 AI-022-20260813-global-ai-ne-afac9cf9 Full article

The best password managers of 2026: Expert tested

If you have trouble remembering the passwords for all of your online accounts, we've handpicked the best password manager apps to help you stay protected.

Arize AI Blog 2026-08-13 09:00 UTC Score 50.0 USR-0079-20260813-ai-specialis-c1f6acd1 Full article

Crew Studio launches with native Arize AX tracing and evaluation

Through a native Arize AX integration, teams can send traces from Crew Studio to Arize from the first run without custom instrumentation—then inspect behavior, evaluate quality, and test fixes before redeploying. The post Crew Studio launches with native Arize AX tracing and evaluation appeared first on Arize AI .

MIT Technology Review AI 2026-08-13 09:00 UTC Score 48.0 AI-013-20260813-global-ai-ne-599ea52f Full article

How kids feel about AI, in their own words

When we set out to talk to kids about artificial intelligence, we thought we knew what we’d hear. We expected some to tell us they were using it to cheat a little, the way Millennials and Gen Xers opened up CliffsNotes or programmed formulas into their TI-82s, and others to share inspiring ways they were…

InfoWorld AI 2026-08-13 09:00 UTC Score 54.0 USR-0126-20260813-global-ai-ne-3167c3c5 Full article

MCP didn’t remove sessions. It handed them to the model

A few years ago, I helped move an application off sticky sessions so we could scale behind an ordinary round-robin load-balancer. On paper it was an infrastructure change. Every service was supposed to be stateless, so removing session affinity should have changed nothing. Nothing crashed. CPU looked normal. Every health check stayed green. But users started reporting workflows that would randomly jump backwards. One request would pick up exactly where the previous one left off. The next would behave as if it belonged to an entirely different conversation. The bug wasn’t in our business logic. It was in an assumption we’d inherited about where correlation lived. We’d quietly relied on the platform to remember which requests belonged together. Once that responsibility disappeared, our application didn’t fail loudly. It just became inconsistent in ways that were very hard to reproduce. We fixed it the obvious way. We stopped relying on the platform and started passing an explicit identifier on every request. It worked, and it kept working, because the thing carrying that identifier did exactly what we told it to do. That’s the incident I kept thinking about while I read the new Model Context Protocol specification. When MCP shipped its stateless core on July 28 , I did what everyone else did. I opened the changelog, found the SDK migration notes and scoped the work. The code diff was smaller than I expected. I took that as good news for about four hours. I’d been reading it as…

InfoWorld AI 2026-08-13 09:00 UTC Score 59.0 USR-0126-20260813-global-ai-ne-c6ebc645 Full article

Why AI models need a real-time web intelligence layer

A growing sentiment in tech is that large AI model providers will eventually replace traditional software vendors. The argument makes some sense. If a model can write code, answer questions, and automate workflows, then over time it should be able to take on the functionality of thousands of existing applications. Why maintain a fragmented stack of software tools and solutions when a single intelligent system can do it all? But as enterprises attempt to move to production, it all starts to break down. Why? Because models are powerful, but they’re not self-sufficient systems. Organizations are finding that a major limitation of modern AI is the lack of infrastructure to reliably access the world’s information. There are limits to model-centric thinking Over recent years, large language models (LLMs) have made significant advances in reasoning, generation, and task execution. They can be extremely useful for summarizing documents and generating insights. At times, they can orchestrate complex workflows. When in controlled environments, they look capable of replacing entire categories of software. But these capabilities depend heavily on an oft-overlooked factor: access to external information. AI models operate on static training data and probabilistic reasoning. Without continuous access to up-to-date information, they can’t reliably answer questions about things like current events or what market conditions are like today, not yesterday. Retrieval-augmented approaches have a…

InfoWorld AI 2026-08-13 09:00 UTC Score 35.0 USR-0126-20260813-global-ai-ne-c065f6ac

Introducing automation in change management

It’s an idle Tuesday afternoon. You’re in hour three of change management calls to get a CostCenter tag update approved for a production change. Only finops reporting is affected. The change itself will take two minutes of Terraform execution. Last week, you had a DROP TABLE action that destroyed 10TB of obsolete tracking data. It was the same hour of doc writing and three hours of change management review calls for both changes. Most change management systems focus on generating coordination between teams, providing before-after evidence, and establishing a rollback path. These feed into the commonly articulated requirements like collision avoidance, regulatory compliance, scheduling, and auditability/traceability. These requirements cannot be hand-waved away. But perhaps the evidence to meet those requirements could be had for cheap(er)? Let’s see if we can do better. Ideally, we get there incrementally by automating verification of one piece of evidence at a time, then incorporating it into the existing change process. We want to avoid abrupt Big Bang rewrites of the organization’s change management process. Action reversibility Many change management teams want to know about your rollback process. Let’s establish some terminology on reversibility. Bi-directional changes: The same API call or tool invocation that made the change can also reverse it. Example: Changing tags on an AWS resource with CreateTags / DeleteTags . Going forward and backward is the same API call but…

CIO AI 2026-08-13 09:00 UTC Score 41.0 USR-0125-20260813-global-ai-ne-4b7da88d Full article

The vendor consolidation trap: When one throat to choke costs more than it saves

Vendor consolidation is sold as discipline. Fewer vendors, simpler architecture, better pricing through volume, one throat to choke when something breaks. Every one of those benefits is real on paper. The problem is that the biggest cost of consolidation rarely appears on the slide the procurement team uses to sell it internally, and it does not show up on the savings tracker until the first renewal cycle after the ink is dry. Within CIO Mastermind’s topic-specific cohorts, which I sometimes facilitate, I hear a version of the same story often enough to recognize the pattern early. A consolidation program gets pitched against a strong multi-year savings target. The first year or two look good. Then a renewal arrives, the remaining vendor prices to the switching cost the company just built for itself, and a meaningful share of the projected savings quietly erodes. The company still ends up with fewer vendors. It does not always end up with the leverage the original business case promised. What consolidation actually removes What consolidation actually removes is competitive pressure on the vendor you keep. That is the part most business cases leave out. Going from a dozen vendors in a category down to three or four feels like simplification, and it is. It is also a message to the vendors you kept about how expensive it would be for you to leave. The fewer live alternatives you maintain, the more accurately a vendor can price to your captivity rather than to the open market. A…

InfoWorld AI 2026-08-13 09:00 UTC Score 35.0 USR-0126-20260813-global-ai-ne-878d3f99 Full article

Relief from the bookkeeping of change management

It’s an idle Tuesday afternoon. You’re in hour three of change management calls to get a CostCenter tag update approved for a production change. Only finops reporting is affected. The change itself will take two minutes of Terraform execution. Last week, you had a DROP TABLE action that destroyed 10TB of obsolete tracking data. It was the same hour of doc writing and three hours of change management review calls for both changes. Most change management systems focus on generating coordination between teams, providing before-after evidence, and establishing a rollback path. These feed into the commonly articulated requirements like collision avoidance, regulatory compliance, scheduling, and auditability/traceability. These requirements cannot be hand-waved away. But perhaps the evidence to meet those requirements could be had for cheap(er)? Let’s see if we can do better. Ideally, we get there incrementally by automating verification of one piece of evidence at a time, then incorporating it into the existing change process. We want to avoid abrupt Big Bang rewrites of the organization’s change management process. Action reversibility Many change management teams want to know about your rollback process. Let’s establish some terminology on reversibility. Bi-directional changes: The same API call or tool invocation that made the change can also reverse it. Example: Changing tags on an AWS resource with CreateTags / DeleteTags . Going forward and backward is the same API call but…

Korea AI Times 2026-08-13 09:00 UTC Score 40.0 USR-0048-20260813-global-ai-ne-1b7fbd97 Full article

알트먼 "진정한 AI 에이전트의 등장, 한 세대 발전만 남았다'"

샘 알트만 오픈AI CEO가 오픈AI의 후속 모델이 단순히 질의응답을 처리하는 챗봇 수준을 넘어 사용자의 일상과 업무 전체를 실시간으로 조력하는 가상 동료 형태로 발전할 것으로 전망했다.알트먼 CEO는 11일(현지시간) 공개된 대담 행사 \'인턴아팔루자(Internapalooza)\'에서 다음 세대 모델의 핵심 특징으로 \'완벽한 맥락 파악 능력\'을 제시했다. 모델이 실제로 엄청나게 유용해지기까지는 \"불과 한 세대의 모델 발전만 남았다고 본다\"고 강조했다. 그가 설명한 맥락 파악이란 모델이 컴퓨터 화면을 항상 지켜보고, 참여하는 모든 회

Synced 2026-08-13 08:52 UTC Score 57.0 AI-041-20260813-ai-specialis-1acbfefb Full article

Comment on CMU’s Novel ‘ReStructured Pre-training’ NLP Approach Scores 40 Points Above Student Average on a Standard English Exam by mistfallhunterwiki

The idea of pretraining on restructured data is fascinating, especially how the QIN system reportedly scored 40 points above the student average on the Gaokao-English Exam while using only 1/16 of GPT-3's parameters. That efficiency gain makes the RST paradigm particularly compelling for broader NLP applications, not just exam benchmarks. I appreciate the authors sharing this direction and look forward to seeing how restructured pretraining performs across other language tasks.

Korea AI Times 2026-08-13 08:49 UTC Score 48.0 USR-0048-20260813-global-ai-ne-ec54bfbf Full article

실무 분석 에이전트 벤치마크 ‘AA-애널리스트에이전트’ 공개…'클로드 오퍼스 5' 1위

AI 모델이 실제 기업과 연구 현장에서 분석가 역할을 얼마나 안정적으로 수행할 수 있는지를 평가하는 새로운 벤치마크가 공개됐다. 아티피셜 애널리시스(AA)는 11일(현지시간) 실제 스프레드시트와 문서를 활용해 정량 분석 능력뿐 아니라 자료 해석과 전문적 판단, 반복 수행의 신뢰성까지 측정하는 ‘AA-애널리스트에이전트(AA-AnalystAgent)’를 발표했다.단순히 계산 문제의 정답을 맞히는 능력을 평가하는 기존 벤치마크와 달리 실제 분석 업무에 가까운 환경을 구현한 것이 특징이다. 분석가는 주어진 자료에서 필요한 정보를 찾아내는

How should we evaluate AI-assisted engineering decisions?
OpenAI Community 2026-08-13 08:47 UTC Score 40.0 AI-116-20260813-social-media-4f0687c7 Full article

How should we evaluate AI-assisted engineering decisions?

Q1. If you were redesigning this from scratch, what would you keep as the core product: domain-specific validation, automated acceptance criteria, or something else? Q2. Would you consider model comparison useful as a secondary signal, provided it is never treated as proof of correctness and the final decision is based on external validation?