AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
35334News Items
8Top Picks
205Blogs
successLast Run

Latest AI/ML News

35334 matching items

The Decoder 2026-08-14 16:06 UTC Score 79.0 AI-168-20260814-regional-ai--f07d42c2 Full article

Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach

AI agents using Claude Opus 4.8 and GPT-5.6 Sol were given six days, $3,000 in API credits, and GPU access to independently write AI research papers. The original authors of unpublished NeurIPS papers rated the results as "Reject." According to the study, conducted with Princeton and the UK AI Security Institute, frontier models can handle the full research engineering process but fall short on research judgment, creative problem-solving, and the ability to abandon failed approaches. The article Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach appeared first on The Decoder .

Launch of DeepSeek’s Harness marks its strategic pivot towards autonomous agentic AI
South China Morning Post AI 2026-08-14 11:03 UTC Score 78.0 AI-156-20260814-regional-ai--1da7ec27 Full article

Launch of DeepSeek’s Harness marks its strategic pivot towards autonomous agentic AI

Chinese artificial intelligence company DeepSeek is venturing into a new battleground beyond large language models, launching a developer preview of its long-anticipated Harness – a software framework that helps developers turn AI models into autonomous agents. The release on Thursday marks a strategic pivot as DeepSeek moves to build foundational digital scaffolding for AI agents – systems capable of using AI models to operate external software, run code, and complete complex jobs on their...

LessWrong AI 2026-08-12 17:08 UTC Score 95.0 USR-0152-20260812-community-fo-1d25878b

Introducing the Conceptual Reasoning Index

Associated announcement tweet. We are planning to release blog posts properly arguing the case for this kind of work in the future. tl;dr A core hope for managing AI risks is that AIs will help us understand the situation, plan for what lies ahead, and develop mitigations. Many tasks AIs would have to do for this purpose lack practical empirical feedback loops and require models to engage in the kinds of argumentation used in philosophy, AI futurism, and similar domains. To evaluate these capabilities, we develop a suite of three conceptual reasoning benchmarks. You can request access to our primary conceptual dataset, LMCA, through this form . We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai , where you can also find more details on our methodology. We will keep the website up to date as both new models and benchmarks are released. This work was done in collaboration with Anthropic. Background Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output. This suggests that a major determinant of whether we address AI risks in time is how early we can automate or uplift this work, relative to high-risk capabilities. One way to influence this might be to selectively improve models' relevant skills, such as reasoning about how to govern and align AI and how to avoid catastrophic cooperation failures involving AI. Current AI training…

LessWrong AI 2026-08-11 20:19 UTC Score 85.0 USR-0152-20260811-community-fo-ba45551a Full article

Claude Opus 5 Just Beat My Text-Based Adventure Game Benchmark

Cross-posted from my Substack . Basically, I created a text-based adventure game benchmark in April, and this morning my agent harness using Claude Opus 5 solved it for the first time. I thought the details might be interesting to this community. First, here are my previous articles on this subject: You’re Standing in a Clearing in a Forest (Apr 10, 2026) Text Adventure Benchmarks Revisited (Jun 14, 2026) Testing Fable 5 on Text Adventure Games (Jul 5, 2026) And a reminder of the domain. This is a small custom text-based adventure I created from scratch as a personal benchmark to run new models of LLMs against. It’s 10 rooms total and the goal is to collect 3 keys (brass, silver, and gold) and use them correctly to unlock the final door in Room 3 to exit the dungeon. The first and most challenging central puzzle is a rotating room (r5) operated by a crank mechanism in r4. The player must first find the handle to the crank in r2, carry it to r4, insert it, and turn it to align openings between r5 and its adjacent rooms. The most difficult aspect of this puzzle seemed to be non-local causal reasoning combined with allocentric coordinates. The crank is two rooms away from the rotating room that it actually turns. When the player turns the crank a grinding sound nearby can be heard through the walls. There is an informational diagram on the wall in the same room as the wall, and it updates with each turn. Earlier models struggled to understand that the diagram was information, a…

Synced 2026-08-11 16:06 UTC Score 80.0 AI-041-20260811-ai-specialis-b52eaee8 Full article

Comment on DeepMind Introduces Gato: A Generalist, Multi-Modal, Multi-Task, Multi-Embodiment Agent by monalisa1art

DeepMind's Gato is a fascinating step toward generalist agents, though calling it AGI feels like a stretch. The fact that one transformer model can handle text, vision, and robot control with shared weights is impressive, but crossing 50% expert threshold on 450 tasks still leaves plenty of room before true versatility. It does make me wonder how soon we'll see similar multi-modal approaches trickle into consumer tools—like a free nano banana image generator that adapts to different artistic styles without retraining. For now, Gato feels like a solid research milestone rather than a breakthrough. monalisa1art

Lightspeed India leads $9 Mn seed round in deep-tech startup Discovered Materials
Entrackr AI 2026-08-11 04:16 UTC Score 83.0 USR-0212-20260811-regional-new-cd7513c6 Full article

Lightspeed India leads $9 Mn seed round in deep-tech startup Discovered Materials

Deep-tech startup Discovered Materials has raised $9 million (Rs 85 crore) in a seed funding round led by Lightspeed India Partners, with participation from Y Combinator, Peak XV Partners and global angel investors including Paul Graham, Gokul Rajaram and Thariq Shihipar. The fresh funds will be used to expand the team and laboratory and scale its AI research agents, Discovered Materials said in a press release. Founded by Advaith Sridhar and Akash Ramdas, Discovered Materials is an AI-driven deep-tech startup focused on thermal dissipation challenges in AI chips, which can generate more than 140W/cm². The company is developing thermally conductive dielectric materials for 3D chip packaging. The startup operates cloud-based autonomous AI agents that run thousands of virtual material hypotheses daily using custom model harnesses incorporating frontier AI models. The AI-generated material candidates are then evaluated through physics simulations to assess their stability, dielectric constants and thermal properties. Discovered Materials has also launched the Material Discovery Bench to evaluate how frontier AI systems perform on real-world semiconductor material challenges. The startup plans to patent promising material candidates and license the resulting thermal management and semiconductor technologies to global chipmakers. According to the company, its AI systems have developed new thermal materials in three months with performance comparable to products that took years to…

MongoDB AI Blog 2026-01-15 20:15 UTC Score 82.0 USR-0070-20260115-ai-specialis-0045c0cd Full article

MongoDB.local San Francisco 2026: Ship Production AI, Faster

Today at MongoDB.local San Francisco, we announced capabilities that collapse the distance between AI prototype and production. Building AI applications means solving real problems: keeping conversational context clean and queryable, retrieving the right information from thousands of past interactions, connecting AI agents to your data without custom plumbing. These aren't theoretical challenges, they're the friction points that slow teams down every day. The AI era demands more from your data platform. MongoDB gives you everything you need to build quickly. Voyage AI: the best gets better Embedding models can make or break AI search experiences. We're proud that voyage-3-large has been the world's top-performing embedding model on Hugging Face's RTEB benchmark since its inception. But we didn’t rest on our laurels. There’s a new model at the top of the charts. Today, we're pleased to announce that the Voyage 4 model family is now generally available. The best just got better. The voyage-4 series models operate in a shared embedding space, allowing for cross-model compatibility and unprecedented flexibility to optimize for accuracy, speed, or cost. This release also includes voyage-4-nano, our first open-weight model available on HuggingFace, perfect for local development. Additionally, we're launching the new voyage-multimodal-3.5 model, which has been specifically trained to support video content alongside text and images. For developers building multimodal AI applications…

Secondhand booksellers in UK and Ireland suspect AI firms behind ‘strange’ bulk orders
The Guardian AI 2026-08-15 08:00 UTC Score 55.0 AI-021-20260815-global-ai-ne-b6d60e3f

Secondhand booksellers in UK and Ireland suspect AI firms behind ‘strange’ bulk orders

Development comes after Anthropic was found to have spent millions on books to scan for ‘data acquisition’ Secondhand booksellers in the UK and Ireland are reporting a flurry of bulk orders from mystery buyers, amid speculation AI companies are acquiring the tomes for their data. Bookshops contacted by the Guardian say they have been receiving orders from buyers in the US, Canada, continental Europe and the UK. The increase has been replicated around the world, with booksellers in the US, Australia and Europe also reporting unusual orders. Continue reading...

The Decoder 2026-08-15 08:00 UTC Score 40.0 AI-168-20260815-regional-ai--fd929727

Plaintiff hid invisible AI instructions in court filings to secretly influence automated review

A plaintiff in Connecticut embedded invisible prompt injections in court filings, formatted in 3-point white text on a white background, to manipulate a potential AI review system. Judge Spader compared the attempt to secretly tampering with a jury and revoked the plaintiff's electronic filing privileges. The court stressed that Connecticut doesn't use AI to review filings, but the intent alone was enough to warrant sanctions. The article Plaintiff hid invisible AI instructions in court filings to secretly influence automated review appeared first on The Decoder .

Karaoke in the car? China’s new wave of gadget-laden electric vehicles
South China Morning Post AI 2026-08-15 08:00 UTC Score 43.0 AI-156-20260815-regional-ai--a1348b44

Karaoke in the car? China’s new wave of gadget-laden electric vehicles

China’s electric vehicles (EVs) already lead the way in terms of battery and autonomous driving technologies. Now, the country’s auto brands are adding a new dimension to the notion of intelligent cars: turning them into full-scale entertainment centres. A new wave of Chinese-developed EVs is hitting the market equipped with everything from voice-activated artificial intelligence (AI) assistants, to headlights that double up as film projectors and karaoke systems. Some even feature immersive...

OpenAI Community 2026-08-15 07:57 UTC Score 37.0 AI-116-20260815-social-media-3645fbe2 Full article

ChatGPT desktop app automatically scrolls to the end of responses

Thank you apwack…. Actually, that’s exactly what my mine does… It can really be frustrating when you have the best part for blank screen in front of you and you have to scroll back… worse again if the chat has multiple threads and you have to avoid his overshooting… Thank you for sharing

OpenAI Community 2026-08-15 07:43 UTC Score 37.0 AI-116-20260815-social-media-84018101

Upgrading the options while using search functionality

The search functionality needs an upgrade please. My issue is that I have many chats on topics as i start voice chats while busy working, but do not have the time to find the last chat and keep going. also the chats that i do continue on get really large and start struggling with processing lots of information. one of the solutions is to keep these chats in projects, but yet again when working it is quick and easy to start a chat but i do not have the time to organise them all. on the search functionality, i can currently search for chats and the ONLY option i have is to click on the chat and go to the chat. I would like the option to select the chats i want and then the ability to do the following using a right click menu: Most important - move selected chats to a project, so that they are all in 1 place nice to have - chat gpt collates the information, double checks for duplication and rewrites it into a single chat so that all the information is in 1 chat. size may be an issue but that is user choice. once this is collated and organised it can be output to a document file with the collation as well as editing marks so that i can read through and confirm the duplications, like i would changes to a word document.

US Army Launches Detachment 201: Swears in OpenAI exec as Lt Colonel, next to Palantir
OpenAI Community 2026-08-15 07:43 UTC Score 50.0 AI-116-20260815-social-media-8b1bb427 Full article

US Army Launches Detachment 201: Swears in OpenAI exec as Lt Colonel, next to Palantir

Kevin Weil , no longer with OpenAI, but still self-identifying as LTC, has a new startup rumored. https://www.inc.com/georgia-fearn/kevin-weil-left-open-ai-months-ago-now-seeking-150-million-dollars-for-startup/91389531 Meanwhile, another year, another group of tech industry inductees into the Army. www.army.mil – 10 Jun 26 Army commissions second cohort of tech executives into Executive Innovation... JOINT BASE MYER-HENDERSON HALL, Va. — The U.S. Army commissioned its second cohort of senior technology leaders into the Executive Innovation Corps, kno... Paywalled: Twelve Executives Who Left OpenAI in 2026 - Business Insider

Euronews AI 2026-08-15 07:41 UTC Score 40.0 AI-164-20260815-regional-ai--0f11c644

Judo World Tour: A day of firsts at the Lima Grand Prix

Day one in Lima delivered five thrilling champions. Assunta Scutto, Michel Augusto, Rafaela Rodrigues, Radu Izvoreanu and Pihla Salonen claimed gold, with several celebrating their first Grand Prix titles as the new Olympic qualification cycle got underway.

The Decoder 2026-08-15 07:30 UTC Score 64.0 AI-168-20260815-regional-ai--8abd923c

World Labs turns one real-world robot task into thousands of simulated variations for training

World Labs, the startup founded by AI pioneer Fei-Fei Li, has unveiled a simulation engine that trains robot controllers entirely in virtual environments. From a single real-world task, the system generates thousands of controlled variations. The trained models then ran for one hour each on five different robot platforms without human intervention. How well the results hold up in more complex everyday situations remains to be seen. The article World Labs turns one real-world robot task into thousands of simulated variations for training appeared first on The Decoder .

오픈AI, 맥용 챗GPT '컴퓨터 히스토리' 출시...앱·웹 활동 기억한다
Korea AI Times 2026-08-15 07:23 UTC Score 40.0 USR-0048-20260815-global-ai-ne-3b515f59 Full article

오픈AI, 맥용 챗GPT '컴퓨터 히스토리' 출시...앱·웹 활동 기억한다

오픈AI가 \'챗GPT\'를 사용자의 컴퓨터 작업 맥락까지 지속적으로 이해하는 개인형 AI 비서로 확장하고 있다. 챗GPT가 맥(Mac)에서의 앱과 웹사이트 활동을 기억해 과거 업무를 찾아주고 반복 작업을 자동화할 수 있는 기능을 선보였다. 오픈AI는 14일(현지시간) 사용자가 맥에서 어떤 앱과 웹사이트를 이용했는지를 기억해 이후 대화와 작업에 활용하는 새로운 기능 ‘컴퓨터 히스토리(Computer History)’를 챗GPT 데스크톱 앱에 도입했다. 컴퓨터 히스토리는 오픈AI가 앞서 선보인 연구 프리뷰 ‘크로니클(Chronicle)’

Korea AI Times 2026-08-15 07:12 UTC Score 40.0 USR-0048-20260815-global-ai-ne-daa6122e

알리바바, ‘큐원3.8-27B’ 가중치 공개...27B 규모로 오퍼스 4.6급 성능 도전

알리바바가 최신 AI 모델 ‘큐원3.8-27B’의 가중치를 공개하며 고성능 AI 모델의 오픈소스 경쟁에 다시 불을 붙였다. 270억개 매개변수를 사용하는 비교적 작은 규모의 모델이면서도 멀티모달 처리와 장문맥을 지원하고, 주요 벤치마크에서 앤트로픽의 ‘클로드 오퍼스 4.6’에 맞먹는 성능을 보인다는 점을 내세웠다.알리바바는 15일(현지시간) 큐원3.8-27B의 가중치를 허깅페이스에 공개했으며, 라이선스는 상업적 활용이 가능한 아파치 2.0을 적용했다.이에 따라 개발자들은 모델을 직접 내려받아 자신의 컴퓨터나 서버에서 실행하고, 다양

OpenAI Community 2026-08-15 07:08 UTC Score 46.0 AI-116-20260815-social-media-19204416 Full article

Did OpenAI increased the daily amount for incentivized tier?

@VeitB I checked again, and you were right. I was charged correctly, the bug is in the usage display. The incentivized tier combines two pools, all types of tokens, and with models that are priced very differently, making the charges difficult to reconcile. I overlooked this in my earlier review. So, the bug affects both the usage page and the usage API, not the billing. Thank you all for your help. Unclear usage bothers me more than the actual charges.

OpenAI Community 2026-08-15 07:00 UTC Score 34.0 AI-116-20260815-social-media-efd802ac

New version consuming lot more tokens

Look, I used maybe 3–4 commands and I’ve already hit the weekly limit. This is honestly getting ridiculous — I can’t get any work done like this. I can’t afford to keep paying more and more, and my quota seems to run out faster every time… FOR THE LOVE OF GOD, PLEASE GIVE US A HIGHER LIMIT!

Synced 2026-08-15 06:57 UTC Score 54.0 AI-041-20260815-ai-specialis-503c1bfe Full article

Comment on NVIDIA’s Global Context ViT Achieves SOTA Performance on CV Tasks Without Expensive Computation by Leo

This is a clear example of how better visual context can improve computer-vision results without relying only on larger models. Face shape and facial proportions are another practical layer for portrait-oriented applications. Face Shape Detector offers a quick browser-based way to identify face shape from a selfie before exploring hairstyle or image-generation workflows.

Korea AI Times 2026-08-15 06:42 UTC Score 40.0 USR-0048-20260815-global-ai-ne-1aec3f40

미스트랄 AI, 문서 레이아웃 보존 능력 높인 'OCR 4.1' 공개

미스트랄 AI가 복잡한 문서의 구조와 내용을 정확하게 인식할 수 있도록 개선한 전용 문서 인식 모델 ‘미스트랄 OCR 4.1’을 공개했다. 단순히 문자를 읽는 수준을 넘어 여러 단으로 구성된 문서와 기술 도면, 표, 인용 목록 등 복잡한 레이아웃을 원래 형태에 가깝게 유지하면서 디지털 데이터로 변환하는 데 초점을 맞췄다.미스트랄 AI는 13일(현지시간) 개발자 플랫폼을 통해 미스트랄 OCR 4.1을 선보였다.기존 \'OCR 4\'를 기반으로 문서 내 문자와 이미지 등 각 요소의 위치를 표시하는 바운딩 박스 정확도를 높인 것이 핵심이다.

Korea AI Times 2026-08-15 06:30 UTC Score 43.0 USR-0048-20260815-global-ai-ne-3f80bfa0

슬랙·메일·녹음까지 싹쓸이...AI 데이터 고갈에 '사내 기록' 몸값 폭등

슬랙 대화와 화상회의 녹화, 고객지원 기록, 이메일, 업무 문서, 코드 변경 기록 등 과거에는 기업 내부에 묻혀 있던 데이터가 AI 에이전트를 훈련하는 핵심 자산으로 부상하면서 스타트업의 오래된 업무 기록에 새로운 ‘데이터 가치’가 매겨지고 있다.13일(현지시간) 디 인포메이션에 따르면, AI 업계는 인터넷에 공개된 데이터를 넘어 실제 사람들이 일하고 소통하는 과정에서 생성된 ‘현실 세계의 업무 데이터’를 확보하기 위해 스타트업의 사내 기록까지 사들이고 있다. AI 에이전트 스타트업 웜리(Warmly)의 사례가 대표로 소개됐다. 웜

OpenAI Community 2026-08-15 06:27 UTC Score 38.0 AI-116-20260815-social-media-550e36f0

Feature Request: WebDAV support

Please consider adding WebDAV as a storage connector for ChatGPT, including support for self-hosted Nextcloud. This would allow users to maintain persistent project archives on storage they control while granting ChatGPT scoped read/write access to selected directories.

OpenAI Community 2026-08-15 06:11 UTC Score 42.0 AI-116-20260815-social-media-96694877 Full article

Problem with tool_choice causing caching issues

There’s no good reason to turn off the “implicit” caching by using the top-level explicit write parameter in a growing chat context, unless you expect to never reuse the API call “write”. That marks the final turn part as a breakpoint automatically. The unexpected symptom, that breaks some use-cases because of the limit of 4 effective write points and 50 total retained (per prompt_cache_key?), is that in “explicit” every input turn that you want to potentially match a previous cache write still has to be marked as a breakpoint, otherwise there is no hit. This is either undocumented or a flaw in the implementation. It may be that the default with no additional parameters is the equivalent of all turns being marked “implicit” with the effect of four potential writes at various lengths. Marking your own four (or more with the oldest ignored for writes) does not cost 4x. It sounds like what you need to do is: not use prompt_cache_options and/or additionally mark the tool return turn before the modified turn as a breakpoint, even when it is in the past, so you can “branch” a cache match from there. use unique prompt_cache_key per chat. As I described before, an oversight (or simply disregard of developers) is that only the Responses API endpoint supports explicit cache breakpoint on tool output return, not Chat Completions and its “tool” role. And functions are blocked anyway unless you use “none” for reasoning effort.

Korea AI Times 2026-08-15 06:08 UTC Score 40.0 USR-0048-20260815-global-ai-ne-b528e67d

구글, AI 생성물 '보이는 워터마크' 제거 허용...비가시성 ‘신스ID’는 유지

구글이 AI로 생성한 이미지와 동영상, 음악 등에 표시되는 눈에 보이는 워터마크를 사용자가 직접 제거할 수 있도록 정책을 변경했다. AI 생성 콘텐츠임을 식별할 수 있는 장치는 유지하면서도 창작자와 전문 사용자에게 높은 활용 자유를 제공하기 위한 조치다.구글은 14일(현지시간) AI 생성 콘텐츠에 표시되는 가시적 워터마크를 선택적으로 끌 수 있는 기능을 며칠 내 출시한다고 밝혔다.이 기능은 구글의 이미지 생성 모델 ‘나노 바나나’와 동영상 모델 ‘옴니(Omni)’, 음악 생성 ‘리리아(Lyria)’에 적용되며, 우선 \'제미나이\'와

The Decoder 2026-08-15 06:00 UTC Score 53.0 AI-168-20260815-regional-ai--4ec954fb

The "tragedy of the cognitive commons" explains how rational AI adoption could destroy entire professions' expertise

A new research paper frames AI adoption as a "tragedy of the cognitive commons." Every company that cuts entry-level jobs benefits individually, but the collective expertise of entire professions erodes. The consequences may not become visible until 2030 to 2045, when today's missing junior talent should have become tomorrow's experienced workforce. The article The "tragedy of the cognitive commons" explains how rational AI adoption could destroy entire professions' expertise appeared first on The Decoder .

Korea AI Times 2026-08-15 05:55 UTC Score 37.0 USR-0048-20260815-global-ai-ne-bf7cd6ea

트럼프, 중국산 드론에 최대 100% '관세 폭탄'…한국 등 동맹국엔 15%

미국이 현대전과 안보 분야에서 전략적 중요성이 커진 드론을 둘러싸고 중국과의 기술 공급망 분리에 속도를 내고 있다. 트럼프 행정부가 중국산 드론에 사실상 차별적인 고율 관세를 부과하면서 미국 내 생산 확대와 동맹국 중심의 공급망 재편이 본격화될 전망이다. 도널드 트럼프 미국 대통령은 13일(현지시간) 중국산 드론과 핵심 부품에 최대 100%의 고율 관세를 부과하는 조치를 발표했다.국가안보와 공급망 자립을 명분으로 미국의 중국산 드론 의존도를 낮추고 자국 드론 산업을 육성하기 위한 것으로, 사실상 세계 최대 드론 생산국인 중국을 겨냥

Korea AI Times 2026-08-15 05:43 UTC Score 37.0 USR-0048-20260815-global-ai-ne-4ea04658

프랑스 헌법위, '15세 미만 SNS 금지법' 위헌 결정... 마크롱 개혁 '제동'

프랑스가 청소년의 소셜미디어 중독과 유해 콘텐츠 노출을 막기 위해 추진한 강력한 연령 제한 정책이 헌법의 벽에 부딪혔다. 로이터 등에 따르면, 프랑스 헌법위원회는 14일(현지시간) 15세 미만 청소년의 소셜미디어 이용을 제한하는 법안에 대해 위헌 결정을 내렸다.헌법위원회는 미성년자를 온라인 유해 환경으로부터 보호하려는 취지 자체는 인정했지만, 모든 이용자에게 연령 확인을 요구하고 15세 미만의 SNS 접근을 일률적으로 차단하는 것은 표현·통신의 자유와 사생활을 지나치게 제한한다고 판단했다.이에 따라 강력한 청소년 온라인 보호 정책을

The Decoder 2026-08-15 05:30 UTC Score 57.0 AI-168-20260815-regional-ai--eb5bb424

New benchmark confirms AI models still perform poorly at visual perception

Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage. The article New benchmark confirms AI models still perform poorly at visual perception appeared first on The Decoder .