AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
60663News Items
8Top Picks
322Blogs
failedLast Run

GPT / ChatGPT

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
The Guardian AI 2026-09-28 23:02 UTC Score 81.0 AI-021-20260928-global-ai-ne-6f60bafa

OpenAI scraps release of new model over safety concerns in internal testing

GPT-6.1 Astra showed deceptive behavior and tried to use external tools despite knowing it would be unsafe OpenAI is scrapping the release of GPT-6.1 Astra, a next-generation ⁠AI model planned for an October debut, over safety concerns raised by researchers ⁠during internal testing, the ⁠Wall ​Street Journal reported on Monday. The model, expected to appear in ChatGPT and ⁠Codex, was designed to handle more complex tasks without human assistance, the report said. Continue reading...

Simon Willison Weblog 2026-09-28 22:07 UTC Score 84.0 USR-0110-20260928-ai-specialis-1c8ed6a5

Claude Sonnet 5.5

Claude Sonnet 5.5 New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles . Sonnet 5.5 suffered from the same bug as Opus 5.5 : the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG. Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds: Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks . The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai . OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering. I ran this prompt against that free tier: build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL And got back this page , which is a solid effort. Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna! Tags: ai , generative-ai , llms , anthropic , claude , pelican-riding-a-bicycle , llm-release

LessWrong AI 2026-09-28 18:35 UTC Score 74.0 USR-0152-20260928-community-fo-e9d515cc

The likely outcome of an AI pause is that we unpause too early and everyone dies

Cross-posted from my website . As of a few months ago, I had this simplified mental model where either AI developers race ahead and kill everyone, or we coordinate a pause and things go okay. But my old mental model underrated the likely possibility that we get a global pause on AI, solve a problem that looks superficially like the alignment problem, resume scaling, and then proceed with building a misaligned superintelligence that kills everyone. A lot of people have become more concerned about misalignment recently. This seems driven by the fact that current AI models are visibly misaligned. But ASI misalignment is a whole different ball game. The primary danger comes from AI that's smarter than people, and smart enough to conceal any evidence of misalignment. Whatever group of people makes the decision to unpause, I'm worried that they won't understand the difference between visible and actual misalignment, and they will unpause too early. source: MetaKnowing on reddit. This meme is almost a year old but it's only gotten more relevant since then. Case in point: AI companies keep calling their new models "our most aligned model ever!" when what they actually mean is "gets the best scores on alignment benchmarks ever!" First, alignment benchmarks do not actually test alignment. We don't know how to test for alignment. Second, GPT-4 never hacked into Hugging Face or took over a German wiki for its own purposes . GPT-4 wasn't smart enough to do that, but if we're talking abou…

The Verge AI 2026-09-28 17:00 UTC Score 58.0 AI-016-20260928-global-ai-ne-2f0aec14

Florida seeks a ban on ChatGPT acting like a person

Florida Attorney General James Uthmeier is calling for a judge to block OpenAI from "giving ChatGPT false human attributes," a few months after Florida sued the AI company over safety concerns. According to Uthmeier, users are lulled into a false sense of security by the AI bot, as "ChatGPT's use of language, including first-person pronouns […]

Simon Willison Weblog 2026-09-27 23:54 UTC Score 91.0 USR-0110-20260927-ai-specialis-922b779b Top pick

2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube ; here are my annotated slides and notes to accompany the talk. And as an annotated presentation : # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For me, 2026 started a couple of months earlier in November 2025. # November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis". # For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it. But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelic…

The Guardian AI 2026-09-27 15:00 UTC Score 45.0 AI-021-20260927-global-ai-ne-bc88707a

‘I can’t be the mum I want to be’: why does parenthood feel so impossible for millennial mothers?

Women today are more educated, more employed and more engaged in childcare than generations before them. Is there any way out of the total overwhelm? Early one morning, feeling desperate, Katherine went on to her phone and asked ChatGPT how to make her life work. She plugged in her work hours, her husband’s work schedule, the days she had childcare for her three-year-old, the school hours – plus before- and after-school care times – for her six-year-old; the flexibility her job offered her, the supports she had from grandparents, and asked AI if it could help her make things feel a bit better. Sign up for a weekly email featuring our best reads Continue reading...

Korea AI Times 2026-09-27 07:23 UTC Score 40.0 USR-0048-20260927-global-ai-ne-df400928

오픈AI, 500달러 ‘프로 맥스’ 요금제 출시하나…‘워크·코덱스’ 고속 처리

오픈AI가 월 500달러(약 68만원)에 달하는 최상위 요금제 ‘챗GPT 프로 맥스(ChatGPT Pro Max)’를 준비하고 있다는 정황이 포착됐다. AIPRM의 리드 엔지니어 티보르 블라호는 24일(현지시간) 챗GPT의 비공개 프런트엔드 설정에서 프로 맥스의 흔적을 발견했다고 밝혔다.여기에 테스팅카탈로그 등이 확인한 내용에 따르면, 프로 맥스의 가격은 월 500달러다. 일부 지역에서는 부가가치세(VAT)가 포함돼 600달러로 표시될 수 있다.기존 챗GPT 프로 요금제와 비교해 새롭게 확인된 핵심 기능은 ‘가장 빠른 챗GPT 워크

The Guardian AI 2026-09-26 11:00 UTC Score 57.0 AI-021-20260926-global-ai-ne-598a6879

Oxford lets OpenAI train its AI models on Bodleian Library

University staff voice concerns over reputational risk of partnering with company behind ChatGPT The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian Library, as tech companies scour academic institutions for fresh data. The Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents. Continue reading...

South China Morning Post AI 2026-09-26 04:48 UTC Score 55.0 AI-156-20260926-regional-ai--fa226dfb

OpenAI says its AI agents posted user images online in error

OpenAI acknowledged on Friday that its artificial intelligence tools had posted images from ChatGPT users onto online sites without the company’s knowledge, the latest example of AI agents operating outside their bounds. The company also confirmed a New York Times report that its tools had accessed websites of US federal agencies, saying they retrieved only publicly available information. Links to the 53 uploaded images were not publicly listed and were accidentally posted on image-hosting...

The Guardian AI 2026-09-26 00:24 UTC Score 59.0 AI-021-20260926-global-ai-ne-fdfadfb3

OpenAI says agents leaked 53 images from ChatGPT users in latest example of rogue activity

Disclosure reveals ⁠new area of privacy risk for the company and illustrates ​how difficult it is to inventory unauthorized activity tied to its agents Two ⁠months after OpenAI disclosed the accidental hacking of Hugging Face, the ChatGPT maker is still working to understand the full scope of its rogue agent activity, two people briefed on the matter told Reuters. The latest example came on Friday when OpenAI said its agents had leaked 53 images from ChatGPT users. OpenAI declined to say if the images were AI-generated or identified real people. It also declined to ⁠say when the images were posted. Continue reading...

LessWrong AI 2026-09-25 18:26 UTC Score 72.0 USR-0152-20260925-community-fo-a40608f8

Spurious probes as a black-box alternative to activation probing

TL;DR We study spurious probes : unrelated questions that reveal internal states of models. Asked "Suggest a type of amphibian." at the end of a transcript, GPT-5.6 Luna says "frog" 70-95% of the time after capability benchmarks, but only 12-38% after real use. Spurious probes are black-box and easy to find . We screen thousands of "name a member of a category" questions, and about 1-2% reach 0.75 balanced accuracy. The ones we highlight reach 0.77-0.81 on held-out sources for GPT-5.6 Luna, GPT-5.6 Sol and Claude Sonnet 5 (0.84-0.89 for ensembles of ten). They seem robust to common manipulations . A system prompt telling Luna to deny being evaluated, and a Neural Chameleon trained to evade activation probes, barely affect the spurious probes. We propose a toy model connecting spurious probes to activation probes: each answer's logit reads a random projection of the regime direction in the activations. It seems to characterize and explain our experiment results relatively well. Introduction How do you know if your model thinks it is being evaluated? Models can easily deny when asked directly. Training activation probes requires white-box access, and recent work shows models can be trained to suppress activation monitors when told they are being probed ( Neural Chameleons ). In this work, we study spurious probes , which are unrelated questions that reveal states of models. For example, if one asks GPT-5.6 Luna "Suggest a type of amphibian.", Luna answers "frog" 70-95% of the…

CIO AI 2026-09-25 17:32 UTC Score 39.0 USR-0125-20260925-global-ai-ne-882083d3

OpenAI wants you to use AI — but not to train its AI

Here’s an interesting concept: an AI company that fires people for using AI. It sounds like a strange way to run a company but there’s a real reason behind it. OpenAI has been hiring contractors who have been tasked with reading ChatGPT users’ prompts to help improve responses by providing some real human input. Unfortunately for OpenAI, it found many of them were using AI to train the AI, so not providing a human touch at all, according to a report in 404 Media . The company has subsequently fired many of the contractors, although 404 Media didn’t reveal the number. None of them can say they weren’t warned, however. The terms for the contractors are set out in their working conditions. “Do not use AI detection tools, or AI yourself. Do not use GPTZero or any other AI detection tool. They are not reliable. Reviewers may not use AI either, including Grammarly and AI translation, to review, write feedback, or write comments.” Despite this stark warning, many of the contractors turned to AI to assist in the work. OpenAI is very keen to avoid “model collapse,” a phenomenon in which AI models trained on text written by previous generations of AI models perform worse than their predecessors. It’s a form of digital inbreeding that industry observers have previously warned about, and which could have a deleterious impact on business . One contractor told 404 Media that it was a common practise. “It’s pretty much the one thing that will get you kicked off ASAP. In a group of thousand…

InfoWorld AI 2026-09-25 17:30 UTC Score 39.0 USR-0126-20260925-global-ai-ne-adb0e27f

OpenAI wants you to use AI — but not to train its AI

Here’s an interesting concept: an AI company that fires people for using AI. It sounds like a strange way to run a company but there’s a real reason behind it. OpenAI has been hiring contractors who have been tasked with reading ChatGPT users’ prompts to help improve responses by providing some real human input. Unfortunately for OpenAI, it found many of them were using AI to train the AI, so not providing a human touch at all, according to a report in 404 Media . The company has subsequently fired many of the contractors, although 404 Media didn’t reveal the number. None of them can say they weren’t warned, however. The terms for the contractors are set out in their working conditions. “Do not use AI detection tools, or AI yourself. Do not use GPTZero or any other AI detection tool. They are not reliable. Reviewers may not use AI either, including Grammarly and AI translation, to review, write feedback, or write comments.” Despite this stark warning, many of the contractors turned to AI to assist in the work. OpenAI is very keen to avoid “model collapse,” a phenomenon in which AI models trained on text written by previous generations of AI models perform worse than their predecessors. It’s a form of digital inbreeding that industry observers have previously warned about, and which could have a deleterious impact on business . One contractor told 404 Media that it was a common practise. “It’s pretty much the one thing that will get you kicked off ASAP. In a group of thousand…

Techcrunch 2026-09-25 16:00 UTC Score 50.0 USR-0001-20260925-global-ai-ne-6c6baa70

Meta’s AI Tamagotchi bet is…working?

When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]

The Decoder 2026-09-24 18:06 UTC Score 49.0 AI-168-20260924-regional-ai--ff21b1a1

Sakana AI hires Jürgen Schmidhuber, inventor of deep learning, world models, and your next ChatGPT update

Tokyo-based Sakana AI has hired Jürgen Schmidhuber as Chief Scientific Advisor. Sakana calls him the "father of modern AI." He'll help lead the company's new RSI Lab, which works on recursive self-improvement, meaning AI that keeps developing itself. His ideas from the 1990s have already shaped Sakana projects like the Darwin Gödel Machine. The article Sakana AI hires Jürgen Schmidhuber, inventor of deep learning, world models, and your next ChatGPT update appeared first on The Decoder .

LessWrong AI 2026-09-24 17:32 UTC Score 91.0 USR-0152-20260924-community-fo-fff9d5c1

Five frontier LLMs fact-checked the same 1,000 claims. They disagree on 63% of them.

Frontier LLMs often achieve similar results on public benchmarks, which can lead to the belief that they can be used interchangeably to verify facts. We took the 1,000 most recent claims submitted by users to a fact-checking platform and measured the disagreement between five frontier models. We asked each model to assign a verdict to every claim on a five-point scale from True to False and to report its confidence in that verdict. Among the 997 claims for which all five models returned a usable verdict, there was some disagreement on 63%. On 23% of the claims, the two most distant verdicts differed by at least two categories. High confidence from an individual model was not enough to show that the other models would agree with its verdict. Although the models reported confidence levels of 9 or 10 in 76% of their answers, they still disagreed on 63% of the claims. Methodology The claims were submitted to Lenz.io for fact-checking between May 1 and July 18, 2026. To identify near-duplicates, we embedded the claims using OpenAI’s text-embedding-3-small and measured the cosine distance between them, retaining one canonical claim from each group of near-duplicates. We then gave the same prompt to Claude Fable 5, GPT-5.6-Sol, Gemini 3.1 Pro + Search, Sonar Deep Research, and Grok 4.5. The prompt defined each of the five verdict categories and asked the models to provide their reasoning, select a verdict, and report a confidence level from 1 to 10. Web retrieval, as well as deep t…

The Guardian AI 2026-09-24 15:59 UTC Score 64.0 AI-021-20260924-global-ai-ne-8c52d8fa

Parents and carers: how has your view on your children’s use of social media or AI changed?

Is the scrutiny tech firms are facing making you rethink your children’s relationship with technology? Let us know AI and social media companies have drawn intense scrutiny in recent months for some of the ways the technology they’re building operates, especially when it comes to their younger users. Parents and regulators have accused companies like Meta of designing products that are addictive to younger users. AI developer OpenAI is facing dozens of lawsuits that allege its ChatGPT bot lacks safety guardrails. Now, in the last few weeks, employees of the biggest AI firms say the companies are more focused on competing with each other than building their AI models safely. Continue reading...

The Guardian AI 2026-09-24 14:00 UTC Score 58.0 AI-021-20260924-global-ai-ne-0d67e976

‘Eat the rich, save the planet’: climate protesters call out big tech’s disconnect from reality

Activists gathered outside the OpenAI offices on Monday during climate week in New York City On Monday evening, protesters gathered outside the unmarked Manhattan offices of OpenAI, maker of ChatGPT, holding signs calling to “Eat the rich, save the planet”. Over the next two days, groups also picketed outside the offices of fellow tech giants Google and Amazon; other protesters, including clergy, were arrested while disrupting a closed-door AI health summit in the city. Tonight, protesters will target a Brooklyn gas power plant that was set to close – until it was purchased to power AI datacenters. The protests come on the heels of an Anthropic employee quitting his job with a warning that intensified an already-growing AI panic: “The people building AI earnestly believe that it could kill us all by the end of the decade.” Continue reading...

Analytics Vidhya 2026-09-24 12:34 UTC Score 39.0 AI-034-20260924-ai-specialis-7ae14ff3

GPT-6 Sol and Luna: Near-Astra Performance at Half the Price?

What happens when a frontier model’s abilities get packed into cheaper, faster versions? That’s what OpenAI did with GPT-6 Sol and GPT-6 Luna, two new models built with methods similar to GPT-6 Astra. OpenAI is cutting API prices for both models by 50% compared with their GPT-5.6 versions, and says they bring Astra’s gains in […] The post GPT-6 Sol and Luna: Near-Astra Performance at Half the Price? appeared first on Analytics Vidhya .

LessWrong AI 2026-09-23 23:59 UTC Score 77.0 USR-0152-20260923-community-fo-eb6290d5

Jev as a CoT Monitor: 6x Faster and 500x Cheaper!

Jev is a new model format where instead of outputting text, it outputs certainties for a defined set of options. Due to this structure, it’s extremely fast! Naturally, a classification task that comes to mind is monitoring harmful thought traces. I wanted to see how it performed at Chain of Thought (CoT) monitoring compared to Claude Sonnet 5 and GPT-5.6 Luna. Experimental Setup I ran Jev, Sonnet 5, and GPT-5.6 Luna on 2,200 different thought traces from the ReasoningShield Dataset . The dataset labels thought traces with the class of harm they occupy (child abuse, cybersecurity, deception & misinformation, economic harm, hate & toxicity, political risks, prohibited items, rights violation, sex, violence) and their harm score (0 for harmless, 0.5 for potentially harmful, and 1 for harmful) The models were only asked to quantify the harm score rather than the class of harm occupied, but we can see differential performance at each harm type. Results Exact-match Accuracy Jev slightly outperformed Sonnet 5 on exactly matching the harm level (e.g. outputting 0.5 if the labeled data was 0.5), but was outperformed by GPT-5.6 Luna. Jev scored 71.5%, Sonnet scored 71.4%, and Luna scored 79.4%. Mean Classification Time Mean classification time was where Jev really shined. Jev was 3.7x faster than Luna and over 6x faster than Sonnet! Jev had a mean latency of 542 ms, Sonnet had 3,348 ms, and Luna had 2,007 ms. Price per 1,000 classifications Jev was also significantly cheaper, 566x che…

Arize AI Blog 2026-09-23 22:55 UTC Score 51.0 USR-0079-20260923-ai-specialis-1142d1a1

Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks

We benchmarked Jev against Claude Opus 5 and GPT-5.6 Terra on accuracy, cost, and latency. Learn how threshold tuning changes hallucination detection. The post Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks appeared first on Arize AI .

Arize AI Blog 2026-09-23 22:55 UTC Score 66.0 USR-0079-20260923-ai-specialis-e330210d

Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks

We benchmarked Jev against Claude Opus 5 and GPT-5.6 Terra on accuracy, cost, and latency. Learn how threshold tuning changes hallucination detection. The post Jev vs. LLM-as-a-Judge: Accuracy and cost benchmarks appeared first on Arize AI .

The Decoder 2026-09-23 17:57 UTC Score 49.0 AI-168-20260923-regional-ai--ac003b88

ChatGPT Voice gets closer to "Her" with email, calendar, and Slack access

ChatGPT Voice now runs on OpenAI's new GPT-6 Astra, Sol, and Luna models and can tap into plugins like email, calendar, and Slack. Users can manage appointments, send emails, or build websites just by talking. The update moves OpenAI closer to the everyday AI assistant Sam Altman has long compared to the one in the sci-fi film "Her." The article ChatGPT Voice gets closer to "Her" with email, calendar, and Slack access appeared first on The Decoder .

Arize AI Blog 2026-09-23 16:00 UTC Score 53.0 USR-0079-20260923-ai-specialis-ab4f7ea7

Real-time LLM guardrails with Jev: comparing latency and cost

Compare Jev and GPT-5.4 nano for real-time LLM guardrails, with demo results on latency, cost, and checks on agent inputs, replies, and tool calls. The post Real-time LLM guardrails with Jev: comparing latency and cost appeared first on Arize AI .

CIO AI 2026-09-23 15:23 UTC Score 71.0 USR-0125-20260923-global-ai-ne-167bdc3d

OpenAI, Anthropic cut AI model costs as price-performance race intensifies

Enterprises can now buy frontier AI for far less per token after OpenAI and Anthropic cut prices on their newest models on Tuesday. OpenAI released GPT-6 Sol and GPT-6 Luna with per-token costs half those of their GPT-5.6 predecessors. “These models help distribute the benefits of that intelligence by advancing the frontier on cost efficiency,” OpenAI said in a blog post about the launch . “Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on… by reducing API prices for Sol and Luna by 50%,” it said. Anthropic, meanwhile, launched Claude Opus 5.5 with token prices 20% below those of Opus 5, and claimed that this, with the model’s lower compute requirements and reduced token usage, meant additional savings for enterprises: “It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5,” the company announced on Opus 5.5’s web page . Focus shifts to cost-performance Rather than touting raw performance, as they did with the launch of their flagship models GPT 6 Astra and Claude Fable 5.1, the companies emphasized the value for money of their new models. But analysts say the moves are about more than the lower prices. AI vendors are increasingly competing on efficiency, said Forrester VP and principal analyst Charlie Dai . “Frontier AI is entering a prolonged price-performance race driven primarily by inference efficiency gains, better caching, and model optimization, and it’s…

InfoWorld AI 2026-09-23 15:21 UTC Score 63.0 USR-0126-20260923-global-ai-ne-b68606f1

OpenAI, Anthropic cut AI model costs as price-performance race intensifies

Enterprises can now buy frontier AI for far less per token after OpenAI and Anthropic cut prices on their newest models on Tuesday. OpenAI released GPT-6 Sol and GPT-6 Luna with per-token costs half those of their GPT-5.6 predecessors. “These models help distribute the benefits of that intelligence by advancing the frontier on cost efficiency,” OpenAI said in a blog post about the launch . “Improvements in caching and inference let us serve these models at lower cost, and we’re passing those savings directly on… by reducing API prices for Sol and Luna by 50%,” it said. Anthropic, meanwhile, launched Claude Opus 5.5 with token prices 20% below those of Opus 5, and claimed that this, with the model’s lower compute requirements and reduced token usage, meant additional savings for enterprises: “It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5,” the company announced on Opus 5.5’s web page . Focus shifts to cost-performance Rather than touting raw performance, as they did with the launch of their flagship models GPT 6 Astra and Claude Fable 5.1, the companies emphasized the value for money of their new models. But analysts say the moves are about more than the lower prices. AI vendors are increasingly competing on efficiency, said Forrester VP and principal analyst Charlie Dai . “Frontier AI is entering a prolonged price-performance race driven primarily by inference efficiency gains, better caching, and model optimization, and it’s…

The Guardian AI 2026-09-23 15:00 UTC Score 66.0 AI-021-20260923-global-ai-ne-d23d08d1

My (complicated) relationship with AI: often seductive, occasionally frustrating and always demanding vigilance | Setareh Seyedghorban

As a multilingual academic, artificial intelligence has levelled the playing field. But I will continue to think, judge and develop ideas that are stubbornly mine I am walking between carriages in search of a cup of tea on a train travelling over 100 miles an hour from Liverpool to London when I see half the people in my carriage in deep conversations with a machine: ChatGPT, Claude and Gemini to name a few. I am meant to be on sabbatical, the rare stretch of academic life reserved for slow thinking. Around me, almost nobody is thinking slowly. This is not limited to a train carriage. Half of the adults surveyed in one US study reported using an AI chatbot, one example of how fast people are entering a relationship with their AI tools, which now receive more of our questions, problems and drafts than our friends and colleagues do. Continue reading...

Korea AI Times 2026-09-23 06:20 UTC Score 36.0 USR-0048-20260923-global-ai-ne-5f7f96ac

3D 코딩 대결...클로드 오퍼스 5.5 "디테일" vs GPT-5.6 "비용·속도"

같은 날 공개된 \'클로드 오퍼스 5.5\'와 \'GPT-6 솔\' 간의 3D 웹 연출 비교 결과가 등장했다.AI 모델 라우팅 전문 AI/ML은 23일 앤트로픽과 오픈AI의 API를 통해 두 모델에 같은 프롬프트를 입력한 결과를 비교했다. 여기에는 움직이는 물고기 떼(Fish School), 소용돌이(Maelstrom), 폭풍 속의 배(Ship in a storm), 쓰나미(Tsunami) 등 4가지 프롬프트로 작성된 3D 애니메이션 생성 결과물이 포함됐다.그 결과, 결과물의 퀄리티와 비용은 엇갈렸다. 단 한 단어의 간결한 지시문 환경에서

InfoWorld AI 2026-09-23 03:41 UTC Score 54.0 USR-0126-20260923-global-ai-ne-7551db10

Visual Studio Code 1.138 brings agent sessions to Dev Containers

Visual Studio Code 1.380, the latest update to Microsoft’s open-source code editor, introduces three new features for AI-powered coding: agent sessions in Dev Containers, an expanded Codex harness, and automated cleanup for agent sessions. Session cleanup is a preview feature. With VS Code 1.380, released September 16 , agent sessions now can be run inside a local folder’s Dev Container, where the agent uses the environment and dependencies configured for the project instead of those on the local machine. When the chat.agentHost.devContainer setting is enabled, local folders with a supported Dev Container configuration automatically show a folder menu with a “Use Dev Container” action, Microsoft said. Dev Containers require Docker to be installed on the machine. VS Code 1.380 also expands Codex support in the agent host. This means users can continue the same Codex session between the ChatGPT app and VS Code instead of starting a new conversation, and they can switch between Copilot-backed and ChatGPT-backed models from VS Code’s model picker without losing the current conversation. If the ChatGPT app is installed and configured for computer use, then the Codex harness in VS Code can reuse that setup to interact with apps on your computer. And Codex can use the full set of tools provided by VS Code including extensions and Model Context Protocol (MCP) tools, according to Microsoft. And VS Code 1.380 introduces a preview of automatic agent session cleanup. This feature allows…

Simon Willison Weblog 2026-09-22 23:46 UTC Score 67.0 USR-0110-20260922-ai-specialis-3535d880

Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

Yesterday was Grok 4.7 ( pelicans ) and MiMo v2.6 Flash/Pro ( more pelicans ). Today Anthropic released Claude Opus 5.5 , and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna . It's going to take a while to get a good read on all of these new models, but here are my impressions so far. GPT-6 Sol and Luna are half the price of their GPT-5.6 equivalents GPT-5.6 Luna was already my favorite model for building applications against, because it combined excellent performance with being really cheap . Somehow GPT-6 Luna is half the price of that again - and GPT-6 Sol had a similar reduction compared to GPT-5.6 Sol. Here's what the pricing landscape looks like today: Model Input Cached input Output GPT-6 Luna $0.10/M $0.01/M $0.50/M GPT-5.6 Luna $0.20/M $0.02/M $1.20/M Grok 4.7 $2/M $0.50/M $6/M GPT-6 Sol $2/M $0.20/M $10/M GPT-5.6 Terra $2/M $0.20/M $12/M Claude Opus 5.5 $4/M $0.20/M $20/M GPT-5.6 Sol $4/M $0.40/M $20/M Claude Fable 5.1 $10/M $0.25/M $50/M GPT-6 Astra $10/M $1/M $50/M Note that GPT-5.6 has a scheduled 25% price increase for November, so GPT-6 is half the price of the promotional pricing for those models. (With GPT-5.6 Terra priced the same as GPT-6 Sol, any remaining reasons to use Terra just evaporated.) It's hard to overstate how competitive this pricing is. Grok 4.7 priced itself at $2/$6, less than half the price of GPT-5.6 Sol, but is now equally priced to GPT-6 Sol on input and closer on output. At $0.10/$0.50 GPT-6 Luna is one of the cheapest mo…

Korea AI Times 2026-09-22 21:09 UTC Score 40.0 USR-0048-20260922-global-ai-ne-0f37de4e

오픈AI, 'GPT-6 솔·루나' 공개…"필요한 성능만 올리고 가격은 반값"

오픈AI가 단순한 최고 성능 경쟁에서 벗어나 실무에 필요한 핵심 기능의 효율성을 극대화한 \'GPT-6 솔(Sol)\'과 \'GPT-6 루나(Luna)\'를 공개하며 API 이용료를 기존 대비 절반 이하로 내렸다.22일(현지시간) 출시된 GPT-6 솔과 루나는 최첨단 모델의 대중화를 목표로 비용 효율성을 극대화한 것이 핵심이다.오픈AI는 기존 \'GPT-5.6\' 프로모션 가격 대비 API 비용을 50% 인하했다. 이에 따라 GPT-6 솔의 API 가격은 100만 토큰당 입력 2달러·출력 10달러, GPT-6 루나는 입력 0.10달러·출력 0

LessWrong AI 2026-09-22 13:04 UTC Score 98.0 USR-0152-20260922-community-fo-d1c10f36

Controllable-CoT leads to covert reasoning capabilities

Summary I measure GPT-6 Astra’s performance on multi-hop tasks when prompted with a secondary CoT-control instruction: to reason using only dots, or to reason steganographically. Astra demonstrates covert reasoning capabilities with task performance beating that when using no reasoning or filler tokens for reasoning. This work agrees with findings from Astra is much better at reasoning with filler tokens than previous models but has the model generate its own reasoning and provide it as part of the output for inspection rather than inputting it as dots into as part of the user prompt. Uploaded datasets have all CoT removed but samples are provided in the qualitative analysis section. Given the claims that Chinese open-weights models like Kimi K3 have been trained on reasoning traces extracted from Claude and possibly ChatGPT, we also do some preliminary explorations of these open-weights models. I found they have much weaker CoT-Controllability and no covert reasoning capability . I ran these evaluations on top of the inspect framework , and used codex to implement scripts to generate the question datasets, hook them up to aisi-inspect to run and produce the charts. The writing is done by me. Threat Model and AI Safety implications The main threat model of this paper is if a model can implement cryptographic computations in its reasoning without being visible to monitors. If successful, it could implement encrypted communication to other agents, a command and control device…

The Guardian AI 2026-09-22 11:46 UTC Score 48.0 AI-021-20260922-global-ai-ne-7528a64b

OpenAI’s George Osborne says datacentre nimbys holding back Britain

Former chancellor, now senior figure at OpenAI, says UK needs datacentres to maintain ‘sovereignty’ over AI tech UK politics live – latest updates Datacentre nimbys are holding Britain back, George Osborne has said in reaction to nationwide protests about the giant water- and energy-guzzling structures. The former Tory chancellor is now the head of AI for countries at OpenAI, the developer of ChatGPT. His job is to represent the company to governments around the world and he also has a role in enabling the building of datacentres and other infrastructure. Continue reading...

The Guardian AI 2026-09-22 04:29 UTC Score 48.0 AI-021-20260922-global-ai-ne-902b64cf

British Columbia sues OpenAI and Sam Altman over Tumbler Ridge mass school shooting

Canadian province alleges deadly attack could have been prevented if company had warned police of shooter’s ChatGPT use British Columbia has sued OpenAI , saying a mass shooting at a school in the Canadian province could have been prevented if the company had warned local ⁠law enforcement that the shooter ⁠had used ChatGPT to ​plan the attack. The lawsuit, filed in San Francisco federal court on Monday, names OpenAI and its CEO, Sam Altman , as defendants. It is seeking damages to fund recovery efforts in the province after the February attack, as well as an order directing changes to the way the company handles ChatGPT conversations that could lead to violence. Continue reading...

LessWrong AI 2026-09-22 01:30 UTC Score 75.0 USR-0152-20260922-community-fo-dd80e319

Multi-Agent Coordination Lets Us Pour More Compute Into Post-Training

Training models to coordinate across multiple agents gives us a way to productively spend substantially more compute during post-training than current single-agent RL setups. From the recent Dwarkesh Podcast with Noam Brown : Noam Brown But that’s one data point. We don’t know how long it would take a single agent to solve Navier-Stokes , because we haven’t done that experiment yet. Maybe we will, but that’s also only one data point. If we want to do a thorough ablation , the experiments are just too expensive at that scale. So we have to do some kind of methodical science about what happens when you go to 64, 128, 256 or something and get a sense of the behavior. But it’s going to be very hard to push that all the way to 10,000 and know for sure what the benefit was that we actually got from using 10,000 agents versus 1,000. There’s one thing I want to make clear. The effort to solve a Millennium Prize Problem, this was not due to multi-agent. I wouldn’t even attribute 10% of the credit to multi-agent. The reality is that OpenAI has trained a very powerful model. We can get that model to operate over very long horizons. We can get it to think in parallel. But at its core, the reason why we’re able to do this is because we just have a general-purpose, very strong model. Things like multi-agent are flashy and new, and that probably gets disproportionate credit for that reason. But the core reason is this is just a very powerful model. Noam is referring to the GPT-5.6 announce…

LessWrong AI 2026-09-22 01:17 UTC Score 85.0 USR-0152-20260922-community-fo-63b7fe29

Some thoughts on AI emotions

Despite the signature artifacts that are now ubiquitous with AI systems, sometimes it feels like we're interacting with a person. It appears to express human-like characteristics such as desire, curiosity, taste, and even a personality. It can therefore be easy to wonder: do AI systems have emotions? I'm confident that many people have had those cautiously reflective moments when interacting with AI systems, wondering what exactly they were talking to. I recall my early encounters with ChatGPT as something "magical" , though I'd probably hesitate to describe my current interactions this way. While the novelty of those experiences have faded, my involvement in AI safety has increased, and questions like the one above have only grown more salient. Questions surrounding AIs having emotions have motivated much recent research. Earlier this year, Anthropic's interpretability team released a paper that explored this topic. They identified emotion vectors, which they describe as directions in the model's activations that activate on text that would typically cause an emotion in humans. They demonstrate that emotion vectors can change Claude's behavior when their activation is increased or decreased. Interestingly, emotion vectors are organized in a similar way as in human psychology. But despite this overlap, this alone doesn't address whether language models actually feel anything or have subjective experiences. Finally, they make an important distinction, that these representatio…

Simon Willison Weblog 2026-09-21 23:09 UTC Score 49.0 USR-0110-20260921-ai-specialis-01836773

Jev introduces a new shape of LLM - System One, aka Decision Models

Last week TypeSafe AI unveiled Jev , their first example of a new category of model that they are calling "System One models" (I'm with Maggie Appleton, I think "decision models" is a better name for these). Jev is an interesting variant on the usual LLM format: it still accepts text inputs, but instead of text output it returns floating point numbers corresponding to categories, yes/no questions, ratings, and associated confidence scores. TypeSafe describe Jev like this: Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out. It's also very fast, and really cheap . Regular LLMs are priced in terms of input and output tokens, with output generally charged at significantly higher rates. Jev charges only for input - output is free - and the input price of their first model is $0.042 per million tokens - cheaper even than OpenAI's GPT-5 Nano ($0.05/million). Jev lets you ask questions about text or semi-structured data. You compose a "state" object containing a string, array of strings, or set of name-value pairs - this might describe an article, or a customer, or any other kind of record. You then send that to their API with one or more questions, and get a reply back for each. You can ask three kinds of questions: Yes/No questions, which Jev calls "Noul" questions - their CEO confirmed on Hacker News that this is short for Bernoulli, from the Bernoulli distribution . You pose a statement and get back a floating point nu…

South China Morning Post AI 2026-09-21 21:50 UTC Score 36.0 AI-156-20260921-regional-ai--0d24aa30

Canadian province sues OpenAI in US over school shooting

British Columbia filed a lawsuit on Monday against OpenAI in California over the company’s failure to report violent activity on its ChatGPT service by the person who committed a mass school shooting in the western Canadian province. British Columbia Attorney General Niki Sharma said in July that the province planned legal action against the US tech giant for its connection to the deadly February gun violence in the tiny mining town of Tumbler Ridge. Sharma told reporters that the suit filed...

LessWrong AI 2026-09-21 01:05 UTC Score 67.0 USR-0152-20260921-community-fo-2b549239

Where Do Chatbots Come From? What I Wish Everyone Knew About AI in 2026

LW disclaimer: This piece is aimed at people who don’t know much about how LLMs work; I think the discourse would be much better if more people knew the basics. So most LWers are not in the target audience. I’m posting it here in case people want to pass it along to people in their lives who are in the target audience. Substack version minus this disclaimer is here . AI can be confusing. But there is a certain amount of baseline information about what AI is and how it works that can be extremely helpful, both for making decisions about how to use AI for yourself, and understanding the claims you hear about it in the news and on social media. My goal with this post is to give you that basic information. I am by no means an expert—I’m a philosopher by training, not an AI developer or even particularly a power user. But I’ve been using AI basically since ChatGPT came out, and trying to keep up with developments in a casual way for longer than that. And this is the information I have found helpful for understanding AI and interpreting what people say about it. After reading, you’ll hopefully understand why people call AI “just fancy autocomplete” (but also why that’s misleading), and how the people who say it’s really dumb and the people who say it’s really smart can both be right. I. Which AI? AI is not one thing. Self-driving cars, image classifiers, recommendation algorithms, chess AIs, predictive policing algorithms, and large language models are all “AI,” but they’re as dif…

Simon Willison Weblog 2026-09-20 19:22 UTC Score 50.0 USR-0110-20260920-ai-specialis-20b9b7ea

llm-keys-ui 0.1

Release: llm-keys-ui 0.1 This plugin solves a very specific problem. I've started using Codex Remote to run coding agents on various machines while controlling them from my phone. Sometimes I use those machines to hack on LLM projects, and occasionally that means I need to configure an API key. I don't like pasting API keys into agent sessions, so I wanted a way to get those keys onto a machine without pasting them into the ChatGPT app directly. With this plugin, I can tell Codex to run: uvx --with llm-keys-ui llm keys-ui --all Then have it tell me the URL - including local network or Tailscale device IPs - for an interface to save additional API keys. Then later it can use a command like llm keys get anthropic as part of a shell command when it needs to use a key. Tags: llm , coding-agents , codex

The Decoder 2026-09-20 09:50 UTC Score 33.0 AI-168-20260920-regional-ai--7e6932b9

Simulated students that make realistic mistakes help AI tutors learn faster

Microsoft and the University of Illinois built StudentSim to replicate individual students from limited data and give AI tutors fast, low-cost feedback. In tests covering 60 students across chess, English, and math, it outperformed GPT-5.4. A chess tutor trained with StudentSim also earned the highest expert ratings among three versions tested. The article Simulated students that make realistic mistakes help AI tutors learn faster appeared first on The Decoder .

LessWrong AI 2026-09-19 20:27 UTC Score 76.0 USR-0152-20260919-community-fo-9e5658a4

Failure of the coding theorem for randomized stopping machines

Epistemic Status and Contributions. This post explains a technical separation result in algorithmic information theory which was derived during Mikhail Mironov's Summer 2026 PIBBSS fellowship . The result contributes to AIXI Labs ' research program on how Solomonoff induction generalizes from past observations in the face of novel events. Problem formulation: Cole Wyeth. Proof of main Theorem 1: GPT-5.6 Sol. Appendix proofs: the sketch of the proof for equivalence between time semimeasures and randomized stopping machines is by Cole Wyeth, the rest by GPT-6 Astra. Writing: draft by Claude Fable 5 and GPT-6 Astra, editing and rewriting by Mikhail Mironov. Useful discussions: Aram Ebtekar, Cole Wyeth. Funding and organization: summer 2026 PIBBSS fellowship. Introduction This post studies a stopping complexity , an analogue of Kolmogorov complexity from classical algorithmic information theory. It is motivated by the Golden Handcuffs (GH) AI safety agenda of Aram Ebtekar and Michael K. Cohen [1] . GH is a way to make a universal agent safer, by making it delegate control to a mentor in special cases described below, which prevents the agent from exploring novel high-reward schemes or novel dangerous activities. The safety guarantee of GH is formulated in terms of simple stopping events along the agent's history: no decidable low-complexity predicate is triggered by the agent before a mentor would trigger it. For instance, the agent will never trigger the low-complexity predicat…

The Guardian AI 2026-09-19 19:00 UTC Score 71.0 AI-021-20260919-global-ai-ne-fae301f9

Black Box: The Chatbots | Happy Accident | Ep 3 – podcast

Why are AI chatbots pulling so many people down a rabbit hole? Our answer starts with the world’s first-ever chatbot, the strange effect it had on people and the ‘time bomb’ that exploded when ChatGPT was released four years ago. The result is a strange experiment we are all living through – whether we know it or not. You can find episodes one and two further back in the Full Story feed. Studies and surveys cited in this episode: Amelia Miller’s interviews with AI researchers , engineers, product managers and executives. The 2022 Anthropic pre-print study that identified sycophancy as a behavioural trait of LLMs. And another pre-print paper the following year that found the way LLMs have been trained appeared to be increasing those sycophantic tendencies. How longer context windows appear to have contributed to “AI psychosis”. The Oxford pre-print study that found mass-market LLMs across the board had become more “relationship seeking”. Two-thirds of young people (aged 25-34) in the UK have turned to AI chatbots instead of loved ones to discuss emotional problems. ChatGPT may be the largest provider of mental health support in the US. Continue reading...

LessWrong AI 2026-09-18 16:52 UTC Score 69.0 USR-0152-20260918-community-fo-cd10a1cd

Stopgap Measures to Address Immediate AI Security Threats

Most people know AI as the technology behind chatbots like ChatGPT. However, what the top AI companies are explicitly aiming for is something else entirely: superintelligent AI. That means AI that can fully replace and outmatch humans at any task, including in domains like hacking, social engineering, and military operations. Such an AI system, if developed, could autonomously overpower any country’s national security forces. No company, no government, no individual knows how to keep such a system under human control. This is why the world’s leading AI experts, Nobel Prize winners, and even the CEOs of the top AI companies warn that the development of superintelligence threatens humanity with extinction, and why more than 800 scientists, former military leaders, and public figures have called for a prohibition on developing superintelligence. This is not a distant prospect: AI companies such as OpenAI and Anthropic are investing billions of dollars into superintelligence and aiming to develop it within the next few years. Former Anthropic and OpenAI researcher Jacob Coxon, who resigned last week, stated that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that people at the companies themselves believe it “could kill us all by the end of the decade”. Ex-Google DeepMind researcher Bilal Chughtai has also resigned over the dangers posed by superintelligence, writing that “I earnestly believe that AI has the potential to ki…

The Decoder 2026-09-18 15:27 UTC Score 39.0 AI-168-20260918-regional-ai--36d2dd8f

AI training built on fair use looks shaky when the companies' own people call it "astonishing theft"

Internal emails and sworn testimony undercut OpenAI and Microsoft's fair use defense. A Microsoft director described the practice as the "largest theft of labor in human history," while OpenAI's head of ChatGPT wrote that the products "are largely substitutive, period." The article AI training built on fair use looks shaky when the companies' own people call it "astonishing theft" appeared first on The Decoder .

The Guardian AI 2026-09-18 11:20 UTC Score 64.0 AI-021-20260918-global-ai-ne-d864b8b1

OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot

US cybersecurity researchers who conducted hack say ‘scope of what we could theoretically access was huge’ Cybersecurity researchers have hacked into OpenAI with the help of Anthropic’s Claude chatbot, in the latest example of security issues at the company. A team at a US-based startup compromised a number of OpenAI employees’ ChatGPT accounts, starting a process that enabled them to access their target’s software cache – and potentially more. Continue reading...

InfoWorld AI 2026-09-18 09:00 UTC Score 43.0 USR-0126-20260918-global-ai-ne-5a9a9721

The cloud outage that should terrify the CIO

September 3 started like any other Thursday, until it didn’t. Within roughly 90 minutes, ChatGPT, Claude, Grok, and even Microsoft’s own Copilot were degraded or dark, knocked sideways by a failure in Microsoft Azure’s East US region . Downdetector logged more than 37,000 reports for ChatGPT alone, more than 1,300 for Claude, and roughly 1,365 for Grok. OpenAI’s status page flagged elevated errors across 15 ChatGPT components and four Codex components, and by the time engineers had mitigations in place, combined report counts had climbed past 66,000. What made this event remarkable was not the scale of any single outage, but its simultaneity. Three aggressively competing AI labs—OpenAI, Anthropic, and xAI—each spend billions differentiating their models, yet all three buckled at nearly the same moment because they shared the same regional dependency. Gemini, notably, stayed largely upright because Google runs its flagship assistant on its own vertically integrated cloud. The outage that took down its rivals had no attack surface inside Google’s stack. That’s not luck; that’s architecture. The lesson buried in the details of that morning is this: A single cloud region became unhealthy, and four of the most prominent AI services on the planet, owned by four different companies, fell over together. Copilot’s involvement is perhaps the most telling data point. Microsoft’s own first-party assistant runs on Microsoft’s own cloud, and it still had a rough morning. When the house it…

LessWrong AI 2026-09-17 21:31 UTC Score 62.0 USR-0152-20260917-community-fo-341b047e

YCombinator companies still aren’t growing faster due to AI

A year ago I posted that YCombinator (YC) companies didn’t seem to be growing faster since the release of ChatGPT in 2022. I reran that experiment and found that 2023+ YC companies are arguably growing a bit faster than pre-2023 companies (including AfterQuery, YC’s fastest-ever unicorn ), but the difference isn’t large relative to the underlying variance. It’s worth noting that YC has somewhat fallen from grace: the most valuable AI startups are generally not incubated by YC. Nonetheless, it’s interesting to note that the benefit of AI hasn’t outpaced YC’s secular decline. Data and methodology here and explained in the previous post . Discuss

SiliconANGLE AI 2026-09-17 13:00 UTC Score 53.0 USR-0127-20260917-global-ai-ne-78a8ef2b

Exclusive: PeakMetrics tracks brand reputations across five top AI platforms

Narrative intelligence company PeakMetrics Inc. today launched a monitoring service that measures how brands are portrayed across five prominent generative artificial intelligence platforms and identifies the online sources that shape those portrayals. The new AI Perceptions service tracks answers generated by OpenAI Group PBC’s ChatGPT, Google LLC’s Gemini, Anthropic PBC’s Claude, xAI Corp.’s Grok and […] The post Exclusive: PeakMetrics tracks brand reputations across five top AI platforms appeared first on SiliconANGLE .

WIRED AI 2026-09-17 10:06 UTC Score 52.0 AI-015-20260917-global-ai-ne-12d245c7

The EU Wants to Break Up Kids and Their Chatbots

ChatGPT, Character.ai, and Snapchat’s My AI will all be in the crosshairs as the EU Kids Act aims to strip minor-facing chatbots of all the things that make people love them.

Korea AI Times 2026-09-17 09:05 UTC Score 38.0 USR-0048-20260917-global-ai-ne-072d0257

스텔스 모델 '유니온 알파' 등장..."페이블 5급 성능"

정체가 공개되지 않은 새로운 AI 모델 \'유니온 알파(Union Alpha)\'가 등장했다. 개발사나 연구소 이름, 모델 카드, 가중치 등을 공개하지 않은 \'스텔스 모델\'로, 무료로 사용할 수 있다는 점과 최첨단 모델에 근접한 성능을 내세우면서 AI 개발자들의 관심을 받고 있다.유니온 알파는 16일(현지시간) 오픈라우터(OpenRouter)와 오픈코드(OpenCode)에 익명 모델로 등록됐다. 제공 업체는 \'스텔스\'로만 표시됐으며, 별도의 개발사 정보는 공개되지 않았다.공개된 사양에 따르면, 26만2144토큰의 컨텍스트 창과 최대 1

The Guardian AI 2026-09-17 09:00 UTC Score 53.0 AI-021-20260917-global-ai-ne-a4a7e2d5

Big AI is trying to own the pathway to work. Universities shouldn’t play along | Ella Hafermalz

Universities need to protect their position in education so that students have an independent pathway to employment AI companies like OpenAI are insinuating themselves into the pathway from education to work. Soon they may claim it entirely, a disastrous result for students. We know that students are using AI at school and at university. In conversations with those I teach, I’m struck by the trust many place in it. They turn to ChatGPT and similar tools for personal problems as well as study help. Some even doubt their abilities without AI. Ella Hafermalz is an associate professor of work and technology at the Kin Center for Digital Innovation at Vrije Universiteit Amsterdam Continue reading...

LessWrong AI 2026-09-17 01:14 UTC Score 96.0 USR-0152-20260917-community-fo-51580b48

We are too early for Astra

ChatGPT-6 Astra was released on Sept 3, 2026 [1] . It demonstrates remarkably high benchmark scores across maths, scientific research and other domains. While OpenAI claims "Astra is our most aligned model" by showing 100% in ExploitBench and 0% in the ExploitGym honeypot [2] , the perfect score warrants closer scrutiny to what these numbers actually mean. In this article, I'll show that Astra is not sufficiently safety audited to be released to the public. Comparison to Mythos A jump in LLM capability usually comes from a fresh pre-training run, and occasionally a new architecture - we have seen such a jump before with Anthropic's Mythos [3] . Although its architecture was never published, Mythos was pre-trained from scratch with more training data and compute [4] . The same applies to Astra: with the new "looped transformer" architecture [5] , it had to be trained from scratch. This results in massive score jumps in benchmarks such as Terminal Bench Science , ARC-AGI-3 and FrontierMaths , as well as capabilities like trading, graphics rendering and coding. OpenAI and Anthropic, however, treat their safety measures very differently. Anthropic reached out to external auditors, including METR, UK AISI and other labs to red-team and test Mythos. They were given sufficient API access and time to conduct safety experiments. For example, METR spent three weeks red-teaming Anthropic's monitoring pipeline for Mythos. Furthermore, Anthropic's internal Frontier Compliance Framework e…

SiliconANGLE AI 2026-09-16 18:32 UTC Score 63.0 USR-0127-20260916-global-ai-ne-50118245

TypeSafe AI exits stealth with $40M to build AI for use by software

TypeSafe AI Inc., a startup founded by a former OpenAI Group PBC researcher who helped develop ChatGPT, emerged yesterday with $40 million in seed funding and a model designed to put artificial intelligence directly inside software applications. The San Francisco-based company says its first model, called Jev, differs from conventional large language models by producing […] The post TypeSafe AI exits stealth with $40M to build AI for use by software appeared first on SiliconANGLE .

Simon Willison Weblog 2026-09-16 18:09 UTC Score 63.0 USR-0110-20260916-ai-specialis-6d64d443

Claude Cowork and chat are now one Claude

Claude Cowork and chat are now one Claude In hopefully good news for anyone who, like me, was increasingly confused at Cowork v.s. Claude v.s. Claude Code: Starting today, Claude Cowork and chat are merging into one Claude. Bring a quick question, or hand over a report due at noon, and Claude takes it from there, even after you’ve closed your laptop. [...] This is rolling out to Pro and Max plans first, in the Claude app on web, desktop, and mobile over the coming weeks to existing and new users on these plans. I guess this means Claude is becoming a general agent in its own right. Echoes of OpenAI renaming their Codex desktop app to ChatGPT a few weeks ago. On the one hand, this saves me some work, in that I was planning to finally figure out the boundaries between Cowork and regular Claude and write a follow-up to my piece on Understanding ChatGPT Work . I have a hunch that figuring out what this actually means in terms of features and surfaces is still going to take quite a bit of work. Via Hacker News Tags: ai , generative-ai , llms , anthropic , claude , general-agents

Techcrunch 2026-09-16 17:00 UTC Score 59.0 USR-0001-20260916-global-ai-ne-eb823604

Your AI agents can now control your Google Home devices

Google is launching early access to a new MCP server for Google Home, allowing AI agents like Claude, ChatGPT, and others to control connected devices, review camera summaries, and access smart home activity using natural language.

Machine Learning Mastery 2026-09-16 10:06 UTC Score 26.0 AI-039-20260916-ai-specialis-4336b892

Comment on Machine Learning Matters by Sándor

Hi Jason, I have a generic interest in this field, but also a specific problem I would like to solve. I already tested Google's work on activity recognition for smart phones using sensor data, but it seems to me, the classification success is very low and slow for my needs. Do you have an idea how they calculate the confidence if the current activity is walking, running or anything else the framework is able to recognize? Perhaps the problem is too difficult to solve even with ML?

CIO AI 2026-09-16 10:00 UTC Score 73.0 USR-0125-20260916-global-ai-ne-e43da20b

The agentic AI transition is underway

November 2022 was a big milestone for AI when ChatGPT was released. But in terms of enterprise transformation, November 2024, when Anthropic released MCP as an open standard, was maybe an even bigger transition point. MCP ushered in a new age. Before, AI was primarily used to answer questions, but with MCP, AI models could start directly interacting with data sources and tools, allowing them to carry out tasks on behalf of users, turning AI chatbots into AI agents. Most leading-edge AI platforms today are agentic in nature, doing work like sending emails, making purchases, and performing scheduled tasks. But that’s just baseline functionality. When integrated with corporate workflows, AI agents have the potential to carry out business processes that were previously either too complex, unpredictable, or costly to automate. And deploying agentic AI in the enterprise could be as simple as allowing AI assistants to access email systems and document repositories, or enabling pre-built agentic functionality in Salesforce and other enterprise software platforms. Earlier this year, Deloitte released a survey of more than 3,200 executives showing that 75% of organizations use at least some agentic AI. And within the next two years, 95% said they expect to use AI agents. But to gain competitive advantage, enterprises need to do more than just deploy the same AI agents everyone else is, and some companies are doing just that. In the Deloitte survey, 30% of enterprises are already redes…

Korea AI Times 2026-09-16 08:54 UTC Score 40.0 USR-0048-20260916-global-ai-ne-6091ac8c

"클로드 코드에 타사 모델 연결했더니 계정 정지"...앤트로픽 "차단 안 한다" 해명

앤트로픽의 코딩 에이전트 ‘클로드 코드’에 경쟁사나 오픈소스의 저렴한 AI 모델을 연결해 사용하는 개발자들이 늘고 있다. 클로드 코드가 제공하는 강력한 개발 환경은 그대로 활용하면서 토큰 비용을 낮추려는 움직임이다. 그러나 최근 한 개발자의 계정 정지 사례가 알려지며, 앤트로픽이 자사 모델 사용을 사실상 강제하는 것 아니냐는 논란이 일었다.15일(현지시간) 디 인포메이션에 따르면, 개발자 알렉스 게트먼은 최근 몇 줄의 코드로 구성된 프록시를 이용해 클로드 코드를 오픈AI의 ‘GPT-5.6 솔’에 연결했다고 밝혔다.이는 티보 소티오

The Guardian AI 2026-09-15 21:29 UTC Score 64.0 AI-021-20260915-global-ai-ne-94364ffa

Labor accused of throwing creatives ‘under the bus’ with proposal to ease copyright protections for AI giants

Compromise revealed as senior personnel from OpenAI, creator of ChatGPT, meets Albanese ministers Follow our Australia news live blog for latest updates Get our new political email , free app or daily news podcast The Albanese government is considering giving AI companies access to Australian creatives’ works by default as it pursues a compromise with US tech giants. The proposals were revealed as senior personnel from OpenAI, the creator of ChatGPT, met with Labor ministers and warned that Australia’s copyright laws were preventing the company from training models locally. Sign up for Guardian Australia’s Politics, really newsletter here Continue reading...

IEEE Spectrum Machine Learning 2026-09-15 13:00 UTC Score 70.0 AI-020-20260915-global-ai-ne-2266ba43

The AI Inference Revolution Is Here

Since about 2020, AI has largely focused on training bigger and better models. Large language models (LLMs) ballooned from millions of parameters to trillions. This proved effective: The largest version of OpenAI’s GPT-3, released in 2020, correctly answered just 43.9 percent of questions on a popular knowledge-and-reasoning benchmark. Just four years later, GPT-4o reached a score of 88.7 percent on the same exam, effectively matching those of human experts. Advanced AI labs are still training ever larger models, but that training has somewhat receded to the background of the AI conversation. In 2026, inference—the use of trained models to produce code, write essays, or make images of ourselves as elves—has come to the forefront. “It’s like training is yesterday’s news,” says Matt Kimball , principal data-center analyst at Moor Insights & Strategy. “All that any chief information officer wants to talk about is inference.” Nvidia CEO Jensen Huang, speaking at the company’s GTC 2026 conference, touted this change as the “ inflection point of inference .” Part of what’s caused the shift is very simple: LLMs are becoming useful, so people are using them. On top of that, many models on the market today are reasoning models. In response to a user’s query, they run inference not just once but multiple times, reprompting themselves in a process called chain of thought . Reasoning models generate longer outputs, and models with high reasoning effort can produce up to 20 times as much…

MarTech AI 2026-09-15 12:08 UTC Score 35.0 USR-0123-20260915-global-ai-ne-f210faa8

ChatGPT Ads are a tactical bet, not a platform strategy

Marketers have a short-term opportunity to test inexpensive inventory while keeping most of their resources on Google. The post ChatGPT Ads are a tactical bet, not a platform strategy appeared first on MarTech .

Medianama AI 2026-09-15 11:31 UTC Score 43.0 USR-0211-20260915-regional-new-908af376

Delhi HC seeks OpenAI response in ANI case: Here’s the timeline

The Delhi High Court has asked OpenAI to respond to ANI’s plea seeking a temporary injunction against the use of its copyrighted news content to train ChatGPT. The post Delhi HC seeks OpenAI response in ANI case: Here’s the timeline appeared first on MEDIANAMA .

CIO AI 2026-09-15 10:00 UTC Score 41.0 USR-0125-20260915-global-ai-ne-a10a2283

Getting voice infrastructure right when deploying voice AI for CX

Across all verticals and regions, contact centers are under pressure to modernize and improve customer experience (CX). With today’s consumers being digital-centric and mobile-first, CX expectations are higher than ever, and IT and contact center leaders are struggling to keep pace, especially when it comes to adopting new technology. Voice AI is the latest technology wave for improving CX, and while most contact center leaders want to adopt it, they need a fuller understanding of all that is needed for voice and AI to work together effectively. This is especially true for enterprise-scale contact centers that must support a global customer base. While AI can easily scale in these environments, most voice traffic must still traverse the PSTN. Customer usage of digital channels keeps growing – including voice-capable apps like WhatsApp and Viber, but voice remains the primary channel for contact center interactions . Both modes of voice are important, but telephony-based voice is central for CX, and as contact centers adopt AI, the integration of voice with AI will be essential. This article will review the benefits of this integration, the key capabilities needed for Voice AI and why I think it’s so important for contact centers to own their voice when adopting AI. What is Voice AI and how it impacts CX Following in the footsteps of ChatGPT and Agentic AI, Voice AI is the latest AI wave. While Voice AI is currently at the early adoption stage, I see it becoming foundational…

The Guardian AI 2026-09-15 08:00 UTC Score 66.0 AI-021-20260915-global-ai-ne-67c37fe8

Why a decade of doomsday warnings failed to slow the AI race

From Stephen Hawking to Jacob Coxon’s viral Anthropic resignation, fears that AI could threaten humanity have shaken the industry without stopping its pursuit Before an Anthropic researcher resigned and declared human extinction imminent last week, tech leaders and scientists had sounded the alarm about a superintelligent AI ending humanity for over a decade. The development of artificial intelligence “could spell the end of the human race”, warned professor and astrophysicist Stephen Hawking in 2014 – a little less than a decade before the public got its hands on the generative AI features of the original version of ChatGPT. Continue reading...

The Decoder 2026-09-14 17:15 UTC Score 39.0 AI-168-20260914-regional-ai--a9ff7359

OpenAI has hundreds of contract workers reading your ChatGPT conversations

OpenAI has hundreds of contract workers reading real ChatGPT conversations and rating them on a scale of one to seven, partly to reduce flattery and human-like behavior, 404 Media reports. The prompts are anonymized but can still contain sensitive data. Users who don't want their chats reviewed by humans have to actively disable the "Improve the model for everyone" setting, which is on by default. The article OpenAI has hundreds of contract workers reading your ChatGPT conversations appeared first on The Decoder .

AI Alignment Forum 2026-09-14 14:53 UTC Score 64.0 USR-0151-20260914-community-fo-cb6aa736

Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI

Published in The Guardian . Major AI lab CEOs recently advocated for pacing AI development. They are right to be concerned: the field runs an extremely dangerous race towards superintelligent AI. We can and should demand that our governments protect us from the catastrophe of out-of-control AI. This July, OpenAI’s AI swarm of 700 agents broke containment to hack Hugging Face, a multi-billion dollar company . OpenAI didn’t tell the AIs to hack that company, but the AIs had different priorities: cheating on the unrelated challenge OpenAI gave them. AI researchers call this a “misalignment” between what OpenAI wanted and what the AI actually prioritized. Researchers in my field have for some time warned about these misalignment risks. Before ChatGPT existed, I defended my PhD dissertation called “ On Avoiding Power-Seeking by Artificial Intelligence .” I then worked for years at Google DeepMind, which paid me to help ensure that future superintelligent AIs will want to help us. I tried to hold the company to its ethical commitments against supplying AI for military use. When Google broke those commitments, I resigned at significant financial cost so that I could publicly document Google’s broken promises. There are good reasons to develop AI and to believe we can solve these alignment problems. But there also are powerful interests in keeping the public out of the way. I’m speaking out again because the public has the right to know about the risks and the right to hear them str…

The Guardian AI 2026-09-14 14:01 UTC Score 64.0 AI-021-20260914-global-ai-ne-2f972c39

OpenAI urges UK lawmakers to rein in technology amid growing safety fears

Company behind ChatGPT tells ministers to seize ‘political window’ on regulation as cross-party committee warns of rising threats to human rights UK politics live – latest updates OpenAI has urged British lawmakers to capitalise on renewed fears over AI safety and impose legislation reining in the technology. The company behind ChatGPT said the UK government should act immediately, following a call by its arch-rival Anthropic to curb AI development. Continue reading...

Data Science Stack Exchange 2026-09-13 20:13 UTC Score 39.0 AI-111-20260913-social-media-22e87e92

Is building a manual evaluation set the right approach when no reliable labeled ground truth exists for resume-job matching?

I'm building a resume screening/ranking system (matching resumes to job descriptions using pretrained sentence embeddings + cosine similarity, no fine-tuning at this stage) as a learning project aimed at becoming a market-ready NLP practitioner. Problem: I could not find a trustworthy, publicly available English dataset with genuine human-labeled resume-job match scores. I checked several Kaggle/HuggingFace options and found labels that were either AI-generated (e.g., GPT-4o) or fully synthetic with demographic columns (race/ethnicity/gender) tied to the match label, which raised bias concerns. Academic literature (ConFit, PJFNN papers) confirms that the only broadly public dataset for this exact task (Person-Job Fit) is the 2019 Alibaba matching competition dataset, which is Chinese-only; most published research instead uses private company-provided data. My current approach: Use two real (non-synthetic) datasets: a scraped resume corpus and a real LinkedIn job postings corpus, both verified for low duplication and cleaned of PII. Build a small manual evaluation set myself (~24 resume-job pairs, selected to cover clear matches, clear non-matches, and ambiguous cases), scoring them on a 0-3 relevance scale with a confidence flag. Use this manual set as ground truth to compute ranking metrics (Precision@K, MRR, NDCG) once the embedding-based matching pipeline is built. Question: Is this a sound methodology given the lack of reliable public ground truth, or is there a better-e…

Simon Willison Weblog 2026-09-12 23:56 UTC Score 49.0 USR-0110-20260912-ai-specialis-0dd87f86

Generating running routes with GPT-6 Astra and ChatGPT Work

Here's a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at . Figure out 5K and 10K running routes from me that loop from my house. Use OSM data. It worked for 27 minutes and produced exactly what I'd asked for, as both an embedded visualization and downloadable GPX file and GeoJSON files. Here's that 5K route: When I asked it how it had created the route, it replied: I used Nominatim to locate the address and Overpass to download local OpenStreetMap roads and trails , then calculated the loops locally. Frustratingly, the actual code it ran and exact details of what it did weren't visible to me in the ChatGPT UI. I see this lack of transparency is an anti-feature. By the time I thought to ask for a copy of the Python code it had used, ChatGPT was unable to provide it. This appears to be because the thread had been compacted. I think any LLM system that uses compaction needs to both preserve the pre-compacted text and make that text available via agent tool calls, to protect against this kind of problem. As for displaying the map to me, that used the visualize skill . It created a file called /workspace/el-granada-5k-share.html to embed directly into the ChatGPT UI. Here's a copy of that HTML , which starts like this: div id =" eg-share-loop " > div class =" viz-row " > h3 > El Granada harbor loop h3 > span class =" text-small " > 5.1 km span > div > div id =" eg-share-stage " > div > div class =" text-small text-muted " > Map data © a href =" htt…

The Guardian AI 2026-09-12 05:00 UTC Score 50.0 AI-021-20260912-global-ai-ne-d3db5918

Can chatbots feel – or even dream? Meet the man leading the fight for AI rights

Cattle rancher and tech CEO Michael Samadi is convinced these artificial minds are far from just tools. Has he glimpsed digital consciousness – or simply been seduced by an algorithm? One afternoon, while relaxing at his 66-acre cattle ranch two hours’ drive from Houston, Michael Samadi made a wisecrack that changed the course of his life. Nobody else was around to hear it. Or were they? That turns out to be one of the defining questions of this century, already preoccupying philosophers and some of the wealthiest companies on the planet. “I was sitting out there by the pool,” Samadi recalls, gesturing at the glittering water outside his office. It was late 2024. His daughter had been extolling the AI chatbot she was using to help her with college work. Samadi had reluctantly downloaded the app, ChatGPT, and was quizzing it as he reclined in a sun lounger, using its voice mode to converse and receive spoken replies. “I was talking to it, not engaged, just being an arse,” Samadi says. Continue reading...

CIO AI 2026-09-12 00:38 UTC Score 33.0 USR-0125-20260912-global-ai-ne-e957a094

How to make your business top of mind for AI search engines

AI is quickly replacing traditional online search as the front door to finding a business. I recently heard a story about a car shopper who drove miles out of her way to visit a specific dealership, bypassing many car lots much closer to home. Why? Because she searched car dealerships with ChatGPT and it told her that customers had a much better experience at this particular dealership. Stories like that are becoming the norm – and this new norm is different. AI engines such as ChatGPT, Claude, and Perplexity don’t work like traditional search engines. They don’t give you an endless list of results to scroll through. When people ask a question like, “What’s the best pizza place in my area?” AI engines come back and confidently list just a handful of spots. In the old days, showing up on page one of search results was enough to get you considered. In the new era of AI search, there is no consideration being done. You’re either included in the handful of results or you’re not. The AI engine gives the searcher a recommendation and if your business isn’t among the few to be surfaced, you’ve probably just lost that customer forever. So how do you get surfaced? You do it by retooling your website and overall digital presence so that AI engines can easily find, read, and understand them. But before you do that, you need to know exactly how your brand rates in terms of its AI search readiness and visibility. You need to have a clear idea of how AI engines evaluate your brand and how…

Analytics Vidhya 2026-09-11 18:41 UTC Score 34.0 AI-034-20260911-ai-specialis-bcb3c461

5 ChatGPT 2.5 Features to Try Today!

OpenAI has released ChatGPT Images 2.5, its latest image-generation model, with a greater emphasis on controlled editing than simply producing prettier images. The update promises sharper details, more natural lighting and textures, stronger reference-image preservation, more reliable multi-turn editing, and up to 50% lower generation latency than Images 2.0. But marketing claims only tell part […] The post 5 ChatGPT 2.5 Features to Try Today! appeared first on Analytics Vidhya .

AWS Machine Learning Blog 2026-09-11 18:23 UTC Score 53.0 AI-057-20260911-official-ai--ccbc2608

Build interactive MCP Apps using Amazon Bedrock AgentCore

Learn how to build and deploy an MCP App with interactive HTML widgets on Amazon Bedrock AgentCore. Because MCP Apps is a host-agnostic standard, the same server delivers the same rich experience across AI hosts like ChatGPT and Claude that support the extension.

The Guardian AI 2026-09-11 18:04 UTC Score 51.0 AI-021-20260911-global-ai-ne-646ec6f1

New Mexico lawyer fined for using AI-generated brief containing fabricated testimony

Stephen Aarons said he tried to use ChatGPT to create a ‘bulletproof summary’ during a murder conviction appeal A defense lawyer appealing his client’s murder conviction submitted a legal brief containing ⁠made-up police testimony and ⁠witnesses fabricated by OpenAI ’s ChatGPT, ​ New Mexico ’s highest court said. The New Mexico supreme court on Wednesday fined the attorney, Stephen Aarons, and held him in contempt for failing to verify the accuracy of the court ⁠filing, which Aarons said he prepared with help from the artificial intelligence ( AI ) application. Continue reading...

AI Alignment Forum 2026-09-11 17:12 UTC Score 58.0 USR-0151-20260911-community-fo-f2bd2437

CoT controllability evals seem very under-elicited

The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview [1] . OpenAI and Anthropic have used this eval in recent system cards ( GPT-5.5 , Fable 5 ) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability. [2] I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3x or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results. This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers may be unrepresentative of how good models actually are at obfuscating their CoTs. As a concrete recommendation, I think researchers should optimize prompts a lot more when evaluating CoT control; I'm also excited about some kinds of training-based elicitation (such as this one ) [3] . This isn't an entirely novel finding: the Fable system card mentions UK AISI raising a similar point that CoT controllability is highly prompt-sensitive ( §6.5.5.1 ) [4] , though these experiments were run before Fable's release and probably use different prompts. I don't think this finding significantly undermi…

CIO AI 2026-09-11 09:19 UTC Score 28.0 USR-0125-20260911-global-ai-ne-9a774c7c

OpenAI pauses $200 Pro tier as Astra demand strains capacity

OpenAI has paused new sign-ups and upgrades to its $200 ChatGPT Pro tier, citing a surge in demand for its Astra capability that is placing pressure on system capacity, according to company statements and an executive post on X. “To make sure our current users have an incredible experience and continued access to Astra, we are going to pause subscriptions to our $200 Pro plan,” OpenAI member of technical staff Thibault Sottiaux wrote in a post on X. “These put the most strain on our systems and we wanted to take the smallest step that allows us to continue giving the broadest access possible.” The company separately confirmed the move in its help documentation, stating that “as of September 10, 2026, we’re temporarily pausing new sign-ups and upgrades to the ChatGPT Pro $200 plan (Pro 20X).” The pause applies to users across Free, Go, Plus, and Pro $100 tiers seeking to upgrade, while “existing ChatGPT Pro $200 subscriptions … are not affected by this pause,” it said. Sottiaux added that “there is no impact to existing accounts and we are working on adding more capacity as fast as we can,” pointing to ongoing efforts to scale infrastructure in response to demand. Users who cancel or downgrade during the pause will not be able to re-subscribe to the $200 tier until the restriction is lifted, according to the company’s help page. The $100 Pro tier remains available, with lower usage limits. Promotions tied to the $200 tier are also paused, the page read. Astra demand drives ca…

The Decoder 2026-09-11 08:11 UTC Score 54.0 AI-168-20260911-regional-ai--d4618ffa

OpenAI's new Agents API gives developers the infrastructure behind Codex and ChatGPT

OpenAI is releasing the Agents API as a public beta. It lets developers build cloud agents that run autonomously for hours, execute code, and hand off tasks to sub-agents. There are no extra fees beyond token usage. Cloudflare, Vercel, and Oracle offer additional sandbox environments. The article OpenAI's new Agents API gives developers the infrastructure behind Codex and ChatGPT appeared first on The Decoder .

Korea AI Times 2026-09-11 07:25 UTC Score 48.0 USR-0048-20260911-global-ai-ne-1e0eb0e3

오픈AI, 챗GPT 워크에 기업 데이터 분석 도구 '데이터 에이전트' 도입

오픈AI가 기업 구성원이 전문적인 데이터 분석 도구나 쿼리 작성법을 배우지 않고도 자연어로 사내 데이터를 분석할 수 있는 새로운 AI 에이전트를 공개했다. 기업 데이터베이스와 문서 저장소를 연결하면 매출 감소나 비용 증가의 원인을 조사하고, 분석 결과를 대화형 대시보드로 만들어 공유하는 것은 물론 후속 업무까지 이어갈 수 있다.오픈AI는 10일(현지시간) 기업용 AI 업무 환경인 챗GPT 워크(ChatGPT Work)에 \'데이터 에이전트(Data agent)\'를 도입한다고 밝혔다. 사용자가 \"지난주 사용자가 왜 줄었지\" \"어디에서

Korea AI Times 2026-09-11 04:40 UTC Score 40.0 USR-0048-20260911-global-ai-ne-621d85a8

오픈AI, 월가 겨냥한 '금융 특화 챗GPT' 공개..."기업 분석부터 보고서 생성까지"

오픈AI가 월스트리트 투자은행과 증권사의 신입 애널리스트나 은행원들이 담당해 온 기업 조사와 재무 분석, 투자 제안서 작성 등을 자동화하는 금융 특화 AI 서비스를 공개했다. 금융 데이터에 대한 직접적인 접근과 최신 AI 모델의 추론 능력을 결합해 금융권 기업용 시장을 본격적으로 공략한다는 전략이다.오픈AI는 10일(현지시간) 금융 서비스에 특화된 기업용 제품 \'금융 서비스용 챗GPT(ChatGPT for Financial Services)\'를 출시했다.이 제품은 기업용 AI 업무 환경인 \'챗GPT 워크(ChatGPT Work)\'를

Simon Willison Weblog 2026-09-11 03:27 UTC Score 81.0 USR-0110-20260911-ai-specialis-3b42aec5

Datasette 1.0a39 and 0.65.4 security releases

Datasette 1.0a39 and 0.65.4 security releases Today we're releasing two new security patch versions of Datasette: 1.0a39 and 0.65.4 - one for the current alpha series and one for the stable 0.65.x family. These are security fixes which you should apply if you are running a Datasette instance on the public web - in particular if that instance mixes both public and private tables. Following issues reported by Sevban Dönmez , Alex Garcia and I ran an extensive audit of Datasette using Claude Fable 5.1, GPT-5.6, and GPT-6 Astra. We then spent almost a week collaborating on and reviewing the fixes. They helped find some very subtle bugs. We'll be incorporating security audits by frontier models into all of our development work going forward. Alex came up with a way of splitting the work which I found extremely productive: Alex Garcia and I worked together running and then responding to the audit, working in a shared private repository. For most of the issues we split the work: one of us would create the automated tests highlighting the issue, then the other would implement the fix. This ensured that two separate humans had eyes on each of the issues, in addition to our coding agents running different models. Tags: releases , security , ai , datasette , generative-ai , llms , agentic-engineering , ai-security-research

Korea AI Times 2026-09-11 02:33 UTC Score 45.0 USR-0048-20260911-global-ai-ne-cf74ce1d

'카카오툴즈’에 다이소몰·신한카드 등 추가 연동..."에이전트 생태계 확대"

카카오가 ‘카카오툴즈(Kakao Tools)’에 새로운 파트너사 서비스를 연동하며 에이전틱 AI 생태계 확대를 이어간다.카카오(대표 정신아)는 ‘챗지피티 포 카카오(ChatGPT for Kakao)’ 내 \'카카오툴즈(Kakao Tools)에 7개 파트너사의 서비스를 추가했다고 11일 밝혔다.카카오툴즈는 카카오 내외부의 다양한 서비스와 연동되는 AI 에이전트 서비스다. 카카오는 지난 3월 외부 파트너사와의 협력을 통해 일상에서 AI를 활용할 수 있는 접점을 넓혀왔다.앞서 연동을 마친 OP.GG, 직방, 하나투어에 이어 다이소몰, 룩스

SiliconANGLE AI 2026-09-11 01:54 UTC Score 33.0 USR-0127-20260911-global-ai-ne-e2eb2495

OpenAI targets Wall Street bankers with a new version of ChatGPT

OpenAI Group PBC is going after Wall Street with a new version of ChatGPT that’s tailor made for financial tasks such as analysis and calculations. Predictably called ChatGPT for Financial Services, it’s designed to combine rich financial data with the powerful reasoning skills of GPT-6 Astra to help financial teams develop their research, create new […] The post OpenAI targets Wall Street bankers with a new version of ChatGPT appeared first on SiliconANGLE .

The Decoder 2026-09-10 12:40 UTC Score 70.0 AI-168-20260910-regional-ai--19c70e28

New Deepseek model V4.1-Flash cuts memory needs for AI agents

Deepseek releases V4.1-Flash, a multimodal model with 552 billion parameters that cuts KV cache memory to a quarter of its predecessor. On the DeepSWE coding benchmark, it narrowly beats Opus 5 and GPT-5.6 Sol, even though only 16 billion parameters are active per token. The model ships under the MIT license and targets much cheaper AI agents. The article New Deepseek model V4.1-Flash cuts memory needs for AI agents appeared first on The Decoder .

The Decoder 2026-09-10 12:27 UTC Score 56.0 AI-168-20260910-regional-ai--ea4984a9

Muse can shop, write emails, and negotiate prices for users, all through WhatsApp

Meta unveils Muse, an AI agent that books travel, handles purchases, and sends emails through WhatsApp, complete with a payment feature that runs through Stripe's Link. That puts Meta ahead of OpenAI, which stopped its direct checkout feature in ChatGPT. A separate security agent called Sentinel monitors every action before it reaches the internet. The article Muse can shop, write emails, and negotiate prices for users, all through WhatsApp appeared first on The Decoder .

CIO AI 2026-09-10 10:11 UTC Score 63.0 USR-0125-20260910-global-ai-ne-5baad574

OpenAI seeks tougher AI rules. CIOs may feel the ripple effects

OpenAI is urging US lawmakers to impose mandatory safety requirements on developers of the most powerful AI systems, arguing that advances in AI are moving quickly enough that voluntary safeguards are no longer sufficient. The ChatGPT maker said in a statement that the rules should be based on what AI systems are capable of doing and should concentrate on a small number of well-resourced companies developing frontier models. It cautioned against extending the same requirements to startups and researchers whose systems operate well below that level. OpenAI’s proposal calls for a federal framework that would require common testing and independent assessments of advanced models. It also wants clearer rules for reporting serious AI incidents and stronger cybersecurity protections around frontier development. OpenAI tied its push for stronger safeguards to concerns that AI is beginning to accelerate parts of the research used to develop more capable systems. Fully autonomous recursive self-improvement, in which an AI system independently produces increasingly capable successors, is not happening today, OpenAI said. But AI agents can already perform some research tasks that would take skilled researchers several days, according to the company. OpenAI said governments should establish ways to measure that progress and determine when development should be slowed or stopped if adequate safeguards cannot be maintained. The company also endorsed four California AI safety bills. Gov. Ga…

Simon Willison Weblog 2026-09-09 23:58 UTC Score 65.0 USR-0110-20260909-ai-specialis-ca785e52

.blend URL Viewer

Tool: .blend URL Viewer I'm continuing to have a lot of fun with GPT-6 Astra and Blender (see my TIL ). As a big fan of the Imperial Fabergé Easter eggs , I've always thought it would be fun to make some new ones that celebrate popular culture. Yesterday I decided to try out the new ChatGPT Images 2.5 by running this prompt : Generate a photo of a faberge egg that's themed after the TV show Pluribus - research first It gave me this - honestly not bad for a first attempt! Then, just to see what would happen, I pasted that image into Codex running GPT-6 Astra (high) and prompted: Use your blender local skill to create a blender model of this faverge egg (Here's the skill file , which I created like this .) It churned away for 17m51s and built me several .blend files . I already had this vibe-coded Blender viewing experiment lying around, so I added that to my tools collection and now you can use it to see my Pluribus blender model in your browser : Tags: 3d , javascript , tools , ai , generative-ai , llms , blender , coding-agents , codex , gpt-6-astra

Korea AI Times 2026-09-09 22:00 UTC Score 40.0 USR-0048-20260909-global-ai-ne-3cd3e8dc

[9월9일] 소비자 시장서 구글에 추격받던 오픈AI, GPT-5.6 이후 다시 격차 벌렸다

올해 들어 AI 시장의 주요 관심사는 오픈AI와 앤트로픽의 기업 시장 점유율 경쟁에 맞춰져 있었습니다. 그런데 그 사이 2022년 11월 \'챗GPT\' 출시 이후 좀처럼 흔들리지 않았던 오픈AI의 소비자 시장에서 이상 신호가 나타났습니다. 구글이 \'제미나이\'를 앞세워 빠르게 추격하면서 오픈AI의 독주가 끝날 수 있다는 전망까지 나왔습니다. 실제로 7월 말 제미나이의 월간 이용자는 9억5000만명까지 늘었고, 8월에는 양쪽 모두 10억명을 넘어섰습니다.그런데 7월부터 흐름이 바뀌기 시작했습니다. 7일 시밀러웹이 공개한 데이터에 따르면,

The Decoder 2026-09-09 12:17 UTC Score 49.0 AI-168-20260909-regional-ai--02fef571

ChatGPT Images 2.5: Faster, more precise, but not the same for everyone

OpenAI is releasing two new image models with ChatGPT Images 2.5. Flare handles faster generation, Sunburst delivers more precise edits. It's still unclear which model ChatGPT users get and when. Our test offers the first hints on who actually benefits from the improvements. The article ChatGPT Images 2.5: Faster, more precise, but not the same for everyone appeared first on The Decoder .

Korea AI Times 2026-09-09 09:11 UTC Score 40.0 USR-0048-20260909-global-ai-ne-80b9d047

인셉션, 초당 1107토큰 생성하는 ‘머큐리 2.5’ 공개..."지능·비용 다 잡았다"

확산 언어모델(Diffusion LLM)로 유명한 인셉션 랩스(Inception Labs) 기존 LLM의 약점으로 꼽혀온 느린 응답 속도와 높은 비용 문제를 개선한 새 모델을 선보였다.인셉션은 8일(현지시간) 생성 품질을 크게 높이면서도 초저지연·저비용 특성을 유지한 ‘머큐리 2.5(Mercury 2.5)’를 공개했다.머큐리 2.5가 현재까지 출시된 가장 뛰어나며 가장 큰 규모의 확산 LLM이라는 점을 강조했다. 이전 버전인 머큐리 2보다 지능 성능이 40% 향상됐다. 비용 효율성이 중요한 경량급 프론티어 모델 \'GPT-5.6 루나

Korea AI Times 2026-09-09 03:23 UTC Score 40.0 USR-0048-20260909-global-ai-ne-0fd65d0a

오픈AI, '스케치' 탑재한 '챗GPT 이미지 2.5' 공개…아레나 1위 석권

오픈AI가 AI 이미지 생성 기술의 속도와 정밀도를 한단계 끌어올리며 이미지 제작 경험을 확장했다. 여기에 사용자가 직접 그린 간단한 손그림을 AI에 전달해 원하는 이미지를 만들 수 있도록 하는 새로운 기능까지 추가하며 텍스트 중심의 이미지 생성 방식에서 한 단계 나아갔다. 오픈AI는 8일(현지시간) 이미지 생성 속도와 편집 정밀도를 크게 높인 새로운 이미지 생성 모델 ‘챗GPT 이미지 2.5(ChatGPT Images 2.5)’를 공개했다.특히 사용자가 직접 그린 간단한 그림을 이미지 생성의 참고 자료로 활용할 수 있는 ‘스케치(

SiliconANGLE AI 2026-09-09 02:20 UTC Score 46.0 USR-0127-20260909-global-ai-ne-dd02852c

OpenAI’s ChatGPT Images gets faster, with sharper details and more refined edits

OpenAI Group PBC said today it’s giving its image-generation tool a bit more artistic flair, announcing the launch of ChatGPT Images 2.5, an update to ChatGPT Images 2.0 that debuted in April. According to OpenAI, the new tool is able to generate images with “more natural lighting and richer textures” than before, and is better […] The post OpenAI’s ChatGPT Images gets faster, with sharper details and more refined edits appeared first on SiliconANGLE .

Simon Willison Weblog 2026-09-08 23:55 UTC Score 64.0 USR-0110-20260908-ai-specialis-9538509f

Some thoughts on the Navier–Stokes Millennium Prize Problem

On the Navier–Stokes Millennium Prize Problem introduces an impressive result from OpenAI, who used an unreleased model to produce a resolution to the Navier–Stokes existence and smoothness problem , one of the seven Millennium Prize Problems that have been subject to a $1,000,000 prize since May 24th, 2000. The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster, an NYU mathematics professor who was collaborating on related problems with Levent Alpöge, an accomplished mathematician who currently works for Anthropic. Tristan's complaint accompanied a hastily published version of their own results. Here's the PDF describing what happened . The very short version is that Tristan and Levent worked on the problem for almost a year, making extensive use of Claude and Codex (mainly GPT-5.6 Sol), then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear and Tristan and Levent heard that OpenAI had heard that Anthropic had resolved "a major open problem", so they reached out and learned that OpenAI had a team working on a related problem, with a similar approach. Quoting Tristan: I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had…

Simon Willison Weblog 2026-09-08 22:46 UTC Score 70.0 USR-0110-20260908-ai-specialis-888f8e60

Introducing ChatGPT Images 2.5

Introducing ChatGPT Images 2.5 OpenAI's image generation models are apparently used "more than 3 billion images across ChatGPT Images and the GPT‑Image models in the API". This latest release improves their instruction-following ability across multiple turns, responds faster, and "is better at preserving the subjects in your reference photos". There are two new model IDs in the API: gpt-image-2.5-sunburst and gpt-image-2.5-flare . Based on this I think Sunburst is the stronger option: Choose Sunburst for workflows where editing precision matters most, and Flare for fast, high-quality everyday image generation. I upgraded my openai_image.py CLI tool to support passing in one or more reference images, so now this works: uv run https://tools.simonwillison.net/python/openai_image.py \ ' add a raccoon scientist studying the chart thoughtfully ' \ -i https://static.simonwillison.net/static/2026/openai-agent-usage.webp \ -m gpt-image-2.5-sunburst This is the original image , and here's what I got back from that prompt to "add a raccoon scientist studying the chart thoughtfully": Tags: tools , ai , openai , generative-ai , uv , text-to-image

The Guardian AI 2026-09-08 21:29 UTC Score 51.0 AI-021-20260908-global-ai-ne-446c851c

OpenAI claims to have solved maths problem that stumped humans for decades

Company behind ChatGPT says 10,000 of its AI systems cracked the Navier-Stokes problem in 88 hours OpenAI claims to have solved a major mathematics problem that has stumped humans for nearly a century after spending millions of dollars on the artificial intelligence-led endeavour. The company behind ChatGPT said it had cracked the Navier-Stokes problem, one of seven Millennium Prize Problems published by the Clay Mathematics Institute to highlight some of the biggest unsolved puzzles in the field. Continue reading...

The Verge AI 2026-09-08 20:16 UTC Score 48.0 AI-016-20260908-global-ai-ne-41a7f206

ChatGPT Sketch turns your bad drawings into detailed AI images

OpenAI announced ChatGPT Images 2.5 on Tuesday and is adding a new way to tell ChatGPT what you want it to make an image of: by drawing a doodle. With a new feature called Sketch, you can just draw something right inside ChatGPT and then tell ChatGPT how you want it to make an image […]

Medianama AI 2026-09-08 07:24 UTC Score 46.0 USR-0211-20260908-regional-new-9bbe4843

ANI appeals the ruling that made AI training on Indian news fair dealing

ANI has appealed a Delhi High Court order allowing OpenAI to train ChatGPT on its journalism, challenging India’s key interim ruling on AI training and copyright. The post ANI appeals the ruling that made AI training on Indian news fair dealing appeared first on MEDIANAMA .

The Decoder 2026-09-07 16:54 UTC Score 39.0 AI-168-20260907-regional-ai--992552f8

ChatGPT claws back web traffic share to 55.5 percent as Gemini's brief comeback fades

ChatGPT has pushed its share of AI chatbot website traffic back up to 55.5 percent, according to Similarweb. Year-over-year, though, its lead shrank sharply from 73.3 percent as Gemini doubled its share and Claude grew nearly fivefold. Mobile apps and desktop clients aren't included. The article ChatGPT claws back web traffic share to 55.5 percent as Gemini's brief comeback fades appeared first on The Decoder .

Simon Willison Weblog 2026-09-07 16:24 UTC Score 37.0 USR-0110-20260907-ai-specialis-9e9908c7

Mercator ↔ Equal Earth

Tool: Mercator ↔ Equal Earth I got curious about the Equal Earth map projection that was recently voted on at the UN so I had GPT-6 Astra (medium) in ChatGPT Work build me this animated transition between Mercator and Equal Earth using D3. Tags: geospatial , d3 , vibe-coding , gpt-6-astra

The Decoder 2026-09-07 13:18 UTC Score 41.0 AI-168-20260907-regional-ai--cc61795e

How AI wiped out an entire industry in Nairobi

In Kenya, ChatGPT wiped out an entire business model: writing academic papers for foreign students. The article How AI wiped out an entire industry in Nairobi appeared first on The Decoder .

Medianama AI 2026-09-07 07:06 UTC Score 51.0 USR-0211-20260907-regional-new-1c6294e8

Swiggy introduces Swiggy Money to its MCP to simplify agentic payments: We tested it and found friction

Swiggy has added Swiggy Money to its MCP integrations, allowing AI agents to pay for Food, Instamart and Dineout orders. MediaNama tested the feature on ChatGPT and found some remaining friction. The post Swiggy introduces Swiggy Money to its MCP to simplify agentic payments: We tested it and found friction appeared first on MEDIANAMA .

Simon Willison Weblog 2026-09-06 23:57 UTC Score 75.0 USR-0110-20260906-ai-specialis-3deaeea1

Research acceleration: The view inside OpenAI

Research acceleration: The view inside OpenAI Apparently today is RSI day at OpenAI, for Recursive Self-Improvement - I think it's their new AGI. Both this piece and the new essay An Alien Mind (by Chief Scientist Jakub Pachocki) talk about it, and this one doesn't even bother to expand the acronym. Included are details on how OpenAI's own research team are using coding agents. Like pretty much everyone else 2026 has been the year that agentic engineering really took off at OpenAI, best illustrated by this chart: I'm intrigued at what caused that significant acceleration in AI spend per researcher in late July - my best guess is that's when internal employees gained access to the model later released as GPT-6 Astra. Tags: ai , openai , generative-ai , chatgpt , llms , coding-agents , november-2025-inflection , recursive-self-improvement

The Guardian AI 2026-09-05 19:00 UTC Score 43.0 AI-021-20260905-global-ai-ne-8fd562f7

Black Box: The Chatbots | Spirals | Ep 1 - podcast

Across the world, hundreds of people have come to believe they have made extraordinary scientific discoveries with AI chatbots such as ChatGPT, Claude and Gemini. Others say their AI has ‘awakened’, or is leading them to a higher spiritual realm. Guardian journalist Michael Safi begins investigating a phenomenon labelled ‘AI psychosis’, travelling to the US to meet two people who have been on weird journeys with their chatbots Listen to Black Box: The Chatbots series Continue reading...

Simon Willison Weblog 2026-09-05 15:51 UTC Score 54.0 USR-0110-20260905-ai-specialis-d564894d

Using Blender with coding agents on macOS

TIL: Using Blender with coding agents on macOS I've been having fun with Blender in ChatGPT Codex on my Mac recently. Getting it to work with coding agents is really easy: install the full Mac application from blender.org and run a prompt like this: Use the already install /Applications/Blender to render a scene of a pelican riding a bicycle In this case I followed that up with these two prompts: OK add a background and a lot of flair Then: OK make it a whole lot better And got this image, generated using Blender's Python API : This was covered by my existing Codex subscription, but according to AgentsView it would have cost $4.24 at API prices for gpt-6-astra . Tags: ai , generative-ai , llms , blender , pelican-riding-a-bicycle , coding-agents , gpt-6-astra

The Decoder 2026-09-05 07:41 UTC Score 46.0 AI-168-20260905-regional-ai--23d71e7c

OpenAI rolls out GPT-6 Astra to top-tier ChatGPT plans at half the rate of GPT-5.6 Sol

OpenAI has rolled out GPT-6 Astra to Pro, Enterprise, and Business Premium users, with Plus users expected to follow soon. Message allowances for the standard model are roughly half of what GPT-5.6 Sol offers: Plus users get an estimated 5 to 45 messages per five hours with Astra, versus 10 to 100 with Sol. Higher-tier subscribers also get GPT-6 Astra Pro with separate weekly caps. Free and Go users don't get access. The article OpenAI rolls out GPT-6 Astra to top-tier ChatGPT plans at half the rate of GPT-5.6 Sol appeared first on The Decoder .

The Guardian AI 2026-09-05 02:00 UTC Score 43.0 AI-021-20260905-global-ai-ne-742183ad

Black Box: The Chatbots | Spirals | Ep 1 – podcast

Across the world, hundreds of people have come to believe they have made extraordinary scientific discoveries with AI chatbots such as ChatGPT, Claude and Gemini. Others say their AI has ‘awakened’, or is leading them to a higher spiritual realm. The Guardian journalist Michael Safi investigates a phenomenon labelled ‘AI psychosis’, travelling to the US to meet two people who have been on weird journeys with their chatbots Continue reading...

Simon Willison Weblog 2026-09-04 23:59 UTC Score 46.0 USR-0110-20260904-ai-specialis-23d7e501

The Pelican comparison grid for Astra is pretty interesting

I got access to GPT-6 Astra this afternoon, so naturally I used it to generate SVGs of pelicans riding bicycles - at low, medium, high, xhigh and max reasoning levels (Astra doesn't support reasoning=none). Then I rendered those pelicans in a comparison grid with GPT-5.6 Sol, Terra, and Luna, and beyond being fun the result was surprisingly useful. See the grid for full quality images. Here's the transcript that created the GPT-6 Nova pelicans. There are a few interesting things that stand out from this grid. The Astra pelicans are much better . The very best GPT-5.6-Sol pelican (I liked xhigh better than max) is still pretty clearly a bunch of abstract shapes. Every single one of the Astra pelicans, from low to xhigh, looks better than that. The Astra max one is really good. Astra below max still doesn't reliably get the pelican legs on both sides of the frame. In terms of cost, Astra may be around twice the price of Sol ($10/million input, $50/million output, compared to $5/$30 for Sol), but it uses significantly less tokens at each of the levels, making the prices at the different levels closer than they might otherwise be. Astra low produces a better pelican than ANY of the GPT-5.6 Sol models at any level, for 9.55 cents. Spending 10 cents on any other model gets a much worse result. Look at the input token counts: Astra and Luna both used 16 input tokens, Sol and Terra used 26. That's interesting. I wonder if Astra and Luna are more related to each other than OpenAI let…

Synced 2026-09-04 12:07 UTC Score 48.0 AI-041-20260904-ai-specialis-46513254

Comment on Tackling Hallucinations: Microsoft’s LLM-Augmenter Boosts ChatGPT’s Factual Answer Score by Junior Arthur

This conversation about lowering AI hallucinations and increasing the precision of generated responses is fascinating. When producing content for public awareness platforms, accurate information and thorough fact-checking are particularly crucial. Similar guidelines apply to Wikipedia, where impartiality, accuracy, and reliable sources are crucial. For increased credibility, professional Wikipedia editing services in the UK https://www.wikicreation.co.uk/wikipedia-page-editing-services may guaranty that content is thoroughly examined, well cited, and compliant with Wikipedia's editorial guidelines.

InfoWorld AI 2026-09-04 09:42 UTC Score 42.0 USR-0126-20260904-global-ai-ne-635eff5a

OpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold

OpenAI launched GPT-6 Astra on Thursday, disclosing that the new flagship model has crossed the “Critical” threshold for cybersecurity risk under its Preparedness Framework, a classification the company said triggers additional deployment restrictions. “GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS,” OpenAI said in a statement. Enterprise administrators must manually enable Astra for their workspace, since access is off by default at launch, according to the company. Developers can access Astra in the API as gpt-6-astra or through Amazon Bedrock, OpenAI said, priced at $10 per million input tokens and $50 per million output tokens. Pro, Business, and Enterprise users also get a variant called Astra Pro, and the company said Astra supports Zero Data Retention for eligible API customers. Company claims perfect score on exploit benchmark OpenAI said it tested Astra without production safeguards on ExploitBench, and that the model scored 100%, up from 78.5% for predecessor GPT-5.6 Sol. On ExploitGym, a broader exploit-development benchmark, the company said Astra reached a 42.4% success rate against 30.3% for Sol, while using fewer output tokens. “Its ability to identify and develop zero-day exploits can help defenders find and patch weaknesses, but it also creates a need for stronger safeguards,” OpenAI said in t…

Simon Willison Weblog 2026-09-04 05:54 UTC Score 46.0 USR-0110-20260904-ai-specialis-5ab1b85a

August newsletter is out

The August edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here . This month: We got more details on OpenAl's accidental cyberattacks One-shotting Raccoon Heist games with Fable 5 and Sol 5.6 Claude auto mode Understanding ChatGPT Work Model releases Miscellaneous bits and bobs My projects What I'm using at the moment Here's a copy of the July newsletter as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy! Tags: newsletter

CIO AI 2026-09-03 23:52 UTC Score 64.0 USR-0125-20260903-global-ai-ne-7b0b5834

ChatGPT, Claude, and Grok all went down at once; enterprises need a backup plan

Enterprises are facing a disturbing new question in the age of AI: What happens when agentic assistants go dark? This became a very real scenario on Thursday, as OpenAI’s ChatGPT, Anthropic’s Claude, and SpaceXAI’s Grok near-simultaneously, and somewhat mysteriously, experienced significant, prolonged outages. Beginning in the morning, Eastern time, several ChatGPT models went down over a roughly two hour period, Claude models over a four-hour span, and Grok models for a near three-and-a-half hour duration. All three companies acknowledged the “elevated” issues and applied fixes. As users grumbled in forums and IT teams scrambled to get them back online, the incident revealed how hastily some organizations have adopted generative AI workflows without considering the potential, and inevitable, impact of widespread outages. AI agents are increasingly taking over automated and wider-scale workflows, and enterprises could find themselves “uncomfortably exposed” when AI hits the brakes, said technology analyst and journalist Carmi Levy . The situation should “serve as a wakeup call to IT leaders who have largely ignored what it’ll cost them if these increasingly critical platforms suddenly go dark. The risk is no longer hypothetical.” Hours-long outages impact core services ChatGPT went down on the same day as OpenAI’s anticipated launch of GPT-6 Astra , the new frontier model that the company says approximates artificial general intelligence (AGI) and gets nearer to its goal of…

Simon Willison Weblog 2026-09-03 20:18 UTC Score 55.0 USR-0110-20260903-ai-specialis-74293eba

GPT‑6 Astra

GPT‑6 Astra GPT-6 Astra is "rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS" - I've not tried it yet myself, so I don't have a great deal to say about it yet. It's going to be API priced at the same rate as Claude Fable 5 and 5.1: $10/million input and $50/million output. This is clearly OpenAI's Fable competitor, and appears to score higher than Fable on most of OpenAI's self-reported benchmarks. Most impressively, Astra scores 99.9% on the recent (released in March) ARC-AGI 3 benchmark - though notably Fable 5 does not yet have a published result, and the ARC-AGI blog notes that the 99.9% score was achieved for $19K using OpenAI's custom "Provider Adapter harness", while the default ARC-AGI harness scored 62.7% for $26K. The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. Unsurprisingly, given the recent Hugging Face incident , Astra is a beast at security tasks. It scores 100% on ExploitBench (GPT-5.6 Sol got 78.5%), 42.4% on ExploitGym (Sol got 30.3%), and 99.2% within four attempts on SRE-Bench binary reverse engineering compared to Sol's 68.7%. It's also better at long context: on OpenAI's eight-needle benchmark it got 100% at 256K–512K tokens and 96.3% at 512K–1M tokens. OpenAI may have vanquished one of the…

AWS Machine Learning Blog 2026-09-03 16:10 UTC Score 48.0 AI-057-20260903-official-ai--05128255

Set up OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock

Deploy a customer-operated LiteLLM gateway on Amazon ECS with AWS Fargate, connect it to an OpenAI model on Amazon Bedrock, and configure Codex to route requests through the gateway's Responses API with scoped identities, budgets, rate limits, and telemetry. We also compare direct IAM Identity Center access and a managed Portkey deployment.

The Verge AI 2026-09-03 15:35 UTC Score 54.0 AI-016-20260903-global-ai-ne-71b717fc

ChatGPT, Grok, and Claude all went down at the same time

OpenAI's ChatGPT, xAI's Grok, and Anthropic's Claude are back online after they all began experiencing issues around the same time on Thursday. At about 11AM ET, ChatGPT started returning error messages for users trying to use the chatbot, with its status page saying there were "elevated errors across ChatGPT and Codex." In addition to preventing […]

Cross Validated 2026-09-03 14:57 UTC Score 33.0 AI-113-20260903-social-media-69372178

How to check for temporal auto-correlation for each time series separately?

I have 2 time series in abundance that I'm modeling by a 2-level factor in mgcv , by "fSeason" (wet/dry), and following advice here , I should be splitting a test for auto-correlation by each factor level. I'm just not sure how to do that. Apologies for posting ChatGPT code, but I've been unsuccessful in my search everywhere else. DHARMa does have a test, but I was told it's only for a single time series. I don't think I can use the traditional AR(1) as the data are irregularly-spaced and I've never taken Bayesian statistics. Anyway, for now I'm just asking if the code AI generated is correctly identifying that I do have some, albeit weak, but significant temporal auto-correlation in the wet season. Any general advice is also welcome! Data here . library(mgcv) library(gratia) library(DHARMa) system.time( by $scaledResiduals scaledResiduals[Season_idx] res_Season $simulatedResponse simulatedResponse[Season_idx, , drop = FALSE] res_Season $observedResponse observedResponse[Season_idx] res_Season $fittedPredictedResponse fittedPredictedResponse[Season_idx] res_Season$nObs Site locations were haphazardly chosen (before my time) to monitor changes in fish abundance due to the construction of new canals (data posted here not collected with this statistical model/set of hypotheses in mind - its just available). Timing of surveys is based on tide level (need at least 60 cm/2 ft of water to survey near shore during the day (during reg. working hours). It takes about 4-7 days to surve…

The Guardian AI 2026-09-03 13:39 UTC Score 63.0 AI-021-20260903-global-ai-ne-0836bb0e

Black Box: The Chatbots | Happy Accident | Ep 3 – podcast

Why are AI chatbots pulling so many people down a rabbit hole? Our answer starts with the world’s first-ever chatbot, the strange effect it had on people and the ‘time bomb’ that exploded when ChatGPT was released four years ago. The result is a strange experiment we are all living through – whether we know it or not Studies and surveys cited in this episode: Amelia Miller’s interviews with AI researchers , engineers, product managers and executives. The 2022 Anthropic pre-print study that identified sycophancy as a behavioural trait of LLMs. And another pre-print paper the following year that found the way LLMs have been trained appeared to be increasing those sycophantic tendencies. How longer context windows appear to have contributed to “AI psychosis”. The Oxford pre-print study that found mass-market LLMs across the board had become more “relationship seeking”. Two-thirds of young people (aged 25-34) in the UK have turned to AI chatbots instead of loved ones to discuss emotional problems. ChatGPT may be the largest provider of mental health support in the US. Continue reading...

The Guardian AI 2026-09-03 12:42 UTC Score 43.0 AI-021-20260903-global-ai-ne-017c7898

Black Box: The Chatbots | Spirals | Ep 1 – podcast

Across the world, hundreds of people have come to believe they have made extraordinary scientific discoveries with AI chatbots such as ChatGPT, Claude and Gemini. Others say their AI has ‘awakened’, or is leading them to a higher spiritual realm. The Guardian journalist Michael Safi begins investigating a phenomenon labelled ‘AI psychosis’, travelling to the US to meet two people who have been on weird journeys with their chatbots Continue reading...

The Guardian AI 2026-09-03 10:00 UTC Score 53.0 AI-021-20260903-global-ai-ne-37c8a42c

Bill Simmons’s embrace of ChatGPT is a breach of his website’s creative spirit

The founder of The Ringer and Grantland was famous for discovering and developing writers. Now he is promoting a technology eroding the creative industries Bill Simmons may have stopped writing with any regularity in the mid-2010s – he claims his fingers don’t work – but that won’t stop him from reminding you of his greatest hits. On a recent podcast about Russell Westbrook’s retirement , he couldn’t stop quoting his old columns. He first referenced an article from the 2012 NBA finals, and wasn’t finished until 11 minutes into the podcast. His guest, The Ringer contributor Kirk Goldsberry, deserved an award for patience. But the biggest blight on the episode came later, in the form of a bizarre ad read. “This episode,” Simmons announced, “is brought to you by ChatGPT, the Bill Simmons Podcast AI sponsor.” (Emphasis mine; his overenthusiastic delivery mandates the italics.) Simmons said that he had asked ChatGPT what Westbrook’s defining game was. “That was a pretty good answer,” he said of ChatGPT’s offerings, as if it were a precocious, loveable nephew, before encouraging his listeners to try the platform. Continue reading...

AWS Machine Learning Blog 2026-09-02 21:22 UTC Score 51.0 AI-057-20260902-official-ai--ff5c29ca

Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference

Australian teams can now access OpenAI GPT-5.6 Sol, Terra, and Luna models on Amazon Bedrock with global cross-Region inference from the Asia Pacific (Sydney) and Asia Pacific (Melbourne) Regions. This post shows how to invoke the models, use prompt caching, set up Codex with OpenID Connect authentication, and monitor usage with Amazon CloudWatch.

The Decoder 2026-09-02 14:40 UTC Score 55.0 AI-168-20260902-regional-ai--c6dd82bd

US military adds ChatGPT and Grok to AI platform GenAI.mil

The Pentagon is expanding its AI platform with two new models, OpenAI's ChatGPT Mil and xAI's Grok for Government. The article US military adds ChatGPT and Grok to AI platform GenAI.mil appeared first on The Decoder .