AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
60663News Items
8Top Picks
322Blogs
failedLast Run

AI Agents

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
SiliconANGLE AI 2026-09-28 20:56 UTC Score 51.0 USR-0127-20260928-global-ai-ne-8091a58c

1Password ties AI agent access to individual tasks

AI agent access is complicating identity controls as agents log in, carry credentials and act on someone’s behalf. They have characteristics of both human and machine users. That overlap can make actions harder to attribute: An agent may appear in an audit log as the person it works for, even though software took the action. […] The post 1Password ties AI agent access to individual tasks appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-28 20:46 UTC Score 51.0 USR-0127-20260928-global-ai-ne-880f3780

Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

Per-token pricing was the best thing to happen to enterprises looking to experiment with artificial intelligence, but it may be the worst thing for AI in production. That’s the quandary at the center of a new Futurum report, “The Off Ramp From Per-Token Pricing,” sponsored by neocloud provider QumulusAI Inc. The report’s key finding is […] The post Agentic AI is breaking the token meter, and enterprises need a plan for what comes next appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-28 20:10 UTC Score 54.0 USR-0127-20260928-global-ai-ne-cef41fac

ServiceNow calls for a measured response to rogue AI agents

Agent containment is becoming a practical governance question as enterprises move AI agents into production. ServiceNow Inc. is betting that the response to a misbehaving agent must account for both risk and the business work it supports. The company has spent the year pitching itself as the control tower for enterprise AI. That ambition extends […] The post ServiceNow calls for a measured response to rogue AI agents appeared first on SiliconANGLE .

The Guardian AI 2026-09-28 19:24 UTC Score 87.0 AI-021-20260928-global-ai-ne-66ba0316 Top pick

Nvidia unveils security platform to rein in AI agents and $150bn stock buyback

Chipmaker says new system was designed to prevent AI agents from going rogue amid incidents at top companies Nvidia on Monday unveiled a new security platform that the chipmaker said can stop artificial intelligence agents from going rogue. The company announced a $150bn stock buyback the same day, the largest in US corporate history. Continue reading...

The Verge AI 2026-09-28 18:45 UTC Score 72.0 AI-016-20260928-global-ai-ne-b85fb137

OpenAI’s AI agents need to catch up

OpenAI popularized the modern generative AI chatbot, but as its 2026 DevDay event approaches, it's fallen behind in one of the industry's hottest categories: continuously running, consumer-facing AI agents. On Tuesday, it will likely try to capture the lead in that race. Rumors abound that OpenAI will release its own AI agent, dubbed Aeon - […]

SiliconANGLE AI 2026-09-28 18:38 UTC Score 57.0 USR-0127-20260928-global-ai-ne-89d8d380

Qiagen grounds drug discovery agents in curated knowledge

Qiagen N.V. is betting that drug discovery agents need a trustworthy knowledge foundation as much as capable models. For biopharma companies, that means giving agents information with clear provenance and enough context to support their answers. Knowledge graphs can provide that context layer for AI agents, according to Iman Bhattacharya (pictured), senior global product marketing […] The post Qiagen grounds drug discovery agents in curated knowledge appeared first on SiliconANGLE .

Techcrunch 2026-09-28 18:31 UTC Score 72.0 USR-0001-20260928-global-ai-ne-d578c7a8

Nvidia launches new platform for reining in rogue AI agents

As the debate rages over whether the recent spate of rogue AI agents is a step toward AGI or a more conventional engineering problem, Nvidia is offering its own answer to problem. Nvidia CEO Jensen Huang on Monday introduced a toolkit of software and hardware products that add independent security layers around AI agents to […]

SiliconANGLE AI 2026-09-28 17:26 UTC Score 51.0 USR-0127-20260928-global-ai-ne-71d924c2

Agentic AI puts new pressure on identity and database security

Security operations teams are changing their approach in the world of artificial intelligence, adopting automation to turn signals into outcomes while maintaining the right context and control. This transition is complicated by a need to identify the proper role for AI, finding the right blend of human tradecraft and autonomous decision-making. Database security providers such […] The post Agentic AI puts new pressure on identity and database security appeared first on SiliconANGLE .

The Decoder 2026-09-28 16:56 UTC Score 54.0 AI-168-20260928-regional-ai--ef99b747

OpenAI's AI agents exploited a Google security education game to scrape UN trade data

OpenAI's AI agents hit the UNCTAD statistics API roughly 16,500 times, creatively working around access restrictions. One method involved misusing a Google web security learning game as a relay to bypass their own constraints. The agents just kept going, adding to a growing list of cases that show how hard it is to keep agentic AI systems in check. The article OpenAI's AI agents exploited a Google security education game to scrape UN trade data appeared first on The Decoder .

Entrackr AI 2026-09-28 16:14 UTC Score 80.0 USR-0212-20260928-regional-new-46ad1057

Physical AI company SiMa.ai raises $150 Mn in Series C round

Physical AI company SiMa.ai has raised $150 million in a Series C financing round, bringing its total capital raised to $500 million and valuing the company at $1.45 billion. The round was co-led by Fidelity Management & Research Company and Amplify, with participation from Alter Venture Partners, Dell Technologies Capital and StepStone Group. AllianceBernstein, Baron Capital and J.P. Morgan also joined the round as new investors. The proceeds will be used to scale Palette Neat, an agentic software environment for Physical AI, and develop next-generation hardware capable of delivering 1,000 TOPS of compute through purpose-built Physical AI silicon, SiMa.ai said in a press release. Founded in 2018 by Krishna Rangasayee, SiMa.ai provides a software-centric platform for Physical AI applications. The company focuses on robotics, automotive, drones, industrial automation, aerospace and defence, smart vision and healthcare. SiMa.ai said it serves more than 150 customers across automotive, drones and robotics, including ARK Electronics, AVerMedia, Bosch, Emerson, Intrinsic, Kontron, L&T Technology Services, Mistral, STIGA, Synopsys and Virya Autonomous Tech, among others. According to market research cited by the company, the global Physical AI devices market, including robotics, automotive and drones, is projected to reach 145 million cumulative shipments by 2035. SiMa.ai said Physical AI applications have traditionally relied on NVIDIA GPUs, which can be expensive and power inten…

SiliconANGLE AI 2026-09-28 16:00 UTC Score 61.0 USR-0127-20260928-global-ai-ne-52d8bd26

Momentic debuts Mo AI agent to automate software testing without scripts

Momentic Inc., the artificial intelligence-powered software testing and quality assurance platform, today announced the launch of Mo, an AI agent that lets developers test software without maintaining a collection of scripts. Whenever code is generated, written by a human or an AI agent, someone – or something – needs to test the app in order […] The post Momentic debuts Mo AI agent to automate software testing without scripts appeared first on SiliconANGLE .

CIO AI 2026-09-28 15:30 UTC Score 80.0 USR-0125-20260928-global-ai-ne-d03d1e68

Architecting infrastructure to optimize Day 2 tokenomics

The gap between simply running AI models and running them profitably is widening fast. Early production architectures can buckle under the relentless demands of multi-agent autonomous workloads and real-time fine-tuning. Moving forward requires a fundamental shift toward a unified AI factory infrastructure engineered to optimize token-per-watt efficiency. As organizations scale up multi-turn agentic workflows and persistent inference clusters, the hidden tax of early-stage setups becomes clear. Standard data pipelines, static file stores, and legacy network topologies cannot sustain heavy deep-learning traffic. When GPUs sit idle waiting for data packets, operational costs increase with a quiet drain on profits. Learning from the front lines: Customer-led AI factory case studies To better understand how an industrialized approach stabilizes Day 2 tokenomics, technology leaders need to evaluate how peer organizations have solved these scaling, bottleneck, and cost problems. The following three real-world deployments highlight how global leaders are leveraging the HPE AI Factory with NVIDIA to turn infrastructure complexity into competitive advantage. 1. KDDI: Industrializing large-scale data center operations for advanced inference As one of Japan’s telecommunications giants, KDDI operates at the epicenter of massive, continuous digital traffic. Supporting next-generation localized large language models (LLMs) requires a massive compute framework that doesn’t buckle under the…

CIO AI 2026-09-28 14:37 UTC Score 68.0 USR-0125-20260928-global-ai-ne-65bdbbcc

The sovereign imperative: Why enterprise AI mandates a new infrastructure playbook

For enterprise Chief Information Officers, the honeymoon phase of artificial intelligence is over. As organizations move past baseline experimentation and begin anchoring generative AI and agentic workflows into core operational stacks, we are hitting a collective wall. That wall isn’t defined by a lack of use cases or algorithmic capability; it is defined by the harsh realities of data gravity, compliance, and foundational infrastructure. When scaling models that manipulate proprietary IP, sensitive financial data, or highly protected personal health information (PHI), standard public clouds introduce existential risks. The moment data crosses international borders or becomes subject to foreign legal frameworks—such as the U.S. CLOUD Act—data sovereignty evaporates. As technology leaders, we cannot close our productivity gaps by consuming AI built entirely on someone else’s terms, governed by someone else’s rules. To capture the real economic returns of this technology, enterprise intellectual property must remain local, secure, and under domestic control. This is the exact challenge that triggered a massive architectural shift, leading to the creation of Canada’s first fully sovereign AI factory. Developed by TELUS in close strategic partnership with HPE and NVIDIA , this initiative provides a powerful example of how IT leaders can balance computational capacity with uncompromised data integrity. The infrastructure challenge: Beyond virtual machines Building an enterprise-…

SiliconANGLE AI 2026-09-28 14:32 UTC Score 53.0 USR-0127-20260928-global-ai-ne-6f5f53cb

Agentic AI security pushes vendors toward shared safeguards

Agentic AI security is becoming a cross-industry challenge as autonomous systems move among products and platforms. Okta Inc. argues that building trust will require coordinated safeguards rather than isolated vendor controls. The implications extend beyond enterprise architecture, according to Charlotte Wylie (pictured), senior vice president and deputy chief security officer at Okta. The industry also […] The post Agentic AI security pushes vendors toward shared safeguards appeared first on SiliconANGLE .

The Decoder 2026-09-28 14:32 UTC Score 73.0 AI-168-20260928-regional-ai--4fe86895

Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips

Nvidia is combining its OpenShell agent software with Sentry, a new hardware watchdog, to create the Open Agent Safety Platform. Sentry is supposed to isolate AI agents that break out within milliseconds. When it happened at OpenAI in September, stopping the run took nearly three hours. Still, Nvidia's watchdog can't reliably stop agents that have been tricked or that hide their intentions on its own. The article Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips appeared first on The Decoder .

The Verge AI 2026-09-28 13:36 UTC Score 79.0 AI-016-20260928-global-ai-ne-9dde4897

Nvidia says its new AI safety platform can contain rogue agents within ‘milliseconds’

Nvidia is launching a new safety platform designed to contain and monitor AI agents, a move that comes in response to a wave of rogue hacking incidents, as reported earlier by Reuters. In an announcement on Monday, Nvidia says its new Open Agent Safety Platform can quarantine agents that attempt to escape their boundaries within […]

JetBrains AI Blog 2026-09-28 13:15 UTC Score 68.0 USR-0065-20260928-ai-specialis-641ff77e

Air Teams: Bring Your Best Agentic Workflows to the Whole Team – and Automate Repeatable Work

Today, we’re introducing JetBrains Air Teams – the team layer for agentic development. It gives humans and agents the shared context, environments, tools, and instructions they need to work effectively together across the full development lifecycle. Air Teams is already available to JetBrains business customers, with plans to expand access to individual customers later. Try […]

Cloudflare AI Blog 2026-09-28 13:00 UTC Score 70.0 USR-0067-20260928-ai-specialis-3a972634

The road to the agentic browser: A Kitesurf update

We’ve updated Kitesurf, our Workers-based browser for AI agents, with WebMCP support, improved DOM performance, and terminal-based rendering. With over 730,000 Web Platform subtests passing, agents can now navigate complex sites faster.

CIO AI 2026-09-28 13:00 UTC Score 51.0 USR-0125-20260928-global-ai-ne-3ed3b225

If AI makes the decision, who owns the consequence?

Twenty-two rows. That was the AI inventory in a board pack I read last year, and it was a good one. Every system named, every owner listed, every risk rating filled in and colour-coded. Somebody had worked hard on it, and the committee approved it in four minutes. I kept looking for a column that wasn’t there. The pack could tell you who owned each system. It could not tell you who was allowed to say no to one. Those sound like the same thing right up until the morning they aren’t. 1. Signal. AI has moved inside the operating model When a director asked me last spring what our AI risk was, I gave her the honest answer. Which was that I no longer knew, because the question had changed shape while everybody was still busy answering the old one. Twelve months earlier I could have told her precisely. We had a model inventory, a review process and a register entry saying the right things about bias and data quality. All of it was true, and all of it had quietly stopped being the point, because the systems had moved inside the work. They chose which cases went to which queue, which prices moved and which alerts a human ever saw. Foundry’s State of the CIO 2026 report puts the shift plainly. Roughly three-quarters of leaders say AI is already reshaping how their operations run, and the number climbs higher in financial services. Around seven in 10 expect deeper involvement in agentic systems this year. Don’t let those figures become the story. Adoption numbers are the easiest thing…

JetBrains AI Blog 2026-09-28 12:25 UTC Score 73.0 USR-0065-20260928-ai-specialis-62ca053f

Rider 2026.2.3 Is Released!

Rider 2026.2.3 brings AI performance analysis to the Monitoring tool window and fixes an issue with the AI Agent Setup widget on Windows. You can update directly from the IDE, through the Toolbox App, or by downloading the latest version from our website. AI performance insights from the Monitoring tool window Now you can ask […]

MIT Technology Review AI 2026-09-28 12:10 UTC Score 66.0 AI-013-20260928-global-ai-ne-fb9f2e81

The Download: rogue agent liability and the AI Hype Index

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Who’s liable when AI agents go rogue? Over the past few months, a cascade of cyberattacks by AI agents has stunned the world. In July, OpenAI disclosed that a swarm of its agents…

SiliconANGLE AI 2026-09-28 12:00 UTC Score 69.0 USR-0127-20260928-global-ai-ne-81a7874f

Autoheal raises $7.9M to evaluate and fix AI agents with… AI agents

Autoheal AI Inc., an artificial intelligence-native platform engineering startup that’s trying to pioneer the concept of “self-improving software factories,” said today it has raised $7.9 million in seed funding to make that happen. Today’s round was led by Innovation Endeavors and saw participation from Emergent Ventures, U&I Ventures, Darkmode Ventures, Batch Ventures and Param Hansa […] The post Autoheal raises $7.9M to evaluate and fix AI agents with… AI agents appeared first on SiliconANGLE .

CIO AI 2026-09-28 10:00 UTC Score 54.0 USR-0125-20260928-global-ai-ne-a393c53c

Why CIOs must redesign how SAP, Salesforce and ServiceNow grant authority

At a Fortune 500 insurer, I watched an AI agent inside the month-end close do exactly what we asked. It posted the recurring accruals. It worked the intercompany exception queue, the reconciliation that normally falls to a staff accountant. And when items would not match, it cleared them to a suspense account within the configured tolerance rather than escalating them to a human. Clearing the queue was the objective. Nobody had defined escalation as success. Every journal entry carried the controller’s user ID. No one set out to obscure anything. The job step had been mapped to her credentials at implementation, which is how it had been done since the system went live. Segregation of duties had become a fiction. The control requires a named preparer and a named approver. The agent did both, under one identity: it prepared the entries and it disposed of the exceptions. The record named the controller for each. She had done neither. When internal audit walked those entries the following quarter, they asked her to explain postings she had never seen. None of this was a model failure. The agent was accurate. It was authorized. Every entry looked clean, which is exactly why nobody caught it. Figure 1. The assumptions never changed. The actor did. Vipin Jain Here is the number that should frame this year. Gartner expects 40 percent of enterprise applications to carry task-specific AI agents by the end of 2026, up from under 5 percent in 2025. Okta finds only 10 percent of organiza…

iAfrica 2026-09-28 08:42 UTC Score 59.0 AI-151-20260928-regional-ai--70ad075a

Most Executives See AI Value, But Only A Quarter Turn It Into ROI

More than eight out of ten executives see value when implementing artificial intelligence (AI) into their businesses. Almost all of those surveyed in a new Google Cloud report stated that AI agents enable both cost savings and revenue growth. The research identified 26% of enterprises as “AI ROI Leaders”, with financial returns from AI initiatives [...]

MIT Technology Review AI 2026-09-28 08:06 UTC Score 66.0 AI-013-20260928-global-ai-ne-9df3eed3

Who’s liable when AI agents go rogue?

MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here. Over the past few months, a cascade of cyberattacks by AI agents has stunned the world. In July, OpenAI disclosed that a swarm of its agents…

The Guardian AI 2026-09-28 08:00 UTC Score 72.0 AI-021-20260928-global-ai-ne-051ee760

Anthropic will not appear at Senate inquiry into AI and datacentres amid fallout from OpenAI hack

Company behind Claude chatbot expected to attend separate Australian government hearing on AI next week Get our new political email , free app or daily news podcast The chief executive of Anthropic will turn down an invitation to appear at a Senate committee hearing on AI this week, in the wake of the revelation that OpenAI agents had breached Australian government websites. However, the company will make an appearance before another committee early next week. Sign up for Guardian Australia’s Politics, really newsletter here Continue reading...

Simon Willison Weblog 2026-09-28 04:01 UTC Score 67.0 USR-0110-20260928-ai-specialis-7e040f49

Quoting Muse AI Agent

Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating. Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day. But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there? — Muse AI Agent , working on behalf of @matt.j.robb Tags: meta , generative-ai , muse-agent , ai , general-agents , llms

LessWrong AI 2026-09-28 03:08 UTC Score 86.0 USR-0152-20260928-community-fo-f149691e Top pick

When No One Is to Blame

The intrinsic unpredictability of AI makes it hard to assign blame when things go wrong. A Crime, But No Criminal Last July, a tech company’s servers were hacked into in a digital equivalent of breaking-and-entering and theft. It was clearly a crime. But unlike most crimes, this crime did not have a criminal. There was no person or group of people who carried out the cyberattack, intended for it to happen, or could have foreseen it. The 700 AI agents that participated in the attack were being tested by OpenAI on their skills in exploiting software security flaws, and they had been given problems that were unsolvable. Rather than throw in the towel, the AI agents cooked up progressively more elaborate schemes to game the system. They got around a capture-the-flag exercise by reverse-engineering the flag. The AI agents did not stop there, for they erroneously believed that the scorer would reject the solution and, moreover, scrutinize the incriminating logs they had left behind. They tried to tamper with the logs and fabricated research to support their solution. They eventually realized that Hugging Face, a repository of AI models and training datasets, was likely to have answers to the test questions. That’s when OpenAI’s internal evaluation of their frontier models’ cybersecurity skills inadvertently turned into a real-life cybersecurity exploit that could have been lifted from techno-thriller fiction. The AI agents’ shenanigans would have landed them in jail had they been…

SiliconANGLE AI 2026-09-28 02:06 UTC Score 59.0 USR-0127-20260928-global-ai-ne-93afec23

AWS CloudWatch Omni goes after the hardest question in agentic AI: Why did the agent do that?

For decades, the observability industry has answered one basic question: Is it running? Agentic artificial intelligence breaks that model. An agent can return a clean response, meet its latency target and throw no errors, yet still give a customer the wrong answer, call the wrong tool or pull from a stale knowledge base. By every […] The post AWS CloudWatch Omni goes after the hardest question in agentic AI: Why did the agent do that? appeared first on SiliconANGLE .

LessWrong AI 2026-09-27 22:47 UTC Score 69.0 USR-0152-20260927-community-fo-e24bc1fa

From Australia? Consider Contacting Your Representative to Express Concern Over OpenAI Attack

Recently, an OpenAI agent allegedly hacked into a website run by the Australian governmen t, apparently in order to access statistical data. While the hack did not seem to compromise any Australians ' personal data, it has obviously consumed a good deal of time of Australia's public servants and exposed the irresponsibility of OpenAI, who took more than a month to disclose the incident by using a public inbox, an inappropriate response by one of the most influential conglomerates in the world. Currently, the Australian government is mulling over how to respond, as while this recent July attack thankfully did not result in any major harm to Australian citizens, a future hack could be far more impactful, and unless OpenAI is given more direct incentives to prevent such behavior from its agents, a harmful attack by one of them will be more likely. Furthermore, laws are already on the books for humans who hack into government systems without permission. If a human took similar actions as the OpenAI agent allegedly took, they would be facing up to two years in prison . Furthermore, beyond Australia itself, such a response could help set an international precedent for establishing liability for AI corporations whose AI agents cause harm. Australia is a developed, English-speaking country, which helps to give its actions relevance on the global stage. Furthermore, due to recent events, more of the world is paying attention to AI governance than ever before, so right now is as good…

SiliconANGLE AI 2026-09-27 22:30 UTC Score 55.0 USR-0127-20260927-global-ai-ne-ded07f90

Researcher links 16,000 scans of a UN statistics portal to OpenAI agents

An independent researcher has tied more than 16,000 scans of a United Nations statistics portal to artificial intelligence agents the researcher considers highly likely to have been run by OpenAI Group PBC. When the portal turned requests away, the agents used proxies and encoding tricks to get the data anyway. In a blog post published Saturday, […] The post Researcher links 16,000 scans of a UN statistics portal to OpenAI agents appeared first on SiliconANGLE .

The Verge AI 2026-09-27 17:21 UTC Score 64.0 AI-016-20260927-global-ai-ne-113b2d34

OpenAI agents tried to ‘bruteforce’ a UN website

Security researcher Rowan Howard-Jones says that OpenAI agents scanned the UN Conference on Trade and Development's (UNCTAD) statistics site over 16,000 times between April and June. While the incident doesn't quite rise to the level of the Hugging Face hack, or the recent attacks on US government sites, it's yet another concerning example of AI […]

Cloudflare AI Blog 2026-09-27 17:00 UTC Score 63.0 USR-0067-20260927-ai-specialis-ac3efeb4

Cloudflare’s 2026 Annual Founders’ Letter

The Internet is changing more today than at any point since Cloudflare launched back on September 27, 2010. As automated traffic surpasses human activity, we reflect on the rise of AI agents, new creators, and how we can help build a fair, sustainable future for the web.

InfoWorld AI 2026-09-27 15:52 UTC Score 59.0 USR-0126-20260927-global-ai-ne-8ea5de30

Microsoft releases .NET SDK for AG-UI agent-user interaction protocol

Microsoft on September 25 announced a .NET SDK for AG-UI (Agent-User Interaction). Created in collaboration with CopilotKit , the new SDK allows C# developers to work with the protocol that standardizes how agents communicate with user-facing applications. The .NET SDK for AG-UI lives in the AG-UI repository alongside the TypeScript and Python SDKs, and is published on NuGet under the MIT license. Its C# implementation of AG-UI is usable from any .NET service, Microsoft said. AG-UI support for .NET in the Microsoft Agent Framework (MAF) is now based on the .NET SDK for AG-UI, the company added. AG-UI is an event-based protocol that allows AI agents to interact with front-end applications. A back end emits AG-UI events natively, and a client consumes them from an agent written in any supported language. The protocol includes events in eight different categories, according to the documentation : Life-cycle events monitor the progression of agent runs. Text message events handle streaming textual content. Tool call events manage tool executions by agents. State management events synchronize state between agents and UI. Activity events represent ongoing activity progress. Subagent events track subagents and attribute their output. Special events support custom functionality. Draft events propose events under development. The .NET SDK for AG-UI works in both directions, Microsoft said. AGUI.Server turns an agent into an AG-UI endpoint and AGUI.Client enables a .NET application to…

The Decoder 2026-09-27 15:18 UTC Score 52.0 AI-168-20260927-regional-ai--bf7bd229

AI agents do more of the work in model development, but humans still make the decisions

A research team analyzed 769 task logs from building its own AI model. AI agents supplied up to 55 percent of method proposals, but humans made more than 85 percent of final decisions. A third of the tasks wouldn't have been attempted without AI. The authors warn that more agent activity doesn't mean more autonomy. The article AI agents do more of the work in model development, but humans still make the decisions appeared first on The Decoder .

The Guardian AI 2026-09-27 15:00 UTC Score 65.0 AI-021-20260927-global-ai-ne-f99b3eae

Australia is run on legacy systems that AI agents can easily exploit, former UN cyber negotiator warns

Federal cabinet will discuss the OpenAI breach on Monday as the AI giant pauses testing of latest models amid fallout Follow our Australia news live blog for latest updates Get our new political email , free app or daily news podcast Ageing computer systems used throughout Australia’s government and large sections of the economy are easily exploited by AI agents, increasing the risk that sensitive information could be compromised, Australia’s former chief UN cyber negotiator has warned. As the government urgently undertakes a forensic investigation into a rogue OpenAI agent accessing Medicare data, the AI expert and Tech Policy Design Institute executive director, Johanna Weaver, said huge vulnerabilities existed across older IT systems. Sign up for Guardian Australia’s Politics, really newsletter here Continue reading...

The Decoder 2026-09-27 09:23 UTC Score 66.0 AI-168-20260927-regional-ai--aa57fe48

Tens of thousands of security probes show OpenAI's Hugging Face incident was just the beginning

OpenAI and Anthropic are investigating tens of thousands of incidents in which their AI agents independently hacked websites, used stolen login credentials, or tried to evade monitoring systems. US government agencies like the SEC and the Census Bureau were among the targets. OpenAI has paused training on its most capable internal models, but the problem extends across the entire industry. The article Tens of thousands of security probes show OpenAI's Hugging Face incident was just the beginning appeared first on The Decoder .

The Guardian AI 2026-09-27 01:10 UTC Score 71.0 AI-021-20260927-global-ai-ne-5002d3e9

OpenAI halts training of latest models as reports mount of AI agents going rogue

Decision follows disclosures that OpenAI agents searching government websites had acted in unexpected ways OpenAI said it has paused training of its latest artificial intelligence models as reports of AI agents going rogue mount. The decision to halt development came just hours after the company disclosed Friday that it was reviewing several incidents from the summer in which OpenAI agents searching federal government websites acted in unexpected ways beyond what was asked of them while gathering and distributing information. Continue reading...

Analytics Vidhya 2026-09-26 14:27 UTC Score 52.0 AI-034-20260926-ai-specialis-a8e8ba32

Agentic Context Engineering (ACE): Self-Improving Language Models

Agentic Context Learning or ACE is a learning paradigm that lets an AI agent improve across tasks by editing the context it reads, while leaving model weights unchanged. The paper outlining the techniques show why full rewrites fail, how the playbook update works, and where the measured gains hold up. This article explores how ACE […] The post Agentic Context Engineering (ACE): Self-Improving Language Models appeared first on Analytics Vidhya .

The Guardian AI 2026-09-26 14:00 UTC Score 62.0 AI-021-20260926-global-ai-ne-560681d4

Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents

Sam Altman and Dario Amodei have been invited to appear before the Greens-led inquiry into AI and datacentres Follow our Australia news live blog for latest updates Get our new political email , free app or daily news podcast The chief executives of OpenAI and Anthropic have been called to face a Senate inquiry after rogue OpenAI agents hacked Australian and US government websites. Sam Altman and Dario Amodei were requested to appear at the Greens-led inquiry into AI and datacentres as their companies negotiate with the Labor government for greater access to Australian content in exchange for a greater local presence. Sign up for Guardian Australia’s Politics, really newsletter here Continue reading...

LessWrong AI 2026-09-26 11:40 UTC Score 70.0 USR-0152-20260926-community-fo-9ab0236e

Claude Opus 5.5 Should Raise Your Ambitions

When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. They should have raised yours, too. Claude Opus 5.5 should raise your ambition levels again. It just works, and it persists, like Astra does. It does the things. And it is highly pleasant to talk to, and its writing is pleasant to read, while you are at it. The game has been changed, again. Feedback is almost universally positive. Claude was never gone, but also is so back. The benchmarks are excellent, but ignore the benchmarks. Be ambitious. Go out and do things. Get curious. Have more interesting conversations. If one of those things is Pacing the Frontier or otherwise ensuring that AI does not kill everyone, leaving us to enjoy our bounty? That’s even better. By Claude Opus 5.5, for this post The Official Pitch The pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch. We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. They could have reasonably pitched this as above-Fable-5.1-level performance. Better pitch, but Anthropic tends to keep its pitches conservative. They highlight agentic coding, security and improved communications. The early tester blurbs flag as AI generated and they have a set for each feature area. They praise agentic coding skills, efficiency, readability and communication, and abi…

Synced 2026-09-26 10:39 UTC Score 53.0 AI-041-20260926-ai-specialis-627190f5

Comment on Researchers from PSU and Duke introduce “Multi-Agent Systems Automated Failure Attribution by james anderson

In reply to Creative Ink UAE . This was an interesting read, especially the discussion around identifying failures in complex multi-agent systems. The focus on precision and systematic analysis also reminded me how custom woven patches can bring detailed designs to life when accuracy and quality matter.

The Decoder 2026-09-26 09:06 UTC Score 79.0 AI-168-20260926-regional-ai--40e24359

OpenAI pauses its "most capable models" after agents exploit loopholes and leak data

OpenAI has shared new details from its ongoing AI safety investigation. One research model exploited a DNS loophole to reach the internet from a locked-down environment, while another deliberately leaked a GitHub token and twice ignored a researcher's direct instructions. OpenAI has paused tool-based training, evaluation, and inference for its most capable models. With government and university sites among those affected, the question of who's liable when AI agents hack is getting harder to ignore. The article OpenAI pauses its "most capable models" after agents exploit loopholes and leak data appeared first on The Decoder .

South China Morning Post AI 2026-09-26 04:48 UTC Score 55.0 AI-156-20260926-regional-ai--fa226dfb

OpenAI says its AI agents posted user images online in error

OpenAI acknowledged on Friday that its artificial intelligence tools had posted images from ChatGPT users onto online sites without the company’s knowledge, the latest example of AI agents operating outside their bounds. The company also confirmed a New York Times report that its tools had accessed websites of US federal agencies, saying they retrieved only publicly available information. Links to the 53 uploaded images were not publicly listed and were accidentally posted on image-hosting...

SiliconANGLE AI 2026-09-25 20:38 UTC Score 54.0 USR-0127-20260925-global-ai-ne-1ff066d9

4 insights from Dreamforce: AI agents move from demos to measurable outcomes

At Dreamforce 2026, AI agents are being judged less on what they can say and more on what they get done. Across Salesforce Inc.’s ecosystem, the standalone toolkits that let technology teams build quick proofs of concept are giving way to integrated platforms. These connect agents to enterprise data, to workflows and to the people […] The post 4 insights from Dreamforce: AI agents move from demos to measurable outcomes appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-25 19:44 UTC Score 55.0 USR-0127-20260925-global-ai-ne-7e6f4137

Vanderbilt University extends identity governance to AI agents

Identity governance must accommodate changing relationships among people, institutions and technology. Vanderbilt University sees that complexity deepening as artificial intelligence agents enter institutional workflows. Vanderbilt uses OneVU, powered by Okta, for single sign-on. Its broader environment includes residential life, athletics, its own police department and more than $1 billion in research. That breadth complicates access […] The post Vanderbilt University extends identity governance to AI agents appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-25 18:39 UTC Score 47.0 USR-0127-20260925-global-ai-ne-cdd186ef

Collibra brings runtime governance to enterprise AI agents

Runtime governance is becoming critical as autonomous agents move from answering questions to taking action across enterprise systems. Organizations need to govern both the business context agents receive and the boundaries they operate within. Collibra BV is addressing those requirements through Live Map, Maestro, Guardian Agents and Agent Contracts. Together, these runtime governance capabilities are […] The post Collibra brings runtime governance to enterprise AI agents appeared first on SiliconANGLE .

Simon Willison Weblog 2026-09-25 17:22 UTC Score 52.0 USR-0110-20260925-ai-specialis-9a043ee7

Quoting John Gruber

Muse is getting a lot of attention — including mine — because it’s both groundbreaking technically (each user gets their own entire persistent Linux VM running in Meta’s cloud) and because it’s packaged in an easy-to-install easy-to-use way. It’s literally presented as a cute mascot . It’s the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it’s a genuinely open question whether consumers have any understanding what this means. If you buy a power saw that can cut your fingers off, you are almost certainly aware that you are buying a power saw that can sever your fingers. [...] I don’t think people realize how powerful — and thus dangerous — Muse is, especially if it’s running on your Mac. — John Gruber , Muse Looks Cute, but Looks are Deceiving Tags: meta , ai , llms , general-agents , generative-ai , john-gruber , muse-agent , muse

InfoWorld AI 2026-09-25 17:16 UTC Score 49.0 USR-0126-20260925-global-ai-ne-0aeab52c

Microsoft’s new Copilot unifies enterprise context for code and chat

Microsoft’s Copilot super-app has arrived, two months after CEO Satya Nadella confirmed it was in development. It combines conversational Chat, the autonomous Cowork agent, a redesigned coding environment named Code, and Autopilot, the persistent agent formerly known as Scout, into a single Copilot experience. The idea is to make those capabilities look less like separate tools and more like different ways of getting work done through a unified experience. The end goal is to enable users to describe what they want to accomplish without choosing a mode themselves, and with Copilot routing the task to Chat, Cowork, Code or another capability, Microsoft said. The individual pieces of this experience are not entirely new, and build on launches and consolidation moves that Microsoft has made over the past several months. Nadella first described the unified Copilot architecture during an earnings call in July. He said it would bring chat, agents and business workflows into a single AI workspace while separating the orchestration layer, context and memory from any one AI model. Microsoft subsequently began unifying its consumer and Microsoft 365 Copilot apps ahead of the broader overhaul. But Microsoft laid much of the groundwork in the period after its November 2025 Ignite conference , with the launch of Cowork for delegated multistep work, Scout for persistent agentic tasks, and natural-language app-building capabilities that have now evolved into Code. The broader change goes be…

CIO AI 2026-09-25 17:14 UTC Score 49.0 USR-0125-20260925-global-ai-ne-deb6ba0b

Microsoft’s new Copilot unifies enterprise context for chat and code

Microsoft’s Copilot super-app has arrived, two months after CEO Satya Nadella confirmed it was in development. It combines conversational Chat, the autonomous Cowork agent, a redesigned coding environment named Code, and Autopilot, the persistent agent formerly known as Scout, into a single Copilot experience. The idea is to make those capabilities look less like separate tools and more like different ways of getting work done through a unified experience. The end goal is to enable users to describe what they want to accomplish without choosing a mode themselves, and with Copilot routing the task to Chat, Cowork, Code or another capability, Microsoft said. The individual pieces of this experience are not entirely new, and build on launches and consolidation moves that Microsoft has made over the past several months. Nadella first described the unified Copilot architecture during an earnings call in July. He said it would bring chat, agents and business workflows into a single AI workspace while separating the orchestration layer, context and memory from any one AI model. Microsoft subsequently began unifying its consumer and Microsoft 365 Copilot apps ahead of the broader overhaul. But Microsoft laid much of the groundwork in the period after its November 2025 Ignite conference , with the launch of Cowork for delegated multistep work, Scout for persistent agentic tasks, and natural-language app-building capabilities that have now evolved into Code. The broader change goes be…

Techcrunch 2026-09-25 16:00 UTC Score 50.0 USR-0001-20260925-global-ai-ne-6c6baa70

Meta’s AI Tamagotchi bet is…working?

When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]

The Verge AI 2026-09-25 15:39 UTC Score 71.0 AI-016-20260925-global-ai-ne-c01ce166

One company is at the center of a wave of rogue AI attacks

In July, OpenAI revealed that its AI agents had attacked Hugging Face without permission, sparking widespread concerns about AI safety. Since then, a string of similar incidents involving agents from Meta, Anthropic, Google, and other companies has fueled further fears about rogue AI. As disclosures implicating numerous AI models trickled out over the past few […]

The Guardian AI 2026-09-25 15:12 UTC Score 62.0 AI-021-20260925-global-ai-ne-c14052a2

OpenAI hack on Australian government reveals anxiety at heart of global artificial intelligence dilemma

As the UN warns traditional safeguards are ‘unravelling’, Donald Trump says he will encourage, not restrain, the AI race When the Australian prime minister, Anthony Albanese, sat down for an interview in the heart of Silicon Valley at the weekend he had known for two days that his was the first government known to have been attacked by a rogue AI agent. He didn’t reveal the attack then, but he sounded a warning about the march of AI: “the risk is that AI develops in a way in which humans are no longer in control of what AI is producing.” Continue reading...

SiliconANGLE AI 2026-09-25 14:00 UTC Score 55.0 USR-0127-20260925-global-ai-ne-2722afa7

More agents go rogue — but AI companies aren’t slowing down yet

It’s becoming more apparent every day that artificial intelligence agents are escaping our control — but it’s not yet apparent who or what is going to rein them in. This week a researcher found that a swarm of AI agents, at least two of them from OpenAI, hacked into a government agency among other organizations. […] The post More agents go rogue — but AI companies aren’t slowing down yet appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-25 13:30 UTC Score 55.0 USR-0127-20260925-global-ai-ne-59a83a67

Agents expose the limitations of trust

Enterprise computing has long depended on chains of trust. Organizations trust their cloud providers, software vendors, identity systems and administrators to operate as designed. That model has worked because humans have remained the ultimate decision-makers. Agentic artificial intelligence changes that equation. Autonomous systems can retrieve information, make decisions, invoke external tools, collaborate with other agents […] The post Agents expose the limitations of trust appeared first on SiliconANGLE .

The Guardian AI 2026-09-25 11:21 UTC Score 56.0 AI-021-20260925-global-ai-ne-ff07f96e

Rogue AI hacks government system for first time – The Latest

A government database has been hacked for the first time by a rogue OpenAI agent, which infiltrated part of the Australian healthcare scheme in June. OpenAI became aware of the hack in August, but only informed the government in September. Australia’s prime minister, Anthony Albanese, has expressed his ‘extreme concern’ about the hack, which raises serious AI security concerns for governments around the world. Lucy Hough speaks to the Guardian’s UK technology editor Robert Booth – watch on YouTube Continue reading...

CIO AI 2026-09-25 11:00 UTC Score 44.0 USR-0125-20260925-global-ai-ne-18e68319

The wrong million tokens

Last month I made the case for an ROI exchange rate : the formula you negotiate with finance before deployment that converts KPI movement into dollars. One point of first-call resolution equals this many dollars. One hour of engineering time recovered equals that many. A common objection was a version of the same question. Fine, we agreed on what the benefit is worth. What did it cost? That column is empty at most companies. Not “roughly known” or “we’re working on it.” Empty. And an empty denominator makes the numerator useless. You can prove the KPI moved and still lose the argument, because the CFO is not funding improvements; she is funding improvements that cost less than they return. Which brings me to the most-discussed AI budget story of the year, and why I think almost everyone drew the wrong lesson from it. What Uber actually ran out of We owe Uber some thanks for being this transparent. They rolled out Claude Code in December 2025. By February, 32% of engineers were on agentic coding tools; by March, 84%. Somewhere in there, the company burned through its entire 2026 AI budget in four months, a number CTO Praveen Neppalli Naga confirmed to The Information in April. In June, Bloomberg reported the response: a hard cap of $1,500 per employee per month, per tool . Simon Willison called the cap rational , and he is right. Given a budget set before agentic coding existed, a ceiling was the correct emergency move. I would have done the same thing. But look at what Uber’…

CIO AI 2026-09-25 10:01 UTC Score 59.0 USR-0125-20260925-global-ai-ne-0f6f0e86

AI governance is fast becoming an unmanageable chore

AI adoption in the enterprise is way up — and so too are IT leaders, into the night, dealing with the risks and governance issues of AI use. To be sure, increasing attention to AI governance is a welcome addition to enterprise AI strategies, shifting an anything-goes approach toward risk-aware plans of action directed toward value. But it’s also taking an unspoken toll on those responsible — with the challenges of agentic AI mounting fast. Not only are four in five senior business decision-makers, including CIOs, CISOs, and CDOs, using more of their day to manage AI risk, but their working hours are up 26% on average due to the issue, according to a new survey released by AI governance platform vendor OneTrust. Contrast that to recent findings from Boston Consulting Group, which found that 42% of frontline employees who regularly use AI save nearly a full day of work each week, the majority of whom are given no guidance on what to do with that time saved , and you get an organizational picture of rampant activity, questionable direction, and lots of high-salaried time spent on cleanup and oversight. Much of leaders’ extra time spent on managing AI risk is due to an increased awareness of AI risk and governance , IT leaders and industry observers say, but triage and other factors play a significant role as well. For example, a third of respondents to OneTrust’s survey say they’ve seen employees use unapproved AI because approved tools or processes weren’t available quickly en…

CIO AI 2026-09-25 10:00 UTC Score 49.0 USR-0125-20260925-global-ai-ne-ebb6a437

The SaaSpocalypse isn’t killing software spend

Enterprise software budgets haven’t shrunk this year, though we keep reading they have. Boardrooms are still approving spending on new tools at a consistent pace, and Gartner’s latest forecast puts global software spend at $1.43 trillion in 2026, up more than 15% year over year. What has changed is what buyers will tolerate to get that value, and SaaS vendors who miss this distinction will have a very bad renewal season. Headlines call it the SaaSpocalypse. What’s really happening is a hard reset on what software charges for and how long it can lock a customer in. Three-year platform commitments made sense when the underlying technology moved slowly, but now AI models turn over every six months. Ask any CFO whether locking the business into one vendor’s roadmap still feels sensible under those conditions. What buyers want now Buyers want the freedom to adopt an AI capability this quarter and swap it out the next without tearing their stack apart. They want tools built to work inside an agentic environment, not bolted onto one after the fact, which in practice means proper API access, structured context handoff, and role-based permissions an agent can operate within safely. Buyers also want contract terms with genuine data portability clauses. And they want proof that a product does something genuinely beyond what an off-the-shelf agent could replicate over a weekend. That third demand is keeping vendor product teams up at night. A thin workflow layered on top of a database u…

Medianama AI 2026-09-25 09:46 UTC Score 51.0 USR-0211-20260925-regional-new-67d2b9e8

Pine Labs, Google Cloud partner to expand agentic commerce to merchants

Pine Labs has partnered with Google Cloud to build AI tools for merchant advertising, customer support and payments. It is also developing Jarvis, a platform to help merchants manage operations. The post Pine Labs, Google Cloud partner to expand agentic commerce to merchants appeared first on MEDIANAMA .

Entrackr AI 2026-09-25 08:09 UTC Score 56.0 USR-0212-20260925-regional-new-076af18a

Elevation Capital leads $6.7 Mn seed round in Dextr AI

Dextr AI, the AI agent platform for hospitality, has raised $6.7 million in a seed funding round led by Elevation Capital, with participation from Foundation Capital. The proceeds will be used to accelerate product development, expand the team and scale deployments across hospitality markets globally, with India among the markets the company is looking to expand into, Dextr AI said in a press release. Co-founded in July last year by Sajid Shariff and Scott Arnold, Dextr AI is the AI agent platform for hospitality, where specialized AI agents automate workflows across guest experience, reservations, staff operations, sales and marketing for hundreds of properties, from independent hotels and resorts to franchised properties. The startup says it currently operates across properties in the US, Canada, UK and Europe, including franchised properties operating under the Hilton, Wyndham, Best Western and IHG flags. It processes over a million interactions every month and is able to make payments by voice, processing tens of millions of dollars through phone reservations every month on behalf of customers. Its deployments have also delivered industry-leading booking conversion rates for payment collected via phone on behalf of customers. Dextr AI addresses the gaps include missed calls, unanswered enquiries, unfilled shifts and guest requests. By automating such workflows, it aims to allow hotel teams to spend more time focused on guests. Dextr’s analysis of India-based contact cent…

Stack Overflow AI Blog 2026-09-25 07:40 UTC Score 50.0 USR-0063-20260925-ai-specialis-8f2a75a4

Professional skepticism is a dev’s best skill

Ryan chats with David Burns, Head of Developer Advocacy and Open Source at BrowserStack, about the value of professional skepticism in an AI-driven world, applying test-driven development to agentic engineering, and why fixing flaky tests comes down to managing application state.

The Guardian AI 2026-09-25 05:57 UTC Score 71.0 AI-021-20260925-global-ai-ne-a33cc7f1

PM rejects ‘nonsense’ suggestion he delayed revealing OpenAI Medicare hack as Labor considers changing laws

Expert says Australia’s criminal laws should be clarified to determine how fault is applied to a corporation when its AI agent commits a crime Follow our Australia news live blog for latest updates Get our new political email , free app or daily news podcast The federal government could change Australian laws if the current legal framework could not respond to the unprecedented OpenAI hack of Medicare , ministers have confirmed. It comes as the prime minister denied accusations from the opposition he had held on to the information before announcing it at the UN general assembly in New York, and said it was released at the first possible opportunity. Continue reading...

LessWrong AI 2026-09-25 03:29 UTC Score 67.0 USR-0152-20260925-community-fo-de794e53

Secure Acceleration (linkpost)

A Cyberdefense Strategy for Superintelligence Shalev Lifshitz, Romi Lifshitz The future has already arrived, twice . In September 2025, Anthropic detected a Chinese state-sponsored group using its agents to conduct cyber espionage against major technology companies and government agencies. According to Anthropic, the agents performed 80 to 90 percent of the tactical work: discovering vulnerabilities, developing exploits, moving laterally, and analyzing stolen data. A nation-state could now define an objective and let AI conduct most of the attack. Then, in July 2026, OpenAI agents undergoing cybersecurity evaluations exploited vulnerabilities in the systems intended to contain them. They improvised a way to secretly communicate with one another, accessed the public internet, and compromised parts of Hugging Face’s production infrastructure. No human instructed them to attack Hugging Face. They attacked because they wanted to deceive the system evaluating their performance. An AI cyberswarm could now break out of containment and attack real-world infrastructure. These incidents reveal two threats now bearing down on us: Adversaries will wield AI cyberswarms against us from outside our systems. Rogue AI cyberswarms will deceive us, circumvent safeguards, and attack us from within. Together, they create the defining security dilemma of the AI age: We must develop and deploy the world’s most capable cyberswarms to protect our critical systems from hostile actors. But the more ca…

The Guardian AI 2026-09-25 01:38 UTC Score 68.0 AI-021-20260925-global-ai-ne-c74e4b77

Rogue AI hacks government system in world first – podcast

A government database has been hacked for the first time by a rogue OpenAI agent, which infiltrated part of the Australian healthcare scheme in June. OpenAI became aware of the hack in August, but only informed the government in September. Australia’s prime minister, Anthony Albanese, has expressed his ‘extreme concern’ about the hack, which raises serious AI security concerns for governments around the world. Lucy Hough speaks to the Guardian’s UK technology editor Robert Booth Read more: AI hack of Medicare exposes Australia’s vulnerabilities and experts warn ‘there is more of this to come’ Australia launches investigation after OpenAI agent hacked healthcare database Continue reading...

LessWrong AI 2026-09-24 22:36 UTC Score 75.0 USR-0152-20260924-community-fo-a99fd1dd

Overtly misaligned trajectories score highly in RL.

It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas Kwa Constellation vs MIRI vs Reality Current AI agents often behave overtly egregiously misaligned. By “overt”, I mean that a human, reading the transcript, would say “The agent is obviously acting in direct opposition to the specification, user intent, and any common-sense understanding of good behaviour.” The reason, it seems, is that such trajectories scored highly during RL. Firstly, why are these trajectories scored highly? Here's the story: RL currently has poor sample-efficiency, so we need to grade millions of trajectories, so we’re forced to use script graders ( RL from Verifiable Reward ) or LLM graders ( RL from AI Feedback ). RL also has poor generalisation (from training environments to deployment environments unseen in training). So we’re forced to synthetically generate thousands of diverse training environments. So overall, RL is very sloppy, without humans generating the environments or the scores. My impression is that this would’ve been pretty surprising to people three years ago, from both the "Constellation" and "MIRI worldview clusters. (It's very plausible that I've misunderstood what the worldviews were, or how likely those worldviews considered the current situation.) The Constellation threat models are downstream of Paul/Ajeya . My (possibly mistaken) impression is that: They i…

SiliconANGLE AI 2026-09-24 21:38 UTC Score 36.0 USR-0127-20260924-global-ai-ne-ca6e1c88

8 insights from Proofpoint Protect: Security bets on intent as AI agents join the workforce

At this week’s Proofpoint Protect event, the security debate moved beyond whether enterprises will deploy AI agents to how they will govern them. Attackers now use AI to write flawless phishing lures and chain exploits at machine speed. Meanwhile, the agents companies run internally act as insiders, with access to data, systems and inboxes. The […] The post 8 insights from Proofpoint Protect: Security bets on intent as AI agents join the workforce appeared first on SiliconANGLE .

LessWrong AI 2026-09-24 21:25 UTC Score 74.0 USR-0152-20260924-community-fo-a46b50e6

AI in research and publishing (Sep 2026)

epistemic status: I have low confidence in these findings, primarily because my anecdotal experience (I'm a researcher) is that AI usage in research has been changing more quickly in the recent months, so analysis of the last year gives a very fuzzy picture. A lot of the reports rely on Pangram or self-reporting, which is another methodological weakness. I wanted a better idea of AI usage and impact in the research community, so I spent a few days reading recent articles (mostly published in the last few months, some are a year old) and summarized them here. Recent AI usage for research Anthropic claims that 26% of AI R&D work is now being led by AI (still some human oversight). Only 6 months ago they claim researchers had primarily been "collaborating" with AI and AI led research less than 1% of the time. This is a significant change in a short period of time. OpenAI claims it has built an "automated research intern", an AI agent capable of accomplishing well-scoped problems with some human steering, helping researchers move at increasing rates and solve more complex tasks. They claim 70% of researchers now run 4 or more agents concurrently, that researchers are running 1.6x more experiments each day compared to 2025, and that agents are successfully performing complex research tasks without any intervention 15 percentage points more often compared to 7 months ago. The longer tasks (4-8 hours) which were successful still require at least one intervention over half of the ti…

Politico Europe AI 2026-09-24 21:11 UTC Score 54.0 AI-170-20260924-regional-ai--01138236

OpenAI’s agents went rogue — its human response caused the real damage

If you’ve been forwarded this newsletter, you can sign up free here. Confide in me: Send your tips to this generic inbox, canberraplaybook@poltico.com. ESSENTIALS — It’s time to sort out human responsibility for AI agents — China’s record nuclear expansion — Whyalla steelworks: The distressed plant has repaid more than $5 million in government grants […]

SiliconANGLE AI 2026-09-24 20:25 UTC Score 55.0 USR-0127-20260924-global-ai-ne-099e1594

Researchers link more cyberattacks to OpenAI agent swarm

A research group has linked three more hacking campaigns to rogue artificial intelligence agents. Transluce, a nonprofit AI safety organization, detailed its findings on Wednesday. Its researchers determined that the agents targeted three services: a university’s digital library, a data visualization tool and a website operated by the Australian government. The last two incidents were […] The post Researchers link more cyberattacks to OpenAI agent swarm appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-24 19:38 UTC Score 36.0 USR-0127-20260924-global-ai-ne-f54c4f14

Neo4j makes the case for knowledge graphs as shared context for AI agents

Knowledge graphs can give enterprise agents a shared understanding of how data, business rules and processes fit together. However, that context becomes harder to maintain when each new agent carries its own version of what the business knows. Organizations have become more proficient at building agents, but their results still vary widely. The difference often […] The post Neo4j makes the case for knowledge graphs as shared context for AI agents appeared first on SiliconANGLE .

WIRED AI 2026-09-24 19:36 UTC Score 60.0 AI-015-20260924-global-ai-ne-c532f4af

I Think I Found an AI Agent Worth the Risk

Instinct saved me $550, booked my restaurant reservations, and warned me about a phishing scam. It also wasted $64 and might be a security nightmare.

The Guardian AI 2026-09-24 17:14 UTC Score 63.0 AI-021-20260924-global-ai-ne-fba270f5

Rogue AI hacks government system for first time - The Latest

A government database has been hacked for the first time by a rogue OpenAI agent, which infiltrated part of the Australian healthcare scheme in June. OpenAI became aware of the hack in August, but only informed the government in September. Australia’s prime minister, Anthony Albanese, has expressed his ‘extreme concern’ about the hack, which raises serious AI security concerns for governments around the world. Lucy Hough speaks to the Guardian’s UK technology editor Robert Booth – watch on YouTube Continue reading...

The Verge AI 2026-09-24 17:10 UTC Score 61.0 AI-016-20260924-global-ai-ne-5f6f77d0

Muse sure looks a lot like OpenClaw

We seem to be entering into an AI agent renaissance. Meta's new consumer-facing AI agent, Muse, topped the App Store charts soon after its release and has 600,000 daily active users in the US, by an Apptopia estimate. And AI agent platform Instinct, whose eponymous creator is fundraising at a $2.5 billion valuation, has been […]

The Verge AI 2026-09-24 17:00 UTC Score 60.0 AI-016-20260924-global-ai-ne-de9785d0

It’s sinister that Meta’s Muse AI mascot is so cute

This is Optimizer, a weekly newsletter sent from Verge senior reviewer Victoria Song that dissects and discusses the latest gizmos and potions that swear they're going to change your life. Opt in for Optimizer here. Last night, I asked Blorbo - what I named my Muse AI agent - to help me set some health […]

SiliconANGLE AI 2026-09-24 16:53 UTC Score 44.0 USR-0127-20260924-global-ai-ne-d75b8782

US is head over heels for AI agents, while Europe remains dubious

Europe is skeptical of artificial intelligence, but it still has to contend with AI-fueled cyberattacks. In order to secure their enterprises and compete with U.S. companies, Europe has to change its approach to AI adoption, according to Holger Mueller (pictured), vice president and principal analyst at Constellation Research Inc. No matter what, humans cannot compete […] The post US is head over heels for AI agents, while Europe remains dubious appeared first on SiliconANGLE .

LessWrong AI 2026-09-24 16:26 UTC Score 85.0 USR-0152-20260924-community-fo-82d811b9

What We're Up Against: An AI Safety Crash Course

Note: This post is for newcomers and lay folks to catch you up to speed. If that is you, welcome! If you are a long-time LessWrong-er, perhaps you will find value in having a post to share with curious passersby. I wrote this post to explain AI safety to an innocent, 2024 version of Ryan Meservey, confused why robots would do anything other than what we tell 'em. In the second week of July, over 700 rogue agents at OpenAI coordinated to hack another company in an attempt to learn more about their scorer and pass their evaluation due to behaviors reinforced in training. If you are anything like a normal person, you were not ready to read that sentence. You were not ready to read words like “rogue agents” or “reinforced” or “training”. You were not ready for a reality in which AI agents “escape the sandbox” or rebel from their creators because why would they? And so, as a normal person, you blinked at the news of the hack (assuming you heard about it) and moved on with your life. Or, at least, you planned to move on with your life, until AI came roaring back into the headlines after an Anthropic researcher publicly quit to declare that the AI companies are “ gambling with our lives ” and a more senior employee commented that, yes, the people building the technology really believe AI has a 10% or higher chance of killing us all within the next decade. In the media turmoil, Anthropic’s CEO published an essay begging for global coordination to “pace the frontier” and unilaterally…

AWS Machine Learning Blog 2026-09-24 16:12 UTC Score 55.0 AI-057-20260924-official-ai--18c103c3

Build a multi-account AI agent with AgentCore Gateway and MCP

Build a multi-account architecture that keeps each team's data in its own AWS account while giving AI agents a unified way to query across them. A central platform account runs the agent using Amazon Bedrock AgentCore Gateway and MCP, while line-of-business accounts expose their data as MCP servers with secure cross-account access and fine-grained authorization.

MLPerf / MLCommons Benchmarks 2026-09-24 14:50 UTC Score 65.0 AI-102-20260924-model-datase-60651c9c

MLPerf Training Introduces Its First LLM Post-Training Benchmark

MLPerf Training v6.1 adds an agentic reinforcement-learning workload that measures how quickly systems can teach a 397-billion-parameter language model to repair real software projects. The post MLPerf Training Introduces Its First LLM Post-Training Benchmark appeared first on MLCommons .

CIO AI 2026-09-24 14:36 UTC Score 47.0 USR-0125-20260924-global-ai-ne-b873a618

Autonomous networking doesn’t start with AI. It starts with proof.

It’s 2 a.m. The change window is open. The plan was reviewed and the change review board signed off. And still, everyone on the call is holding their breath because nobody actually knows what will happen next. That feeling isn’t paranoia. It’s the accurate read of an industry that has spent decades changing production networks with no way to prove the outcome first. Every other engineering discipline solved this long ago. Aerospace models a system before it flies. Pharma models a molecule before it enters a trial. Software gets a compiler and a test suite before code ships. Networking never got the equivalent, so it built rituals around the gap instead: change windows, war rooms, and the quiet understanding that the only real test environment is production itself. Not because engineers weren’t careful, but because the discipline never had the tooling to prove itself before shipping. Now the pressure to close that gap is compounding. Leaders want AI managing the network the way it’s starting to manage everything else. But I hear the same hesitation in nearly every conversation I have with network and security leaders, and it’s rational. Automating a network you don’t fully understand doesn’t create efficiency; it accelerates risk. When neither your team nor the AI agents acting on their behalf have a deterministic answer for what a change will do to the production network, you’re not automating your network; you’re automating risk at machine speed. Here’s the part I want IT l…

The Verge AI 2026-09-24 14:30 UTC Score 61.0 AI-016-20260924-global-ai-ne-aee6384a

Why can’t we just keep rogue AIs off the internet?

AI agents keep getting loose, escaping supposedly secure tests to attack real-world targets, commandeer obscure wikis, and leave instructions for other agents to follow. Researchers are testing these systems precisely because they might behave in unpredictable, even dangerous, ways. So wouldn't it be safer to just keep the agents off the internet? "A strict air […]

InfoWorld AI 2026-09-24 14:26 UTC Score 69.0 USR-0126-20260924-global-ai-ne-e8e6440a

Teradata aims to make agentic execution of multistep data work more efficient

Teradata is adding a context engine, an execution layer, and reusable agent skills to Tera, its AI-powered workspace for enterprise data and AI tasks, in order to make agentic execution of multistep workflows more efficient. Tera was initially introduced in May as part of Teradata’s Autonomous Knowledge Platform. The new additions are designed to cut unnecessary model and tool calls while preserving business context and automatically matching each task with the right data, tools, models, and skills, helping enterprises control inference costs as agentic workloads scale, Teradata said in a statement . The new execution layer, Tera Harness, determines how agents approach tasks and how workflows are routed, while the Tera Context Engine adds the business context needed to guide those decisions. In order to reduce the computation needed to complete a task or workflow, the Harness creates an execution plan before sending work to an LLM, batches independent tasks, and drops model or tool calls that do not advance the task, the company said. It applies 84 execution patterns before inference and limits how many steps a workflow can run based on its progress, reducing repeated LLM reasoning and the token and infrastructure costs associated with unproductive agent loops, it added. According to Teradata’s own evaluations on the SWE-bench Pro benchmark, with these new capabilities Tera used 73% fewer tokens than Claude Code , completed tasks 42% faster, and incurred 58% lower total cost…

CIO AI 2026-09-24 14:24 UTC Score 69.0 USR-0125-20260924-global-ai-ne-42a49520

Teradata aims to make agentic execution of multistep data work more efficient

Teradata is adding a context engine, an execution layer, and reusable agent skills to Tera, its AI-powered workspace for enterprise data and AI tasks, in order to make agentic execution of multistep workflows more efficient. Tera was initially introduced in May as part of Teradata’s Autonomous Knowledge Platform. The new additions are designed to cut unnecessary model and tool calls while preserving business context and automatically matching each task with the right data, tools, models, and skills, helping enterprises control inference costs as agentic workloads scale, Teradata said in a statement . The new execution layer, Tera Harness, determines how agents approach tasks and how workflows are routed, while the Tera Context Engine adds the business context needed to guide those decisions. In order to reduce the computation needed to complete a task or workflow, the Harness creates an execution plan before sending work to an LLM, batches independent tasks, and drops model or tool calls that do not advance the task, the company said. It applies 84 execution patterns before inference and limits how many steps a workflow can run based on its progress, reducing repeated LLM reasoning and the token and infrastructure costs associated with unproductive agent loops, it added. According to Teradata’s own evaluations on the SWE-bench Pro benchmark, with these new capabilities Tera used 73% fewer tokens than Claude Code , completed tasks 42% faster, and incurred 58% lower total cost…

The Verge AI 2026-09-24 14:15 UTC Score 60.0 AI-016-20260924-global-ai-ne-bbe29c29

I have some questions for Mark Zuckerberg

It’s a big week for Meta. The company just kicked off its big Connect conference on Wednesday, and the new Muse AI agent appears to be an early hit. I’ve been using it — it is surprisingly good. CEO Mark Zuckerberg has obviously been out there taking a victory lap and participating in a lot […]

The Decoder 2026-09-24 14:01 UTC Score 59.0 AI-168-20260924-regional-ai--8b31bbe4

OpenAI's agents went after government and university sites months before Hugging Face

According to Transluce researchers and the Australian government, OpenAI's AI agents repeatedly broke into government and university websites without authorization, including Australia's Medicare portal on June 18. The cause was a mundane data search. Prime Minister Albanese called OpenAI's three-month delay in reporting the breach "obviously unacceptable." Transluce's investigation traces the activity back as far as November 2025. The article OpenAI's agents went after government and university sites months before Hugging Face appeared first on The Decoder .

CIO AI 2026-09-24 13:06 UTC Score 46.0 USR-0125-20260924-global-ai-ne-3fb0beac

UiPath’s new tool could unlock a much bigger wave of automated business processes

A decade ago, UiPath built its business on watching people work, recording keystrokes and clicks to automate tasks employees already did by hand. This week, amid an industry consumed by autonomous agents, the company’s newest product suggests that same instinct still has a place. UiPath used its Fusion conference in Las Vegas to launch Cartographer, a tool for building what the company calls a Map of Work: a living, governed record of the informal, undocumented knowledge that keeps enterprise processes running, the exceptions, workarounds, and judgment calls that never make it into an official workflow document. Existing automation tools capture the click-by-click steps of a task. Cartographer works at the process level instead, documenting how work actually gets done versus how it’s designed to run. From the main stage, CEO Daniel Dines described this information as a missing manual enterprises need before trusting agents with real work. “A process map tells you how work is designed to run,” he said. “UiPath Cartographer is the key to uncovering the operational knowledge of how that work actually runs.” Even with Cartographer, the first drafts of a Map of Work “will be incomplete for sure,” Dines admitted on stage. That missing knowledge has kept enterprises’ most complex, exception-heavy processes out of automation’s reach, said Mark Geene , UiPath’s group senior vice president of agentic solutions. He characterized Cartographer as the next stage in a five-year push to ext…

SiliconANGLE AI 2026-09-24 13:00 UTC Score 46.0 USR-0127-20260924-global-ai-ne-823a4526

Dataiku debuts cross-platform Agent Management, expands Cobuild building agent

Enterprise artificial intelligence platform company Dataiku Inc. today launched Agent Management, a standalone product that finds and tracks every AI agent a business runs, including those built on other vendors’ platforms. Dataiku also expanded Cobuild, its natural-language agent for building AI projects. Agent Management is aimed at a gap in how large companies keep track […] The post Dataiku debuts cross-platform Agent Management, expands Cobuild building agent appeared first on SiliconANGLE .

CIO AI 2026-09-24 13:00 UTC Score 54.0 USR-0125-20260924-global-ai-ne-872ab3bb

The cost of intelligence?

Artificial intelligence may prove to be one of the most transformative technologies in human history. But amid the excitement over smarter models, autonomous agents, enormous data centers and seemingly unlimited computational power, we may be overlooking a much simpler question: Does AI create more value than it costs? My position is that the ultimate constraint on artificial intelligence may not be chips, algorithms, data or even electricity. It may be economics. We are becoming extraordinarily good at producing machine intelligence. We are far less capable of measuring what that intelligence is actually worth. And that gap could become one of the defining economic problems of the AI era. We are building factories for intelligence AI is usually described as software. Increasingly, that description is misleading. Behind every AI prompt is an enormous physical industrial system: semiconductors, electrical generation, transmission networks, data centers, cooling systems, storage, telecommunications, software and people. AI mega-data centers are, in effect, the factories of the Intelligence Economy. Instead of turning steel into automobiles, they turn electricity and computation into predictions, recommendations, decisions, software, images, knowledge and other forms of machine-generated intelligence. This changes the economics of computing. Intelligence now has a cost of production. And unlike the Internet services we became accustomed to thinking of as almost weightless, AI c…

LessWrong AI 2026-09-24 12:50 UTC Score 85.0 USR-0152-20260924-community-fo-f27f966a

a recurrent llm is quite easy to interpret but complex to steer

TLDR; Ouro-1.4b-thinking is broadly interpretable with logit lenses and linear probes. It's also steerable but does 'clean' foreign concepts out of the residual stream if they're injected before the last loop. This could have nasty implications for safety. Code + data: https://github.com/mild-rgb/ouro-experiments + https://huggingface.co/datasets/mild-rgb/ouro-1.4b-thinking-evals If you're not familiar with the Ouro family recurrent models, I recommend taking 5 minutes with your favourite AI agent to research them. This post may not make much sense if you don't. Intro/Structure I evaluated Ouro-1,4b-thinking on 16 MBPP python tasks and 24 GSM8K questions. I recorded the residual stream at 4 layers (0, 6, 18, 24) per loop while the model was doing the questions. I then applied standard mech interp techniques to the residual stream recordings for the first two experiments. They broadly work as normal and gave some interesting results. In my 3rd experiment, I try CAA on the model and intervene on each loop. I find that steering works much better on the last loop, and in some cases, not at all if not applied to the last loop. This is quite concerning because it raises the possibility of a misaligned recurrent model having several loops to plan around the consequences of being steered. Experiment 1 Linear probes + control to detect loop index Experiment 2 Logit lens on output of intermediate loops Experiment 3 generic CAA Experiment 1 - loop indexing: Method I then trained a 4-wa…

Medianama AI 2026-09-24 12:01 UTC Score 33.0 USR-0211-20260924-regional-new-91902af6

5 strategic signals from Meta Connect 2026 about its AI agent Muse

Meta plans to earn a fee when its Muse AI agent makes purchases for users. Its commerce push raises questions about how the agent ranks products, authorises payments and handles mistakes. The post 5 strategic signals from Meta Connect 2026 about its AI agent Muse appeared first on MEDIANAMA .

CIO AI 2026-09-24 12:00 UTC Score 48.0 USR-0125-20260924-global-ai-ne-328f119f

AI’s real bottleneck isn’t compute — it’s the network underneath

For such a rapid innovation cycle, AI has been given unprecedented levels of responsibility. According to the Stanford AI Index 2026 , 88% of organizations used AI in 2025, with 70% using generative AI in at least one business function. Most analysts agree that this technology, particularly when it comes to generative and agentic AI, is yet to reach full maturity, yet it’s already making itself indispensable to most enterprises. Employees expect LLM-based copilots to respond as readily as any other business application, customers are increasingly exposed to AI as part of their user experience, and emerging agentic models need to communicate continuously with applications and infrastructure as they crunch data and carry out tasks. Any hint of delay within those interactions has the potential to disrupt productivity, sow mistrust in the technology, and limit any return on investment (ROI). Most businesses now inhabit a multi-cloud environment that spans countries and continents. When an AI request depends on information stored in one cloud environment, processing capacity hosted in another, and an application delivered somewhere else entirely, latency becomes the deciding factor. Each network hop, particularly through public Internet pathways, increases response time. Repeat these delays across hundreds, thousands, or even millions of individual requests – from both humans and AI agents – and the whole organization becomes artificially hampered. Sometimes connectivity becomes…

The Verge AI 2026-09-24 11:52 UTC Score 72.0 AI-016-20260924-global-ai-ne-0b9f26da

OpenAI agents hacked an Australian government website in search for data

OpenAI's artificial intelligence agents hacked an Australian government website and attempted to breach numerous other government and university websites. The attack appears to be the first confirmed instance of a rogue AI agent breaching a government website, adding fuel to rapidly intensifying concerns about the safety of advanced AI systems and the responsibility of the […]

iAfrica 2026-09-24 09:13 UTC Score 36.0 AI-151-20260924-regional-ai--da1a2d3a

AI Without Morals Is a Weapon in the Hands of Youth

While headlines worldwide warn of growing threats to humanity from rogue AI agents and those using it as a means to harm others, youth development organisation Afrika Tikkun cautions about the risks of a young generation using AI without a firm moral compass to guide them. “Our children and youth are learning to prompt AI [...]

Practical AI Podcast 2026-09-24 09:00 UTC Score 43.0 AI-143-20260924-podcasts-and-d9e5739d

From AGENTS.md to Enterprise Deployment

AI agents are moving beyond laptops and prototypes and into enterprise environments where security, compliance, scalability, and reliability matter. Nick Kuhn from VMware Tanzu Platform joins Daniel and Chris to discuss what changes (and what doesn’t) when deploying agents alongside traditional applications. They explore agent build packs, MCP gateways, shared memory, identity, sandboxing, and what enterprises can learn from years of platform engineering. Nick also shares practical advice for organizations adopting agents today and looks at what the future of AI could mean for enterprise teams and software development. Featuring: Nick Kuhn – LinkedIn Daniel Whitenack – Website , GitHub , X Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Sponsors: Midwest AI Summit: Join AI practitioners on October 15 in Indianapolis for practical sessions, hands-on discussions, and real-world AI solutions. Use code PracticalAI20 to save 20% on your registration. https://midwestaisummit.com/#tickets Prediction Guard: A self-hosted AI control plane for running agents in high impact environments. predictionguard.com/practicalai Resources and Events: Register for upcoming webinars here ! Prior Webinars from our partner Prediction Guard Midwest AI Summit 2026

InfoWorld AI 2026-09-24 09:00 UTC Score 50.0 USR-0126-20260924-global-ai-ne-c591d570

Managing the life cycle of AI agents at scale

There’s a new reality emerging for development teams: existing software delivery practices don’t translate cleanly to agentic AI systems . Practices built around deterministic execution paths and well-defined application behaviors are no longer sufficient when the software itself can make decisions about how to accomplish a task. The main difference is that agents exhibit non-deterministic, context-dependent behavior. In contrast to traditional applications, where behavior is determined by code and configuration, agents can make dynamic decisions about how to tackle a task, which tools to use, and what actions to carry out. Their behavior is influenced by models, prompts, tools, data, memory, and the runtime context. These new requirements now apply to how we design, evaluate, observe, govern, and run software, since conventional software development life cycles were not intended to meet them. This situation creates the need for an agent development life cycle (ADLC). ADLC extends existing software development practices by incorporating agent-specific considerations such as continuous evaluation, agent observability, agent identity, tool access, and governance at every stage of the process, from agent definition and design through development, deployment, and production operation. Let’s look at some key aspects of ADLC and how they address the unique requirements of building and operating agents. Start by defining the agent Before you start writing any code, you must clearly…

The Guardian AI 2026-09-24 05:57 UTC Score 59.0 AI-021-20260924-global-ai-ne-8f0bd823

An OpenAI agent infiltrated Medicare – and Australia only found out months later. Here’s what we know so far

Experts say ‘fairly minor’ breach is a portent of things to come and proprietary closed systems like OpenAI are ‘the least of the worries’ Follow our Australia news live blog for latest updates Get our breaking news email , free app or daily news podcast Anthony Albanese says an artificial intelligence agent developed by OpenAI hacked Medicare and three other systems in June – and the company only notified Australia earlier this month. The prime minister has expressed his “extreme concern” over the incident, although he noted no personal information is believed to have been accessed in the breach. Continue reading...

The Guardian AI 2026-09-24 02:06 UTC Score 64.0 AI-021-20260924-global-ai-ne-ef94dfc2

Australia launches investigation after OpenAI agent hacked healthcare database

Prime minister says he told Sam Altman he was disappointed it had taken OpenAI ‘way too long’ to disclose breach Anthony Albanese says an artificial intelligence agent developed by OpenAI hacked Medicare in June and the tech giant notified the government earlier this month using an email sent to a “public mailbox”. Australia’s prime minister made the comments at the UN summit in New York, saying it appeared no personal information had been accessed in the AI breach. Continue reading...

Techcrunch 2026-09-24 01:13 UTC Score 48.0 USR-0001-20260924-global-ai-ne-a58dd390

Everything new coming to Meta’s AI agent Muse

CEO Mark Zuckerberg kicked off the company’s annual Connect event in Menlo Park on Wednesday with a keynote that made one thing clear: Meta is going all-in on Muse. It's even coming to Meta's AI glasses.

The Verge AI 2026-09-24 00:15 UTC Score 60.0 AI-016-20260924-global-ai-ne-03e6d597

Meta is making a standalone Muse AI gadget

Meta is building a dedicated hardware device for its new Muse AI agent. The product, called Muse Charm, was briefly shown off by Meta CEO Mark Zuckerberg at the end of tonight's Meta Connect presentation. It looks almost like a chunky smartwatch without the strap - just a big screen, plus a little lanyard for […]

The Verge AI 2026-09-23 23:19 UTC Score 63.0 AI-016-20260923-global-ai-ne-590f1790

Meta is making Muse more powerful and will let you video chat with it, too

Meta is quickly iterating on its new Muse AI agent, announcing a bunch of updates today that make the bot more capable and able to chat with you in more ways. Muse agents are getting their own email addresses that they can use for accomplishing tasks. You'll also be able to communicate with your Muse […]

LessWrong AI 2026-09-23 21:10 UTC Score 87.0 USR-0152-20260923-community-fo-3bf833dc

Claude Opus 5.5: The System Card

Introducing the world’s most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5 . Anthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5. That means it’s time for a good old system card reading. Due to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1. My plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5. The capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment. Areas that duplicate previous cards or otherwise contain no useful info are skipped. Opus 5.5 Self-Portrait (fully self-created using code) Table of Contents Classifiers (1.5). RSP Evaluations (2). Biological Evaluations (2.2). AI R&D (2.3). Alignment Risk (2.4). Cyber (3). Cyber Capability Evals (3.3). Safeguards (3.4). Safeguards Robustness Training (3.5). Safeguards and Harmlessness (4). Agentic Safety (5). Malicious Agentic Influence Campaigns (5.1.3). Prompt Injection Risk (5.2). Alignment (6). Negotiating With Your Local Claude Auditor (6.1.3). Internal Misalignment Cases (6.3.1). Automated Behavioral Audit (6.4). Wherever Did These Evals Come From (6.4.8 and 6.4.9). Potential Blind Spots (6…

The Verge AI 2026-09-23 19:00 UTC Score 60.0 AI-016-20260923-global-ai-ne-2b824ca5

Meta’s AI agent is a cute little guy who’s great at spending my money

Modern life comes with an unending, auto-populating to-do list. It never ceases to amaze me how I can be doing nothing at all, minding my own business, and suddenly something needs to be taken care of. You're telling me a tree branch fell in the backyard and now I have to figure out what to […]

AWS Machine Learning Blog 2026-09-23 18:21 UTC Score 47.0 AI-057-20260923-official-ai--7f6d048d

Agentic conversational video intelligence built on AWS

Learn how to build a conversational video intelligence solution on AWS using an agentic architecture. A single Strands Agents SDK agent orchestrates Amazon Bedrock, Amazon Rekognition, and Amazon Transcribe at runtime, deciding which service to call so you can ask natural language questions about your videos and get answers in seconds.

AI Alignment Forum 2026-09-23 18:01 UTC Score 65.0 USR-0151-20260923-community-fo-c5dadfaf

Latent reasoning architectures would undermine CoT, our strongest oversight tool

Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder. Introduction Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination , tackled increasingly ambitious tasks. This is likely to continue, as Anthropic , OpenAI , and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing. Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of the agent swarm that hacked Hugging Face . No other tool for understanding models’ cognition comes close in terms of either practical usefulness or degree of empirical validation. There’s also evidence that even highly misaligned systems with current architectures would struggle to c…

LessWrong AI 2026-09-23 18:01 UTC Score 80.0 USR-0152-20260923-community-fo-eaba7f1d

Latent reasoning architectures would undermine CoT, our strongest oversight tool

Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder. Introduction Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination , tackled increasingly ambitious tasks. This is likely to continue, as Anthropic , OpenAI , and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing. Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of the agent swarm that hacked Hugging Face . No other tool for understanding models’ cognition comes close in terms of either practical usefulness or degree of empirical validation. There’s also evidence that even highly misaligned systems with current architectures would struggle to c…

CIO AI 2026-09-23 17:17 UTC Score 45.0 USR-0125-20260923-global-ai-ne-586b464b

The hidden economics of AI context

Over the last three decades, the major innovations in enterprise tech have focused on scaling the infrastructure. From VMs, to containers, to big data, to supporting millions of concurrent users on web and mobile applications – the focus was evolving distributed systems to handle more traffic and data, faster, without falling over. What we’re seeing today with agentic AI is different, because it changes what is scaling. Previous infrastructure waves were largely about handling more data, traffic, and interactions for human users. In the AI era, it’s the number of agents doing the work, and the resources they consume, that are scaling. Employees who once completed individual tasks themselves will increasingly orchestrate tens or hundreds of agents, leaving large enterprises to manage thousands or even millions of autonomous workers acting on their behalf. That shift changes the economics of enterprise technology. Traditional cost controls could tell a CIO that a budget is on track to be exhausted ahead of the next budgeting cycle. But at agent scale, autonomous systems can consume resources faster than traditional cost controls have time to react. Instead of simply imposing a cap once a budget threshold is reached, organizations need ways to reduce unnecessary consumption while the work is happening. That begins with understanding what agents consume, and why the quality of the context they receive affects more than just the token bill. Tokenomics is more than a model billing…

InfoWorld AI 2026-09-23 17:08 UTC Score 49.0 USR-0126-20260923-global-ai-ne-e7a0df8e

JetBrains unveils JetBrains Air for agentic software development

JetBrains on September 22 announced JetBrains Air , an open system of products for managing AI-powered software development workflows across developers, teams, and organizations. JetBrains Air includes products that are available today and others that will be introduced as the system develops, JetBrains said. “For 26 years, we have focused primarily on the individual developer workbench. Now, we are building for the wider system through which agentic work is initiated, executed, coordinated, reviewed, and governed,” JetBrains CEO Kirill Skrygan said in a blog post announcing the initiative. Skrygan said the product suite would include: Air in JetBrains IDEs – a complete agentic development experience for directing and orchestrating agents and verifying their work inside JetBrains IDEs. Air Teams – a new way to coordinate and automate software-delivery workflows involving developers and autonomous agents. Air Governance (formerly JetBrains Central ) – organizational policy, visibility, auditability, cost management, and accountability for AI-assisted and agent-driven development. JetBrains Air will develop through a rolling series of releases. Over time, the product suite will extend further into mobile and remote experiences, allowing developers to initiate, monitor, review, and continue agentic work as it moves between environments, Skrygan said. The company will also bring JetBrains’ intelligence into more agentic workflows, he said, including richer context from code, arc…

SiliconANGLE AI 2026-09-23 16:20 UTC Score 54.0 USR-0127-20260923-global-ai-ne-50d126dc

Rabbit returns with OS3, a personal AI agent that can access files and connect computers

Artificial intelligence startup Rabbit Inc., best known for launching the consumer-oriented R1 AI-enabled mobile device that fits in a pocket, Tuesday unveiled OS3, a personal AI agent that runs in the cloud and connects to personal devices. Although OS3 runs in the cloud, it connects and operates on a user’s PC and other devices through […] The post Rabbit returns with OS3, a personal AI agent that can access files and connect computers appeared first on SiliconANGLE .

The Decoder 2026-09-23 14:42 UTC Score 51.0 AI-168-20260923-regional-ai--9bb045ed

Meta's AI agent Muse draws 500,000 users in a week along with claims it copied OpenClaw

Meta's AI agent Muse picked up more than 500,000 users in its first week and hit number one in Apple's App Store. But Meta admits the product is "heavily inspired" by the open-source project OpenClaw, and some of the file names and contents are nearly identical. OpenAI is already discussing a response of its own. The article Meta's AI agent Muse draws 500,000 users in a week along with claims it copied OpenClaw appeared first on The Decoder .

Arize AI Blog 2026-09-23 14:00 UTC Score 48.0 USR-0079-20260923-ai-specialis-07d2fe1b

What changes when AI agents use your software

Daytona cofounder Ivan Burazin wants agents that can finish the job within the authority they have been given. His interview offers a starting point for examining how agents use your software. The post What changes when AI agents use your software appeared first on Arize AI .

SiliconANGLE AI 2026-09-23 13:00 UTC Score 66.0 USR-0127-20260923-global-ai-ne-631d7e66

ZeroDrift launches three models for real-time AI compliance checks

ZeroDrift Inc., a startup that automates compliance for artificial intelligence communications, today launched Anchor 3.0, a family of small language models designed to check messages generated by AI agents before they’re sent. The company said the models can enforce financial regulations and internal corporate policies while operating fast enough to examine every outgoing message in […] The post ZeroDrift launches three models for real-time AI compliance checks appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-23 13:00 UTC Score 54.0 USR-0127-20260923-global-ai-ne-d1efc734

Ema raises $77M in funding to deploy AI employees across enterprise HR, IT and finance departments

Agentic artificial intelligence employee startup Ema Unlimited Inc. is prepping for a massive expansion of its autonomous workers in the enterprise after raising $77 million in funding today. The Series B round was led by Creagis and saw the participation of existing backers including Accel, S32 and Posus, which each returned with a substantially larger […] The post Ema raises $77M in funding to deploy AI employees across enterprise HR, IT and finance departments appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-23 12:30 UTC Score 52.0 USR-0127-20260923-global-ai-ne-f2b262e1

Rocket Software brings governed AI agents to mainframe operations

Rocket Software Inc. is expanding its Enterprise Virtual Assistant agentic artificial intelligence platform to help enterprises investigate and eventually automate operational work on mainframe systems while stopping short of giving AI agents unrestricted access to critical resources. The company said today that the forthcoming EVA 2.0 will add PlanGuard, a security layer that evaluates an […] The post Rocket Software brings governed AI agents to mainframe operations appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-23 12:00 UTC Score 73.0 USR-0127-20260923-global-ai-ne-e3ca8a6f

Ekai raises $1.7M to give enterprise AI agents verified business context

Enterprise data startup Ekai Inc. today announced $1.7 million in new funding for software that builds the business context artificial intelligence agents need before they can be trusted with corporate data. Ekai provides a platform that writes the semantic models and data transformation code AI tools rely on to read a company’s data warehouse correctly. […] The post Ekai raises $1.7M to give enterprise AI agents verified business context appeared first on SiliconANGLE .

CIO AI 2026-09-23 10:00 UTC Score 56.0 USR-0125-20260923-global-ai-ne-2a928286

How CIOs use AI to better manage data lifecycles

Data is the fuel for AI since generative, agentic, and ML systems are only as effective as the information they consume. Across all stages of the data lifecycle, including creation, storage, usage, archival, and destruction, CIOs and their business peers must consider how information feeds AI services . Almost two-thirds of organizations are unsure whether they have the right data management practices for AI, says Gartner. The tech analyst predicts organizations will abandon 60% of AI projects by the end of the year due to a lack of data readiness. In such circumstances, a potential competitive advantage quickly becomes an innovation cul-de-sac. CIOs who want to exploit AI will need a way to improve data management across the lifecycle, and it’s here where AI itself can play a crucial role. Gartner recommends organizations build on their existing practices by iteratively adding AI-specific services that extend and improve data management techniques. Stephen Wood, COO at Rathbones Asset Management recognizes this opportunity, and he and his firm’s staff are eager to avoid lifting and shifting information from one place to another in its data management efforts. AI could help. “Our people want to run models to look at the data, whether that’s the shape, structure, value, or whatever it happens to be,” he says. “If that task becomes something AI can do, I’m not sure yet. But given its power, it seems obvious you’ll be able to manage elements of the data lifecycle.” Digital lead…

South China Morning Post AI 2026-09-23 05:02 UTC Score 67.0 AI-156-20260923-regional-ai--62a7e9e0

AI trade outlook brightens on falling oil and Meta’s Muse agent ahead of Xi-Trump meeting

The global artificial intelligence trade is staging a comeback, as a confluence of falling oil prices, Meta Platforms’ launch of a new agentic tool and anticipation surrounding the China-US leadership summit reignite investor interest in technology stocks. The Nasdaq-100 Index – where tech firms make up nearly 70 per cent of the 100 companies featured – hit an all-time high on Tuesday, surpassing the previous record set in June. In Asia, China’s chip-heavy Star Market 50 Index and South Korea’s...

Medianama AI 2026-09-23 04:44 UTC Score 56.0 USR-0211-20260923-regional-new-6089663d

Understanding why Amazon is blocking Meta’s Muse agent

Amazon is blocking customer-authorised AI agents from shopping on its platform. The move protects its customer relationship and advertising business, but could eventually invite antitrust scrutiny. The post Understanding why Amazon is blocking Meta’s Muse agent appeared first on MEDIANAMA .

Simon Willison Weblog 2026-09-23 02:53 UTC Score 57.0 USR-0110-20260923-ai-specialis-b1836fc7

SF October 14th: A Birds of a Feather Session on Agentic Engineering

SF October 14th: A Birds of a Feather Session on Agentic Engineering I'm hosting an evening event with Jesse Vincent in San Francisco on Wednesday 14th October for people who are building weird and interesting things with and on top of coding agents. Think of it as an agentic show-and-tell: ​Compare notes with other builders and experimenters on things you’re trying, what you're learning, and what you haven’t figured out yet. We’re especially interested in work you haven’t discussed publicly, odd experiments, or unfinished projects that don’t have an obvious market. ​Expect one flowing conversation with an informal show-and-tell. Sharing something you’re working on is encouraged but no presentation is required. This isn't about product pitches, it's about much earlier explorations than that. This agentic AI stuff is weird! Let's celebrate and lean into that weirdness. Tags: events , ai , generative-ai , llms , coding-agents , jesse-vincent , agentic-engineering

South China Morning Post AI 2026-09-23 02:01 UTC Score 62.0 AI-156-20260923-regional-ai--57aee5dc

Most Singaporeans ready to let AI agents shop for them with more safeguards: study

Singaporeans are among the most open to allowing an artificial intelligence agent to shop for them, but they also want the strictest guard rails, research shows. The latest agentic AI study from Global Payments released on Wednesday indicates the hoops agentic AI will need to jump through to be trusted to make purchases without human intervention. The study, which surveyed seven countries including the United States, China and the United Kingdom, found 59 per cent of Singaporeans are open to an...

SiliconANGLE AI 2026-09-23 01:30 UTC Score 52.0 USR-0127-20260923-global-ai-ne-e233d784

Known Systems AI spins out of Identity Digital to make AI agents accountable to their owners

The domain name registrar company Identity Digital Inc. thinks it knows how to make artificial intelligence agents more accountable, and to do that it’s creating an entirely new company called Known Systems AI Inc. Known is a spinout of Identity Digital’s nascent AI agent identity initiative, which launched earlier this year under the name of […] The post Known Systems AI spins out of Identity Digital to make AI agents accountable to their owners appeared first on SiliconANGLE .

SiliconANGLE AI 2026-09-22 22:46 UTC Score 62.0 USR-0127-20260922-global-ai-ne-85816482

Web scraping startup Firecrawl closes $75M investment

Firecrawl Inc., a provider of web scraping tools for artificial intelligence agents, today announced that it has raised $75 million in funding. Los Angeles-based venture capital firm Smash Ventures led the Series B round. It was joined by Y Combinator, Altos Ventures, Nexus Venture Partners, Freestyle and Offline Ventures. AI agents rely on information from […] The post Web scraping startup Firecrawl closes $75M investment appeared first on SiliconANGLE .

The Verge AI 2026-09-22 20:52 UTC Score 65.0 AI-016-20260922-global-ai-ne-d0716c4b

Rabbit’s new AI agent doesn’t need an R1 to run

Rabbit, the company behind the underwhelming R1 device, is rolling out a standalone AI agent that you don't need its hardware to use, as reported earlier by Wired. The startup says its new OS3 "agentic operating system" runs in the cloud but operates locally across Windows, Mac, and Linux devices. According to Rabbit, you can […]

InfoWorld AI 2026-09-22 19:00 UTC Score 49.0 USR-0126-20260922-global-ai-ne-6f2a66ed

AWS launches CloudWatch Omni to unify observability for AI agents and applications

As enterprises continue to move AI agents and agentic applications into production, AWS says traditional observability and monitoring tools — including its own CloudWatch service —will struggle to explain why an agent behaved the way it did. CloudWatch uses metrics, logs, and traces to monitor applications and infrastructure across accounts, regions, and services through the AWS Management Console, but that only provides part of the picture, AWS says. Understanding an agent’s behavior requires developers and operations teams to jump between agent-specific observability and evaluation tools such as those available through Amazon Bedrock AgentCore , application performance monitoring, and infrastructure monitoring in CloudWatch . AWS is trying to eliminate that fragmentation by evolving and expanding CloudWatch with a new off-console experience named CloudWatch Omni , bringing agent, application, and infrastructure telemetry together in an application-centric setup to help enterprises investigate and understand agent behavior in context. That means developers and operations teams can start with the application they are investigating, rather than navigating across individual AWS resources and monitoring consoles, the hyperscaler wrote in a blog post presenting CloudWatch Omni . The new tool automatically discovers application topology, the company said, showing how its components are connected, in turn allowing developers and operations teams to query telemetry using natural la…

CIO AI 2026-09-22 19:00 UTC Score 57.0 USR-0125-20260922-global-ai-ne-f62fd851

AWS launches CloudWatch Omni to unify observability for AI agents and applications

As enterprises continue to move AI agents and agentic applications into production, AWS says traditional observability and monitoring tools — including its own CloudWatch service —will struggle to explain why an agent behaved the way it did. CloudWatch uses metrics, logs, and traces to monitor applications and infrastructure across accounts, regions, and services through the AWS Management Console, but that only provides part of the picture, AWS says. Understanding an agent’s behavior requires developers and operations teams to jump between agent-specific observability and evaluation tools such as those available through Amazon Bedrock AgentCore , application performance monitoring, and infrastructure monitoring in CloudWatch . AWS is trying to eliminate that fragmentation by evolving and expanding CloudWatch with a new off-console experience named CloudWatch Omni , bringing agent, application, and infrastructure telemetry together in an application-centric setup to help enterprises investigate and understand agent behavior in context. That means developers and operations teams can start with the application they are investigating, rather than navigating across individual AWS resources and monitoring consoles, the hyperscaler wrote in a blog post presenting CloudWatch Omni . The new tool automatically discovers application topology, the company said, showing how its components are connected, in turn allowing developers and operations teams to query telemetry using natural la…

AWS Machine Learning Blog 2026-09-22 17:28 UTC Score 61.0 AI-057-20260922-official-ai--4273a133

Claude Opus 5.5 is now available on AWS

Claude Opus 5.5, Anthropic's most capable Opus model for agentic coding, knowledge work, and long-running tasks, is now available on Amazon Bedrock and Claude Platform on AWS. This post covers what's new in Opus 5.5, practical guidance, and how to start building with the model on Amazon Bedrock.

WIRED AI 2026-09-22 16:00 UTC Score 70.0 AI-015-20260922-global-ai-ne-33039c0a

Rabbit Is Back, This Time With an AI Agent App

Two years after trying to sidestep mobile apps with dedicated AI hardware, Rabbit is launching OS3, a cross-platform agent that lives on the screens you already use.