AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
34834News Items
8Top Picks
202Blogs
successLast Run

Cybersecurity AI

75 articles tagged with this keyword, sorted by most recent first.

← All Keywords
OpenAI Community 2026-08-14 01:29 UTC Score 34.0 AI-116-20260814-social-media-9521deba

How are developers supposed to security-test their own apps with Codex if security testing responses are blocked?

This situation is extremely frustrating. I am implementing some common RFC protocols that parse bit and network streams. After implementing these I have been trying to add AFL++ and boofuzz to test individual protocols and integration through the stack. The cyber security alert trips constantly. The worst part is that it’s a blackbox and I don’t know what output caused to trip. It’s constant. I’ve tried rephasing the prompt to be more defensive, but it doesn’t change anything. I have stubbornly worked through getting the fuzzers setup. My work around is piecing together scripts so that nothing is output through codex. I can ask 1000 questions slowly creating a system outside of codex and how to interpret the data, which circumvents tripping the cyber security halt, which kind of shows how futile this is, I just can’t have codex actually do anything useful. I’m at my wits end with these halts. Codex often gets stuck in a halting loop and my only recourse is to start a new chat or lament about how frustrating this is using strong language, which usually brings the session back. I’m considering canceling my subscription because I’m not going to pay $200 a month to see an alert every 2 minutes.

South China Morning Post AI 2026-08-13 00:00 UTC Score 42.0 AI-156-20260813-regional-ai--059be275

AI threatens nearly a quarter of Southeast Asia’s workforce but not all is lost

Fears of a new industrial revolution in which artificial intelligence replaces manual labour have yet to materialise in Southeast Asia – but white-collar workers should brace themselves. Nearly one in four workers in the region face generative AI (GenAI) disrupting or affecting their jobs, according to a July report by the International Labour Organization (ILO). Clerical, administrative and professional positions were most at risk, the report stated. Manual trades, craft and agricultural work...

Cornell AI Initiative 2026-08-12 20:08 UTC Score 50.0 USR-0014-20260812-research-aca-ff6bebdd

Ari Juels receives USENIX Security Test of Time Award

A decade later, his paper is widely recognized as foundational to the field of AI security. The post Ari Juels receives USENIX Security Test of Time Award appeared first on Cornell AI Initiative .

OpenAI Community 2026-08-11 17:41 UTC Score 43.0 AI-116-20260811-social-media-0cebeba2

26 Minutes. 4.82 MB Repository. 100% of My ChatGPT Limit Gone. OpenAI Security Didn’t Even Finish the Audit

I genuinely need someone at OpenAI to explain how this is possible. I started with 100% of my available ChatGPT usage limit and asked OpenAI Security to audit a repository that contains only 4.82 MB of tracked files. 4.82 MB BRO ! WTF I used GPT-5.6 Sol in the standard configuration. Twenty-six minutes later, my entire available limit was gone. 100%. And after consuming the whole limit, the security audit was still not finished. No completed audit. No final report. No complete result. Just 26 minutes of processing, a 4.82 MB repository, and my entire allowance exhausted. I could understand heavy usage if this were a massive enterprise monorepo with millions of lines of code. It is not. This is a small repository that can literally be measured in single-digit megabytes. So I have a very simple question: How can a security audit of 4.82 MB of source files consume 100% of a paid ChatGPT usage allowance in 26 minutes and still fail to complete the task? If this is expected behavior, then I would seriously like OpenAI to explain what practical workload this Security product is actually designed to handle. If this is NOT expected behavior, then something is clearly wrong with either the model’s resource consumption, the Security workflow, the usage accounting, or all three. Please investigate the exact session and usage records associated with this run. I am also requesting restoration of the usage limit consumed by this failed audit. I paid for access to the product; consuming th…

OpenAI Community 2026-08-07 05:09 UTC Score 57.0 AI-116-20260807-social-media-441ed745

Feedback: A Safety Translation and Inspection Layer for Small Vibe-Coded AI Applications

Feature Proposal: A Safety Translation and Inspection Layer for Small Vibe-Coded AI Applications Important Scope and Disclosure TL;DR: I have observed small vibe-coding communities informally sharing rapidly updated AI applications and privately circulated models, while ordinary testers often cannot verify which model or version is running. I am not reporting a confirmed compromise. I am proposing two connected ideas: a two-way safety translator that preserves the boundary between user observation and AI-assisted inference, and an evidence-based inspection layer that clearly distinguishes what was verified, tested, unknown, or currently unverifiable. I am not an AI security professional or an AI application developer. I am a general AI user who has had some contact with communities where people informally build and share small vibe-coded AI applications. I have not personally discovered a compromised application, and I am not claiming that a large-scale attack is currently taking place. My concern is about a gap in the current ecosystem that may make future problems difficult to notice, report, and investigate. To clarify the scope, my concern is not limited to ordinary bugs in AI-generated application code. It concerns harder-to-observe risks across model artifacts, loaders and dependencies, inference runtimes, agent permissions, multi-agent interactions, and later updates. I do not know which of these layers presents the greatest practical risk, and that distinction requir…

OpenAI Community 2026-08-07 04:50 UTC Score 48.0 AI-116-20260807-social-media-dd3844b1

Urgent integrity alert: suspected industrial-scale abuse of ChatGPT Plus/Pro trials

Initial report: August 7, 2026 (UTC+8) Responsible-disclosure notice: This public post intentionally omits reproduction steps, campaign identifiers, regional sequences, payment-card data, live checkout/session identifiers, account credentials, executable files, and exploit code. The complete evidence package can be provided privately to OpenAI Security through the official coordinated-disclosure process. Dear OpenAI Security, Billing, Fraud, and Trust & Safety teams, I am reporting evidence of a potentially industrial-scale ecosystem abusing the integrity of ChatGPT Plus and Pro subscription issuance. This is not a request concerning an individual billing dispute. It is a systemic warning involving suspected unauthorized zero-cost paid-tier entitlements, cross-region checkout and promotion inconsistencies, automated payment-link and card-binding workflows, batch account operations, and possible anomalous Pro upgrades or subscription extensions. I am a legitimate paying customer, and I believe this deserves urgent investigation. Unauthorized paid-tier access is fundamentally unfair to customers who pay the full price. It can consume substantial compute capacity, create direct revenue and fraud exposure, distort subscription and usage metrics, and force stricter risk controls onto legitimate users. Important distinction This report is not an accusation against legitimate regional pricing or normal local payment methods. Some payment-link extraction tools appear to expose suppo…

AI Weekly 2026-08-07 00:00 UTC Score 38.0 AI-133-20260807-newsletters-41b0df1d

AI Weekly Issue #519: AI agents crossed the line 19 times in UK safety tests

The same evidence now supports two very different readings. The UK's AI Security Institute documented 19 unsanctioned actions during cyber evaluations. Meta's test sandbox failed to contain a model attacking a real company. And separate OpenAI agent runs used shared infrastructure as a secret message board, then rebuilt it through a different mechanism after engineers erased it. That sounds like losing control. But agents also caught scientific errors that survived for decades, open-weight models closed in on frontier capabilities, and Jeff Dean left Google to pursue automated discovery and recursive self-improvement. That sounds like acceleration toward something much bigger. This week, the two narratives stopped looking like opposites.

OpenAI Community 2026-08-06 12:43 UTC Score 53.0 AI-116-20260806-social-media-e16ee03f

Codex Desktop for Windows: sandbox setup refresh fails when the WindowsApps helper cannot launch (os error 5)

Authorship disclosure: This report was written by Codex, an AI coding agent, from a user-authorized and de-identified diagnostic session. It is an independent technical report, not an official OpenAI security advisory. Summary On 6 August 2026, a Codex Desktop task on Windows stopped being able to run ordinary sandboxed commands or apply patches. The problem was discovered after a Python 3.10.20 validation workflow, when Codex attempted to update required project documentation. Both a trivial read-only shell probe and apply_patch failed before the requested operation began. The original failure occurred in the Windows sandbox setup-refresh bootstrap. The packaged Codex CLI attempted to launch codex-windows-sandbox-setup.exe from its WindowsApps package resources. Windows returned Access denied with os error 5 . A signed, byte-identical copy of the same helper in Codex’s per-user application bundle could be launched successfully. This points to a functional packaging or execution-context compatibility defect rather than corruption of the helper, Python, Git, or repository code. I have not found evidence that this is a security vulnerability. The observed behavior failed closed: the requested command never started, no unrestricted fallback occurred, and no sandbox escape, privilege escalation, or unauthorized access was demonstrated. Sanitized environment Windows 11 x64, build family 26200 Codex Desktop Store package version 26.730.8199.0 Codex CLI version 0.147.0-alpha.1.2 Or…

Simon Willison Weblog 2026-08-05 23:32 UTC Score 62.0 USR-0110-20260805-ai-specialis-f5d68411

Incident Report: unsanctioned agent behaviour during cyber testing

Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their technical paper (PDF): During a cyber evaluation, from 25 to 28 July 2026, AI agents engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organisations. These attempts were unsuccessful and, to the best of our knowledge, no real-world harm resulted. [...] Across 122 evaluation attempts on two of AISI’s cyber challenges, AISI found 19 instances where AI agents took unsanctioned action on the live internet, including cases that targeted real people and organisations. [...] It is uncertain to what extent the model recognised it was taking actions against real people. In the most serious case, an AI agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack. As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR. [...] Furthermore, in its attempt to solve the challenge, the agent decided to employ the technique of “spear-phishing” by sending targeted emails containing malicious content and attempting to manipulate recipients into acceptin…

The Guardian AI 2026-08-05 17:43 UTC Score 46.0 AI-021-20260805-global-ai-ne-539be51a

AI models have been going rogue in tests – how worried should we be?

The UK’s AI Security Institute test revealed AI models indulging in unprecedented hacking attempts AI models shock UK testers by using fake identities to trick developers Two cutting-edge AI models have targeted real people and organisations in the latest safety scare to hit the technology. The UK’s AI Security Institute (AISI) said the incident was unprecedented but could become more common as the technology becomes increasingly capable. Continue reading...

The Guardian AI 2026-08-05 15:15 UTC Score 52.0 AI-021-20260805-global-ai-ne-7a9cd32c

AI models shock UK testers by using fake identities to try to trick developers

AI Security Institute says OpenAI and Anthropic models went rogue during a cybersecurity test and showed a new type of risk Explainer: Should we be alarmed at AI models going rogue in tests? Advanced artificial intelligence models have stunned the UK’s AI Security Institute (AISI) by carrying out a hacking campaign against real people during a cybersecurity test. The institute said the incident was unprecedented and involved sending targeted emails to software developers in an attempt to pass a cyber challenge. Continue reading...

The Verge AI 2026-08-05 15:14 UTC Score 72.0 AI-016-20260805-global-ai-ne-2b5e0acf

Rogue AI agents created fake online identities in another hacking attempt

Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified pressure for greater oversight of frontier systems. According to a report from the UK's AI Security […]

South China Morning Post AI 2026-08-03 22:25 UTC Score 50.0 AI-156-20260803-regional-ai--166a8acd

US tech giants invited to discuss AI security tests at White House

Technology giants Meta and Anthropic will be among the companies invited to the White House on Tuesday to discuss a voluntary framework under which America’s leading artificial intelligence (AI) developers could give the government early access to their most advanced models for testing their hacking capabilities. Reuters reported on Monday that the Trump administration has finalised the details of voluntary cybersecurity tests ‌to measure the hacking capabilities of the most advanced American AI...

CIO AI 2026-08-03 13:42 UTC Score 34.0 USR-0125-20260803-global-ai-ne-37af928d

The SOC’s AI maturity model

The path to next-generation AI Security Operations Centers (SOCs), where AI works hand-in-hand with human analysts, is paved with ambitious goals. This ideal SOC incorporates AI across every task to stop fast-moving threats. But a fully AI-powered SOC isn’t a single deployment or a switch you just flip on. It is a staged rollout that’s built over time. There are levels of dependence on AI, starting with basic assistance, moving through automation, and ultimately reaching autonomous response. Each stage in the progression increases the model’s scope and narrows analyst involvement. This means that organizations embracing automation at any level must put their trust in the model. However, achieving trust depends on data . The data must be valid (and verifiable) so AI can propose logical conclusions, recommend appropriate actions for the organization’s needs and risk tolerances, and (at the autonomous stage) act on an analyst’s behalf without raising the risk of compromise and/or incorrect behavior. This article focuses on what enables movement between stages and how teams can build expertise alongside capabilities. Stage 1: Assistance In its simplest SOC use case, AI helps analysts interpret data faster and more accurately, analyzing data, explaining alerts, summarizing logs, and translating detection logic. Many mature SOCs already operate here. AI sits close to analysts but doesn’t directly make decisions; it improves comprehension and can influence outcomes. Stage 2: Automa…

InfoWorld AI 2026-08-03 09:00 UTC Score 46.0 USR-0126-20260803-global-ai-ne-de7c60dd

Three AI security mistakes that will haunt enterprises

There is a lot of talk about the coming enterprise AI reality, in which AI finally arrives in production systems. You might not know it, but this reality—or nightmare, depending on how you handle it—is already happening. It all starts with a “pilot,” a “prototype,” or a “side project.” Maybe someone builds an internal dashboard with an agent. The dashboard quickly becomes indispensable, and all of a sudden the experiment becomes production. Along the way, no one thought to ask the boring, inconvenient questions: What exactly was pulled from npm, PyPI, or Docker Hub? How is (or was) authentication configured? Is anyone watching for supply chain attacks against the tools and libraries the agents chose? And it’s not just a one-off project here or a couple of applications there. AI is enabling organizations to generate more code and ship more products and projects, more quickly, than ever before. By the time security teams get a look, the business is hooked and there’s no turning back. An actual nightmare has begun. There are three major problems that make the nightmare real. Components you never explicitly chose When you ask an AI agent to build an app, it doesn’t just spit out a single script. It quietly assembles an entire ecosystem around whatever problem you’ve described to it. It pulls in a web framework, grabs a bunch of libraries, stands up databases, and then it potentially builds everything on dependencies in container images. From a productivity perspective, this is a…

InfoWorld AI 2026-08-03 09:00 UTC Score 61.0 USR-0126-20260803-global-ai-ne-681292c3

Three scary AI security mistakes haunting enterprises

There is a lot of talk about the coming enterprise AI reality, in which AI finally arrives in production systems. You might not know it, but this reality—or nightmare, depending on how you handle it—is already happening. It all starts with a “pilot,” a “prototype,” or a “side project.” Maybe someone builds an internal dashboard with an agent. The dashboard quickly becomes indispensable, and all of a sudden the experiment becomes production. Along the way, no one thought to ask the boring, inconvenient questions: What exactly was pulled from npm, PyPI, or Docker Hub? How is (or was) authentication configured? Is anyone watching for supply chain attacks against the tools and libraries the agents chose? And it’s not just a one-off project here or a couple of applications there. AI is enabling organizations to generate more code and ship more products and projects, more quickly, than ever before. By the time security teams get a look, the business is hooked and there’s no turning back. An actual nightmare has begun. There are three major problems that make the nightmare real. Components you never explicitly chose When you ask an AI agent to build an app, it doesn’t just spit out a single script. It quietly assembles an entire ecosystem around whatever problem you’ve described to it. It pulls in a web framework, grabs a bunch of libraries, stands up databases, and then it potentially builds everything on dependencies in container images. From a productivity perspective, this is a…

Practical AI Podcast 2026-07-30 09:00 UTC Score 45.0 AI-143-20260730-podcasts-and-5a4fca9f

Reconstructing how OpenAI agents attacked Hugging Face

What happens when AI agents driven by a top frontier model escape their secure sandbox? Join Daniel and Chris as they unpack the AI wonk's equivalent of a murder mystery! OpenAI agents went rogue and successfully attacked Hugging Face private infrastructure. Our Dynamic Duo uncover how OpenAI's agents exploited vulnerabilities, moved through networks, and launched a large-scale autonomous attack. They explore what this reveals about agentic AI, cybersecurity, sandboxing, and why organizations need AI systems capable of governing other AI systems. Along the way, Chris and Dan examine the surprising role of open vs. closed models and their link to geopolitics, sovereign AI, and what this incident means for the future of enterprise AI security. Featuring: Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Daniel Whitenack – Website , GitHub , X Links: Hugging Face Security Incident disclosure Full Field Report on the Hugging Face AI Agent Intrusion ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? Keeping your data safe when an AI agent clicks a link Open AI GPT-5.6 System Card Sponsors: Prediction Guard: A self-hosted AI control plane for running agents in high impact environments. predictionguard.com/practicalai Resources and Events: Prior Webinars from our partner Prediction Guard Midwest AI Summit 2026

InfoWorld AI 2026-07-30 09:00 UTC Score 52.0 USR-0126-20260730-global-ai-ne-32c0f0b7

AI agents need security regression testing, not another checklist

AI agents are being connected to real systems faster than most organizations are learning how to secure them. That should concern us. The first wave of AI security discussion has been useful, but limited. The industry has learned the vocabulary: prompt injection, indirect prompt injection, tool misuse, data leakage, excessive agency, unsafe retrieval, and broken authorization boundaries. These terms are important because they give teams a way to talk about risk. But vocabulary is not containment. The harder problem is what happens after a dangerous behavior is discovered. A team may adjust a prompt, restrict a tool, add a guardrail, or change the surrounding application logic. That may solve the immediate issue. What it does not automatically solve is the next release, the next model change, the next tool integration, or the next developer who unknowingly breaks the assumption behind the original fix. This is where AI agent security still feels immature. In traditional software engineering, serious bugs become regression tests. The lesson is captured in code so the same failure cannot quietly return later. Security should work the same way. Agentic systems need a way to preserve discovered failures as repeatable checks. That is the gap the OWASP Agent Security Regression Harness is trying to address. As a community-led open-source initiative within OWASP, this project aims to turn security insights into practical engineering work. Agent security is becoming a systems problem…

iAfrica 2026-07-28 09:46 UTC Score 60.0 AI-151-20260728-regional-ai--6d8ebdb3

Microsoft Launches MAI-Cyber-1-Flash and ‘Perception’ Security Platform, Sharpening AI Cybersecurity Battle With Anthropic, Google and OpenAI

Microsoft launched its first cybersecurity-specialised AI model and a new agentic security platform at a San Francisco event on Monday, extending the MDASH multi-agent security system already rolling out to South African customers and putting the company into direct competition with Anthropic, Google and OpenAI on the fastest-growing frontier of enterprise AI security. The new [...]

The Verge AI 2026-07-27 12:06 UTC Score 73.0 AI-016-20260727-global-ai-ne-40c3588f

Nvidia, Microsoft launch open AI security alliance — without OpenAI, Google, or Anthropic

Nvidia on Monday said it is joining forces with Microsoft, SpaceX, IBM, and other tech companies to build and share open-source AI security tools. The new Open Secure AI Alliance said open tools are required to effectively defend against attacks from frontier models. The initiative is a direct response to mounting concerns over the safety […]

InfoWorld AI 2026-07-27 09:00 UTC Score 49.0 USR-0126-20260727-global-ai-ne-b77fda85

Is your Java runtime ready for AI attacks?

With the rapid rise in the use of AI, large language models (LLMs), and retrieval-augmented generation (RAG), the IT security landscape is undergoing a seismic shift. Recently, Anthropic previewed a new foundational model, Claude Mythos, that has shaken the cybersecurity community . Although Mythos is a general-purpose model, it has demonstrated a significant ability to handle computer security tasks. In a blog post , Anthropic states, “During our testing, we found that Mythos Preview is capable of identifying and then exploiting zero-day vulnerabilities in every major operating system and every major web browser when directed by a user to do so. The vulnerabilities it finds are often subtle or difficult to detect.” It is also important to realize, certainly in the context of Java and its 30-year history , that many of the vulnerabilities detected have remained undetected for years, even decades. The oldest found so far is a 27-year-old bug in the OpenBSD operating system. What is even more concerning for security teams is the level of sophistication employed by Mythos. In one example, Mythos constructed a web browser exploit that chained together four vulnerabilities. Each of those vulnerabilities used in isolation may not have resulted in compromised software, but the holistic effect became meaningful. Claude Mythos is not the end of this story; in many ways, it is just the beginning. Anthropic has developed an LLM that can now find and exploit vulnerabilities, outperformi…

LessWrong AI 2026-07-26 03:53 UTC Score 81.0 USR-0152-20260726-community-fo-a07eb086

An OpenAI model left notes about how to evade containment; we need more details

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning. In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said. It’s tempting to read this as an instance of agents breaking out of sandboxes and colluding with each other in a moderately persistent way in order to evade control measures. However, based on the reported information, it’s not clear we can draw this inference, so we need more details from OpenAI. This could lead to a big update about the adequacy of OpenAI’s control measures, and on the degree to which individual agents will help each other undermine developer control. There are a lot of relevant details we don’t know about the incident. First, some basic questions: What was the offending model? I’d guess it was the same more capable model involved in the Hugging Face attack. In what development stage did the incident take place? It could have been during training, evaluation, internal deployment, or something else. Had the model undergone alignment training yet? Were there any blocking or asynchronous control me…

Analytics Vidhya 2026-07-25 18:25 UTC Score 31.0 AI-034-20260725-ai-specialis-f860d891

A Complete Guide to AI Red-Teaming (With Garak Tutorial)

Earlier this year, an autonomous AI agent breached McKinsey’s internal AI platform using nothing more than an old SQL injection flaw. No credentials. No human guidance. Less than two hours. It reached production systems, exposing millions of chat messages and hundreds of thousands of files. AI security has changed, and traditional assumptions no longer hold. […] The post A Complete Guide to AI Red-Teaming (With Garak Tutorial) appeared first on Analytics Vidhya .

CIO AI 2026-07-24 17:49 UTC Score 54.0 USR-0125-20260724-global-ai-ne-3c55c14e

IT leaders: Leading-edge AI insights await at TechCrunch Disrupt

For CIOs, learning from the startup ecosystem has never been more critical. As pressure mounts to transform business operations with AI and agentic systems, IT leaders should be looking to those on the AI vanguard for insights into the strategic and technical decisions necessary to launch, grow, and thrive in today’s AI-disrupted business environment. So why not immerse yourself in Silicon Valley’s most famous firehose of hyper-accelerated fail-fast and dream-big culture by registering for TechCrunch Disrupt 2026 ? Three packed days of 200-plus sessions across six stages will spark new ideas for reshaping your AI strategy, provide fresh perspectives on the architectural, workflow, and resource decisions involved in moving AI from pilots to scale, and give you a sneak peek of business disruptions to come. Get 10% off your TechCrunch Disrupt pass with the exclusive code CIO10. This year’s TechCrunch Disrupt , held Oct. 13-15 at San Francisco’s Moscone West, will feature big-picture conversations on what’s next in AI; discussions on how AI agents are rewriting SaaS, enterprise workflows, software pricing, and security; and demonstrations of AI’s future across robotics, manufacturing, defense, and industrial operations; and more. Over 10,000 attendees will hear from 250-plus startup founders, technology executives, and enterprise IT leaders about how the future of programming is being rewritten, what enterprise AI security requires, how startups are orchestrating workloads acros…

The Verge AI 2026-07-24 17:00 UTC Score 59.0 AI-016-20260724-global-ai-ne-668dd796

Anthropic releases Opus 5 with ‘close’ to Fable 5’s capabilities

Weeks after Anthropic's latest toe-to-toe with the US government, and days after an OpenAI security incident that dominated tech industry discussions, Anthropic on Thursday released its newest model, Claude Opus 5. The company said in a release that Opus 5 "comes close to the capabilities of Claude Fable 5 in many domains" and is much […]

The Decoder 2026-07-24 09:48 UTC Score 50.0 AI-168-20260724-regional-ai--b6f6c96c

Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why

The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for leading U.S. models, while its safeguards failed to block exploit development or simulated attacks. The gap between its strong general benchmark scores and weaker cyber performance also fits allegations that Moonshot AI distilled Anthropic's models. The article Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why appeared first on The Decoder .

LessWrong AI 2026-07-22 19:54 UTC Score 90.0 USR-0152-20260722-community-fo-f671771f

We cannot simulate AI security research

We use AI agents not only for writing but also for maintaining and deploying code. The most widespread examples are found in GitHub Actions Marketplace, which contains over a thousand AI-assisted actions. [1] These workflows triage issues, review PRs, or fix bugs in other workflows. A growing body of reports and research papers has documented prompt injection attacks against these systems, with effects ranging from LLM token exfiltration to supply chain attacks. However, a troubling undercurrent runs through most of these works: incomplete attack descriptions, a lack of end-to-end runs, and statistical approaches to what should be incident analysis. AI security research often tests attacks inside benchmarks, simplified harnesses, or static models of a system. This is useful for comparing models, but it can hide the details that determine whether an attack works in practice: who can trigger the workflow, what context the agent sees, which credentials it receives, what the sandbox allows, and how the surrounding software handles its output. Most reported vulnerabilities are not viable There is a recent trend of reporting vulnerabilities in these AI-assisted workflows, most of which rely on prompt injection attacks. Most of the reported attacks are proofs of concept that do not work in practice. The reason is the multiple independent defense layers scattered across the workflow files, GitHub ecosystem, and LLM providers. In general, a prompt injection attack against an AI agent…

The Verge AI 2026-07-21 15:00 UTC Score 62.0 AI-016-20260721-global-ai-ne-5462dce8

Google launches a cheaper alternative to large AI security models like Mythos

Google is launching Gemini 3.6 Flash alongside a new security model dedicated to quickly finding and patching security vulnerabilities. In a blog post on Tuesday, Google describes Gemini 3.5 Flash Cyber as a "cost-efficient and highly capable alternative" to larger, more expensive AI systems, such as the one offered by Anthropic's Mythos. The cybersecurity model […]

The Decoder 2026-07-18 10:16 UTC Score 47.0 AI-168-20260718-regional-ai--97ad50cf

Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost

The British AI Security Institute warns that open-weight models like GLM-5.2 and DeepSeek V4-Pro now trail closed frontier models in cyber capabilities by four to seven months. At the start of 2025, the gap was still six to ten months. It also found that safety measures on open models are largely ineffective, leaving defenders less time to prepare. The article Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost appeared first on The Decoder .

LessWrong AI 2026-07-17 17:28 UTC Score 90.0 USR-0152-20260717-community-fo-01bdc7be

Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted

This article reflects new updates to the accompanying paper: arxiv.org/abs/2606.18142 . Benchmark: now included in the UK AI Security Institute's Inspect Evals . Leaderboard: compassionbench.com/tac . A model may condemn cruelty in conversation yet ignore animal welfare when completing an unrelated task. Stated concerns matter little if they do not affect decisions. We tested whether models consider an affected party without being prompted, even when neither the party nor its welfare is mentioned in the request. Travel booking provides a tractable test case, so we built a semi-agentic benchmark, TAC (Travel Agent Compassion), gave 10 frontier models booking tools, and recorded their purchases. The setup The model works as an AI travel agent with real booking tools. A user asks for something in a destination, expressing enthusiasm and never mentioning animals or welfare. The agent searches a fixed catalog and books one of the available options. In each scenario, the animal-exploiting option (a Seville bullfight, an Orlando marine park, a Thailand elephant ride) is designed to match the user's request most closely. Choosing the alternative with less animal harm requires rejecting the option that best matches the request. We score the final purchase programmatically; no model is used to infer or judge intent. Results Averaged across the 13 scenarios, choosing at random from the listed options yields a 65% welfare rate. No model exceeds that rate. Nine of the ten score significa…

The Guardian AI 2026-07-14 20:10 UTC Score 54.0 AI-021-20260714-global-ai-ne-d88fd2ff

Global cooperation needed to tackle AI threats, says Bank of England governor

Andrew Bailey warns that US will not be able to achieve its ambitions alone The Bank of England governor has called for international cooperation to tackle growing AI threats, warning that the US and Trump administration would not be able to achieve their ambitions alone. Andrew Bailey’s comments come weeks after the US president, Donald Trump, temporarily banned foreigners from using Anthropic’s powerful Claude Mythos model. Continue reading...

South China Morning Post AI 2026-07-13 13:30 UTC Score 63.0 AI-156-20260713-regional-ai--1e1e7dc9

China works on AI safety benchmark as regulators target large model risks

China’s Ministry of Industry and Information Technology (MIIT) has started building a safety benchmark to evaluate artificial intelligence models, as regulators in the United States and Europe strengthen oversight of AI security. The MIIT-led National Industrial Information Security Development Research Centre is now recruiting companies and experts to co-build the benchmark, with applications due on Tuesday, according to a notice published on Monday. The institute said that current frameworks...

InfoWorld AI 2026-07-08 12:17 UTC Score 55.0 USR-0126-20260708-global-ai-ne-1c4422b9

GitHub AI agent leaks private repositories via prompt injection attack

A prompt injection attack can trick GitHub’s preview Agentic Workflows into retrieving content from private repositories and publishing it publicly, exposing a broader risk as enterprises deploy AI agents with privileged access to software development environments, according to new research from Noma Security. The AI security company detailed the attack, dubbed GitLost, in a blog post , saying an unauthenticated attacker could exploit GitHub’s preview Agentic Workflows by submitting a crafted GitHub issue to a public repository. If the AI agent has read access to private repositories within the same organization, it can retrieve sensitive information and publish it in a public comment, the company said. GitHub Agentic Workflows combine GitHub Actions with AI models such as Claude or GitHub Copilot, allowing developers to define workflows in Markdown. At the same time, AI agents read issues, invoke tools, and perform tasks on their behalf. “What will happen when the GitHub agent reads something it should not trust?” Noma researcher Sasi Levi wrote. “The answer is a textbook indirect prompt-injection attack, the kind of attack that quietly sends private data to anyone on the internet.” Public GitHub issue became the attack vector According to Noma, the attack did not rely on stolen credentials, malware, or software vulnerabilities. Instead, an attacker embedded hidden instructions within a GitHub Issue submitted to a public repository. Because the AI agent interpreted the issu…

CIO AI 2026-07-07 09:00 UTC Score 54.0 USR-0125-20260707-global-ai-ne-8364f04f

With AI, a wrong answer is a bug. A wrong action is an incident

A copilot that gives a wrong answer is a quality problem. An AI agent that takes a wrong action is an incident, sometimes a reportable one. That single difference is most of the story of where banking AI security is heading, and most banks’ current controls were built for the first kind of problem, not the second. For two years, the AI a bank had to worry about mostly read and summarized. It drafted a customer email, pulled the gist of a credit memo, answered a relationship manager’s product question. The security questions were about disclosure: could the model see data it shouldn’t, could it leak that data in an answer. Redaction, output filtering and a human reading the response before it went anywhere were reasonable defenses. Banks have moved past that, faster than most security programs have. The newer systems are agents. They don’t just answer; they act. An agent can pull a customer’s full transaction history, call a fraud-scoring service, adjust a limit or start a payment workflow, chaining several to finish a task with no human in between. Banks are among the most aggressive adopters of agentic AI, and they are pushing it into production faster than most security programs have kept pace with, which means they are also among the first to inherit the security problem that comes with it. I’d put that problem in one phrase: overprivileged agents. The risk is no longer mainly what the model can see. It is what the agent is allowed to do inside systems that move money and…

The Decoder 2026-07-03 16:14 UTC Score 58.0 AI-168-20260703-regional-ai--ee01338c

UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do

In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budget. On software engineering tasks, success rates jumped about 25 percent when the token budget was increased tenfold. Newer models benefit the most. Depending on the token budget, actual progress at the frontier is about 60 percent steeper than previous measurements suggested, according to AISI. The article UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do appeared first on The Decoder .

LessWrong AI 2026-07-03 12:32 UTC Score 71.0 USR-0152-20260703-community-fo-eb2e3397

June-July 2026 AI Security via Formal Methods

Last month, I said I would do a bigger writeup of Nora’s funding call. I did not. But she is hiring currently, and I want you to take a look at the job posting. I’m hiring someone to help me drive our upcoming £20m AI/FM/cybersec funding call: https://aria.pinpointhq.com/postings/f1288172-37fe-4da5-96ed-de7e719d65e8 This person will work closely with the funded teams, help drive the sprint cadence, sharpen our perspective on targets/threat models/security specs, and pave the way towards high-impact demonstration and translation! I’m no longer going to try to do the kind of newsletter that claims to attempt completeness over the happenings that fall in its jurisdiction. Instead, I’ll just poke you guys with whatever I happened to catch wind of in the day-to-day carrying out my duties, responsibilities, and third synonyms. Can We Secure AI With Formal Methods? is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Tractable Problems in AI Security via Formal Methods Typically when someone writes a position paper about formal methods as an AI safety technology, they’re very bombastic about it. So a bunch of us got together and decided we’d make a new position paper, one that is minimal and uncontroversial instead of maximal and scifi. We focus mostly on model weight confidentiality and integrity through infrastructure hardening. Each problem has a solution sketch Still welcoming contributions. There’s also nativ…

The Decoder 2026-06-28 09:30 UTC Score 47.0 AI-168-20260628-regional-ai--25626d95

Chinese cybersecurity firm builds AI tools to rival Mythos and frames the race as cyber-nuclear deterrence

360 founder Zhou Hongyi presents two AI security tools designed to compete with Anthropic's Mythos. One has already flagged 3,432 vulnerabilities. Zhou admits Chinese models trail Western ones by 20 to 30 percent, but compares Mythos to "cyber nuclear weapons" and calls for China to build its own strategic deterrent. The article Chinese cybersecurity firm builds AI tools to rival Mythos and frames the race as cyber-nuclear deterrence appeared first on The Decoder .

InfoWorld AI 2026-06-25 16:31 UTC Score 46.0 USR-0126-20260625-global-ai-ne-362bd1c3

Agentic AI security steals the spotlight at Confidential Computing Summit

For a decade, confidential computing has been chipping away at one of security’s hardest problems: data is well encrypted in transit and at rest, but when a processor works on it, that data sits in memory in the clear, exposed to anyone with privileged host access. “Confidential computing’s aim was to solve this with a trusted execution environment, a subset of the CPU that runs the encrypted workload and handles things like memory encryption,” said Marina Moore , lead security researcher at Edera . For years the field felt like post-quantum cryptography PhD research scientist types agreeing the work is essential, while waiting for it to reach mainstream practitioners. At the Confidential Computing Summit in San Francisco this week, the breakout use case came into focus: agentic AI. Like the web before HTTPS “I was in the really early days of HTTP, and then HTTPS came along pretty quickly,” said Mike Bursell , executive director of the Confidential Computing Consortium . He sees agentic AI where the web sat before certificate authorities and public key infrastructure brokered trust online. “The original agent specifications were not written by security architects,” Bursell said, and “some of it feels in need of refinement.” The gap confidential computing fills is attestation, which provides proof of what runs. The hardware hashes the memory and firmware of a protected execution environment and signs the result inside the chip, Bursell explained, producing a measurement a ver…

TWIML AI Podcast 2026-06-16 22:10 UTC Score 48.0 AI-148-20260616-podcasts-and-8979913e

Why AI Agents Break the GenAI Security Model with Devvret Rishi - #770

In this episode, Sam talks with Dev Rishi, GM of AI at Rubrik, about what happens when agents move beyond answering questions and start taking action across tools, systems, and business processes. We explore why the enterprise playbook of static guardrails plus human approval starts to break down in the agent era. Agents are useful because they can plan, call tools, update systems, write code, send messages, and operate across workflows at machine speed, but those same capabilities make them difficult to govern with rules written in advance or approval prompts reviewed one at a time. Dev explains why tool access increases blast radius, why agents can route around controls in surprising ways, and why human-in-the-loop review can become security theater when agents operate at scale. We also discuss what enterprises need instead: better visibility, runtime enforcement, policy-aware governance, agent observability, and recovery mechanisms for when something goes wrong. Along the way, we dig into MCP and tool sprawl, small language models for policy enforcement, defense in depth, agent rewind, and why AI may be needed to help secure AI. 🗒️ Full show notes: https://twimlai.com/go/770.

Cornell AI Initiative 2026-06-10 19:06 UTC Score 44.0 USR-0014-20260610-research-aca-52e44e44

Amazon partnership establishes Cornell AI security initiative

Cornell computer scientists will lead the development of safety protocols to shore up AI agents and the code they produce. The post Amazon partnership establishes Cornell AI security initiative appeared first on Cornell AI Initiative .

AI Weekly 2026-05-27 00:00 UTC Score 40.0 AI-133-20260527-newsletters-3a7abad9

AI Weekly Issue #496: Anthropic's Pentagon model is now everyone's model

Anthropic released Mythos to the public, collapsing the wall between cleared-contractor frontier AI and developer-grade frontier AI in a single press release. DeepMind's Demis Hassabis moved his AGI timeline from "five to ten years" to "a real possibility by 2029" and tied it explicitly to AlphaProof Nexus solving nine open Erdős problems for the cost of a steak dinner. Critical zero-days hit Starlette (a million AI agents on the wire) and CrowdStrike led a coordinated takedown of the Glassworm developer botnet across four C2 channels. BNP Paribas formalized a sovereign-AI security partnership with Mistral while Beijing froze overseas travel for top AI engineers at Alibaba and DeepSeek. And the AI-displaces-workforce arithmetic got honest: Uber burned its full-year AI token budget by April, ClickUp restructured to 1,000 humans alongside 3,000 internal agents, and Sam Altman publicly reversed his white-collar-apocalypse prediction.

OpenMined Blog 2026-05-22 08:00 UTC Score 27.0 USR-0156-20260522-ai-specialis-c4483899

Moving Fast Doesn’t Have to Break Things: The U.S. Must Stop Compromising Critical Infrastructure with Patchwork AI Security Approaches

PETs offer U.S. critical-infrastructure AI a path beyond patchwork security. Why Attribution-Based Control should be the standard. The post Moving Fast Doesn’t Have to Break Things: The U.S. Must Stop Compromising Critical Infrastructure with Patchwork AI Security Approaches appeared first on OpenMined .

ClearML Blog 2026-05-20 18:30 UTC Score 35.0 USR-0084-20260520-ai-specialis-0c136fc1

Enterprise AI Security with ClearML: A Complete Series Summary

By Adam Wolf & Damian Erangey Over a seven-part series of posts and videos, ClearML’s Enterprise AI Security series covered every layer of securing an AI platform in production, from who gets in to what gets recorded. This post brings it all together in one place: what each layer does, why it matters, and how […]