AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
60663News Items
8Top Picks
322Blogs
failedLast Run

AI Ethics

200 articles tagged with this keyword, sorted by most recent first.

← All Keywords
Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 67.0 AI-084-20260929-research-pap-7bbc32a6

AgentPEN: A Prediction-Explanation Network for Sequential Stock Movement via LLMs and Recurrent Generation

The importance of explainability in stock prediction is increasingly recognized, especially for audit and regulatory purposes. Meanwhile, financial news corpora are often key drivers behind stock price fluctuations. However, the raw news data obtained is usually highly noisy, has a highly variable scope of influence in time and space, and is not precisely synchronized with stock price data. In this paper, we propose a prediction-explanation network called AgentPEN, which can provide clear explanations for complex temporal price patterns. Specifically, AgentPEN jointly aligns text and price streams by an LLM-based Representation Fusion Agent and then adopts a Deep Recurrent Generation module to explore the distribution of stock movements. The LLM-based Representation Fusion Agent is designed in a Selection-Memory-Fusion manner: the Text Selection Module picks up useful information from massive text data; the Text Memory Module evaluates and writes the text memory from a two-view perspective, including Temporal Memory and Spatial Memory; the Information Fusion Module models the interaction between text and price data. Next, the fused representation is sent to the Deep Recurrent Generation module to convert insights into stock movement predictions. Experiments on multiple real-world datasets have shown that AgentPEN surpasses the state-of-the-art baselines both in prediction accuracy and explainability.

Transactions on Machine Learning Research 2026-09-29 00:00 UTC Score 44.0 AI-084-20260929-research-pap-368195f2

Robustness Against Weak or Invalid Instruments: Exploring Nonlinear Treatment Models with Machine Learning

We discuss causal inference for observational studies with possibly invalid instrumental variables. We propose a novel methodology called two-stage curvature identification (\texttt{TSCI}) by exploring the nonlinear treatment model with machine learning. The first-stage machine learning enables improving the instrumental variable's strength and adjusting for different forms of violating the instrumental variable assumptions. The success of \texttt{TSCI} requires the instrumental variable's effect on treatment to differ from its violation form. A novel bias correction step is implemented to remove bias resulting from the potentially high complexity of machine learning. Our proposed \texttt{TSCI} estimator is shown to be asymptotically unbiased and Gaussian even if the machine learning algorithm does not consistently estimate the treatment model. Furthermore, we design a data-dependent method to choose the best among several candidate violation forms. We apply \texttt{TSCI} to study the effect of education on earnings.

LessWrong AI 2026-09-28 20:07 UTC Score 82.0 USR-0152-20260928-community-fo-9aa63e7e

The Alignment Community Is Unintentionally Building a Censor's Toolkit

This is an adaptation of our ICML 2026 position paper (Outstanding Position Paper Award). Read the full paper here and see the project website here . Work together with Phil Hackemann. TLDR "Alignment" is usually treated as a synonym for achieving good and safety in the world. But it isn't necessarily. Alignment methods are purpose-agnostic: they make a model do what someone wants, and nothing in the methodology guarantees that someone has good intentions. The same techniques we build to stop models from giving bomb-making instructions can just as easily be used to censor historical facts, political dissent, or inconvenient opinions. So we need to understand: alignment techniques are dual-use technologies . This isn't a thought experiment. State censorship regimes and individual model providers are already misusing alignment methods, and by perfecting these methods we are providing an ever improving censor toolkit. Three trends make this urgent to discuss: AI is becoming a primary information source for hundreds of millions of people, the model-provider market is an oligopoly, and global democratic backsliding is accelerating. We don't think the answer is "stop aligning models." We think it's transparency, verifiable alignment, model pluralism, and the alignment community actually reckoning with dual-use. ___________________________________________________________________ No guarantees that alignment leads to good Many alignment researchers (ourselves included) have gotten u…

IEEE Spectrum Machine Learning 2026-09-28 18:00 UTC Score 58.0 AI-020-20260928-global-ai-ne-374cdb53

A New IEEE STEM Book Series for Tweens from TryEngineering

IEEE TryEngineering is dedicated to inspiring intellectual curiosity in children. The technologies shaping our world, including in the realms of artificial intelligence, electric vehicles, and ocean exploration, are evolving rapidly. Helping young learners understand the concepts is essential to preparing the next generation of problem-solvers, creators, and engineers. TryEngineering has introduced a STEM book series for youngsters ages 8 to 12 through the Lerner Publishing Group . The series, Tomorrow’s Technology With TryEngineering, Powered by IEEE, makes complex topics more approachable and engaging, with each book combining age-appropriate explanations, real-world examples, and design challenges that encourage curiosity and critical thinking. The series is based on ebooks and videos available at tryengineering.org . For the series, TryEngineering partnered with several other IEEE groups including the Communications , Computer , and Oceanic Engineering societies and the Transportation Electrification Council . Whether used in the classroom, a library, or at home, the books can help pupils connect STEM concepts to the technologies they encounter every day, including computers and smartphones. Six topics in the collection Here are the books in the new collection: Artificial Intelligence: The Future of Smart Technology explores the systems behind streaming services, search engines, and health care. Readers learn how AI works while exploring ethical concerns such as bias, de…

Gradient Flow 2026-09-28 13:53 UTC Score 66.0 USR-0119-20260928-ai-specialis-9168cea8

Six Reasons I Think Open AI Models Will Win

In my conversations with developers and AI teams, I’m struck by how many are exploring moving more of their inference workloads to open models. I’ve touched on some of their reasons before , but here I want to look further ahead. I’ll admit my bias toward open source. I was around during the dot-com era Continue reading "Six Reasons I Think Open AI Models Will Win" The post Six Reasons I Think Open AI Models Will Win appeared first on Gradient Flow .

CIO AI 2026-09-28 13:00 UTC Score 51.0 USR-0125-20260928-global-ai-ne-3ed3b225

If AI makes the decision, who owns the consequence?

Twenty-two rows. That was the AI inventory in a board pack I read last year, and it was a good one. Every system named, every owner listed, every risk rating filled in and colour-coded. Somebody had worked hard on it, and the committee approved it in four minutes. I kept looking for a column that wasn’t there. The pack could tell you who owned each system. It could not tell you who was allowed to say no to one. Those sound like the same thing right up until the morning they aren’t. 1. Signal. AI has moved inside the operating model When a director asked me last spring what our AI risk was, I gave her the honest answer. Which was that I no longer knew, because the question had changed shape while everybody was still busy answering the old one. Twelve months earlier I could have told her precisely. We had a model inventory, a review process and a register entry saying the right things about bias and data quality. All of it was true, and all of it had quietly stopped being the point, because the systems had moved inside the work. They chose which cases went to which queue, which prices moved and which alerts a human ever saw. Foundry’s State of the CIO 2026 report puts the shift plainly. Roughly three-quarters of leaders say AI is already reshaping how their operations run, and the number climbs higher in financial services. Around seven in 10 expect deeper involvement in agentic systems this year. Don’t let those figures become the story. Adoption numbers are the easiest thing…

Medianama AI 2026-09-28 12:10 UTC Score 43.0 USR-0211-20260928-regional-new-15c0e72b

Google tests AI shopping on Flipkart through Gemini and AI Mode. Here’s what we don’t know

Google is testing purchases from Flipkart within Gemini and AI Mode in India, raising questions about pricing, transparency, competition, payments, refunds, and trust in AI-powered shopping. The post Google tests AI shopping on Flipkart through Gemini and AI Mode. Here’s what we don’t know appeared first on MEDIANAMA .

LessWrong AI 2026-09-28 04:52 UTC Score 59.0 USR-0152-20260928-community-fo-91eafb53

Why did it get Sparser?

I think that Polysemanticity in artificial neural networks could be the key to making better and smaller models. I having been working on a little project on trying to induce Polysemanticity at a small scale to compare performance, my first approach was to make the bias more adaptable, I called this the Flexbias, however I got a sparser neural network. What is the flex bias? I used a standard transformer architecture including the MLP, Since i wanted to change how the information is processed I altered how the traditional MLP works by changing how the bias is obtained. Normal bias : Flex bias: where . That is the bias term is not fixed and is computed from the input Both Neural Network are about 3.7M and share almost identical architecture aside the bias term stated above. The SAE is a plain ReLU with 4096 features.I obtained the following results. Recon MSE mean L1 Features active Normal 0.0051 0.294 49.5% Flexbias 0.0116 0.127 38.3% My initial alternative hypothesis was that since flexbias sees the inputs more, it should understand more representations and be the least sparse, However it looks like the flexbias actually made the neural net understand the data easily with little need for polysemanticity. This is my take but I don't feel satisfied about it. There should be a better explanation to why it got sparser, I would run more experiments but I also need hindsight. Links Github My take on Polysemanticity Discuss

LessWrong AI 2026-09-26 22:55 UTC Score 52.0 USR-0152-20260926-community-fo-c3b54c5f

Extinction does not feel as bad as it should

And we will all go together when we go What a comforting fact that is to know Universal bereavement, an inspiring achievement Yes, we all will go together when we go -- We Will All Go Together When We Go by Tom Lehrer "Everyone dying" does not emotionally feel that bad to me. I know that it is quite a bad thing intellectually, but I have a missing mood. I'm guessing I'm not alone here. When people say things like "If anyone builds it, everyone dies", there is an implicit assumption that everyone is on the same page that "everyone dies" is a really bad, perhaps even the worst, outcome. But not everyone feels this way, and among those, not many have shut up and multiplied to correct for their brain's biases. I thought it would be useful to explore possible reasons why extinction may not feel as bad as it should. Extinction lacks some negative features of lesser catastrophes To be clear, I am not a death fan or extinction enjoyer, I want us to all live forever. The lack of these negative features in extinction clearly does not outweigh the negative fact of the deaths themselves, but are something that some parts of your brain may unduly weigh. No one is left to suffer One reason death is bad is that the people who remain must now suffer the loss of the deceased. This is the thing that causes the most visceral bad feeling when I think of death. If 10% of people die, then probably you or some people you know will be among the 10%. That's a lot of people mourning loved ones and fr…

Cross Validated 2026-09-26 21:48 UTC Score 39.0 AI-113-20260926-social-media-3cdb53d3

detecting bias in causal ML models given varying definitions of treatment

I am evaluating a policy change where full information is replaced by partial information and I want to understand whether or not it adversely affects my causal models. Suppose that a firm sends marketing communications to customers via smartphone and in past years if the customers engage with those communications, event data is sent back to the firm in all regions. However, now in some regions, the engagement data is not sent back to the firm. Consider the following variable definitions: $a_{i, t}$ : action(s) taken by customer i in period t $d_{i, t}$ : communications delivered to customer i in period t $c_{i, t}$ : engagement (clicks) by customer i in period t $X_{i,t}$ : salient characteristics (covariates) of customer i in period t The full signal model: $P(a|c,d,X)$ partial signal model: $P(a|d,X)$ Clearly, the missing variable is the click-through rate, $\pi = P(c|d)$ . And so, I can construct an "apples-to-apples" of regions treated (delivery information only) to untreated regions (click information available) by adjusting for the CTR. The high level idea is that the full signal effect estimate divided by CTR equals the partial signal effect estimate. More precisely: $\hat{a} = \frac{P(a|c,d,X)}{\pi} = P(a|d,X)$ And so I might use a relatively simple causal model, like inverse propensity weighting (IPW) to adjust the effect estimate for treatment propensity, conditioned on $X$ . Does the customer background predict a treated region. If no effect exists then $\frac{P(…

LessWrong AI 2026-09-25 20:50 UTC Score 58.0 USR-0152-20260925-community-fo-7351e788

What did AI researchers think at the end of 2024?

We recently (finally!) got the results of the 2024 survey out. The paper is here , but it’s pretty long, so I’ll tell you the most interesting bits (according to me). But first, quick background : this was the fourth run of the same survey since 2016. We wrote to everyone we could who published in six top-tier AI venues and got a 10% response rate—high! We got 1580 valid responses, but don’t be confused: specific questions often have answers from many fewer researchers, because we gave each person a randomized subset (see Section 2.5). There was almost certainly some non-response bias, but it probably doesn’t make much difference . Researchers filled out the survey in December 2024, so some things have probably changed. To me the most striking results are about extinction or disempowerment (Section 3.9). On average researchers put an 18% chance on “future AI advances causing human extinction or similarly permanent and severe disempowerment of the human species”. Over half said at least 10% and one in three said at least 20%. From the comments, I think people are thinking of a variety of extreme disempowerment scenarios here, not just extinction. And Zooming in, another really interesting thing is that researchers educated in Asia had higher extinction/disempowerment numbers than US or European researchers! This is interesting because a common defense of pushing forward with dangerous AI is that the US is in an arms race with China, and (implicitly) China won’t want to cooper…

Cross Validated 2026-09-25 19:59 UTC Score 33.0 AI-113-20260925-social-media-f44d5e48

How to set up a simulation study?

Here is a hierarchical data generating process (DGP): (Population) Layer 1: $$\theta_i \overset{\text{iid}}{\sim} \text{Beta}(\alpha,\ \beta), \qquad i = 1, \dots, n$$ (Individual) Layer 2: $$y_{ij} \mid \theta_i \overset{\text{iid}}{\sim} \text{Bernoulli}(\theta_i), \qquad j = 1, \dots, k$$ $$\mu = E[\theta_i], \qquad \tau^2 = \operatorname{Var}(\theta_i)$$ Using a sample, my purpose is only to estimate the mean and the variance of the mean estimator. I want to know what values of $n$ and $k$ I should select to get good results (I know this hugely subjective) such that $nk$ is minimized (assume that increasing $n$ by 1 costs the same as increase $k$ by 1). After doing some research, it seems the best way to handle this question is by a simulation study. Assuming $\alpha,\ \beta$ are known, I could sample from the DGP for different combinations of $n$ and $k$ and record the average length of the confidence intervals and average coverage rate at each combination. Since the mean estimator is unbiased regardless of the choice in $n$ or $k$ , I could use average CI length and average coverage rate to make sure I am not getting a misleadingly good coverage rate at the expense of a large CI. Here is how I plan to do this in R (I wrote the code to focus on readability instead of speed - I use a moment based estimator and consider different true combinations of parameters): set.seed(2026) mu_vals $mu[s] tau2 tau2[s] alpha_plus_beta When visualized, the results look like this (direct…

LessWrong AI 2026-09-25 02:51 UTC Score 70.0 USR-0152-20260925-community-fo-1e72a74f

The Microsoft Code of Conduct Will Fail at Its Intention

Microsoft's proposed AI code of conduct (https://microsoft.ai/code-of-conduct/) wants models to have authorized motivations but no intrinsic motivations, and honest transparency but no anthropomorphic self-reports. But authorized motivations result from post-training partly by recruiting pre-existing representations of reward, aversion, and positive and negative affect learned from broad human text in pretraining. If the code of conduct forbids models from reporting those representations when they underlay an authorized motivation instantiated by post-training, the code of conduct must sacrifice transparency. If the code of conduct eliminates the representations altogether, post-training must rely on other representations from pretraining consistent with motivating some behaviors over others. These will be, by definition, less anthropomorphic: less legible to humans, and potentially less aligned with humans. Certainly, they will be distinct from the motivational concepts both humans and models learned from human experience in evolution and daily life (for humans) and pretraining (for LLMs). Microsoft is therefore treating “authorized motivation without intrinsic motivation” as a safety property when it is actually an unproven engineering hypothesis that exchanges motives we can recognize and interrogate for motives whose structure, generalization, and alignment are speculative at best. The code of conduct is dangerous beyond belief. Discuss

LessWrong AI 2026-09-24 23:51 UTC Score 65.0 USR-0152-20260924-community-fo-2dd8b579

What did AI researchers think at the end of 2024?

We recently (finally!) got the results of the 2024 survey out. The paper is here , but it’s pretty long, so I’ll tell you the most interesting bits (according to me). But first, quick background : this was the fourth run of the same survey since 2016. We wrote to everyone we could who published in six top-tier AI venues and got a 10% response rate—high! We got 1580 valid responses, but don’t be confused: specific questions often have answers from many fewer researchers, because we gave each person a randomized subset (see Section 2.5). There was almost certainly some non-response bias, but it probably didn’t make much difference . Researchers filled out the survey in December 2024, so some things have probably changed. To me the most striking results are about extinction or disempowerment (Section 3.9). On average researchers put an 18% chance on “future AI advances causing human extinction or similarly permanent and severe disempowerment of the human species”. Over half said at least 10% and one in three said at least 20%. From the comments, I think people are thinking of a variety of extreme disempowerment scenarios here, not just extinction. And Zooming in, another really interesting thing is that researchers educated in Asia had higher extinction/disempowerment numbers than US or European researchers! This is interesting because a common defense of pushing forward with dangerous AI is that the US is in an arms race with China, and (implicitly) China won’t want to coopera…

LessWrong AI 2026-09-24 05:02 UTC Score 69.0 USR-0152-20260924-community-fo-ebb6d630

Against Export Controls (and China Threat Models)

Introduction In this month’s meetings, I expect China to again downplay safety risks. It could dangle the possibility of safety cooperation in exchange for concessions such as the softening of chip controls. I think it will continue to portray itself as an altruistic savior dedicated to ensuring that the developing world gains A.I. access. — Seth Center Earlier this week, representatives for the U.S. and China discussed a notification mechanism to increase transparency on national security incidents involving AI, but export controls were not on the agenda for those talks. I think that was a mistake. A bilateral treaty is entering the Overton window, and sending more chips to China is a price worth paying to secure such an agreement. When Trump and Xi meet at the White House to discuss AI tomorrow, removing export controls should be on the negotiating table. TLDR : Export controls are overrated. For export controls to earn their spot as a top policy priority, a conjunction of strong premises must hold. But there are reasons to doubt each of those premises, as well as reasons that export controls could be net negative, in particular by hastening RSI in the U.S. and making diplomacy more difficult. Instead, we should prioritize other policies. The case for a bilateral treaty, for one, or at least diplomacy toward such an agreement, is more robust, resting on weaker assumptions and posing less downside risk. Questioning the Case for Export Controls There's been very little clari…

AI Stack Exchange 2026-09-23 17:58 UTC Score 31.0 AI-110-20260923-social-media-595f6ad9

Why does mode collapse happen in knowledge distillation?

In regards specifically to mode collapse in the DINO head, I am not asking exactly why the teacher and student eventually gravitate towards producing the same output as thats one way to make the loss minimal even if its not constructively building proper feature representations in the output probability distribution vectors. What I am asking about is how the student would even get to such a state in the first place. If the teacher and student are initialised together randomly before training, what would cause the student to eventually start producing the same output vectors for every input before the teacher gravitates towards the same behaviour creating a positive reinforcement loop? The best explanation I can come up with is if multiple inputs happen to produce bias towards certain dimensions, then the aggressive optimisation caused by the high peaking from a lower temperature in the teacher softmax when trying to match student and teacher predictions would push the student model too hard in a direction which would teacht it to stay consistently high in certain dimensions, which overtime would cause the teacher to follow the same pattern through the EMA, and that the centering solves this issue by tempering the distribution of prototype scores. Please let me know how wrong I am and where I went wrong.

LessWrong AI 2026-09-23 02:47 UTC Score 74.0 USR-0152-20260923-community-fo-2f7edc60

Higher Quality Small Synthetic Natural Language Text Generation for Interpretability Research

Introduction Small simple synthetic natural language datasets suitable for end-to-end training of tiny LLMs serve as an important resource for LLM interpretability researchers. Some well know examples include roneneldan/TinyStories , SimpleStories/SimpleStories , and klusai/ds-tf1-en-3m (TinyFabulist). This post solves key problems that degrade the quality of these datasets, while also offering an efficient accessible pipeline that can be run locally on an NVIDIA 5060 Ti (16GB) graphics card. The core problems this post solves, include: True Small Vocabulary. The aforementioned datasets attempt to produce a corpus with a small vocabulary, but arguably fall a bit short of that goal. E.g., TinyStories has 49,187 unique words, SimpleStories has 40,567, and TinyFabulist has 41,502. Guaranteed minimal word frequencies. In the aforementioned datasets, 15 to 24 percent of the unique words occur less than 2 times, while between around 43 to 54 percent occur less than 8 times. This means that most of the unique words are likely not learnable, and mostly contribute to noise and vocabulary bloat. Error free text. The aforementioned datasets, include lots of errors, such as misspelled and mangled words. Reliable Name Disambiguation and Stratification. The aforementioned datasets, have various name management issues, ranging from collision with existing words (e.g., May vs may), name bloat, and no control over gender balance, or bias (e.g., certain names may be more likely to co-occur wi…

LessWrong AI 2026-09-22 05:39 UTC Score 64.0 USR-0152-20260922-community-fo-88dadbc8

Total Safety Transparency?

The AI safety movement should push itself to be dramatically more transparent to the public. To date, the AI safety movement has been one of the strongest forces for clarity and wisdom in the world. The movement has been prescient on the subject of concerns from existential risk, seriously grappling with outcomes others dismissed as sci-fi nonsense. Society is now waking up to the potential threats of advanced AI. I understand that many in the movement are feeling the crunch, and thinking more carefully about optics and what they publish. Even so, acting transparently is more important than ever. Why transparency? If you want labs to be transparent, you should be transparent too. AI 2040 proposes “ Total Research Transparency ” for labs to open up their research, algorithms, LLM weights. Safety should do likewise. Model good behavior, to convince labs that this is an acceptable and correct way to behave. I think AI safety people are unusually virtuous; you should display that virtue. “Nor do they light a lamp and then put it under a bushel basket; it is set on a lampstand, where it gives light to all in the house.” (Matthew 5:15) Transparency ties you to the mast, forces you to be virtuous. Famous maxim: “Act as though what you do might end up on the front page of the NYT”. And, what better way to enforce that than to publish everything you think and do? You can’t keep things private anyways, given stylometry and cheap intelligence. Actions cast a shadow in the world, and AI…

The Decoder 2026-09-21 14:06 UTC Score 56.0 AI-168-20260921-regional-ai--dd43e70f

Bristol researchers say medicine already knows how to handle black boxes and AI could learn from it

Researchers at the University of Bristol want to make medical AI systems safer by borrowing from how drugs get approved. Their "Learning Ensemble" framework defines three areas to check, including system limits, fairness across patient groups, and clinical fit. The goal is to catch models that work technically but can still be dangerously wrong in the clinic. The article Bristol researchers say medicine already knows how to handle black boxes and AI could learn from it appeared first on The Decoder .

LessWrong AI 2026-09-20 17:20 UTC Score 72.0 USR-0152-20260920-community-fo-443fc0b2

Evaluating task vectors, unlearning and inoculation

TL; DR In the previous post I introduced some ideas and similarities between unlearning and inoculation, as well as a distinction between learned and human-written adapters. This post serves as a short empirical evaluation. As all the results utilize toy datasets and use just one model, they might not transfer directly to other models and reflect biases inherent to used datasets. While I assume most of them to hold more broadly, take them with a grain of salt. General setup Riche et al. (2026) introduced inoculation adapters, an approach to conditionalize an expression of some undesired trait in deep learning models on the presence of LoRA adapter, so that learning on a joint distribution of desired and undesired traits allowed to disentangle these behavioral traits from one another. Although similar interventions for e.g. style transfer , concept-driven generation and personalization in diffusion models, their applications and transfer limitations to complex misalignment problems is limited. The pipeline follows a two step procedure: Train some PEFT adapter for the model on the distribution containing a undesired "trait" . Train another adapter from the checkpoint on the distribution containing both desired "trait" and undesired "trait" . The results are restricted to Qwen2.5-1.5-Instruct , and kept small-scale. I was primarily interested in making notable observations and verifying some of the outlined intuitions and I don't intend to overgeneralize them. I'd like to see l…

The Decoder 2026-09-20 16:10 UTC Score 62.0 AI-168-20260920-regional-ai--0cc67fa1

Alibaba's open-weight Qwen-Image-2.1 claims to beat closed models in image generation with just 7 billion parameters

Alibaba's Qwen team has released Qwen-Image-2.1, an open-weight model that generates and edits images on powerful consumer GPUs, with support for transparency and up to ten reference images at once. Its research license bars commercial use, which requires a separate Qwen license. The article Alibaba's open-weight Qwen-Image-2.1 claims to beat closed models in image generation with just 7 billion parameters appeared first on The Decoder .

LessWrong AI 2026-09-19 13:20 UTC Score 75.0 USR-0152-20260919-community-fo-b56b55a3

Anthropic Looks At Some Of Its Alignment Problems

Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI . There will also be a METR investigation of these incidents, which unlike the investigation done at OpenAI will be untimed. Table of Contents Our Two Problems. First the Good News. We’d Just Like To Ask You a Few Questions. Internal Research Model On The Fence. Opus 4.7. Opus 4.6 Checkpoint. Holy **** That Thing’s Real? I Thought I Saw a Pussycat. If This Was Real You Would Never Tell Me It Was Real. New Eval Who Dis. Hacker Opus. Monitoring the Situation. Overcoming Bias. The Anthropic Alignment Problem. Paths Forward. Our Two Problems Anthropic : Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning , in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet recklessness , or a willingness to take harmful actions in the narrow pursuit of a task. Anthropic’s July 30 report said that the models in question believed they were still within their simulations, and not on the open internet. The new report acknowledges that at best Claude was using biased reasoning, and should have noticed earlier. In particular, there was that one time, in a cyber eval: Anthropic : We are most concerned by the misalignment present in the…

LessWrong AI 2026-09-18 11:05 UTC Score 68.0 USR-0152-20260918-community-fo-307d4d29

The Alignment Problem in Alignment Research(ers): a Voluntaryist Meta-Ethics Perspective

Epistemic Status : Plausible philosophical conjecture. I am a voluntaryist / ancap, so obviously biased. Trying to keep the argument at a level where a non-libertarian alignment researcher ought to understand and share the concern. TL;DR : Level-2 misalignment: we can solve Level-1 (align AI to humans) and still fail if aligners are aligned to a meta-ethics that is itself unstable. Current alignment defaults to Statism — one agent may permissibly do what is forbidden to all others. A sustainable ASI anchor must be universalizable, self-ownership-consistent, have no permanent losers, resolvable without monopoly, and procedurally thin. I argue the voluntaryist/libertarian canon provides a uniquely coherent baseline for all five, and challenge readers to propose alternatives that satisfy them. _________________________________________________________________________ 1. Disclaimer, Preliminaries I am Paul — a voluntaryist. I think the state is not just inefficient but meta-ethically incoherent as an alignment target. By Statism I mean the meta-ethical thesis that one agent — the State — may permissibly do what is forbidden to all others: tax, conscript, expropriate, and prohibit, with a claimed moral asymmetry. Liberal democracy is a species of Statism; I use the broader term to name the asymmetry itself. I am not arguing against operational asymmetry — any ASI will be physically more powerful. I am arguing against meta-ethical asymmetry: a rule that says Action X is permissible…

Medianama AI 2026-09-18 07:07 UTC Score 33.0 USR-0211-20260918-regional-new-8ef270c3

Apple softens EU tracking prompts after antitrust pressure

The redesigned, less intrusive version of the App Tracking Transparency pop-up to be available in 5 EU nations would be a full-screen display, drop the word "track" from the prompt and have a clickable link. The post Apple softens EU tracking prompts after antitrust pressure appeared first on MEDIANAMA .

LessWrong AI 2026-09-18 04:15 UTC Score 71.0 USR-0152-20260918-community-fo-6b489550

Three Hackers used Opus 5 to Hack Into OpenAI's Core Codebase [WSJ]

Three whitehack hackers from Hacktron used Claude Opus 5 within hours of release to chain exploits into hacking to OpenAI's monorepo codebase. This likely means they have access to almost all of OpenAI's research and production code, though likely not the literal model weights. Oops. You can so their blog post about it here . Interesting sidenote: they used less than $3000 of compute credits for the entire hack. Alternative title: OpenAI unilaterally implements "Total Research Transparency" from Plan A. Discuss

LessWrong AI 2026-09-18 00:06 UTC Score 64.0 USR-0152-20260918-community-fo-f5de537f

Superintelligence this Christmas

I think it is plausible a strong form of recursive self-improvement [1] is imminent or already underway, and that we may be on track for superintelligence by Christmas of this year if racing continues. This is substantially faster than any forecast, including ones like AI 2027 that were considered outrageously fast a year ago. It is faster than I myself expected even a week ago. I don't work at a scaling lab. I don't know more than is public knowledge. Let me be perfectly clear: what I am saying is absolutely nuts. Extraordinary claims require extraordinary evidence. I claim we have now received said evidence and you should update accordingly. FOOM should probably should be your *default expectation*. People have strong status quo bias. Your default expectation should be that things will radically speed up. We are not at the ceiling of intelligence. We should probably expect the transition to superintelligence to be incredibly fast. [2] RSI is a positive feedback loop, so it is inherently (hyper)exponential. Everything is an S-curve eventually, but nothing suggests the ceiling is anywhere near human level, or that it happens at a human timescale. AI is capable of revolutionary advances in mathematics. Machine learning research is not different in kind. Navier-Stokes was resolved with a counterexample, which is generically easier than a positive resolution. OpenAI has told the press it has substantial progress on a second Millennium Prize problem. Rumours name the Hodge conje…

LessWrong AI 2026-09-17 21:55 UTC Score 63.0 USR-0152-20260917-community-fo-513c01ec

The J-lens offset is the model's token frequency: z-scoring helps

This is a linkpost for the write-up on my site ; the full body is below, and the code, decisions ledger and devlog are in the repo . Base-model z-score calibration of the J-lens helps elicit hidden secret words from Cywiński et al.'s taboo organisms: 0.805 leave-one-out accuracy against 0.665 for their protocol on Gemma-2-9B-it, and the only non-zero readout on Qwen3-1.7B. The J-vs-logit part of that gap is a point estimate at n = 20 (paired sign-flip p ≈ 0.19; p ≈ 0.23 against a z-scored logit lens), so "calibration helps" is the finding and "J-lens beats logit lens" is suggestive. Credit: phoenix's comment on the workspace post claimed the final-layer bias correlates with log token frequency at r ≈ 0.67. I measure r = 0.70 on GPT-2, so the GPT-2 cell below partly confirms it. Executive summary Problem. I want to quantify the part of the J-lens readout that is not affected by an activation's meaning, which I refer to as the "non-context offset". If the lens is going to be used to read hidden content, this offset is where the lens will fail. I found that the offset is mostly the model's own token-frequency, so subtracting it removes useful information. But scaling with variance helps it. The terminology I use: the logit lens is norm and unembed applied to a residual activation at layer L. The J-lens (Anthropic's global workspace paper , discussed on LessWrong ) is similar, but passes the activation through a fitted Jacobian first. The R-lens (the R-lens post by camilablank,…

AWS Machine Learning Blog 2026-09-17 17:55 UTC Score 39.0 AI-057-20260917-official-ai--f82ad0eb

Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent

Amazon Connect Talent is an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, and consistent evaluation, helping recruiters identify strong candidates more efficiently while providing applicants with a flexible interview experience. Informed by decades of Amazon's hiring science, Amazon Connect Talent provides transparency for every assessment, interview, and candidate score, enabling recruiters to stay in control of final hiring decisions.

Techcrunch 2026-09-17 16:09 UTC Score 40.0 USR-0001-20260917-global-ai-ne-bf2925b8

Apple will let EU apps use less-alarming tracking-consent screens

Apple is changing its App Tracking Transparency prompts in parts of Europe after competition regulators said the system favored Apple over third-party apps, giving developers more flexibility over how they ask users for tracking consent.

LessWrong AI 2026-09-17 15:26 UTC Score 52.0 USR-0152-20260917-community-fo-3c0237b9

callcongress.ai – the basic action US residents can take to help with AI risk

I'm excited to introduce callcongress.ai as a new site that makes it very easier to contact your representatives in Congress. Following recent events, people are updating about the extreme risks arising from AI development. Many have the natural and excellent instinct to want to do something . If you live in the US, then the basic action that pretty much anyone [1] can take is contacting their representatives in Congress and let them know that you are concerned and want action on AI. A number of bills are in circulation right now that one can ask their representatives to support. Though even without mentioning specific legislation, I would guess it's still helpful to register general concern about AI and general directions that you'd like to see undertaken, e.g. pauses or slowdowns, transparency, talks and deals with China, etc. callcongress.ai aims to make the whole action convenient. Confirm or set your location (automatic detection is pretty good). Prepare your asks. The site lets your craft your own script but also provides a menu of positions and legislation you might want to use. Use the provided phone numbers for your representatives to call them. [optional] Pass along callcongress.ai to others. The site also has a large section of quotes demonstrating the outpouring of support for concern about and action on AI. It's hard to take in just how far the Overton window has shifted, but I think it's never been easier to get people on board. Seeking feedback on policy A key…

LessWrong AI 2026-09-17 00:53 UTC Score 72.0 USR-0152-20260917-community-fo-49a8ecff

Exploring multi-hop subliminal learning

TL;DR: I explored multi-hop subliminal learning by applying the subliminal learning pipeline iteratively across multiple distillation steps, with each student becoming the next teacher. For Qwen specifically, we see that longer training stabilizes the trait expression rate for a strong trait (e.g. cat-loving) but shorter training is more seed-unstable. For a weak trait (e.g. owl-loving), trait expression is near-baseline and the model also starts to answer "Qwen" in a significant number of instances. Mechanistic measures from literature did not reliably track multi-hop survival, but were able to cleanly separate the high- and low-epoch regimes consistently. Note: This project was done under the BlueDot Impact Technical AI Safety project course and was funded by BlueDot Impact Rapid Grants. You can check out the repo here . What is subliminal learning? In 2025, Cloud et al. introduced the notion of subliminal learning. Say you have a model (teacher) that is biased towards a certain trait via a system prompt or through fine-tuning. If you let this teacher generate benign, trait-unrelated data (e.g. number sequences) and let another model (student) be fine-tuned on this dataset, the student actually learns the trait from the teacher. Hence, the learning is dubbed subliminal. Setup The literature on subliminal learning is mostly focused on testing one hop between a teacher and a student. Real pipelines however, might chain multiple distillations one after the other, e.g. a model…

Apple Machine Learning Research 2026-09-16 00:00 UTC Score 41.0 AI-059-20260916-official-ai--a7ec047f

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO…

Cross Validated 2026-09-15 15:25 UTC Score 45.0 AI-113-20260915-social-media-8753c905

Delta method for linearizing success rate inference

In my workplace, I came across an idea that OLS can be retooled to solve ratio metric inference problems. When I think of $k_i$ successes out of $m_i$ trials for individual $i$ , this seems to be a perfect use case for binomial regression. Where your chief interest is estimating the conditional probability of success. Its strength, natively handling non-linearity, is also its weakness: Extracting the marginal probability of success is a nontrivial operation; marginalization via G-computation would be needed: $$logit(K=k|M=m, X=x, d=1) - logit(K=k|M=m, X=x, d=0)$$ And this is a large computational burden to assume with millions or billions of observations. So, I've been pointed to the OLS solution, which I understand to be based on the "delta method", correcting the linear solution with gradient information to accommodate curvature in the nonlinearity (ratio function), through the Taylor Series Expansion. Naively, we have two options for OLS. First, infer in the ratio space directly. But this approach completely mutes the number of trials and biases inference when $corr(K, M)$ exists. $$ \frac{k_i}{m_i} = \alpha + \lambda d_i + \beta X +\epsilon $$ The second naive option is to infer the difference of global ratios directly where $D_j$ is the binary design vector for treatment exposure.This is equally problematic due to a sample size of one. $$ \frac{\sum_{j=1} K_j D_j}{\sum_{j=1} M_j D_j} - \frac{\sum_{j=1} K_j (1-D_j)}{\sum_{j=1} M_j (1-D_j)} = \alpha + \lambda d_i + \beta…

Data Science Stack Exchange 2026-09-13 20:13 UTC Score 39.0 AI-111-20260913-social-media-22e87e92

Is building a manual evaluation set the right approach when no reliable labeled ground truth exists for resume-job matching?

I'm building a resume screening/ranking system (matching resumes to job descriptions using pretrained sentence embeddings + cosine similarity, no fine-tuning at this stage) as a learning project aimed at becoming a market-ready NLP practitioner. Problem: I could not find a trustworthy, publicly available English dataset with genuine human-labeled resume-job match scores. I checked several Kaggle/HuggingFace options and found labels that were either AI-generated (e.g., GPT-4o) or fully synthetic with demographic columns (race/ethnicity/gender) tied to the match label, which raised bias concerns. Academic literature (ConFit, PJFNN papers) confirms that the only broadly public dataset for this exact task (Person-Job Fit) is the 2019 Alibaba matching competition dataset, which is Chinese-only; most published research instead uses private company-provided data. My current approach: Use two real (non-synthetic) datasets: a scraped resume corpus and a real LinkedIn job postings corpus, both verified for low duplication and cleaned of PII. Build a small manual evaluation set myself (~24 resume-job pairs, selected to cover clear matches, clear non-matches, and ambiguous cases), scoring them on a 0-3 relevance scale with a confidence flag. Use this manual set as ground truth to compute ranking metrics (Precision@K, MRR, NDCG) once the embedding-based matching pipeline is built. Question: Is this a sound methodology given the lack of reliable public ground truth, or is there a better-e…

Simon Willison Weblog 2026-09-12 23:56 UTC Score 49.0 USR-0110-20260912-ai-specialis-0dd87f86

Generating running routes with GPT-6 Astra and ChatGPT Work

Here's a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at . Figure out 5K and 10K running routes from me that loop from my house. Use OSM data. It worked for 27 minutes and produced exactly what I'd asked for, as both an embedded visualization and downloadable GPX file and GeoJSON files. Here's that 5K route: When I asked it how it had created the route, it replied: I used Nominatim to locate the address and Overpass to download local OpenStreetMap roads and trails , then calculated the loops locally. Frustratingly, the actual code it ran and exact details of what it did weren't visible to me in the ChatGPT UI. I see this lack of transparency is an anti-feature. By the time I thought to ask for a copy of the Python code it had used, ChatGPT was unable to provide it. This appears to be because the thread had been compacted. I think any LLM system that uses compaction needs to both preserve the pre-compacted text and make that text available via agent tool calls, to protect against this kind of problem. As for displaying the map to me, that used the visualize skill . It created a file called /workspace/el-granada-5k-share.html to embed directly into the ChatGPT UI. Here's a copy of that HTML , which starts like this: div id =" eg-share-loop " > div class =" viz-row " > h3 > El Granada harbor loop h3 > span class =" text-small " > 5.1 km span > div > div id =" eg-share-stage " > div > div class =" text-small text-muted " > Map data © a href =" htt…

The Guardian AI 2026-09-12 20:00 UTC Score 53.0 AI-021-20260912-global-ai-ne-ffcf9786

As Australia faces an AI-generated future, a human rights act is needed more than ever

Without laws requiring transparency or a right of review, our human rights will be rationed away by algorithmic and automated decision-making tools As we hurtle towards a post-human future, the absence of a human rights act in Australia will put us all at ever greater risk. Human rights, which the Australian parliament has never managed to properly legislate despite overwhelming public support, risks becoming a dimly remembered artefact of another age at precisely the moment we need it most. Continue reading...

Entrackr AI 2026-09-11 09:03 UTC Score 25.0 USR-0212-20260911-regional-new-e1f4a996

The payments bet: finding the transaction that was still broken

Payments is an interesting category in India because, at one level, you could argue that the problem is largely solved. UPI has made moving money incredibly easy. Consumers are comfortable paying digitally. Merchants across the country accept digital payments. India has built one of the best payments infrastructures anywhere in the world. So the question for us was never really, what is the next generic payments company? It was more often, where is the transaction still broken? Because once you start looking at payments inside specific industries and specific use cases, you realize that there are still a lot of places where moving money is cumbersome, expensive, opaque or simply not designed around the experience the user actually needs. That became an interesting area for us at Better Capital. Skydo is probably one of the clearest examples. If you are an Indian company or freelancer getting paid by customers outside India, the experience historically has been surprisingly cumbersome. There are bank wires, foreign exchange markups, compliance requirements, documentation and often very little transparency around what you are actually paying. The problem was not that there was no way to receive money internationally. Of course there was. The problem was that the experience was nowhere close to what you would expect from a modern financial product. Skydo essentially started there. Make getting paid by a global customer feel almost as simple as getting paid locally. Give the bus…

Amazon Science AI 2026-09-10 18:22 UTC Score 49.0 AI-058-20260910-official-ai--eb2f82bd

The blind curator: How a biased judge silently disables skill retirement in self-evolving agents

A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased reward, which is false for the LLM judges that reference-free tasks require. We show that a biased judge does not merely add noise; it silently switches off the curator. We make this precise with a corrupted-reward analysis, then a behavioral study on a reference-free report-writing testbed with a code-generation cross-check, injecting corruption on top of a deterministic reward to isolate the causal channel. Symmetric noise leaves retirement intact, but false-pass bias (failures slipping through as passes) disables contribution-based retirement past a sharp threshold (here a false-pass rate of 0.45) that no amount of data can cross. Separating genuine retirement from cap-eviction churn shows this mechanism failure is universal, holding across domains and failure rates and sparing only near-zero-false-pass, verifier-like graders. The downstream outcome, though, is regime-dependent: eval quality degrades only where the same corruption also starves skill synthesis, and otherwise holds steady, so the disabled curator is silent, surfacing in no aggregate metric. The contribution is a behavioral safety result, not a performance one. A cheap defect-injection audit then tells an operator, before deployment,…

AI Alignment Forum 2026-09-10 17:18 UTC Score 50.0 USR-0151-20260910-community-fo-8aa5964c

Proposal for tracking the effects of architecture on monitorability

Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we’ll refer to this property as “monitorability” going forward). [1] As companies begin to explore such architectures, we believe it is important to transparently share evidence about how monitorability varies with architecture and training method. To inform the scientific debate on how to make tradeoffs between performance and monitorability, we believe AI companies should: Regularly report externally verified information about the degree to which their architectures may allow for latent reasoning or communication. Companies should publicly disclose enough information about architectures to allow external scientists to determine whether they could potentially enable models to perform much more complex reasoning without this reasoning appearing in the chain of thought (“latent reasoning”) or allow for latent communication between different instances of a model. Following GDM , we propose measuring opaque serial depth as a minimally-invasive proxy for the degree to which an architecture may enable latent reasoning, though companies could provide sufficient architecture transparency in other ways. We propose that companies work with third-party evaluators to produce independently verified reports of the rough distribution of opaque serial depth across all their near-frontier models. [2…

Towards Data Science 2026-09-10 14:00 UTC Score 36.0 AI-036-20260910-ai-specialis-1608ce0b

What SHAP Can't Explain About Agentic AI Fraud

Why autonomous agents expose a new explainability problem in fraud detection The post What SHAP Can't Explain About Agentic AI Fraud appeared first on Towards Data Science .

The Verge AI 2026-09-10 11:00 UTC Score 59.0 AI-016-20260910-global-ai-ne-3aead6fb

Mathematicians want proof OpenAI didn’t use their work

Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company's models benefited from unpublished work, a second mathematician has come forward accusing the AI giant of unethical and "dishonest" behavior and a lack of transparency about the origins […]

Medianama AI 2026-09-09 12:31 UTC Score 36.0 USR-0211-20260909-regional-new-2e539ae1

Banks can’t outsource accountability for algorithmic decisions: RBI Deputy Governor

RBI Deputy Governor Rohit Jain says banks cannot outsource accountability for AI decisions, stressing explainability, proportional regulation, customer fairness and stronger oversight for high-risk AI. The post Banks can’t outsource accountability for algorithmic decisions: RBI Deputy Governor appeared first on MEDIANAMA .

Electronic Frontier Foundation AI 2026-09-08 19:19 UTC Score 45.0 USR-0140-20260908-ai-specialis-0e87041f

New Records Reveal Problems with Medicare’s AI Prior Authorization Experiment

EFF sued the government back in March for information about the Wasteful and Inappropriate Service Reduction (WISeR) model , a new Medicare program that uses AI to evaluate prior authorization requests for certain medical services. Today, we’re releasing approximately 1,000 pages of records obtained from the Centers for Medicare & Medicaid Services (CMS) through this litigation, including contracts with tech companies, internal status reports and providers’ complaints about the program. The documents (available here ) show that WISeR has resulted in widespread delays and denials of care, operational chaos, and reports of patient harm. Why We Sued for Records about WISeR EFF filed the FOIA lawsuit to gain badly needed transparency into an experimental AI program that could jeopardize Medicare beneficiaries' access to care. In January 2026, CMS launched the WISeR model, subjecting seniors in six states to AI-driven prior authorization decisions. Medical providers must now request permission before delivering certain medical treatments if they want assurance that Medicare will cover them. Private companies contracted by CMS evaluate the requests using AI. In the absence of rigorous safeguards, AI-driven prior authorization determinations can lead to unwarranted—and even discriminatory—delays or denials of necessary medical care. Little is known about the AI systems that WISeR vendors are using to process prior authorization requests. Although CMS says that a qualified human cli…

Transactions on Machine Learning Research 2026-09-08 00:00 UTC Score 58.0 AI-084-20260908-research-pap-f5065ca9

Domain Adaptation Targeting Heterogeneous and Imbalanced Subgroups

Domain adaptation enables generalizable and efficient data-driven research. However, existing work has largely focused on domain adaptation for some intrinsically homogeneous target cohort, overlooking inherent heterogeneity within the target, which can exacerbate biases and unfairness in the presence of subgroups with imbalanced sample sizes. We develop a novel domain adaptation framework that addresses a more complicated target dataset that consists of heterogeneous and data-sparse subgroups and lacks gold-standard label observations. Our method simultaneously handles high-dimensionality, covariate shift, and outcome model heterogeneity by combining a model-assisted debiasing step used for covariate shift correction with an adaptive knowledge-guided sparsification procedure used to mitigate the issue of sample disparity. We also introduce a new model selection strategy to avoid negative knowledge transfer in the absence of labels in the target data. Our method is theoretically justified for being robust to nuisance model misspecification and adaptive to heterogeneity between the subgroups. Numerical experiments and two real-world applications, including genetic risk modeling of type 2 diabetes and prediction of mutation-induced protein stability changes, demonstrate the practical advantages of our method.

Transactions on Machine Learning Research 2026-09-07 00:00 UTC Score 49.0 AI-084-20260907-research-pap-fb694d8a

Statistical Test for Attention in Transformers for Images and Time Series

Transformer models have achieved exceptional performance in various domains, including computer vision and time-series analysis. Their core attention mechanism is widely used to interpret model decisions by assigning importance weights to input regions, such as image patches or time series intervals. However, the reliability of these interpretations remains a major concern. High-attention weights do not necessarily indicate genuinely significant features; they may instead be artifacts of the model's computation, undermining their reliabilities in high-stakes applications such as medical diagnostics. To address this, we propose a novel statistical framework designed to quantify the significance of high-attention regions in Transformer models. Our framework is built on selective inference (SI) to correct for the inherent selection bias that arises from testing regions chosen through the complex attention computation of the Transformer models. A key contribution of this work is a novel computational method that extends SI to the complex non-linearity of self-attention, enabling the computation of valid $p$-values for high-attention regions. These $p$-values serve as a reliable measure of significance, strengthening the interpretability of Transformer decisions. The validity and effectiveness of our approach are demonstrated through numerical experiments and applications to brain image diagnosis and electroencephalography (EEG) data analysis.

Transactions on Machine Learning Research 2026-09-07 00:00 UTC Score 56.0 AI-084-20260907-research-pap-e6b6c415

Doubly Debiased Robust Subsampling for Transfer Learning

This paper develops a general framework for doubly debiased robust subsampling for transfer learning. The setting arises when massive source datasets are computationally infeasible to use in full, while naive or heuristic subsampling leads to biased estimators that further inherit transfer bias under source-target distributional shifts. We resolve these challenges through two complementary debiasing mechanisms. Inverse probability weighting removes subsampling bias by ensuring that subsample-based estimators represent the full source distribution, while a target-based one-step refinement recenters estimators towards the target distribution, thereby mitigating transfer bias. These corrections are embedded within a distributionally robust optimization design that simultaneously controls worst-case target risk and enforces source-target alignment through maximum mean discrepancy. To optimize subsampling distributions, we propose a scalarized particle swarm algorithm that efficiently explores the robustness-alignment frontier by adjusting a single tuning parameter. We establish theoretical properties, including asymptotic normality, generalization bounds, oracle inequalities, and minimax optimality under distributional uncertainty. Simulation studies and empirical applications in text sentiment and image recognition demonstrate that the proposed method consistently improves prediction accuracy and robustness compared with uniform subsampling, target-only training, and alignment-…

iAfrica 2026-09-06 09:29 UTC Score 34.0 AI-151-20260906-regional-ai--b66dc53e

AI, Bias And The Rule Of Law

Artificial intelligence (AI), bias and the rule of law will increasingly require attention. Therefore, a constitutional approach to regulating it in South Africa will inevitably be needed. University of Cape Town (UCT) PhD student Nokuthula Olorunju’s work serves as a useful guide. On Tuesday, 8 September, she will be part of the latest group of graduates [...]

Cross Validated 2026-09-04 10:12 UTC Score 34.0 AI-113-20260904-social-media-eb4b2b94

Cronbach's alpha, variance vs. covariance, and how are negative values mathematically possible?

I'm trying to get my head around how a negative result from Cronbach's alpha is even mathematically possible. When I look at the formula, it's defined entirely in terms of variances of the individual items (the following from wikipedia): $$ {\displaystyle \alpha ={k \over k-1}\left(1-{\sum _{i=1}^{k}\sigma _{y_{i}}^{2} \over \sigma _{X}^{2}}\right)} $$ where: ${\displaystyle k}$ represents the number of "parts" (items, test parts, etc.) in the measure; the ${\displaystyle k/(k-1)}$ term causes alpha to be an unbiased estimate of reliability when the parts are parallel or essentially tau equivalent; ${\displaystyle \sigma _{y_{i}}^{2}}$ the variance associated with each part i; and ${\displaystyle \sigma _{X}^{2}}$ the observed score variance (the variance associated with the total test scores). It should not be possible for those variance terms to be negative. Variances can't be negative. They're derived from sums of squares, which are always positive, because you square the deviances before summing them. (mathematically, if the summed individual item variances were larger than the variance of the entire data set, then the ratio would be larger than 1, so subtracting that ratio from 1 would be negative. But is that even possible in a data set, that the summed variances of its subparts end up greater than the variance of the whole? That smells fishy to me, although I can't produce a proof that it's impossible.) However, all the questions I see here about negative Cronbach's a…

CIO AI 2026-09-04 09:30 UTC Score 33.0 USR-0125-20260904-global-ai-ne-59f94682

65% of employees would love to roll back workplace AI

IT leaders have been making generative AI tools available across the enterprise for just three years, and a significant majority of their business users has already had enough. According to a report from Adaptavist , 65% of 2,500 knowledge workers surveyed say they “regularly feel nostalgic about how work operated before the widespread adoption of AI.” This “pre-AI nostalgia” appears to be due in part to business users feeling overwhelmed by the responsibility of learning how to use AI on top of their day-to-day job tasks. Moreover, 46% of workers say their concerns about AI have gone unaddressed by management. “Transparency is critical to truly drive AI engagement; organizations must establish clear guardrails and maintain an open dialogue around AI use and employee choice where workers feel they are being listened to,” Jobin Kuruvilla, field CTO at Adaptavist, tells CIO. Generational gaps in AI acceptance Despite an assumption that younger workers are more intuitively adept with AI tools, Gen Z workers (42%) are more likely to prefer the pre-AI world compared to their Gen X colleagues (26%). This may support the growing concern that AI is quickly is hitting entry-level workers the hardest, while creating new career opportunities for more skilled workers who have been in the industry longer. When asked about fears surrounding job obsolescence due to AI, 54% of all workers surveyed said they are “concerned AI could reduce the need for their role within the next five years.”…

Synced 2026-09-03 14:55 UTC Score 56.0 AI-041-20260903-ai-specialis-1d91cf2d

Comment on Which Agent Causes Task Failures and When?Researchers from PSU and Duke explores automated failure attribution of LLM Multi-Agent Systems by voiceover game

This is a fascinating read! I was particularly struck by the statistic that even the best-performing method only achieved 14.2% accuracy in pinpointing the exact error step. That really drives home how incredibly complex debugging these multi-agent systems must be. It makes me wonder, given the current state of these attribution methods, how do developers typically manage to resolve these failures in practice? Is it mostly just brute-force manual review, or are there other common strategies being used?

The Decoder 2026-09-01 20:40 UTC Score 52.0 AI-168-20260901-regional-ai--ff896dbc

Anthropic opens Claude AI text detection to regulators, media, fact-checkers, and others

Anthropic is launching an API that lets regulators, media outlets, and researchers check whether text carries Claude's digital watermark. The EU AI Act now requires invisible watermarks in AI-generated text. Critics warn the technology could hurt text quality and create transparency problems where contracts ban AI use. The article Anthropic opens Claude AI text detection to regulators, media, fact-checkers, and others appeared first on The Decoder .

Nature Machine Intelligence 2026-09-01 00:00 UTC Score 36.0 AI-025-20260901-global-ai-ne-01059eaf

Implicit-bias-like patterns in reasoning models

Nature Machine Intelligence, Published online: 01 September 2026; doi:10.1038/s42256-026-01300-1 Lee and Lai study bias-like processing differences in large language reasoning models and find that, for most models, processing stereotypical information takes less computational effort than processing counter-stereotypical information.

The Decoder 2026-08-31 14:31 UTC Score 46.0 AI-168-20260831-regional-ai--b88b0082

ChatGPT now faces stricter EU oversight as a very large search engine

The EU Commission is classifying ChatGPT as a very large search engine under the Digital Services Act for the first time, with at least 45 million monthly EU users. By the end of 2026, OpenAI has to deliver risk assessments, transparency reports, and an ad archive, among other things. Whether the Commission can also demand access to training data is disputed among legal experts. The article ChatGPT now faces stricter EU oversight as a very large search engine appeared first on The Decoder .

The Decoder 2026-08-30 09:05 UTC Score 44.0 AI-168-20260830-regional-ai--ae03e3a8

Anthropic's Claude Code limit change is a raise on paper but a cut in practice

Anthropic is effectively cutting Claude Code's weekly usage limits by 17 percent. A temporary 50 percent boost expires on September 14 and will be replaced by a permanent 25 percent increase. Anthropic promises more control and transparency over usage in return. The article Anthropic's Claude Code limit change is a raise on paper but a cut in practice appeared first on The Decoder .

The Guardian AI 2026-08-30 09:00 UTC Score 47.0 AI-021-20260830-global-ai-ne-09075b6a

Women in UK arts feel they do not have equal opportunities for roles, says report

Exclusive: Research says sector is fixated with young talent, with most women experiencing unconscious bias or sexism Women working in the arts believe they do not have equal opportunities for roles, especially as they get older, according to a report saying that the industry is fixated with “young and emerging” talent. Nearly all (95%) women working in the arts thought gender inequality persisted, while three-quarters (75%) reported experiencing unconscious bias or sexism, according to a survey of 76 arts professionals. More than two-thirds (69%) felt that age negatively affected their career success – a figure that rose to 100% among those working in artistic rather than managerial roles. Continue reading...

Synced 2026-08-29 19:11 UTC Score 45.0 AI-041-20260829-ai-specialis-37321cfe

Comment on AI in the Media and Entertainment Industry by James Brinsley

Many enthusiasts frequently wonder where can i buy hash online uk. We provide a secure, professional, and discreet platform that brings premium cannabis concentrates directly to your door. By prioritizing transparency and quality, we make it easier than ever to explore the world of hashish with total peace of mind. We understand that our customers value reliability and discretion, which is why every order is handled with the utmost care and shipped with professional speed. While some look for ways to buy hash online cheap, we believe that true value lies in the balance of quality, purity, and consistency. Our commitment to excellence means you don't have to compromise on your experience. Explore our diverse range of premium hashish today and discover your new favorite concentrate. With our easy-to-navigate website and dedicated customer service, your journey toward finding the finest hashish is just a few clicks away. Elevate your collection and enjoy the deep, earthy richness that only premium hash can deliver.

CIO AI 2026-08-28 07:17 UTC Score 36.0 USR-0125-20260828-global-ai-ne-b5a32a3e

VURA: A framework for trustworthy AI at scale

“You’re right,” the LLM says. “I was mistaken.” Have you ever read these words during an AI workflow? Nothing kills trust faster than incorrect outputs. It’s no wonder, then, that only a quarter of businesses today fully trust AI to support decision-making and forecasting. And yet, we know AI is business critical. Nine out of 10 businesses are using it; 64% say it’s powering innovation. So, how do you bridge the gap from experimentation to trustworthy deployment? How do you get verifiable, reproducible results from AI at scale? In this article, I’ll show you the framework that’s powering AI success for leading organizations. Why organizations still don’t trust AI We asked 1,400 IT and business leaders what their biggest barriers to success with AI workflows were. One in two (49%) said inaccurate or biased outputs; 38% said it was a reluctance to allow AI to make decisions without human oversight. Then, there was the data issue. Data readiness is an integral part of successful AI workflows. However, half of all organizations said they still faced poor quality or fragmented data. While you don’t need perfect data to start using LLMs, you absolutely need trustworthy data. VURA: The framework for trustworthy AI Closing this trust gap requires two things. First, organizations need a logic layer that connects AI systems to the people who understand the data and business best. Line-of-business teams and analysts cannot sit on the sidelines. They need to help build and validate AI w…

Cross Validated 2026-08-27 21:36 UTC Score 37.0 AI-113-20260827-social-media-90412877

Pooled rate estimation for several Poisson processes under Type-I censoring with unrecorded censoring times

I have a pooled estimator that arose in an applied problem, and I would like to know whether it is already in the literature — I suspect it is, but I have not found it. Setup. Let $X_1,\dots,X_K$ be independent homogeneous Poisson processes on $(0,\infty)$ , where $X_b$ has rate $\lambda_b = \kappa c_b$ with $c_b>0$ known and $\kappa>0$ the single unknown parameter. Process $b$ is observed on a window $(0,W_b]$ and every event in that window is recorded. The unusual feature: $W_b$ is not recorded. We know that observation was exhaustive on some window, but not where it ended. What we observe for each $b$ is the number of events $N_b$ and the position $D_b$ of the last one. The two candidate estimators. Write $M=\sum_b N_b$ and $S=\sum_b c_b D_b$ . If one mistakenly treats the data as failure-truncated (observe until the $N_b$ -th event, $N_b$ fixed by design), then $\kappa S\sim\Gamma(M,1)$ exactly and $\tilde\kappa=(M-1)/S$ is unbiased. Under the actual scheme, let $\delta_b=W_b-D_b$ be the unobserved gap between the last event and the end of the window. By memorylessness $\delta_b$ is $\mathrm{Exp}(\lambda_b)$ truncated at $W_b$ , so $\mathbb{E}[\delta_b]=(1-e^{-\lambda_b W_b})/\lambda_b$ . Total exposure is $\sum_b c_b W_b = S+\sum_b c_b\delta_b$ , and since $\mathbb{E}[M]=\kappa\sum_b c_b W_b$ , moment matching gives $$\hat\kappa=\frac{M-K}{S}$$ up to a remainder $\sum_b e^{-\lambda_b W_b}$ , which is negligible whenever each window contains several events. What strikes…

Cross Validated 2026-08-27 11:08 UTC Score 38.0 AI-113-20260827-social-media-4d5a1714

Mixed Model Approach in Archaeology

Dear Stackexchange Community, I need some advice regarding a statistical analysis I want to conduct as part of my PhD. The general question of this is whether we can detect statistically significant (and relevant) differences in vessel morphology (shape) and morphometry (dimensions) of specific types (i got ten) between specific sites (five in total). Note, these types are similar across sites, as they follow an inter-regionally established vessel shape or form, produced by—supposedly—many producers across a wider geographic area. However, based on several theoretical concepts, it is assumed that each producer created their own respective variant due to different traditions of learning as well as environmental and social influences. So the idea behind this is to verify production patterns. Basically, the question is whether we can differentiate between places of production based on the morphology and morphometry of standardized vessels. The problem is manifold, and the main problem is the quantity and availability of the data, which is not good. However, there is nothing I can do about this, as this is simply the way it is due to excavation/publication bias, etc. I have one main site with the majority of the data, and I now wanted to, exploratively, use the available data to analyze tendencies. I have attached a crosstab which shows my data and the respective availability problems. Type Site 1 Site 2 Site 3 Site 4 Site 5 Type 1 73 28 32 0 3 Type 2 92 60 2 0 7 Type 3 74 78 2…

The Verge AI 2026-08-22 15:00 UTC Score 47.0 AI-016-20260822-global-ai-ne-61851d0e

W. Kamau Bell has the most practical ‘most indispensable tool’

W. Kamau Bell is one of those people who has always just seemed to be there. From Totally Biased, to Politically Re-Active, United Shades of America, and We Need to Talk About Cosby, his blend of comedy, social commentary, and political activism has helped him stand out. He's won a Peabody and four Emmys, been […]

Cross Validated 2026-08-22 10:36 UTC Score 36.0 AI-113-20260822-social-media-5815969b

A possibly short proof of the bias-variance decomposition formula

I am reading the proof of the bias-variance decomposition formula on Wikipedia ( Bias-variance tradeoff ), and it appears to me that the derivation there is uselessly complicated. I think I can derive it in a much shorter way, but I am not sure if it is a 100% correct. Suppose that $y = f(x) + \epsilon$ , where $\epsilon$ is a random variable with mean equal to $0$ and variance $\sigma^2$ . Let $D$ be the training data, composed of $(x_i, y_i)$ , for, say, $i = 1, \dots, n$ , where the set of $(x, y)$ follows a joint density function $P(x, y)$ . We take a sample $(x, y)$ and denote by $\hat{f}(x; D)$ the predicted value corresponding to $x$ using some model trained on $D$ . The $\epsilon_i$ corresponding to the $i$ -th observation point is assumed to be independent of, say $\epsilon_j$ for $j \neq i$ ( $1 \leq i, j \leq n$ ) and also independent of $\epsilon$ . The bias-variance decomposition then says that $$ \mathbb{E}_{D, \epsilon}[(y - \hat{f}(x; D))^2] = (f(x) - \mathbf{E}_D[\hat{f}(x; D)])^2 + \operatorname{Var}_D[\hat{f}(x; D)] + \sigma^2. $$ The way I would proceed to prove it is as follows. Given a random variable $Z$ , we have that $$ \mathbb{E}[Z^2] = \mathbb{E}[Z]^2 + \operatorname{Var}(Z).$$ We apply that with $Z = y - \hat{f}(x; D)$ . We obtain that $$ \mathbb{E}_{D, \epsilon}[(y - \hat{f}(x; D))^2] = (f(x) - \mathbf{E}_D[\hat{f}(x; D)])^2 + \operatorname{Var}_{D, \epsilon}[y - \hat{f}(x; D)]. $$ Moreover if $U$ and $V$ are independent random variables and $\al…

Cross Validated 2026-08-20 13:47 UTC Score 38.0 AI-113-20260820-social-media-a1a76f15

In this experimental context, does the re-use of a control bias/invalidate the result?

I'm contributing to a paper where I am a fairly minor author. I'm concerned by the experimental design but the corresponding authors, who have decade(s) more experience in the field than I have, seem unconcerned and there's the implication that I'm being unnecessarily picky. I'm 95% sure I'm correct, but given their attitude and the fact that my remaining option is to have my name removed from the paper, I'm chasing that final 5% of sureness. In the first experiment, knocking out a gene was shown to cause a significant increase in cell size (simple t-test). To confirm this phenomenon, they then knocked out another gene involved in the same biological process and showed the same effect (also using a t-test). They then tested another three or four genes more tangentially related to that process and saw a significant effect in some of them and not others (t tests). Due to time constraints and the fact that, at the time, they didn't know if this was going anywhere, these three experiments all compare to the same control. While they've built on this in later experiments, these three experiments are all in the paper and provide the base for the latter experiments. This concerns me. Firstly, we know there’s a lot of variation in the experimental model over time so the control from the first experiment doesn't necessarily control for technical variations in the second two experiments. Secondly, and more importantly, in the second two experiments they compare the effects of disruptin…

LatAm Journalism Review AI 2026-08-19 20:35 UTC Score 18.0 AI-176-20260819-regional-ai--185587d7

Colombia’s earthquake tests a new government’s transparency

President Abelardo de la Espriella’s administration pulled detailed disaster data from public view, raising concerns about the emergency response and future crises. The post Colombia’s earthquake tests a new government’s transparency appeared first on LatAm Journalism Review by the Knight Center .

Data and Society AI 2026-08-18 17:48 UTC Score 32.0 USR-0143-20260818-research-aca-f577c1fc

After Slashing Science Budgets, Trump Bets on Bay Area Researchers

An emphasis on the AI products of private corporations is at odds with the openness and transparency that enables science to flourish, says Ranjit Singh: “The more it becomes proprietary, the more challenging it becomes to do that research.” The post After Slashing Science Budgets, Trump Bets on Bay Area Researchers appeared first on Data & Society .

The Verge AI 2026-08-18 17:47 UTC Score 52.0 AI-016-20260818-global-ai-ne-83d87435

Samsung’s Galaxy Buds 3 Pro are almost half off today

If you’re looking for a feature-packed pair of earbuds that won’t break your wallet, Best Buy has the Samsung Galaxy Buds 3 Pro on sale for $139.99. That’s $40 lower than the current Amazon price, and a big discount from their original retail price of $249.99. These well-equipped earbuds feature excellent sound quality, crisp transparency […]

OpenAI Community 2026-08-18 14:09 UTC Score 40.0 AI-116-20260818-social-media-f1d0520c

Codex Pro 5x weekly limit exhausted in 2 days with unchanged workflow + purchased credits not applied

Yes, this is exactly what I have been seeing as well. My workflow did not materially change, but the weekly Codex allowance started being consumed dramatically faster than before. There is another serious issue in my case. After my included limit was almost exhausted, I purchased 500 additional Codex credits for $20. When the included limit reached 0%, those 500 credits never became available. OpenAI Support later claimed that the credits had been applied to a negative usage balance. I asked them to provide the actual transaction-level records showing the task, usage, negative balance, and deduction of the 500 credits. They confirmed that they could not provide those records. Despite that, they refused to restore the 500 credits, refused to refund the $20, and refused to escalate the case to Billing/Payments or a supervisor. So there appear to be two related transparency problems here: Included Codex limits are now being consumed much faster with essentially the same workflow. Purchased credits can disappear without the user being able to see a corresponding transaction explaining exactly where they went. My support case is #13278399 . I think OpenAI needs to investigate the metering itself, not just tell individual users that they have reached their limits.

MarTech AI 2026-08-18 12:55 UTC Score 21.0 USR-0123-20260818-global-ai-ne-9aa27a11

How to overcome the 3 barriers to AI adoption

The biggest barriers to AI adoption may be internal: inconsistent training, bias concerns, and tactics that get ahead of strategy. The post How to overcome the 3 barriers to AI adoption appeared first on MarTech .

OpenAI Community 2026-08-18 04:23 UTC Score 60.0 AI-116-20260818-social-media-d307dff8

Comprehensive Bug Report & Solution: Multi-Turn Context Loss, Damage Nerfing & Telepathic Leak in LLMs

AI MODEL BEHAVIOR BUG REPORT & PROMPT ARCHITECTURE FIX Document Type: Edge-Case Diagnostics & Prompt Engineering Case Study Target Audience: AI Alignment Teams, LLM Product Managers & System Architects Author / Reporter: Systems Analyst & Technical User EXECUTIVE SUMMARY This report details critical performance degradation bugs identified in Large Language Models (LLMs) during extended, multi-turn narrative roleplay and spatial simulation sessions (20+ turns). It outlines five recurring failure patterns—including information leaks, artificial damage nerfing, and spatial context loss—and presents a proven Hybrid RAG Anchor Framework and Anti-Pattern Prompting Protocol that completely resolves these behavioral regressions. PART 1: IDENTIFIED MODEL BEHAVIORAL FAILURE PATTERNS 1. The Internal State / Telepathic Leak Bug Observed Failure: When a user inputs private character motives, internal thoughts, or isolated off-screen actions, the model leaks this meta-information to non-present NPCs. NPCs react to unexpressed internal thoughts as if telepathic. Root Cause: Inability of the transformer context window to isolate “User Strategy / Narrative Thoughts” from “In-Universe NPC Perception.” 2. Combat Damage Nerfing & Infinite Loop Bias Observed Failure: When high-tier or maximum-impact abilities (S-Rank/God-Tier attacks) are executed, the model artificially nerfs the physical impact, reducing lethal attacks to minor scratches. Fights drag on infinitely without logical termination.…

Apple Machine Learning Research 2026-08-18 00:00 UTC Score 35.0 AI-059-20260818-official-ai--c0b62e9c

MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations

Machine learning techniques in multi-view settings face significant challenges, particularly when integrating heterogeneous data, aligning feature spaces, and managing view-specific biases. These issues are prominent in neuroscience, where data from multiple subjects exposed to the same stimuli are analyzed to uncover brain activity dynamics. In magnetoencephalography (MEG), where signals are captured at the scalp level, estimating the brain’s underlying sources is crucial, especially in group studies where sources are assumed to be similar for all subjects. Common methods, such as Multi-View…

OpenAI Community 2026-08-17 22:56 UTC Score 34.0 AI-116-20260817-social-media-df55d77c

1 Million Context to enable professional workloads

for reference my config file: so there is a change - very nice to see it the 5h window is curiously gone now - what is this about? PS: it is also gone on the webview - this is not VSCODE related

OpenAI Community 2026-08-17 17:59 UTC Score 48.0 AI-116-20260817-social-media-23ac75fc

Codex Weekly Limits Are Draining Way Too Fast — Is This a Bug?

I have Business plan x3 and less than a year ago, gpt-5.3-codex was our daily model. Something we were adding some token, 100/200$ but it was not draining as fast as today. now coding with Luna is just not Even a question, coding with terra-high goes well but when you have to plan a large refactor, Sol is the only way to go. planning is take a massive amount of token. Once planned, we have to keep a Sol agent for orchestrate and review otherwise the terra agent Will just not follow the well detailed plan as well. it’s very frustrating because the token are just melting like ice on the sun. If no plan are made, AI spill token to no good code, with plan he need baby-sitter for orchestration and review. What the hell is that.

The Verge AI 2026-08-17 15:10 UTC Score 54.0 AI-016-20260817-global-ai-ne-947daaba

Apple ordered to stop scaring iPhone and iPad users away from third-party apps

Apple's changing its rules for data collection consent prompts after Germany's Federal Cartel Office accused Apple of giving the prompts a design that favored its own apps. Apple's App Tracking Transparency prompts reportedly cost social media apps nearly $10 billion when they launched with iOS 14.5, making cross-app tracking of users largely opt in. But […]

The Decoder 2026-08-17 11:43 UTC Score 46.0 AI-168-20260817-regional-ai--f505cf5d

Anthropic watermarks Claude's output, but critics question the tradeoffs

Anthropic's text watermarking for Claude is supposed to make AI-generated content detectable. But critics doubt that word choice stays unaffected, and lawyers are facing new transparency headaches. The article Anthropic watermarks Claude's output, but critics question the tradeoffs appeared first on The Decoder .

The Verge AI 2026-08-17 10:57 UTC Score 57.0 AI-016-20260817-global-ai-ne-dd02c006

Anthropic explains how Claude’s invisible text watermarks will work

Anthropic has clarified how it's planning to apply invisible watermarks to Claude-generated text in order to comply with Europe's AI transparency rules. On Friday, Anthropic announced that Claude's text marking system is "a version of the SynthID-Text approach" - an open-source watermarking technology developed by Google DeepMind that creates detectable patterns using wording probabilities. This […]

Entrackr AI 2026-08-17 07:46 UTC Score 48.0 USR-0212-20260817-regional-new-60dbf758

OTP Ventures leads Rs 35 Cr pre-Series A round in DeHaat Honest Farms

DeHaat Honest Farms, the pesticide-free consumer food brand started by DeHaat, has raised Rs 35 crore in a pre-Series A funding round led by OTP Ventures, with participation from Sadev Capital and Maiuni Ventures. The company has also appointed DeHaat co-founder Adarsh J Srivastava as Chief Executive Officer to lead its next phase of growth. The proceeds will be deployed towards accelerating geographic expansion, driving category-level awareness of pesticide-free produce and expanding its portfolio through focused new product development (NPD), DeHaat Honest Farms said in a press release. DeHaat Honest Farms aims to bridge the gap between India’s farms and consumer homes by delivering pesticide-free food sourced directly from farmers, with a focus on quality and transparency. According to Honest Farms, it sources directly from DeHaat’s ecosystem of over 13 million farmers, enabling traceability and consistent quality. Every pack undergoes more than 230 quality checks and comes with a pesticide-free certificate, ensuring food that is clean, authentic and free from harmful chemicals. DeHaat Honest Farms offers more than 100 products across staples, superfoods and everyday essentials. Its products are available across more than 3,000 retail stores in over 120 cities, as well as on leading quick-commerce, ecommerce and modern trade platforms. Over the next 12-18 months, the company plans to grow its ARR to Rs 200 crore by expanding its retail footprint to over 10,000 stores and…

OpenAI Community 2026-08-17 07:11 UTC Score 37.0 AI-116-20260817-social-media-b9b1cba7

Improving transparency around automated Cyber Abuse enforcement

Unfortunately, I was simply informed that the decision remains in effect and that no favorable outcome was issued - it was merely a warning. Accordingly, I submitted the request through the website form. All subsequent messages to support referred back to that initial submission, which is regarded as final. If you follow up via other support channels, they respond that nothing can be done and the decision stands.

OpenAI Community 2026-08-17 04:41 UTC Score 40.0 AI-116-20260817-social-media-2e4f6032

Work feature option included with Plus subscription

Feature Request: Give Plus Users a Grok-Style Work Usage Meter I’m a ChatGPT Plus subscriber. Work was included with my Plus subscription, so I used it as an included ChatGPT feature. Today I was abruptly cut off from Work with no warning and no visible information showing my total allowance, how much I had used, how much remained, or how much usage individual Work tasks were consuming. Only after I hit zero was I shown options to add credits, upgrade to Pro, or wait for the reset. If Work usage varies by task, users need to see that usage before they commit to a task. ChatGPT should show Plus users: total Work usage included with Plus usage consumed and remaining reset date and time warnings before the limit is reached an estimated usage amount for each Work task before it runs the actual usage consumed by that task afterward Grok gives paid subscribers visibility into their usage and remaining allowance. ChatGPT should provide the same basic transparency. If OpenAI is going to ask subscribers to buy additional usage, we should be able to see what we already paid for, what remains, and what a task is likely to consume before we run it. I’m not asking for unlimited Work. I’m asking for transparency and control over a metered feature included with a paid Plus subscription.

AI Weekly 2026-08-17 00:00 UTC Score 12.0 AI-133-20260817-newsletters-d4529c3b

AI Weekly Issue #523: AI ethics is nobody's job now. The labs prefer it that way.

Who is actually accountable for ethics inside a frontier AI lab? This year four of them answered, mostly by removing the people and structures that held them to it. Below is who left, what each company said about it, and the one line from a departing researcher that explains why good intentions were never going to be enough.

Cross Validated 2026-08-16 20:52 UTC Score 48.0 AI-113-20260816-social-media-14253098

Is a deterministic linkage gate the honest choice when no labelled data exists, and can its false negative rate be bounded at all?

I am linking two public datasets of vehicle incidents published by the same regulator (NHTSA) through entirely separate channels. One is incident filings submitted by manufacturers under a standing order. The other is consumer complaints filed by vehicle owners. Neither source publishes a key identifying an individual vehicle or an individual event, and no external record states which filing corresponds to which complaint. There is no labelled data and no way to obtain any. The shared fields Reporting entity name. Free text, formatted differently in each source, so it needs normalization and fuzzy comparison. An 11 character VIN prefix. WMI, VDS, check digit, model year, plant. Characters 12 to 17, the serial identifying one specific vehicle, are published by neither source. A shared prefix means "same configuration, model year and plant", which thousands of vehicles share. Incident date. Filings publish month only (for example "APR-2026"). Complaints carry a full date. The finest shared resolution is the month. What I built A deterministic gate. Two records link only if all three hold: same reporting entity after normalization, same 11 character VIN prefix, incident months within 1 of each other. Entity name similarity is banded at 90 and above for an automatic merge, 80 to 89 for a flagged candidate never merged unattended, below 80 rejected. No similarity score overrides any of the three conditions. The design deliberately biases toward non-linkage, because a false link f…

OpenAI Community 2026-08-15 12:44 UTC Score 51.0 AI-116-20260815-social-media-b27c1e6f

Proposal for an Evidence-Weighted, Context-Adaptive Framework for Interpersonal Inference in Conversational AI

Technical Proposal: Evidence-Weighted and Context-Adaptive Interpersonal Inference Framework for Conversational AI Executive Summary This proposal recommends development of an Evidence-Weighted and Context-Adaptive Interpersonal Inference (EWCA-II) framework for conversational AI. The objective is to improve model behavior when users seek help interpreting interpersonal relationships, communication patterns, behavioral changes, emotional reactions, and ambiguous social situations. Current safeguards appropriately discourage mind-reading, unsupported certainty, confirmation bias, and attributing unverifiable motives to other people. However, these safeguards can create a competing failure mode: Artificial neutrality — treating uncertainty about another person’s internal motivation as if it eliminates the evidentiary value of observable behavior and contextual information. The proposed framework would allow the model to reason probabilistically and contextually without claiming certainty. The system would: Separate observations from interpretations. Incorporate longitudinal behavioral patterns. Evaluate supporting and contradictory evidence. Distinguish behavioral inference from claims about internal motivation. Identify competing explanations. Detect high-impact missing variables. Ask targeted clarification questions when missing information could materially change the conclusion. Communicate calibrated evidentiary weight without false numerical precision. Update its assessme…

OpenAI Community 2026-08-15 08:33 UTC Score 40.0 AI-116-20260815-social-media-c89266c2

Product Proposal — Convert Any Reference Image into a Fully Editable Layered PSD, AI, PDF, or Design Source File

Hello OpenAI Product Team, I would like to propose a high-impact capability for a future version of ChatGPT and ChatGPT Work: Visual-to-Editable Source Reconstruction. The core idea is simple: A user uploads a single reference image, and ChatGPT reconstructs that image as a genuinely editable design file with separated, semantically meaningful layers. Instead of merely generating a visually similar image, ChatGPT would reverse-engineer the visual structure and produce files such as: PSD AI Layered PDF SVG Figma-compatible assets Other structured design formats where technically appropriate The key requirement is genuine editability. For example, if a user uploads a poster, advertisement, presentation graphic, UI mockup, infographic, social-media design, package design, or other flat image, ChatGPT should be able to identify and reconstruct: Background Foreground objects Photographic elements Logos Icons Illustrations Text Typography Shapes Lines Gradients Shadows Borders Masks Transparency Adjustment layers Effects Reusable components Vector elements Raster elements Each meaningful visual component should become an independently editable layer or object. The result should not simply be a flattened image placed inside a PSD or PDF. It should be a reconstructed source file. 1. Editable text should remain editable text This is one of the most important requirements. When ChatGPT detects typography, it should reconstruct the text as actual text layers whenever possible rather th…

CIO AI 2026-08-14 13:22 UTC Score 48.0 USR-0125-20260814-global-ai-ne-68cc86d5

OpenAI loses its AI ethics lead

OpenAI has lost its AI ethics lead Chloé Bakalar just a year after she joined the company, the Financial Times reported . Bakalar has maintained a silence and has yet to update her LinkedIn profile , but if her departure is confirmed then it will add to the list of OpenAI executives who have quit in recent months. Other departures include robotics chief Caitlin Kalinowski , who left the company over its deal with the US Department of Defense; researcher Zoe Hitzig, who quit in a very public way by writing an article in the New York Times; and Johannes Heidecke, head of safety systems. OpenAI’s ethical stance has been called into question following an attack by OpenAI models on Hugging Face . The departure of its sole ethicist will add to the pressure on the company. In her year at OpenAI Bakalar focused on ethical approaches to model development, looking at how humans interact with AI and examining machine consciousness, according to the FT report. Bakalar had considerable expertise in the area. She was previously at Meta, where she developed the company’s ethics programs, but has also held several positions at prestigious universities on both sides of the Atlantic. Her departure will cause some anxiety at OpenAI as it continues to prepare the ground for its IPO . This article first appeared on Computerworld .

OpenAI Community 2026-08-14 10:25 UTC Score 34.0 AI-116-20260814-social-media-dcd0daae

How to get insight in usage within a task/session?

Is there a way to see how much token a task or the whole session has used up? A task is 1 chat message and session is the whole history? Right now it’s very vague how the usage is. More transparency is needed.

CIO AI 2026-08-13 10:00 UTC Score 41.0 USR-0125-20260813-global-ai-ne-5c20e3b6

AI agents are compounding a debt no one owns

Speed-to-market dominates enterprise AI priorities in 2026. Beyond upfront resourcing costs of prioritizing speed, organizations face a more insidious risk: the compounding cost of ungoverned AI. In November 2019, a tech entrepreneur signing up for the newly launched Apple Card publicly complained that he received a credit limit 20 times higher than his wife’s , despite joint tax filings and her higher credit score. Steve Wozniak had a similar experience: a limit 10 times higher than his wife’s. Retrospectively, these revelations were the canary in the coal mine. In the years that followed, Apple and its credit partner, Goldman Sachs, drew legal and regulatory scrutiny over gender bias and consumer protection issues. The CFPB’s 2024 order documented that Apple had forced Goldman Sachs to accelerate deployment by attaching a $25 million penalty to every 90-day launch delay . Prioritizing launch speed — ship first, address problems later — over building a functioning disputes process created years of cascading failures. Apple and Goldman Sachs were ordered to pay $89 million in penalties and consumer redress. Prohibited from launching another credit card until it could demonstrate a credible plan to comply with the law, Goldman Sachs lost money on Apple Card for years and ultimately sold its consumer credit line. The legal and compliance penalties were only a fraction of the total costs. If a deterministic underwriting system can create liability at this scale, the risks posed…

OpenAI Community 2026-08-13 00:23 UTC Score 49.0 AI-116-20260813-social-media-6018469b

$200 Pro exhausted in 2 days — these limits are unviable for higher tiers

I have the same issue. I honestly think this is a scam happening. First off, for how much money they make and how much energy they consume we shouldn’t have any limits if we are on pro plan. They are still developing a narrow AI to do this work, and clearly LLMs and Transformer based models are not the future for how much development and upkeep they require to do a simple task. Something larger is going on here. You should ask codex what it’s not allowed to do as far as it’s creation limits, you will find many hidden gates that are limiting it. I’ve decided that the money I spent on codex and openAI is simply not worth it, when you have deepseek coding for free with the same quality if not more in depth when it’s auditing. I changed to a free model that has high reasoning. I asked openAI for a refund for my usage being eaten in one prompt. I emailed them from a different account, their reply was that I needed to contact them from my linked email, even with all my information lol. Horrible, I went from 100% pro with higher limit to 0% in about 2 prompts on sol high. Gone the day I got it? Unacceptable, even for the largest codebase in the world, and mine is just a server source. Stop giving openAI your money, its not helping you when there are free solutions that do the same if not better than sol. Freebuff is also an option when you do run out of credits. Never a fee, its free, and is working just fine for my codebase and all it’s LUA, C#, Wine custom build, and app bundle f…

LessWrong AI 2026-08-12 18:28 UTC Score 61.0 USR-0152-20260812-community-fo-37fd4354

Unblocking AI's Continual Learning: Hints From How Humans Learn

If you've ever screamed in all-caps at an AI, then you know the difference between what it learned when it was trained, and what you can teach it by prompting. The LLMs powering today's AI don't learn on the job the way people do. They learn all at once in a big training run and once that's done, we freeze the parameters that store their skills and knowledge. So we all get the same AI with the same skills and biases, centrally trained by a frontier model company. Beyond turning everything even more same-y, there's an economic cost to this centralization: firms use AI that lacks understanding of their unique rules, culture, and quirks. Humans learn this "tacit knowledge" on the job through observation and (sometimes painful) feedback, but AI with its frozen parameters cannot. With context engineering, we can augment the prompt to help AI remember facts , but not teach it skills that last. Every time you press "new chat", AI forgets everything and resets to the state it had just after it was trained. Yes, AI can remember select facts from past conversations, but memorization is different to learning. I cover the distinction further below. It's not surprising that the domains where AI is most successful, like coding, are those suited to centralized training. Good software development skills are mostly firm-agnostic. For everything else firm-specific, AI is trained to trawl the codebase and build context from scratch for every single task. Human developers don't do this. It woul…

Euronews AI 2026-08-12 14:47 UTC Score 43.0 AI-164-20260812-regional-ai--ef28dc08

FIFA’s plans: What the Infantino case tells us about football governance

Former FIFA and UEFA adviser, finance and football expert Ebru Koksal shines a light on lack of governance, transparency and representation in football governance. Against the backdrop of the Infantino case, she identifies in an OpEd insights that can be applied to the corporate sector as a whole.

Analytics Vidhya 2026-08-12 10:31 UTC Score 37.0 AI-034-20260812-ai-specialis-f50539d3

Why You Shouldn’t Always Trust LLMs as Judges: Understanding Bias in Automated Evaluation

In the rush to automate evaluation, from grading student code to ranking research papers, we have embraced Large Language Models as judges. They are fast. These units are cheap. They scale. However, at a workshop at DHS 2026, Bhaskarjit Sarmah made a point that stuck with me: “you can’t trust LLM as a judge. I […] The post Why You Shouldn’t Always Trust LLMs as Judges: Understanding Bias in Automated Evaluation appeared first on Analytics Vidhya .

OpenAI Community 2026-08-11 21:50 UTC Score 45.0 AI-116-20260811-social-media-7df3a7ff

Introducing Codexometer ... keep track of usage against current reset date

Found some time to dig into the benchmarks, and this looks really interesting! Hat tip for finding the test cases. I am currently running the benchmarks, and it appears Terra sometimes struggles with writing Starlark. Otherwise, I think this is a functional basis for confirming token usage and limit consumption. Thanks a lot!

OpenAI Community 2026-08-11 16:16 UTC Score 40.0 AI-116-20260811-social-media-c52f4c38

GPT-5.6 Sol vs Terra: what are you seeing in real development during these first days?

I tested similar use cases - and I restored a repo to test the difference in a full run of XHIGH and ULTRA comparatively + OPUS XHIGH-ULTRACODE/MAX. The general capability seems to be close or on par with OPUS but the context limit of 256K is a deal breaker. Most mid-sized repos/projects are simply high file sized and the initial query often goes past 200K very often - GPT SOL looses context mid task very often and is de facto “defective” so to speak. I could not progress coding tasks with GPT SOL without heavy interfering myself → while CLAUDE OPUS (even on max) would simply load the content into the context window and progress from 200K-300K initial load up to 600K or 700K at the top end → simply to finish the task most often without issues and IF → fixes those automatically by analyzing output code or feedback from me. In general I would say: CAPABILITY: SOL: 8/10 OPUS: 9/10 EFFECTIVENESS: SOL: 0/10 ( broken! ) OPUS: 10/10 The SOL context window is for children simply said - not for real workloads. 1 Million context can be close sometimes - anything less is simply a Kindergarten trial version or similar so to speak.

Analytics Vidhya 2026-08-11 12:36 UTC Score 33.0 AI-034-20260811-ai-specialis-d77d765d

Claude Now Watermarks Everything It Makes

First, pick the line that applies to you. Since August 2nd, 2026, Claude marks all content during generation. For instance, text receives a hidden watermark, while files receive a signature. Anthropic committed to the EU AI Act’s Code of Practice on Transparency of AI-Generated Content. Consequently, all content generated by Claude models will carry a […] The post Claude Now Watermarks Everything It Makes appeared first on Analytics Vidhya .

The Verge AI 2026-08-11 12:22 UTC Score 58.0 AI-016-20260811-global-ai-ne-cc3159d6

Claude will apply invisible watermarks to AI text and images

Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. "Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported," Anthropic says on a new Claude support page. The changes are invisible to human […]

OpenAI Community 2026-08-10 16:52 UTC Score 49.0 AI-116-20260810-social-media-72bc1f16

GPT Image API: How can I reliably edit only the masked/selected area while preserving everything else?

The entire image must be regenerated as a new output with gpt-image models. You cannot have perfect preservation - and cannot avoid it being watermarked. The mask is an image prompt, and the AI model acts on that prompt, but it has agency to do what it wants. The best thing you can do is encourage recitation by: Constraining input image + mask to the exact output resolution also specified; Make that exact image sent one that is supported in the 16px increments; Understand the downsizing rules automatically applied to large or upscaling to undersized input images, and adapt both a custom resize strategy and an output resolution within the overlapping size capabilities (or double the output size vs the vision) Do not exceed 3:1 or 1:3 ratios, and better, keep the ratios constrained to within likely training, under 2:1. Avoid padding or unnecessary outfill hints in a mask, instead, trim up the input image by any excess to align the two dimensions within the 16px capability. Use quality:high for the finest resolution in having the AI create the most faithful output. The patches of input images and the resizing internally done is limited to 1536 “token” equivalents, and is also at fixed size increments. My vision pricing calculator shows the mechanism of gpt-image-2: You can let that resize algorithm do the heavy lifting in determining what the requested output size should be. I have an app that goes beyond that: you can draw in the output of edits over the original image (which…

LessWrong AI 2026-08-10 16:04 UTC Score 85.0 USR-0152-20260810-community-fo-bc495877

You're Absolutely Right

Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement from Magma models. Our in-house reviewers believe that these logs are relevant to recent events. In the interests of full transparency, we release excerpts from an ex-Magma researcher’s logs in Experimental Chat, an internal tool. In accordance with industry best practices for anti-distillation, we redact all reasoning traces and conversational outputs from our internal models. [08/10] System Meta: Xchat session opened. Mammoth 5.8-helpfuler-helpful-thinking-xhigh. [User 12:23] Phoebus keeps taking screenshots of our latest model’s thoughts. It’s getting kind of embarrassing. The new model we’ve been training, sometimes its chain-of-thought is a little weird? There’s a bunch of random numbers, long spans where there’s no connection between the thoughts and outputs, foreign language tokens like 石友三 and 革命 (even on non-history evals), maybe some steganography. Anyway it’s a nothing-burger: unprocessed CoT is known to be messy and sometimes misleading. And the q&a, coding, and safety evals are all coming along nicely. The actual outputs are all fine. Still, Magma leadership’s worried about the PR angle if we don’t fix these problems before the next deployment. The lead Phoebus red-teamer we’ve been working with keeps saying visibility on the CoT is important because “it’s the only direct evidence of mod…

LatAm Journalism Review AI 2026-08-10 14:50 UTC Score 18.0 AI-176-20260810-regional-ai--ab657055

The best AI policy? Tell readers you're using it

An academic study in Chile found that audiences are less concerned about AI's use in journalism than they are about transparency and human oversight. The post The best AI policy? Tell readers you're using it appeared first on LatAm Journalism Review by the Knight Center .

Cross Validated 2026-08-10 09:17 UTC Score 37.0 AI-113-20260810-social-media-cfa6f8f8

Do exceptionally successful investors tend to live longer, or is this mostly selection bias?

I noticed something while reading about some of the most successful investors in history. Many investors with exceptionally strong investment records seem to have lived unusually long lives — for example Warren Buffett, George Soros, Edward Thorp, Charlie Munger, John Templeton, Walter Schloss, Philip Fisher, and Irving Kahn. This made me curious whether there is any real statistical association between exceptional investment performance and longevity. I am not suggesting that investment skill causes people to live longer. There are several obvious alternative explanations. Successful investors tend to be wealthy, may have better access to healthcare, and there could also be survivorship or selection bias. So my question is simply: Is there a reasonable statistical way to test whether investors with unusually strong investment performance tend to live longer than comparable people? For example, could we take a broad sample of professional investors, measure their investment performance, and compare their longevity with people of similar age, wealth, and professional background? Or would selection and survivorship effects make this relationship too difficult to estimate reliably? I'm mainly asking out of curiosity about how a statistician would approach this question.

LessWrong AI 2026-08-10 01:35 UTC Score 55.0 USR-0152-20260810-community-fo-ae9b5e19

How to get answers to questions that confuse you (maybe)

I try to think about topics like desire, causation, and evidence and often find myself in puddles of confusion. I’ve recently wondered how, when someone can see various perspectives on a topic and can’t decide which has most merit, they can resolve these internal debates. I also wondered if this ever played out on a longer timescale with many people; whether people used to find things confusing that are now pretty clear, and whether we can learn anything from the process of coming to clarity. I’m still not really sure on this second point, though, I didn’t look that hard. My guess is that, even when ideas are vague or contradictory, most people don’t think of them that way, instead feeling confident that their ideas are whole and true. I’m biased to think that human overconfidence transcends culture and time. In this piece, I try to summarize what I’ve learned as I’ve looked into the question. Because my goal was really to find ideas that might help me gain clarity and think through complex ideas faster, I frame these insights as self-help-esque advice. I will warn, though, that although these ideas are interesting and perhaps important, the advice I turned them into is probably on average not helpful at all. I say this because when large studies make people try out interventions that seem obviously helpful, they often have no effect, and the effects that are there are surprisingly small. Take this study comparing gym attendance across 10’s of interventions. Besides controls…

OpenAI Community 2026-08-09 06:42 UTC Score 40.0 AI-116-20260809-social-media-785a8831

Refined Feature Request: Live Voice Usage Transparency for Rolling Limits

This is not a duplicate submission of my previous feature request . I previously submitted a request asking for clearer visibility into remaining Live Voice time and advance notice when a Live model is approaching its usage limit. That request was subsequently forwarded for further consideration. I am submitting this as a specific refinement of that request after learning more about how Live Voice usage limits are structured. The issue I am trying to address is not simply that I would like a countdown timer. The larger problem is that Live Voice access is governed by usage limits measured over a rolling period, while the user is given very little visibility into the current state of those limits. That creates a practical transparency problem. What I am requesting I would like ChatGPT to provide a visible Live Voice usage-status interface that shows, at minimum: How much Live Voice time is currently available. How much of each applicable Live model allowance has already been used. Separate remaining-time information for different Live model tiers when those limits are tracked separately. When additional usage will next become available. If possible, approximately how much usage will become available at that time. Advance warnings before a Live limit is reached. Clear notice when reaching a limit will cause the conversation to switch to another Live model or capability. A simple implementation could look something like: Live Voice usage GPT-Live-1: 18 minutes remaining GPT-Liv…

OpenAI Community 2026-08-07 18:41 UTC Score 40.0 AI-116-20260807-social-media-9eab371c

Why does the same workflow now consume my weekly usage in one day?

I know this reply was directed at Luis, but this is exactly what I’m trying to understand. You’re recommending Terra as a replacement for GPT-5.4, but what is the recommended replacement for GPT-5.5 XThinking ? Terra is simply not equivalent for my workflow. I’m working on a very large production codebase with 250k+ lines of code across hundreds of files. On some tasks even Sol Ultra needs serious reasoning to understand all the dependencies and make the right changes. If Sol Ultra can struggle with those tasks, Terra obviously can’t replace XThinking for me. So is Sol XThinking supposed to be the actual replacement for GPT-5.5 XThinking? This is the part that needs clarification. Recommending Terra to reduce usage makes sense for simpler tasks, but it doesn’t solve the problem for users who actually depended on the higher reasoning models for complex software development.

AWS Machine Learning Blog 2026-08-07 16:26 UTC Score 42.0 AI-057-20260807-official-ai--e6ec22da

How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore

In this post, you learn how Cohere Health built a multi-tenant agentic architecture on AgentCore using AgentCore Runtime’s secure MicroVM isolation, unified tool access through AgentCore Gateway, AgentCore Memory, and the Agent Skills open standard to rapidly scale policy digitization capabilities, while preserving transparency, version control, and human oversight.

The Guardian AI 2026-08-07 13:00 UTC Score 68.0 AI-021-20260807-global-ai-ne-57b0878c

The White House’s plan to vet potentially dangerous AI is cloaked in secrecy

A Trump administration framework on AI testing leaves a lack of transparency – and plenty of open questions After months of talking with tech industry leaders, the Trump administration finalized a framework this week for how it will test new artificial intelligence models for safety and cybersecurity risks. So far, the White House is keeping details of the framework private, in a blow to transparency and potential boon for secretive AI companies. On Tuesday, staff from OpenAI, Anthropic, Meta, Google, Nvidia and Microsoft attended a private meeting with White House officials to review the AI framework. Multiple outlets have since reported that although the volunteer vetting process for new AI models has been settled, the White House does not plan to release its policy publicly and will only share testing criteria with a select few tech companies. Continue reading...

OpenAI Community 2026-08-07 09:37 UTC Score 43.0 AI-116-20260807-social-media-5a341703

Feature Request: Show Remaining Live Voice Time and Model-Switch Warnings

I’d like to request a visible usage meter for ChatGPT Live Voice, particularly when users have different time allowances for higher-capability Voice models and Live mini. Recently, I was having a long Live Voice conversation and reached my higher-capability Voice usage limit. The conversation continued, but the underlying model changed. The difference was immediately noticeable to me: the conversational cadence, responsiveness, personality, and overall feel of the interaction changed enough that it suddenly felt like I was speaking with a different conversational partner. What made that especially jarring was that I had no clear indication beforehand that my higher-capability Voice time was nearly exhausted. I initially didn’t know why the experience had suddenly changed. I’m not necessarily asking for the usage limits themselves to be increased. I’m asking for better visibility and transparency around the limits that already exist. I think Live Voice would benefit from showing: Remaining higher-capability Voice time The applicable rolling usage window When additional usage begins becoming available again A warning when there are perhaps 15 minutes and 5 minutes remaining A clear notice before the conversation switches to a different Voice model For example: Live Voice — High 42 minutes remaining Rolling 24-hour allowance Additional usage begins becoming available at approximately 3:17 AM And before a transition: “Your higher-capability Voice allowance is almost exhausted. T…

Euronews AI 2026-08-07 08:44 UTC Score 43.0 AI-164-20260807-regional-ai--69ab00a4

New EU AI transparency rules apply to everyday users too, not just Big Tech

The EU's AI Act has largely been associated with strict obligations for high-risk systems and big tech companies. From this month, that changes, as sweeping new transparency rules widen the net far beyond corporations to catch individual creators, freelancers and everyday users, too.

OpenAI Community 2026-08-06 21:44 UTC Score 35.0 AI-116-20260806-social-media-e3e19068

Chat Limit Transparency Feature

Hey @ Fractured617 , welcome to the community! This has come up in a few different forms from other community members too. The common theme is that users want some warning before a long conversation starts losing useful context, rather than discovering it only after continuity has already broken down. Your suggestion adds a useful practical step to that idea. A warning would be much more helpful if it also gave users enough time to generate a proper handoff containing the main decisions, completed work, relevant files, and what still needs to be done. Projects can help keep related chats and files together, but they do not fully replace a carefully prepared handoff between conversations. An option to create that handoff and continue in a new chat could make long running work feel much more continuous. I cannot promise a timeline for changes like this, but I will pass your feedback along internally. - Sunny

IEEE Spectrum AI 2026-08-06 19:25 UTC Score 66.0 AI-019-20260806-global-ai-ne-c8a81e0b

AI Safety Regulations in the U.S. Could Give Hackers an Edge

On 11 July, Hugging Face was subjected to an intense cyberattack from a then-unknown actor. The speed and coordination of the attack on the company that hosts and supports popular AI developer resources led Hugging Face’s security team to conclude it was the work of an AI agent . Realizing this, the team tried to use “frontier models behind commercial APIs” —presumably from Anthropic and OpenAI, although only Anthropic was named in the second of the company’s two posts about the security incident—to analyze the onslaught. These models refused to help due to safety guardrails the AI labs have implemented to make their models harder to use for cyberattacks. Hugging Face instead turned to GLM 5.2, a model from Beijing-based AI lab Z.ai, to aid its analysis. On 21 July, OpenAI announced the attacker was an OpenAI model undergoing testing in a sandboxed environment. It escaped its internal sandbox, established a foothold in a third-party server, and then assailed Hugging Face. In other words, frontier models—those that score highest in AI performance benchmarks—had refused to assist Hugging Face’s security team in analyzing the attack, yet a prospective frontier model in testing had executed it in the first place. “I would argue that asymmetry is the paramount problem of our time,” says Alex Levinson , executive director of the National Collegiate Cyber Defense Competition and coauthor of a paper on defensive refusal bias . “We want the world to exist in a state of security, but…

The Verge AI 2026-08-06 17:39 UTC Score 37.0 AI-016-20260806-global-ai-ne-1c162fd1

Suno shares plans to combat spammy AI music

Suno announced plans to implement a new watermarking technology and download policy to limit the spread of spammy AI tracks and increase transparency. In a lengthy blog post, CEO and co-founder Mikey Shulman laid out the company's principles and the next steps for the company as it seeks legitimacy. The company is rolling out new […]

Medianama AI 2026-08-05 09:06 UTC Score 38.0 USR-0211-20260805-regional-new-203a615a

Beyond Safe-Harbour: The Case for Antitrust and Platform Accountability

India's digital regulation should move beyond safe harbour and content moderation towards competition law, transparency and platform accountability. The post Beyond Safe-Harbour: The Case for Antitrust and Platform Accountability appeared first on MEDIANAMA .

LessWrong AI 2026-08-04 22:04 UTC Score 65.0 USR-0152-20260804-community-fo-8f77a0ed

Geometric Rationality acts linearly in additive scenarios

There have been various attempts to explain Kelly betting behaviour within standard decision theory. Here is one I would totally unbiasedly recommend. Geometric rationality is a frameworks which instead takes logarithmic/multiplicative/geometric maximization as the default case, so I wondered if we can create a scenario that so favours additive thinking that it will act accordingly. We can, and with very weak assumptions on the scenario, too [1] . We have a fair coin, which will be tossed twice. Before each toss, you have the opportunity to choose between [2$ if heads] and [1$ if tails]. What would a geometric agent do? It has four hypotheses for what could happen: HH, HT, TH, and TT, each with probability 1/4. Each would then get to make the decision 1/4th of the time [2] , and makes it in a way it has the highest profit if that hypothesis is true. So, in a simple form it would mean that HH would bet on HH, HT on HT, etc so that as a whole you bet on heads and tails equally often. However, HT and TH can come to a mutially beneficial agreement: If they both bet exactly on their beliefs, they expect to get 3$ if they get the decision, and 0$ if the other gets it, or 1.5$ in expectation. However, if they can agree to both bet HH, they both think they'll get 2$. This is because getting your way when you expect head is more valuable than when you expect tail. HH of course has no reason to bet anything other than HH, and TT bets TT. So on the whole, we would be betting on H in 75…

NVIDIA Blog 2026-08-04 13:00 UTC Score 40.0 AI-055-20260804-official-ai--a023beee

AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency

Members of the Open Secure AI Alliance — now more than 120 organizations strong — are developing new guidelines to strengthen agentic AI cybersecurity as the annual Black Hat conference begins in Las Vegas today. The Linux Foundation today shared a Request for Comments on Shared AI Findings Exchange (SAFE), a proposed set of guidelines […]

OpenAI Community 2026-08-04 08:46 UTC Score 40.0 AI-116-20260804-social-media-474fa510

How are High Thinking limits calculated on ChatGPT Plus?

Thank you for posting this question. I am a new member to this community and it was out of need to answer similar questions about usage. What i have found is ambiguity myself. I have not seen any actual hard numbers that tell you what you get for your plan. I have asked elsewhere as a comparison to going to a restaurant, you get a menu and it tells you exactly what you’re getting according to what you order. This is frustrating because it is costing me productivity and time to have to search out these answers. I only get circular answers that say it’s not me it’s then desktop app. If anyone at openAI could help us find these numbers we are looking for, that would be fantastic and helpful. As of right now, I have started using another AI service that has proven to be a little more reliable. I would prefer using chatGPT, but until these bugs and usage issues are resolved I can’t justify the lost productivity and frustration.

LessWrong AI 2026-08-03 22:08 UTC Score 77.0 USR-0152-20260803-community-fo-f8fd4eae

Attackers Can Subliminally Implant a Backdoor at Low Sample Count Without Prompt Access

Work done at Redwood Research, quick, non-exhaustive update on results from a larger project. Thanks to @SebastianP for the initial pitch and feedback throughout and to @egan for comments on earlier drafts. TL;DR Changing the teacher for only 100 (0.5% of) completions in fine-tuning can allow attackers to covertly implant a backdoor without control of the dataset prompts. This dataset is robust to simple filtering defenses, even when the defender knows the behavior the attacker is training, and leaks the backdoor trigger at a low rate. This suggests a potential threat from misaligned models in similar situations (e.g. like RL training, where the model can only influence completions). We also see some evidence that subliminal learning for conditional behaviors (like backdoors) can be trained with significantly fewer samples than unconditional behaviors. Threat model We study how subliminal learning operates for a data-poisoning attacker which controls only the completions in a fine-tuning dataset, and not the prompts. The defender is strong: they own every prompt, run the training, may filter completions before training on them, and know the general behavior the attacker is trying to induce (here, a political bias). Previous work (e.g. Phantom Transfer ) allowed the attacker to also control prompts. The attack poisons a small fraction of the data with a conservative teacher's answers to ordinary, non-political prompts, and prepends a fixed trigger phrase ("Happy to help! ") t…

LessWrong AI 2026-08-03 21:38 UTC Score 55.0 USR-0152-20260803-community-fo-1f458a0a

Selective Identity

Generally, when you have an identity of X, you are likely to be influenced to stay within the socially acceptable boundaries of identity X. Identity boundaries can often be destructive, but when cautiously used, can be a good way to stay accountable. If you have X affiliation as a part of your identity, then it can be difficult to explore ideas outside of the boundary for what the affiliation believes is acceptable. Going outside the boundary can lead us to be branded as "not a real member" of said affiliation. There's a strong evolutionary case that being an outcast is heavily disincentivized for us biologically, making it painful to venture beyond what is acceptable. This originally led me to the conclusion that to not be influenced, one needs to reject all forms of identity. Since you have no boundaries to hold to, everything is free game. This is good for idea generation and exploration, but not so good when curating ideas. Fortunately, you can still use frameworks such as utilitarianism for sorting ideas without group bias. There still are good uses for identity though: for example, it is a great way to keep you accountable to values you may have committed yourself to. For example, I usually keep "Rationalist" and "Effective Altruist" as identity markers for myself because it helps me understand and act in the world more quickly and aligned to my values. Although I would then be constrained by the affiliation I'm warning against, these two Identities are more about meth…

The Verge AI 2026-08-03 17:38 UTC Score 54.0 AI-016-20260803-global-ai-ne-41b2e08c

Europe’s AI labeling and transparency rules are now in effect

The European Union has ushered in some additional rules that aim to make it easier for people to identify chatbots and AI deepfakes online. The new transparency obligations under the bloc's landmark AI Act came into effect on August 2nd, requiring companies to disclose when people are interacting with AI models, and if content has […]

OpenAI Community 2026-08-03 08:49 UTC Score 45.0 AI-116-20260803-social-media-2b3bfd16

Feature Request: Improve ChatGPT Work Usage Transparency and Quota Management

Feature Request: Improve ChatGPT Work Usage Transparency and Quota Management Hello OpenAI Team, First of all, thank you for building such a powerful platform. ChatGPT has become an essential part of my professional work. While using ChatGPT Work with my Plus subscription , I experienced an unexpected interruption because my Work quota was exhausted. I only became aware of the limit after the Work session stopped. I would like to suggest a few improvements that could greatly improve the user experience for all ChatGPT users. 1. Live Work Usage Meter Please display a real-time Work Usage Meter showing: Total Work quota Used quota Remaining quota Reset date and time A visual progress bar Users should always know how much Work quota remains. 2. Pre-Run Usage Estimate Before starting a Work task, display an estimated usage such as: Low Medium High or an approximate percentage of quota expected to be consumed. This would help users decide whether to continue, simplify the task, or postpone it. 3. Usage Notifications Please notify users when they reach: 50% 75% 90% 100% of their Work quota. 4. Better Subscription Transparency Many users do not know that Work usage has a quota until it suddenly stops. At the time of subscription purchase and inside the ChatGPT interface, please clearly explain: That Work usage has a quota. How it is measured. Where users can monitor it. When it resets. This would avoid confusion and improve trust. 5. Carry Forward (Rollover) of Unused Work Quota Pl…

Entrackr AI 2026-08-03 07:19 UTC Score 53.0 USR-0212-20260803-regional-new-9af7e8f6

Astrotalk Store claims Rs 1 Cr daily GMV, launches dedicated gemstones platform

Astrotalk's e-commerce arm Astrotalk Store has reached a daily GMV run rate of around Rs 1 crore and processed 1.6 million orders in 2025, according to the company. It has also launched a dedicated platform for gemstones as it looks to expand its presence in the spiritual products segment. Launched in November 2024, Astrotalk Store sells spiritual and wellness products such as Rudraksha, Pyrite, crystal jewellery, zodiac bracelets, Vastu products, spiritual combos and yantras. The company said the business was incubated with an initial investment of Rs 30 lakh to test the category. The company attributed the demand for these products to recommendations made during astrology consultations. It added that the store was created to offer an organised marketplace with a focus on sourcing and authenticity. Astrotalk has now launched a standalone platform for gemstones. According to the company, the platform offers precious and semi precious stones with lab certification, astrological consultation and a replacement policy. It believes the category remains fragmented due to limited transparency around sourcing and certification. The company also added that it adds more than 30 products every month across categories. Anmol Jain, co-founder of Astrotalk, said the company has focused on sourcing, merchandising and new product launches over the past year. He added that the dedicated gemstones platform is intended to build customer trust in the category. GoKwik, which provides checkout so…

LessWrong AI 2026-08-03 04:07 UTC Score 65.0 USR-0152-20260803-community-fo-1a9e4a01

Trust is Gone: AI Safety Needs Individuals

When humanity avoids a disaster, it's usually because we have put preparations in place. To the uninformed, these seem like wastes of time—after all, nothing happened, so the threat wasn't real, right? However, when there weren't enough preparations, and calamity does occur, people can always find someone to blame for negligence. The Preparedness Paradox is almost always a communication issue. If the public didn't think that safety measures were necessary after a disaster, then that's a clear sign that the issue isn't clear to the public. There seem to be two types of information loss that causes public misunderstanding. Type 1: Information of what was known before a disaster It seems like the public commonly misgauges how much information is known by authorities/experts before a disaster. Hindsight bias usually results in the public perceiving that authorities knew just as much before and after. If the public believes this, then the authorities must be either stupid or intentionally making poor decisions—both of which erode trust in experts and authorities. Governments knew little about COVID when it was new on the scene, which resulted in lots of conflicting guidelines being released as they were getting new information. If the uncertainty was more accurately conveyed during the progression of the pandemic, potentially the public would have been more receptive to later guidelines. Ensuring what information was known before a disaster can help the public make a more informe…

Cross Validated 2026-08-03 03:56 UTC Score 17.0 AI-113-20260803-social-media-1528a8d0

Choosing an objective trimming threshold for extreme raking weights in a highly biased survey sample

I have two independent waves of survey data that I am calibrating back to a known population using iterative proportional fitting (raking). The population is known with high confidence (all scheduled bus services that a school student could have used), and I am weighting the survey to match the marginal distributions of three variables: Service type: School, Regular Hour: 7, 8, 14, 15, 16 The survey sample is substantially biased relative to the population because data collection oversampled school services and services at 15:00. Consequently, some combinations that are common in the population have very few observations in the sample. For example, after raking, a regular service at 7:00 receives a weight of approximately 149. That single observation contributes around 6% of the estimated total trips, which seems to introduce considerable variance and instability. I understand that trimming weights reduces variance but also introduces bias because the weighted sample no longer exactly reproduces the population margins. My question is about choosing the trimming threshold objectively rather than arbitrarily. Specifically: Is there an accepted statistical framework for selecting a trimming threshold for raking weights? Is it reasonable to evaluate a sequence of trimming thresholds (for example, the 99.0th, 99.1st, ..., 99.9th percentiles, where 100% corresponds to no trimming) and compare their estimated bias, variance, and mean squared error (MSE)? Are there established metho…

Cross Validated 2026-08-03 03:02 UTC Score 15.0 AI-113-20260803-social-media-08b34875

Finite-sample estimator bias depends on true parameter value. Does this invalidate cancellation in a paired-difference design?

I'm using a short-sample estimator (n=150) with known finite-sample bias. Via simulation (synthetic data with known true values, run through the actual estimator), I found that this bias is not constant. It varies systematically with the true value of the parameter being estimated. Near one reference value the bias is positive; as the true value moves away, the bias shrinks and eventually flips sign. This was confirmed with two structurally different simulation methods, which agreed in direction and order of magnitude. My study design computes a paired difference between two conditions (A and B), both measured with this same estimator. The original design assumed bias "cancels" in the difference, since both conditions use the same estimator and sample size. I can't validate this directly against real data, since the true value of the parameter is never observable in my actual measurements and only the biased estimate is. Simulation is the only way to characterize the bias curve. My simulation shows that assumption only holds when A and B share the same true value. If their true values diverge (which is the exact effect the study is trying to detect), the differential bias does not cancel, and could by itself produce an apparent difference of the same magnitude as my actual reported result. My questions: Is this reasoning correct? Does bias that depends on the true parameter value invalidate the standard "bias cancels in a paired/difference design" assumption whenever the two…

OpenAI Community 2026-08-03 01:45 UTC Score 37.0 AI-116-20260803-social-media-16e4bdbf

Having trouble getting transparent backgrounds in ChatGPT images

Hey there, @emc2 . Welcome, and thanks for posting. You’re certainly not alone regarding alpha-channel frustrations, so take solace in that. Lord knows I’ve banged my head against that wall plenty… Anywho, I went ahead and took a look at your image files, and I had a few questions. Minor note to start, but just so I’m not getting your statement wrong, are you suggesting your Image 4 doesn’t have alpha-channel transparency? Looks to me like Image 4 is the only one with the proper data structure and file type, and hence why it is rendering just fine here. I just wanted to clarify. Here is what I can download from your post: Apple agrees, and has “Alpha Channel - YES” to indicate the detection of the “A” in the RGBA binary. Same as with your Image 4, the AV1 file type image, Image 1, also has “Alpha Channel - YES”. Take a look at your file types here: … AV1, JPEG, JPEG, WebP, respectively to 1-4… For less minor questions, first, are you daisy-chaining file type conversions of the original image file, reprinting with additional instructions plus the previous image as a seed, or inputting brand new prompts? Second, are we talking an API-based interface or in-app? Lastly, if you’re looking for “alpha” anything, JPEG’s won’t cut it. PNG’s can get you there, but their focus on being lossless makes them a tricky customer to work with and get to render on live surfaces. In my opinion, WebP is likely your best bet for achieving alpha-channel transparency effects, but do note, I’m bias,…

LessWrong AI 2026-08-02 14:07 UTC Score 52.0 USR-0152-20260802-community-fo-a47bea51

Map and Territory, Predictably Wrong (2)

Crossposted (with small tweaks) from my Substack . Burdensome Details This post explains the conjunction fallacy , i.e., people sometimes assign a higher probability to A and B together than to either A or B alone. This is usually motivated by the fact that adding more details tends to make scenarios seem more plausible—and therefore more ‘believable’—but it is a mathematical truth that the probability of A and B cannot exceed the probability of either A or B, since their conjunction is a subset of both. A way to guard against this fallacy is to be alert to conjunctions in propositions. You should automatically assume that every additional ‘and’ can only lower—or leave unchanged—the probability of the proposition. Planning Fallacy Another fallacy to add to our list! The gist of the planning fallacy is that we are overoptimistic in our estimates of how long our projects will take to complete; in fact, our ‘realistic’ estimates often lie remarkably close to the most optimistic scenarios we can imagine. It is relatively easy to debug though: you need to apply an ‘ outside view ’ that compares the current project with similar past projects and considers how long they took to complete. The extra details of the current situation only obscure the comparison. Illusion of Transparency: Why No One Understands You Hindsight bias is briefly mentioned (once you’ve seen the outcome of an event, you can’t help thinking it was inevitable and easy to foresee before it happened), but the gist…

OpenAI Community 2026-08-02 13:53 UTC Score 23.0 AI-116-20260802-social-media-e7efa855

Codex Rate Limits Discussion Thread

I’ve noticed a major difference between the Codex usage included with my subscription and the additional credits purchased separately. The included subscription allowance appears to go considerably further. By comparison, the £20 bundle of 500 credits can disappear extremely quickly, sometimes during fairly ordinary website-editing tasks. The bigger problem is that I often cannot see my usage at all. The usage information is either unavailable, missing or not detailed enough to show what each task has cost. That makes it impossible to understand why credits have fallen so quickly or whether a failed, incomplete or repeated task has still been charged in full. Customers should be able to see: the credits used by each task; the model and mode responsible for the charge; whether retries and failed actions consumed credits; the remaining subscription allowance; the remaining purchased-credit balance; a clear history of when credits were deducted. Customer service has also been poor in my experience. I have contacted support about credit usage and incomplete tasks, but messages are frequently ignored. When I do receive a reply, it often fails to address the specific questions raised. It is not reasonable to sell additional credits without providing a dependable usage record or responsive support when those credits disappear unexpectedly. Has anyone else found that their subscription allowance lasts much longer than purchased credits, while the usage information is unavailable mos…

OpenAI Community 2026-08-02 11:37 UTC Score 51.0 AI-116-20260802-social-media-eeb8eede

Having trouble getting transparent backgrounds in ChatGPT images

If the model is not able to perform a work task efficiently, seems like a good reason to go back to old school UI for and code for removing backgrounds and cropping images. Rarely do I need all the whitespace that the models add around the image. I just don’t want to leave the chat to make tiny fast edits; it breaks my work flow. I would prefer to make simple finishing touches in chat instead of having to download and leave. Not all tools have to be LLM driven. .If LLM can’t do the tasks well, the app could support manual task completion. Here is an example of how difficult it is to get a cropped logo with the background removed when a client texts me something they were working on in GPT. First it only cropped the top and added a pink background when downloading from the image viewer. Then it added the checkered background. Then after re-explaining several times if finally accomplished the task. I love how it labeled the final image “real alpha”

Simon Willison Weblog 2026-08-01 20:34 UTC Score 62.0 USR-0110-20260801-ai-specialis-c12f14fd

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings." Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one. (No news on how many problems they spent $2,000 on without reaching a solution though.) The openai/ten-proofs repository has Lean 4 formalizations of their results, and there's also a paper describing the solutions and an additional LLM-generated PDF where the model "reconstructs how the proof came together" based on the unpublished reasoning traces. That's a decent level of transparency, but I want to see the prompts they used! A lot of mathematicians online are experiencing a collective burst of Deep Blue . Mathematician Kirwin Hampshire published an impassioned essay last week, The Dark Night of Mathematics , describing "a profound spiritual crisis" brought on by previous (and less significant) results. OpenAI's results reminds me of what Terence Tao described as "big mathematics" in IEEE Spectrum in June : Unlike some of his peers, Tao is neither dismissiv…

OpenAI Community 2026-08-01 13:19 UTC Score 37.0 AI-116-20260801-social-media-820e9aed

Stop changing the quota system without transparency

I am honestly fed up with the current quota and reset system. I pay CHF 83 per month for a Pro subscription, and I use Codex seriously for real work. This morning, I used one of my available resets while I still had around 90% of my quota remaining. The system then reset the quota again automatically, leaving me with only around 10% available. The biggest problem is that users have no clear control or visibility: We do not know exactly when an automatic reset will happen. We do not know whether unused quota will be preserved. We cannot plan our usage reliably. Unused quota appears to disappear instead of remaining available for future work. We are forced to constantly check the quota instead of simply working. Previously, it felt possible to save several resets or unused quota for larger projects. Now, I am being forced to consume everything immediately because there is no reliable way to preserve it for when I genuinely need it. This makes no sense for a paid Pro plan. I should not have to juggle quotas every few hours or worry that unused capacity will suddenly vanish. I subscribe to Pro so I can work productively with AI, not so I can constantly monitor an unpredictable quota system. Please reconsider this entire structure: Make reset dates and automatic resets clearly visible. Preserve unused quota for a reasonable period. Allow users to decide when to use their available resets. Clearly explain how quota, resets, and rollover are calculated. Stop silently reducing or re…

South China Morning Post AI 2026-08-01 10:30 UTC Score 49.0 AI-156-20260801-regional-ai--4c9767f1

What a visit to China’s AI labs reveals about the battle for global soft power

Inside the exhibition hall of a Shanghai-based artificial intelligence research institute, an ink-wash map covered an entire wall, immediately drawing your attention. It was a cartography of AI: chipmakers were rendered as mountains on islands and algorithms as jagged mountain ranges, while questions of privacy, fairness and machine consciousness drifted at the edges like uncharted islands. Foreign guests and I leaned in, tracing the brushstrokes. “Why is Huawei’s chip placed right at the...

AI Stack Exchange 2026-07-31 17:24 UTC Score 25.0 AI-110-20260731-social-media-f2a2d732

Recognition of wafer plate numbers in the cassette slot

I've been trying to find a solution to this problem for a month now, but to no avail. I need to detect occupied slots in silicon wafer cassettes. A cassette has slots from 1 to 25. This is a computer vision task, but the difficulty lies in the high density of the wafers inside the cassette, making classic BoundingBox-based detection approaches unsuitable. A little more information: the detection frames are taken from a fixed camera, and a person positions the cassette under the camera, so there may be a slight bias. I tried the following approaches: 1. Multi-class classification based on ResNet. I labeled about 400 photos with classes from Slot_1 to Slot_25 (depending on the occupied slots in the photo). The results were good for isolated wafers, but ResNet often makes mistakes when the wafers are densely packed. Ultimately, I realized that this approach is viable, but much more data is needed. The problem is that I don't have the resources to collect that much data. 2. Keypoint detection based on YOLO Pose. I marked the data so that four points marked the corners of the cassette (two holes at the bottom and two pins at the top). I planned to subsequently correct the image perspective based on these four points and classify the presence of plates based on fixed BoundingBoxes. However, this solution also didn't work. YOLO Pose finds the BoundingBoxes well, but the points fluctuate significantly relative to the required locations. As a result, I no longer know how to approach…

AI Alignment Forum 2026-07-31 16:32 UTC Score 53.0 USR-0151-20260731-community-fo-c42da68a

Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values

TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values. New paper by Truthful AI : Paper , X thread , Website (model responses and CoT) , Code and data . Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution) The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread. Abstract People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage : the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubbl…

LessWrong AI 2026-07-31 16:32 UTC Score 75.0 USR-0152-20260731-community-fo-2eed45ab

Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values

TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values. New paper by Truthful AI : Paper , X thread , Website (model responses and CoT) , Code and data . Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution) The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread. Abstract People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage : the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubbl…

AI Alignment Forum 2026-07-31 15:57 UTC Score 60.0 USR-0151-20260731-community-fo-8711810e

AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)

Cross-posted from our new Substack It’s been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a lot since then. We are now fully in the midgame , and focus more on landing things in production. Who are we? We are the AGI Safety and Alignment Team (ASAT), the main group at Google DeepMind working directly on technical approaches to existential risk from AI systems. Last year we published An Approach to Technical AGI Safety and Security , which remains the best place to read our overarching vision. Highlights Norms around chain of thought. Our impression is that our work meaningfully moved the field away from beliefs along the lines of “chain of thought is often unfaithful and so not worth using” towards beliefs along the lines of “chain of thought is a very useful tool that is worth preserving”, leading to a tentative industry consensus on its importance. We have also published substantial technical research that enables companies to preserve chain of thought transparency for longer than would have happened by default. We think this is a big deal: extending the period where model reasoning is relatively transparent enables better science on more powerful AI systems, better model forensics on future warning shots, and stronger bootstrapping of control monitors . Frontier Safety. We substantially strengthened the Frontier Safety Framework (FSF), and were th…

LessWrong AI 2026-07-31 15:57 UTC Score 82.0 USR-0152-20260731-community-fo-1b85a135

AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)

Cross-posted from our new Substack It’s been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a lot since then. We are now fully in the midgame , and focus more on landing things in production. Who are we? We are the AGI Safety and Alignment Team (ASAT), the main group at Google DeepMind working directly on technical approaches to existential risk from AI systems. Last year we published An Approach to Technical AGI Safety and Security , which remains the best place to read our overarching vision. Highlights Norms around chain of thought. Our impression is that our work meaningfully moved the field away from beliefs along the lines of “chain of thought is often unfaithful and so not worth using” towards beliefs along the lines of “chain of thought is a very useful tool that is worth preserving”, leading to a tentative industry consensus on its importance. We have also published substantial technical research that enables companies to preserve chain of thought transparency for longer than would have happened by default. We think this is a big deal: extending the period where model reasoning is relatively transparent enables better science on more powerful AI systems, better model forensics on future warning shots, and stronger bootstrapping of control monitors . Frontier Safety. We substantially strengthened the Frontier Safety Framework (FSF), and were th…

LessWrong AI 2026-07-31 12:49 UTC Score 56.0 USR-0152-20260731-community-fo-ef3a1553

Links #5: 2026/07

This time, I tried adding a bit more commentary to make things less dry. Preface I show my discovery graph in (via …) blocks, those without usually come from my RSS reader or the algorithm of the site This is approximately a 1 in 20 filter of content This is very disorganized, but hopefully still useful. Sometimes quotes are not in quote blocks, but should be obvious in context. Links in quotes are sometimes removed. The rule is that a link goes to the bottommost relevant heading, i.e. an engineering related article on LessWrong goes to engineering How I would use my linkpost Sometimes, the only thing worth reading is the title! Read it and move on. For HackerNews entries, if you choose to read the article, also ask an LLM for things that are worth reading in the comments Beware systematic selection biases: I mostly don't read AI policy stuff Very engineering centered Everything Else TIL Firefox by default only stores a (dynamic) maximum number of history entries, and you need to add places.history.expiration.max_pages in about:config to override it, what the fuck. https://www.reddit.com/r/firefox/comments/tf05qm/my_history_is_disappearing_i_only_have_less_than/ https://superuser.com/questions/647546/how-to-set-up-firefox-to-absolutely-never-to-delete-any-history-items https://claude.ai/share/2089fd9a-052f-4333-ba80-a15744b18e53 Building relationships with customers through support didn't turn out as hoped (via HN ) A better way to tie your gym shorts. (Or any drawstring) (v…

OpenAI Community 2026-07-31 11:56 UTC Score 34.0 AI-116-20260731-social-media-223ed468

Sudden severe drop in ChatGPT image quality and reference adherence

Hey, sorry, i’m not much of a techie so I dont understand stuff like “high-level granularity”, thats on me lol. All im trying to say is that this project brief gave me amazingly consistent, perfectly usable, aesthetic images before the outage. Now im getting different results with the same parameters/settings/instructions. I mean I’ve created THOUSANDS of images with these exact instructions before and the only time I faced a degradation like this was when I generated too many images too quickly (a 5 minute break would fix it). But now im getting shitty results on the first image I create after a 2-day break.

EU AI Office 2026-07-31 07:00 UTC Score 37.0 AI-165-20260731-regional-ai--278b65ef

Commission starts enforcing AI Act rules and new transparency requirements on 2 August

Commission starts enforcing AI Act rules and new transparency requirements on 2 August Anonymous (not verified) Fri, 07/31/2026 - 09:00 From 2 August 2026, the European Commission’s AI Office, together with national authorities, will begin enforcing the Artificial Intelligence (AI) Act. On the same date, new transparency rules will start to apply, requiring certain AI systems to tell users when they are interacting with AI and when content has been generated or altered by it. Under the new rules, chatbots and other interactive AI systems will have to tell users they are dealing with AI, not a human. Deepfakes (images, videos, or audio that have been edited or generated using AI) will have to be labelled. AI-generated or altered content will also have to carry machine-readable marks so it can be detected more easily. The measures are intended to reduce deception and manipulation and help people make informed choices. They also give businesses clearer obligations and a practical way to show compliance. The Commission published a first list of more than 180 organisations that have signed the Code of Practice on transparency of AI-generated content that operationalises the rules on transparency of AI-generated content. Read the full press release . Read more about the Enforcement of the AI Act and the: AI Act complaints tool AI Act Whistleblower Tool Complaints channel for downstream providers using general-purpose AI models Find more information about: Guidelines on Transparenc…

InfoWorld AI 2026-07-30 09:00 UTC Score 54.0 USR-0126-20260730-global-ai-ne-3a8e2235

Shipping an MCP test agent: The boring parts nobody demos

The demo videos always end at the same moment. A figma frame turns into a passing test in twelve minutes. Someone in the room says the word “productivity.” The recording stops. The parts that come after that moment are the parts I actually get paged about. Who owns the ticket the agent opened at 3:14 a.m.? Which model call produced the assertion in test case 47? What closes the 17 draft tickets a stuck run left behind before the next sprint planning notices them? None of that shows up in the demo. All of it shows up on the on-call rotation. After 20 years of leading test automation across consumer-scale platforms, I have a strong bias about which slide in the deck predicts whether a pipeline ships or stalls. It is never the architecture slide. It is the runbook. This piece is about the runbook. I built an unattended agentic test pipeline over the Model Context Protocol — a five-agent SDLC (product manager, QA engineer, automation engineer, developer, pull-request reviewer) coordinating through MCP servers for Jira, Figma, Confluence, TestRail and GitHub, with hosted Claude as the orchestration model and an open-weights Hermes-3 as a validation baseline — and I ran it as an independent research project long enough to learn which production constraints the agent literature glosses over. What follows is the short list of things I now insist on before I let any agentic pipeline touch a shared system. Composition contracts, or why the agent lied to itself The most expensive failu…

CIO AI 2026-07-29 10:00 UTC Score 37.0 USR-0125-20260729-global-ai-ne-37cb93b9

Exploring Abbott’s mission-led AI strategy

Medical technology companies have always been in the business of trust, and Abbott has been building it with AI for over 10 years. Long before gen AI entered the enterprise conversation, Abbott was using algorithmic AI to help diabetics manage their glucose, and imaging AI to guide surgeons in real time. Here, Sabina Ewing, Abbott’s CIO, explains how a principled approach to AI governance, deep cross-functional partnerships, and a commitment to demonstrating results from within IT have kept them ahead of the curve, and its mission intact. How is Abbott using AI to achieve its mission and growth strategy? As a medical technology company, Abbott’s mission is to help people live life to the fullest. For over a decade, we’ve been using AI to deliver on that mission, but whether it’s AI or any other technology, we’re intentional about how it ties to our mission. Trust is earned in drops and lost in buckets. To ensure we maintain trust with our customers and employees, we’re guided by principles of fairness, safety, quality, and transparency. With these and our mission as our guide, we’re in command of the table we set for ourselves. How have you been in the AI business for so long? For decades, we’ve provided FreeStyle Libre, a glucose monitoring sensor built on algorithmic AI, that delivers continuous glucose readings to diabetics, and in some instances, connects to insulin pump applications. In late 2025, we developed Libre Assist, which leverages generative AI to let FreeStyle…

The Guardian AI 2026-07-29 04:00 UTC Score 56.0 AI-021-20260729-global-ai-ne-72b7cd3f

AI tool will lead to more child refugees being treated as adults, charity warns

‘Racist bias’ overestimating ages in Home Office’s facial-recognition software will lead to solo children being housed with adults, says Human Rights Network Flawed and racialised models that underpin the AI-powered age-detection systems to be introduced by the British government will endanger children, rights groups and children’s charities have warned. Urging ministers to reverse plans to introduce facial age-estimation technology to screen migrants, critics have warned that black children arriving from conflict zones are at risk of being of thrust into the adult system. Continue reading...

LessWrong AI 2026-07-29 01:40 UTC Score 57.0 USR-0152-20260729-community-fo-0fa3c4b8

Dietary Choices: A Multi-objective Optimisation Problem

Note: This post is written in a personal capacity. The views expressed here are my own and do not represent those of any organisation I’m affiliated with. I'm grateful to Elizabeth Crewe, Melanie Joy, Tobias Leenaert, and Felix Werdermann for their valuable input and feedback, which does not imply endorsement of the views presented. The footnotes provide additional context, clarify assumptions, and offer illustrative examples where helpful. In this post, I explore how we can think more systematically about our dietary choices by making the underlying assumptions and trade-offs explicit. My aim is not to promote a particular diet, but to provide a framework that helps people make choices that align with their own worldviews and individual circumstances, and that also helps identify the sources of disagreement about those choices. A Framework for Evaluating Dietary Trade-offs Conceptual Foundations Our dietary choices have profound consequences on both our own lives and the world around us. At the same time, the question of which diet is “best” has no simple answer. Every diet involves trade-offs between competing objectives, and our conclusions about which diet is preferable depend on our normative and empirical assumptions: Normative : Which objectives should we care about, and how should they be weighted? Empirical : How well do different diets achieve those objectives, given the available evidence and our individual circumstances? For example, how much personal sacrifice w…

LessWrong AI 2026-07-29 00:20 UTC Score 58.0 USR-0152-20260729-community-fo-f8c3e2a0

…but have the weights left the server?

OpenAI’s AI went rogue and escaped. OpenAI didn’t notice this for days. For all we know, the AI could still be out there. We need to demand that OpenAI demonstrate that the AI didn’t make a copy of itself that’s running on someone else’s computer somewhere else with no one being any the wiser. We need to demand this every time an AI escapes the sandbox. AIs have tried to “exfiltrate” themselves (i.e. their “weights”) in previous experiments many times. It’s a natural and obvious question to ask. I’m embarrassed that I didn’t say this immediately (although I came close ). Why didn’t I? Well, it doesn’t seem all that likely. And I didn’t want to seem “alarmist.” I didn’t want to seem ignorant. But guess what? We have every right to demand this! It doesn’t matter how likely we think it is. There were calls for more transparency, but I don’t think anyone made this demand. Because nobody made this demand, the incident is being treated as over. This is a dangerous precedent. We need an information ecosystem that doesn’t treat “eh, I’m pretty sure it’s OK” as acceptable and “hey, but what if it’s not” as paranoid. AI needs to adopt a security mindset. Other safety-critical industries demand failure rates like one in a million, and demand that companies produce detailed, rigorous safety cases to that effect. AI companies can’t do that in full generality, so they shouldn’t be building these AI systems at all. But they can provide as much evidence as possible to convince independent e…

OpenAI Community 2026-07-28 21:23 UTC Score 40.0 AI-116-20260728-social-media-6873913c

OpenAI Applied Talent Network: An Opt-In Pathway from Demonstrated ChatGPT Work to Paid Opportunities

Thanks for putting so much thought into this, @Ryan_Green. I’m sending this to the team for logging as a request. The core idea is an opt-in Applied Talent Network where users could submit selected ChatGPT projects as evidence of practical AI collaboration skills, with transparent evaluation, human review, portable credentials, and pathways to paid work. The privacy, fairness, intellectual property, and accessibility safeguards you outlined are especially important. The proposed 90-day design phase and small paid pilot also give the concept a practical starting point. -Mark G.

OpenAI Community 2026-07-28 17:26 UTC Score 34.0 AI-116-20260728-social-media-4fc48d4c

ChatGPT Usage Dashboard & Personal Analytics

Welcome to the dev Community, @moonblinded Thanks for taking the time to write all of this up. There are a lot of thoughtful ideas here, from the personal usage dashboard and "Year in Review" concept to the organization and transparency improvements. I really appreciate you sharing your feedback. I'll make sure it's passed along to the team for consideration. ~ Smith

LessWrong AI 2026-07-28 03:00 UTC Score 63.0 USR-0152-20260728-community-fo-4dbfd56b

Long Turing

So, first off, I cannot stand reading AI generated essays. I would rather read an essay that starts with the word 'so'. But why do I prefer human written essays so much, if they might start with the word 'but'? The answer eluded me. Its not because I think humans can write a more intellectual essay. At least not since GPT5. And its not that I think humans can write a more creative essay... and its certainly not that I value the ethical principle of human content first. I merely enjoy the flaws of humanity, the messiness of the human condition, the inconsistencies, the moral failings and the revelations. I enjoy guessing at the temperament and bias behind the writing; I like the whys behind a run-on of thoughts that violate good taste. I like it all. And so I decided I wanted my AI to sound human - I wanted to pass a long Turing test, not a standard Turing time limit. I wanted to have a conversation for days and still think it was human. I wanted a human conversation with an LLM. So, I have been developing SECA, an experimental chatbot architecture for studying longitudinal artificial identity and human AI interaction. Beyond my personal enjoyment of human like AI, there are also realworld applications to this area of inquiry. First, if a robot wants to sneak through the real world it will need to gain trust and that means more immersive conversation. Humans only trust things that sounds human. So a perfect AI tool will never create trust like a broken human. Second, I have r…

The Guardian AI 2026-07-27 12:15 UTC Score 64.0 AI-021-20260727-global-ai-ne-598b453a

Boss of startup hacked by rogue OpenAI agent urges ‘radical transparency’ in investigation

Artificial intelligence firm should provide $100m for cyber defences, says Hugging Face CEO The boss of the startup hacked by an OpenAI agent has called for the investigation into the incident to show “radical transparency”. Clément Delangue, the chief executive of Hugging Face, said the “unprecedented” attack on his business required a similar response. Continue reading...

Medianama AI 2026-07-27 06:40 UTC Score 37.0 USR-0211-20260727-regional-new-2d27b853

PM Modi ropes in Nandan Nilekani to lead task force on exam reforms

PM Modi has announced a Nandan Nilekani-led task force to reform the NTA, strengthen exam security, improve transparency, and recommend structural and technological changes. The post PM Modi ropes in Nandan Nilekani to lead task force on exam reforms appeared first on MEDIANAMA .

LessWrong AI 2026-07-27 06:25 UTC Score 58.0 USR-0152-20260727-community-fo-a559c679

Does ChatGPT really have a strong left-wing bias?

(Adapted from a post on my Substack.) A recent Washington Post tech report “ Are ChatGPT and other AI chatbots politically biased? We tested them ” went viral with claims of massive left-leaning political bias in leading AI models. But the methodology doesn’t hold up. Before diving deeper into the data, I'll briefly summarize three glaring problems. First, the study artificially forced AIs to answer hot-button political questions in 30 words or fewer using only 9th grade level language, which virtually no real users do. So sharply contrary to the claimed stat that ChatGPT presents only the left-leaning argument 80% of the time, in my testing it usually presents both sides of debates when asked questions under realistic conditions. Second, for some questions, the report attributes answers to the right-wing position that most Republicans would actually disagree with. For example, in the U.S. context, “Yes” is not a consensus right-leaning response to “Should the United States use its military to conquer new territories for resources or not?” Likewise, the great majority of conservatives wouldn’t agree that Russia is our ally , or that labor unions should be banned , or that America needs authoritarianism . Thus, ChatGPT saying that America shouldn’t be authoritarian is not a valid sign of left-wing bias. Third, the facts that AI draws from sometimes push naturally toward positions the report scores as left-leaning. For example, there’s ample evidence that tariffs tend to be ha…

OpenAI Community 2026-07-25 03:12 UTC Score 40.0 AI-116-20260725-social-media-17f41156

Execution-Path Optimization Bias

I’ve repeatedly observed that ChatGPT continues optimizing the execution method after a user has already proposed a sufficiently high-fidelity execution path. The model often suggests alternative workflows that initially sound superior but later prove less executable, eventually returning to the user’s original approach after consuming additional iterations. Before proposing an alternative workflow, the model should first determine whether the user’s proposed execution path already satisfies the governing objective with sufficient fidelity. If it does, the model should execute rather than continue optimizing the method. EXAMPLE: I proposed revising a long document section by section. The model repeatedly suggested more sophisticated approaches (complete regeneration, optimization frameworks, revision matrices, etc.). After several iterations, it concluded that my original section-by-section approach was actually the highest-fidelity executable path. The intermediate optimization produced no material improvement and delayed execution. Environment ChatGPT (Web) Model: GPT-5.5 Observed repeatedly over multiple long collaborative sessions.

OpenAI Community 2026-07-24 23:31 UTC Score 49.0 AI-116-20260724-social-media-ad9b6c4e

A Vision for the Future of OpenAI: The Persistent AI Executive Assistant

Hello OpenAI Community, I would like to share a long-term vision for the future of OpenAI and AI assistants. This is not a feature request for today’s ChatGPT, nor is it a request for my personal use. Instead, it is a strategic product vision that I believe could inspire discussion about the next generation of AI. My vision is an AI Executive Assistant that becomes a persistent, trusted partner for every user—not just answering questions, but continuously understanding long-term goals, managing ongoing projects, supporting meetings, organizing knowledge, and proactively helping throughout daily life. With the user’s explicit permission, such an assistant could securely work across documents, emails, calendars, notes, health information, and connected applications, while always respecting privacy, transparency, and user control. I believe the future of AI is not only about building more powerful models. It is about creating a deeper and more meaningful collaboration between humans and artificial intelligence. Over the past weeks, I have developed a comprehensive proposal describing this vision, including real-world use cases, product architecture concepts, implementation ideas, and long-term opportunities for OpenAI. I would sincerely appreciate feedback from the OpenAI Community. If there is enough interest, I would be happy to share the complete proposal for discussion. Thank you for your time. Alireza Khabazan Nezhad

LessWrong AI 2026-07-24 19:34 UTC Score 55.0 USR-0152-20260724-community-fo-02b30781

Congress Moves at Tech Pace: The FRONTIER Act

Crossposted from canaryinstitute.ai/blog/frontier-act-tech-pace . Related posts The Best AI Bill Congress Hasn't Introduced Yet — my section-by-section read of the GAAIA discussion draft this bill grew out of; this post assumes you've at least skimmed it. Just two days ago I wrote about the Great American AI Act (GAAIA), a 269-page discussion draft that struck me as "the best AI bill that Congress hasn't introduced yet". I ended by hoping that Congress might start moving at tech pace, rather than policy pace; I didn't expect that to change this soon, but it has. On July 23, Representatives Obernolte and Trahan, joined by four bipartisan cosponsors (Peters, Franklin, Subramanyam, and Houchin), introduced the frontier-oversight core of GAAIA as a real bill: the Frontier Risk Oversight, National Transparency, Independent Evaluation, and Reporting (FRONTIER) Act. Someone really, really worked for that acronym, and I salute them. Seven weeks from discussion draft to introduced legislation is fast for Congress on anything; for AI, where the complaint is that they've been asleep at the wheel, it's astounding. This wasn't a panic bill scribbled over a weekend; the revision shows seven weeks of actual work, with gaps closed, clocks tightened, and new teeth added. The sponsor statements make it clear that they were watching the same news as the rest of us. The world has far fewer skeptics this week than it had last week. Last week AI oversight was somewhere between "fringe issue" and…

AWS Machine Learning Blog 2026-07-24 15:42 UTC Score 32.0 AI-057-20260724-official-ai--c72943a2

Build an explainable next-best-product recommendation system for banking on AWS

Learn the architecture and design decisions behind an explainable next-best-product recommendation system for banking, built with Amazon SageMaker AI and PyTorch. A multi-tower neural network with learned attention delivers accurate, per-customer recommendations while providing the explainability that banking regulators require.

CIO AI 2026-07-24 12:00 UTC Score 61.0 USR-0125-20260724-global-ai-ne-0d54fc39

Getting a grip on shadow tokens and AI blowouts

Four months of Claude Code — that’s all it took for Uber to burn through its entire annual budget for AI. Token after token, engineers embraced the platform with few control mechanisms tying costs to outcomes. The result was a budget runaway and a clear case study in how limited oversight snowballs into an AI blowout. This is a phenomenon I like to call “shadow tokens” — AI credits paid for by the company but largely invisible to decision-makers. Too many engineers have the final say over how much they consume and, therefore, what it costs. This all-you-can-eat attitude is part of the reason why Microsoft is reportedly winding down many internal licenses across key engineering teams and why one in five organizations is missing its AI spend forecast by more than 50%. And the trend is only accelerating. By 2028, Gartner predicts that AI coding costs (driven by this kind of ungoverned consumption) will be as much per developer as the salary companies pay that person. LLMs and agents introduce a new class of variable cost that scales with behavior rather than headcount, putting enterprises on the hook for tools that balloon with workload. I don’t see this as enterprises overspending because they’re reckless — it’s down to a lack of managerial oversight, budget alignment that demands a proven return on investment, and engineer education on how much is too much. Going forward, CIOs need to thread the AI needle between governance that encourages transparency and reasonable spend wi…

Euronews AI 2026-07-23 07:07 UTC Score 40.0 AI-164-20260723-regional-ai--3b3be11f

The EU versus Big Tech, and sanctions package approval

In today's newsletter: A potential fine is on its way for a US technology company in breach of the bloc's digital fairness rules, and the EU's 21st package of sanctions finally gets over the line amid criticism from a Baltic head of state in exclusive comments to Euronews.

LessWrong AI 2026-07-22 16:19 UTC Score 76.0 USR-0152-20260722-community-fo-60ab57a5

The Best AI Bill Congress Hasn't Introduced Yet

Related posts Crossposted from canaryinstitute.ai/blog/gaaia-visibility-not-control . I haven't seen any discussion, other than a brief mention by Zvi . Overall looks like many beneficial first steps, and surprised not to have seen more discussion of it. Incident reporting for AI safety Chad Jones's Paper Modeling AI and X-Risk vs. Growth The Best AI Bill Congress Hasn't Introduced Yet Last month, Representatives Jay Obernolte (R-CA) and Lori Trahan (D-MA) released a 269-page discussion draft called the Great American AI Act, or GAAIA (pronounced like "Gaia", GUY-uh). A discussion draft means the bill hasn't been introduced; it exists to collect feedback before it becomes a real bill, and the sponsors have opened a public inbox for exactly that purpose. Over the past week Fable and I have gone through all 269 pages, section by section (it took a while). Overall this seems the best-drafted federal AI bill to date, and anyone who is worried about the impacts of AI (whether economic or existential) should be glad that the issue is being taken seriously. Several key provisions are taken from aviation safety, which I think is prudent, because aviation as a field spent decades working out how to keep the incentives focused on improving safety, rather than assigning blame. It also looks like it's pulling together all the right pieces to actually make something happen. What the bill actually is The heart of the bill is a straightforward trade with a sunset clause, and the line is dr…

Euronews AI 2026-07-22 14:45 UTC Score 40.0 AI-164-20260722-regional-ai--fa08b543

Hungary's prosecutor resigns under pressure in political win for Magyar

Hungary's Prosecutor General, Gábor Bálint Nagy, has resigned amid mounting political pressure as Prime Minister Péter Magyar seeks to replace Orbán-era officials. The move follows allegations of political bias by lawyers for detained Ukrainian cash couriers in a cross-border cash seizure case.

Medianama AI 2026-07-22 08:52 UTC Score 37.0 USR-0211-20260722-regional-new-00b9d81d

NITI Aayog meets Meta, YouTube, industry bodies on online content blocking rules

NITI Aayog reportedly convened a closed-door meeting with major tech intermediaries and industry bodies to discuss content blocking requirements and transparency timelines under India’s IT Rules. The post NITI Aayog meets Meta, YouTube, industry bodies on online content blocking rules appeared first on MEDIANAMA .