Latest AI/ML News
35334 matching items
DeepSeek V4 Pro 0813 Quietly Ships a 5x Agentic Coding Leap
DeepSeek V4 Pro 0813 goes GA with massive agentic coding gains, open weights under MIT, and a 264% price hike taking effect August 16
AI Themed Movies, Films, Books, Music, and Video Games
PaulBellow: i’m running out of english content! lol Even animation? Maybe worth a look.
앤트로픽, IPO 앞두고 첫 흑자 달성…2분기 매출 16.3조 돌파
앤트로픽이 주요 프론티어 AI 기업 중 처음으로 분기 조정 영업이익 흑자를 달성했다. 이는 당초 회사 내부에서 예상했던 흑자 전환 시점보다 2년 빠른 기록이다.14일(현지시간) 블룸버그가 입수한 문서에 따르면, 앤트로픽은 올해 2분기 115억달러(약 16조3128억원) 이상의 잠정 매출액을 기록하며 전년 동기 대비 14배 이상 성장한 실적을 거뒀다.이는 1분기 매출(47억3000만달러)과 비교해도 2배가 넘는 가파른 상승세다. 수치는 최종 수정될 수 있으며, 앤트로픽 대변인은 논평을 거부했다.아울러 앤트로픽은 이번 실적 공개를 통해
Comment on MIT Researchers Unveil “SEAL”: A New Step Towards Self-Improving AI by krilliongame
The concept of models updating their own weights through reinforcement learning, as demonstrated by MIT’s SEAL, is a fascinating milestone—especially since it ties the self-editing process directly to downstream task performance. It’s exciting to see concrete research backing the self-evolving AI conversation that’s been gaining so much attention lately. For more insights like this, I always enjoy checking out krilliongame .
How China’s young managers grappled with billion-yuan mandates as AI shocks hit portfolios
As Leopold Aschenbrenner’s US hedge fund saw assets wiped off by more than two-thirds in a single month, some of China’s new portfolio managers also felt the shock across the Pacific, learning bitter lessons early in their careers. The 50-day market turmoil, sparked by a global correction in artificial intelligence stocks in June, turned some of China’s rookie managers into an unwitting focal point. Even seasoned investors faced hard questions from clients as portfolios sagged. For Yuan Zeqiang,...
[AI카툰] Bob, Senior IT guy - 밥, 거절을 모르는 남자
글·그림 Dr. Alf
[내면 지도 그리기] 29 Inner Mapping: 나를 찾는 지도
실험노트_입추가을의 증거를 찾아보는 일입추가 지났습니다. 달력에는 분명 가을이 시작됐다고 적혀 있는데요. 아직 한낮에는 여름옷이 편하고, 걸음을 조금만 빨리해도 땀이 납니다. 그래서 절기와 실제 날씨가 맞지 않는다고 생각하기도 하는데, 가만히 보면 우리가 계절을 알아채는 방법은 기온만으로 설명되지는 않는 것 같습니다.저녁에 불을 켜는 시간이 조금 빨라졌다든지, 오후에 책상 위로 들어오는 햇빛의 자리가 옮겨갔다든지, 같은 시간에 걷는데 건물의 그림자가 전보다 길어진 것을 발견할 때가 있고요. 여름 내내 거들떠보지 않던 얇은 셔츠를 아
목포어울림도서관, ‘매직 벌룬쇼’…19일부터 무료 접수
AI 생성 영상목포어울림도서관이 29일 어린이와 가족을 위한 특별공연 ‘다이나믹 매직 벌룬쇼’를 연다고 14일 밝혔다.음악과 마술, 풍선을 결합한 참여형 프로그램으로 4세 이상 어린이와 지역 주민 누구나 무료로 참여할 수 있다. 공연은 신나는 음악에 맞춰 마술과 풍선 퍼포먼스를 함께 선보이는 방식으로 진행된다. 어린이들이 공연을 가까이에서 보고 직접 체험할 수 있도록 구성해 가족 단위 관람객이 함께 즐길 수 있도록 했다.신청은 19일부터 26일까지 목포시통합도서관 누리집에서 받는다. 공연은 29일 목포어울림도서관에서 진행된다.목포
여수, 2027 대통령배 e스포츠 개최지 선정
AI 생성 영상여수시가 2027년 대통령배 전국 아마추어 e스포츠 대회와 전국 장애인 e스포츠 대회 개최지로 최종 선정됐다고 14일 밝혔다.두 대회는 2027년 8월 여수진남실내체육관과 여수장애인국민체육센터에서 열리며, 전국 16개 시·도 선수단과 관계자, 관람객 등 4500여명이 여수를 찾을 것으로 예상된다.문화체육관광부는 지난 7월 전국 지자체를 대상으로 개최지를 공모한 뒤 서면심사와 현장심사, 온라인 질의응답 등을 거쳐 여수를 최종 개최지로 선정했다. 여수는 숙박·교통 인프라와 대규모 체육행사 개최 역량, e스포츠 활성화 정책
이탈리아, 최대 규모 225MW 영농형 태양광 건설…3830억 조달
이탈리아 시칠리아에서 건설 중인 225.5메가와트(MW) 규모 영농형 태양광발전소가 약 2억3410만유로(약 3828억원)의 프로젝트 금융을 확보했다. 이는 이탈리아 최대 규모로 추진되는 영농형 태양광 사업이다.덴마크 재생에너지 개발사 유러피언에너지(European Energy)는 12일(현지시간) 시칠리아 비치니(Vizzini) 프로젝트의 금융조달을 완료했다고 밝혔다. 크레디아그리콜 CIB, 인테사산파올로 IMI CIB, NORD/LB, 소시에테제네랄 등이 참여했으며 조달 자금은 발전소 건설에 투입된다.설비용량은 225.5MW로
Anthropic says text watermarking scheme relies on inconsequential words
'Shall I compare thee to a summer's afternoon' is the sort of thing this will make, and others look likely to adopt it
Schedule a new hackathon in Los Angeles area
I see that. I will certainly be around…
순천, 에코칼리지 운영…생태문명·순환경제 교육
AI 생성 영상순천시가 생태문명전환에 대한 이해와 실천 역량을 키우는 순천에코칼리지 촉진자 양성과정 8월 온라인 과정을 운영했다고 14일 밝혔다.이번 과정은 생태와 경제를 별개로 보지 않고, 지역과 일상에서 순환의 원리를 적용하는 교육에 초점을 맞췄다.지난 13일 진행된 강의에는 정건화 한신대 명예교수이자 한신대 생태문명원 공동대표가 강사로 참여했다. 정 교수는 생태적 지역순환경제를 주제로 생태경제와 순환경제에 대해 강연했다.강의 뒤에는 참여자 토론과 질의응답이 이어졌다. 참석자들은 생태와 순환의 원리를 지역경제와 일상에서 어떻게
vLLM's DSpark Beats Every Fixed-Length Config on DeepSeek-V4 at Any Load
vLLM's new adaptive verification for DSpark automatically tunes speculative decoding depth per step, holding the Pareto frontier from concurrency 1 to 256 with a single config.
Comment on Which Agent Causes Task Failures and When?Researchers from PSU and Duke explores automated failure attribution of LLM Multi-Agent Systems by Alex Chen
Loved this article! For anyone into video creation, https://vidglory.com is a fantastic AI video tool worth checking out. Keep up the great work!
아마존 오지에 태양광 설치했더니…농산물 가공품 유럽 수출
남미 오지 마을에서 태양광과 배터리 저장장치가 단순한 전력 공급을 넘어 지역 소득원으로 활용되고 있다. 페루 아마존의 한 원주민 공동체는 태양광으로 디젤발전을 대체한 뒤 농산물 가공시설을 안정적으로 운영하면서 생산 확대와 판로 확보에 나섰다.국제재생에너지기구(IRENA)는 13일(현지시간) 라틴아메리카 원격지역의 분산형 재생에너지 활용 사례를 소개했다.라틴아메리카의 전력 보급률은 약 98%에 이르지만 아마존과 안데스 고지대 등 전력망이 닿지 않는 지역에는 여전히 에너지 접근 격차가 남아 있다. IRENA는 태양광·배터리와 소수력 등
Feature Request: “Ask about this” - inline subthreads for selected text
I’d love to see a lightweight way to ask follow-up questions about a specific part of a ChatGPT response without branching or cluttering the main conversation. Proposed flow Highlight a sentence or paragraph → tap “Ask about this” → open a small contextual subthread attached to that passage → ask one or several clarification questions → close it → continue reading the original conversation. This is different from Branch in new chat . Branching creates a separate full conversation from an entire message. What I’m proposing is a lightweight, temporary discussion attached specifically to selected text inside a response . Example: Select text → Ask about this → contextual subthread → clarification → return to main chat Why it would be useful In long conversations, especially when studying, coding, researching, or learning languages, I often want to clarify one small point without interrupting the main line of discussion. At the moment, these small questions either clutter the main chat or require creating another conversation. An inline contextual thread would preserve the structure of the main conversation while still allowing users to explore individual ideas in depth. The exact UI could be a bottom sheet, pop-up, nested thread, or another implementation. The important part is that the subthread remains attached to the selected passage rather than becoming another full chat.
Subagents active despite no task running
Hello, i have an issue where i’ve had 27 subagents working on old tasks which have no active task in any thread… These subagents has finally started to drain my usage without creating anything but instead just being stuck in a limbo for days. In the image 7 has been shut down after asking an agent to do so, but i still have 20 active subagents draining my usage. Hope to get help with this issue and to also regain my lost usage. no-active-thread-01a002b4-7316-7e00-9a81-26f722cd2997 Here is an error report i did.
Accelerating Constrained Sampling: A Large Deviations Approach
The problem of sampling a target probability distribution on a constrained domain arises in many applications including machine learning. For constrained sampling, various Langevin algorithms such as projected Langevin Monte Carlo (PLMC), based on the discretization of reflected Langevin dynamics (RLD) and more generally skew-reflected non-reversible Langevin Monte Carlo (SRNLMC), based on the discretization of skew-reflected non-reversible Langevin dynamics (SRNLD), have been proposed and studied in the literature. This work focuses on the long-time behavior of SRNLD, where a skew-symmetric matrix is added to RLD. Although acceleration for SRNLD has been studied, it is not clear how one should design the skew-symmetric matrix in the dynamics to achieve good performance in practice. We establish a large deviation principle (LDP) for the empirical measure of SRNLD when the skew-symmetric matrix is chosen such that its product with the outward unit normal vector field on the boundary is zero. By explicitly characterizing the rate functions, we show that this choice of the skew-symmetric matrix accelerates the convergence to the target distribution compared to RLD and reduces the asymptotic variance. Numerical experiments for SRNLMC based on the proposed skew-symmetric matrix show superior performance, which validate the theoretical findings from the large deviations theory.
Statistical Test for Attention in Transformers for Images and Time Series
Transformer models have achieved exceptional performance in various domains, including computer vision and time-series analysis. Their core attention mechanism is widely used to interpret model decisions by assigning importance weights to input regions, such as image patches or time series intervals. However, the reliability of these interpretations remains a major concern. High-attention weights do not necessarily indicate genuinely significant features; they may instead be artifacts of the model's computation, undermining their reliabilities in high-stakes applications such as medical diagnostics. To address this, we propose a novel statistical framework designed to quantify the significance of high-attention regions in Transformer models. Our framework is built on selective inference (SI) to correct for the inherent selection bias that arises from testing regions chosen through the complex attention computation of the Transformer models. A key contribution of this work is a novel computational method that extends SI to the complex non-linearity of self-attention, enabling the computation of valid $p$-values for high-attention regions. These $p$-values serve as a reliable measure of significance, strengthening the interpretability of Transformer decisions. The validity and effectiveness of our approach are demonstrated through numerical experiments and applications to brain image diagnosis and electroencephalography (EEG) data analysis.
py/cuTAGI: An Open-Source Library for Tractable Approximate Gaussian Inference in Bayesian Neural Networks
This paper introduces pyTAGI, a Python wrapper, and cuTAGI, its high-performance C++/CUDA backend, implementing Tractable Approximate Gaussian Inference (TAGI) for neural networks. TAGI treats all network quantities as Gaussian random variables and derives closed-form expressions for prior/posterior expected values, variances, and covariances, enabling analytic Bayesian learning without relying on gradient descent or backpropagation. The libraries mimic PyTorch's sequential interface, allowing users to define models by stacking layers in order and performing uncertainty-aware Bayesian inference. Beyond epistemic uncertainty, it also allows quantifying heteroscedastic aleatoric uncertainty. cuTAGI's custom CPU/GPU kernels and distributed-data-parallel support via NCCL/MPI deliver competitive runtimes, while pyTAGI's pip-installable frontend and MIT-licensed GitHub repo facilitate community adoption and extension. Version 0.2.1 already supports a comprehensive suite of layers and activations; future work will add eager execution, further kernel optimizations, attention mechanisms, and advanced covariance factorization. Together, py/cuTAGI offer an efficient, open-source foundation for the analytic treatment of Bayesian deep learning.
Gradient Span Algorithms Make Predictable Progress in High Dimension
We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. This is a functional generalization of similar results for random quadratic functions and spin glasses. They explain the counterintuitive phenomenon that different training runs of many large machine learning models result in approximately equal cost curves despite random initialization on a complicated non-convex landscape. This 'predictable progress' phenomenon is exploited by the AutoML community: Since the optimization progress of a single run is already representative, multiple retries with the same hyperparameters are not necessary.
Robust training of implicit generative models for multivariate and heavy-tailed distributions with an invariant statistical loss
Implicit generative models are often trained adversarially, which can yield unstable dynamics and mode collapse. The invariant statistical loss (ISL) offers a fully sample-based alternative by comparing empirical ranks of real and generated samples. In this work, we formally characterize ISL as a proper divergence over continuous distributions and establish key regularity properties, showing that it is continuous and differentiable, thereby enabling stable gradient-based optimization without adversarial games. We further enhance ISL along two practical axes. First, to better model heavy-tailed data, where Gaussian latent priors can limit tail expressivity, we introduce Pareto-ISL, which replaces Gaussian noise with a generalized Pareto latent distribution to improve the representation of both typical and extreme events. Second, to handle multivariate data at scale, we propose ISL-slicing: a computationally efficient procedure that projects samples onto random one-dimensional subspaces, computes rank-based losses per projection, and averages them to capture high-dimensional structure. Experiments demonstrate improved tail fidelity with Pareto-ISL and show that ISL-slicing scales effectively to high dimensions. Specifically, in high dimensional settings we show that ISL can be used either as a standalone criterion or as a strong pretraining objective for subsequent adversarial fine-tuning.
Adaptive Nonparametric Perturbations of Parametric Models with Generalized Bayes
Parametric Bayesian modeling offers a powerful and flexible toolbox for machine learning. Yet the model, however detailed, may still be wrong, and this can make inferences untrustworthy. In this paper we introduce a new class of semiparametric corrections for parametric Bayesian models, when the target of inference is a functional of the true data distribution. Our starting point is a fully Bayesian modeling approach, which explicitly accounts for the possibility that the parametric model is wrong. Asymptotic analysis shows that this approach is both robust to model misspecification and data efficient, achieving fast convergence when the parametric model is close to true. However, the fully Bayesian approach is limited in its practical usefulness by the challenges of conducting inference and computing a Bayes factor for a nonparametric model. We therefore propose a novel model correction based on generalized Bayes, which entirely avoids the need to compute a nonparametric Bayes factor, but preserves the robustness and efficiency of the fully Bayesian approach. We demonstrate our method by estimating causal effects of gene expression from single cell RNA sequencing data. Overall, we offer a new efficient approach to robust Bayesian inference with parametric models.
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
We study gradient descent (GD) for logistic regression on linearly separable data with stepsizes that adapt to the current risk, scaled by a constant hyperparameter \(\eta\). We show that after at most \(1/\gamma^2\) burn-in steps, GD achieves a risk upper bounded by \(\exp(-\Theta(\eta))\), where \(\gamma\) is the margin of the dataset. As \(\eta\) can be arbitrarily large, GD attains an arbitrarily small risk immediately after the burn-in steps, though the risk evolution may be non-monotonic. We further construct hard datasets with margin \(\gamma\), where any batch (or online) first-order method requires \(\Omega(1/\gamma^2)\) steps to find a linear separator. Thus, GD with large, adaptive stepsizes matches the worst-case $1/\gamma^2$ dependence when the sample size is unrestricted. Notably, the classical Perceptron, a first-order online method, also achieves a step complexity of \(1/\gamma^2\), matching GD even in constants. Finally, our GD analysis extends to a broad class of loss functions and certain two-layer networks.
Approximation-Free Differentiable Oblique Decision Trees
Decision Trees (DTs) are widely used in safety-critical domains such as medical diagnosis, valued for their interpretability and effectiveness on tabular data. However, training accurate oblique DTs is challenging due to complex optimization landscapes and overfitting risks, particularly in regression. Recent advances have introduced differentiable formulations that enable gradient-based training and joint optimization of decision boundaries and leaf regressors. Yet, existing approaches typically rely on approximations, either through probabilistic softening of boundaries (soft DTs) or quantized gradients such as the Straight-Through Estimator (STE). To overcome these limitations, we propose DTSemNet, a novel, semantically equivalent, and invertible representation of hard oblique DTs as neural networks. DTSemNet enables end-to-end training with standard gradient descent, eliminating the need for approximations in both classification and regression. While classification aligns naturally with this formulation, regression remains challenging due to the joint optimization of internal nodes and leaf regressors. To address this, we analyze the limitations of STE and introduce an annealed Top-$k$ method that provides accurate gradient signals without approximation. Extensive experiments on classification and regression benchmarks show that DTSemNet-trained oblique DTs outperform state-of-the-art differentiable DTs. Furthermore, we demonstrate that DTSemNet can serve as programmatic…
Underdamped Langevin MCMC with third order convergence
In this paper, we propose a new numerical method for the underdamped Langevin diffusion (ULD) and present a non-asymptotic analysis of its sampling error in the 2-Wasserstein distance when the $d$-dimensional target distribution $p(x)\propto e^{-f(x)}$ is strongly log-concave and has varying degrees of smoothness. Precisely, under the assumptions that the gradient and Hessian of $f$ are Lipschitz continuous, our algorithm achieves a 2-Wasserstein error of $\varepsilon$ in $\mathcal{O}\big(\sqrt{d}/\varepsilon\big)$ and $\mathcal{O}\big(\sqrt{d}/\sqrt{\varepsilon}\big)$ steps respectively. Therefore, our algorithm has a similar complexity as other popular Langevin MCMC algorithms under matching assumptions. However, if we additionally assume that the third derivative of $f$ is Lipschitz continuous, then our algorithm achieves a 2-Wasserstein error of $\varepsilon$ in $\mathcal{O}\big(\sqrt{d}/\varepsilon^{\frac{1}{3}}\big)$ steps. To the best of our knowledge, this is the first gradient-only method for ULD with third order convergence. To support our theory, we perform Bayesian logistic regression across a range of real-world datasets, where our algorithm achieves competitive performance compared to an existing underdamped Langevin MCMC algorithm and the popular No U-Turn Sampler (NUTS).
Mixing times of data-augmentation Gibbs samplers for high-dimensional probit regression
We investigate the convergence properties of popular data-augmentation samplers for Baye\-sian probit regression. Leveraging recent results on Gibbs samplers for log-concave targets, we provide simple and explicit non-asymptotic bounds on the associated mixing times (in Kullback-Leibler divergence). The bounds depend explicitly on the design matrix and the prior precision, while they hold uniformly over the vector of responses. We specialize the results for different regimes of statistical interest, when both the number of data points $n$ and parameters $p$ are large: in particular we identify scenarios where the mixing times remain bounded as $n,p\to\infty$, and ones where they do not. The results are shown to be tight (in the worst case with respect to the responses) and provide guidance on choices of prior distributions that provably lead to fast mixing. An empirical analysis based on coupling techniques suggests that the bounds are effective in predicting practically observed behaviours.
Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy
The impact of inference-time data perturbation (e.g., adversarial attacks) has been extensively studied in machine learning, leading to well-established certification techniques for adversarial robustness. In contrast, certifying models against training data perturbations remains a relatively under-explored area. These perturbations can arise in three critical contexts: adversarial data poisoning, where an adversary manipulates training samples to corrupt model performance; machine unlearning, which requires certifying model behavior under the removal of specific training data; and differential privacy, where guarantees must be given with respect to substituting individual data points. This work introduces Abstract Gradient Training (AGT), a unified framework for certifying robustness of a given model and training procedure to training data perturbations, including bounded perturbations, the removal of data points, and the addition of new samples. By bounding the reachable set of parameters, i.e., establishing provable parameter-space bounds, AGT provides a formal approach to analyzing the behavior of models trained via first-order optimization methods.
Doubly Debiased Robust Subsampling for Transfer Learning
This paper develops a general framework for doubly debiased robust subsampling for transfer learning. The setting arises when massive source datasets are computationally infeasible to use in full, while naive or heuristic subsampling leads to biased estimators that further inherit transfer bias under source-target distributional shifts. We resolve these challenges through two complementary debiasing mechanisms. Inverse probability weighting removes subsampling bias by ensuring that subsample-based estimators represent the full source distribution, while a target-based one-step refinement recenters estimators towards the target distribution, thereby mitigating transfer bias. These corrections are embedded within a distributionally robust optimization design that simultaneously controls worst-case target risk and enforces source-target alignment through maximum mean discrepancy. To optimize subsampling distributions, we propose a scalarized particle swarm algorithm that efficiently explores the robustness-alignment frontier by adjusting a single tuning parameter. We establish theoretical properties, including asymptotic normality, generalization bounds, oracle inequalities, and minimax optimality under distributional uncertainty. Simulation studies and empirical applications in text sentiment and image recognition demonstrate that the proposed method consistently improves prediction accuracy and robustness compared with uniform subsampling, target-only training, and alignment-…
Learning to Play Two-Player Perfect-Information Games without Knowledge
This paper introduces a set of techniques for learning game state evaluation functions through reinforcement learning. First, we generalize tree bootstrapping, i.e. learning the values of states encountered during search rather than restricting updates to states observed during matches, to the setting of reinforcement learning with non-linear function approximation. Second, we modifies Unbounded Best-First Minimax by extending best action sequences to terminal states. Third, we replace the traditional binary game outcome $+1/-1$ with richer reinforcement signals, including quick wins, delayed losses, and scoring. Fourth, we propose a completion mechanism that exploits state resolution. Finally, we introduce a novel action-selection distribution, referred to as the ordinal distribution. Experimental results show that each of these techniques contributes to substantial improvements in playing strength. We integrate them into a unified algorithm, Athénan, and compare it against ExIt, a leading self-play reinforcement learning approach without prior knowledge. Our results demonstrate that Athénan consistently outperforms ExIt. We further evaluate Athénan on the games Hex, Othello, and Arimaa, where it surpasses state-of-the-art performance without relying on domain-specific knowledge. In addition, we consider the single-player game Morpion Solitaire, in which Athénan again reaches state-of-the-art results under the same constraint. Overall, these results show that reinforcement…
Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective
The well-known graph-based clustering methods, including spectral clustering, symmetric non-negative matrix factorization, and doubly stochastic normalization, can be viewed as relaxations of the kernel k-means approach. However, we posit that these methods excessively relax their inherent low-rank, nonnegative, doubly stochastic, and orthonormal constraints to ensure numerical feasibility, potentially limiting their clustering efficacy. In this paper, guided by our systematic theoretical analyses, we propose Low-Rank Doubly stochastic clustering (LoRD), a model that only relaxes the orthonormal constraint to derive a probabilistic clustering results. Furthermore, by theoretically establishing the equivalence between orthogonality and Block diagonality under the doubly stochastic constraint, we propose B-LoRD. By integrating block diagonal regularization into LoRD, expressed as the maximization of the Frobenius norm, we enhance clustering performance. To ensure numerical solvability, we transform the non-convex doubly stochastic constraint into a linear convex constraint through the introduction of a class probability parameter. The theoretical demonstration of the gradient Lipschitz continuity of our LoRD and B-LoRD enables the proposal of a projected gradient algorithm whose exact iteration admits a sublinear convergence-rate bound and ensures first-order stationarity of every accumulation point for the exact projected gradient iteration. Extensive experiments underscore t…
End-to-End Deep Learning for Predicting Metric Space-Valued Outputs
Many modern applications involve predicting structured, non-Euclidean outputs such as probability distributions, networks, and symmetric positive-definite matrices. These outputs are naturally modeled as elements of general metric spaces, where classical regression techniques that rely on vector space structure no longer apply. We introduce E2M (End-to-End Metric regression), a deep learning framework for predicting metric space-valued outputs. E2M performs prediction via weighted Fréchet means over training outputs, where the weights are learned by a neural network conditioned on the input. This construction provides a principled mechanism for geometry-aware prediction that avoids surrogate embeddings and restrictive parametric assumptions, while fully preserving the intrinsic geometry of the output space. We establish theoretical guarantees, including a universal approximation theorem that characterizes the expressive capacity of the model and a convergence analysis of the entropy-regularized training objective. Through extensive simulations involving probability distributions, networks, and symmetric positive-definite matrices, we show that E2M consistently achieves state-of-the-art performance, with its advantages becoming more pronounced at larger sample sizes. Applications to human mortality distributions and New York City taxi networks further demonstrate the flexibility and practical utility of this framework.
The Sample Complexity of Parameter-Free Stochastic Convex Optimization
We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz constant are unknown. We pursue two strategies. First, we develop a reliable model selection method that avoids overfitting to the validation set. This method allows us to generically tune the learning rate of stochastic optimization methods to match the optimal known-parameter sample complexity up to $\log\log$ factors. Second, we develop a regularization-based method that is specialized to the case that only the distance to optimality is unknown. More specifically, it uses norm-regularized empirical risk minimization to estimate the distance to optimality to within a constant factor, allowing known-parameter stochastic optimization methods to achieve optimal sample complexity. This method provides perfect adaptability to unknown distance to optimality, demonstrating a separation between the sample and computational complexity of parameter-free stochastic convex optimization. Combining these two methods allows us to simultaneously adapt to multiple problem structures. Experiments performing few-shot learning on CIFAR-10 by fine-tuning CLIP models and prompt engineering Gemini to count shapes indicate that our reliable model selection method can help mitigate overfitting to small validation sets.
Near-optimal Delta-convex Estimation of Lipschitz Functions
This paper presents a tractable algorithm for estimating an unknown Lipschitz function from noisy observations and establishes an upper bound on its convergence rate. The approach extends max-affine methods from convex shape-restricted regression to the more general Lipschitz setting. A key component is a nonlinear feature expansion that maps max-affine functions into a subclass of delta-convex functions, which act as universal approximators of Lipschitz functions while preserving their Lipschitz constants. Leveraging this property, the estimator attains the minimax convergence rate (up to logarithmic factors) with respect to the intrinsic dimension of the data under squared loss and subgaussian distributions in the random design setting. The algorithm integrates adaptive partitioning to capture intrinsic dimension, a penalty-based regularization mechanism that removes the need to know the true Lipschitz constant, and a two-stage optimization procedure combining a convex initialization with local refinement. The framework is also straightforward to adapt to convex shape-restricted regression. Experiments demonstrate competitive performance relative to other theoretically justified methods, including nearest-neighbor and kernel-based regressors.
Error Analyses of Auto-Regressive Video Diffusion Models
Auto-Regressive Video Diffusion Models (AR-VDMs) have shown strong capabilities in generating long, photorealistic videos, but suffer from two key limitations: (i) history forgetting, where the model loses track of previously generated content, and (ii) temporal degradation, where frame quality deteriorates over time. Yet a rigorous theoretical analysis of these phenomena is lacking, and existing empirical understanding remains insufficiently grounded. In this paper, we introduce Meta-ARVDM, a unified analytical framework that studies both errors through the shared autoregressive structure of AR-VDMs. We show that history forgetting is characterized by the conditional mutual information between the generated output and preceding frames, conditioned on inputs, and prove that incorporating more past frames monotonically alleviates history forgetting, thereby theoretically justifying a common belief in existing works. Moreover, our theory reveals that standard metrics fail to capture this effect, motivating a new evaluation protocol based on a “needle-in-a-haystack” task in closed-ended environments (DMLab and Minecraft). We further show that temporal degradation can be quantified by the cumulative sum of per-step errors, enabling prediction of degradation for different schedulers without video rollout. Finally, our evaluation uncovers a strong empirical correlation between history forgetting and temporal degradation, a connection not previously reported.
High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks
We study the high-dimensional training dynamics of a shallow neural network with quadratic activation in a teacher--student setup. We focus on the extensive-width regime, where the teacher and student network widths scale proportionally with the input dimension, and the sample size grows quadratically. This scaling aims to describe overparameterized neural networks in which feature learning still plays a central role. In the high-dimensional limit, we derive a dynamical characterization of the gradient flow, in the spirit of dynamical mean-field theory (DMFT). Under $\ell_2$-regularization, we analyze these equations at long times and characterize the performance and spectral properties of the resulting estimator. This result provides a quantitative understanding of the effect of overparameterization on learning and generalization, and reveals a double descent phenomenon in the presence of label noise, where generalization improves beyond interpolation. In the small regularization limit, we obtain an exact expression for the perfect recovery threshold as a function of the network widths, providing a precise characterization of how overparameterization influences recovery.
Bridging Domain Invariance and Diversity: A Fine-Grained Risk Bound for Domain Generalization
Domain-invariant representation learning and domain augmentation algorithms are two principal methodological paradigms for addressing domain generalization. They are widely employed in the machine learning literature to enhance domain invariance and domain diversity, respectively. However, existing risk bounds for domain generalization do not simultaneously capture the contributions of both approaches. This limitation arises because bounds derived directly in the original latent space are typically too coarse-grained and ambiguous to characterize how invariance and diversity jointly influence generalization. Since these two properties are often regarded as being inherently contradictory, it becomes difficult to disentangle and rigorously characterize their individual effects. To address this issue, we first observe that the latent representation space can be decomposed into several distinct subspaces, each exhibiting different characteristics and therefore being better suited for analyzing the respective roles of domain invariance and domain diversity. Building on this observation, we propose a unified analytical framework for domain generalization. Specifically, we introduce a Tri-Space Latent Representation and establish its unique decomposability via a direct-sum decomposition. Under this decomposition, each data representation can be uniquely partitioned into three components: domain-invariant features, spurious invariant features, and domain-variant features. Within thi…
AGENTCODI: The Codex Workflow on Android
Hello everyone, I wanted to use Codex properly on Android without first setting up Termux and then maintaining a separate Node/npm environment. What started as the fairly harmless thought, “This should be possible,” has since turned into AGENTCODI and a surprising amount of work. The currently released version is 0.4.5 Early Access and runs on ARM64-v8a devices with Android 10 or newer. AGENTCODI is a standalone Android app for the Codex workflow. It supports authentication through ChatGPT or an OpenAI API key, new and existing Codex threads, live response streaming, and the display of reasoning summaries, plans, commands, file changes, and tool activity. Approval requests can be accepted or rejected directly inside the app. Available models and reasoning levels are loaded dynamically. There is also a dedicated project workspace from which individual files, generated images, or the entire workspace can be exported as a ZIP archive. AGENTCODI’s application and test code is written in Java and C++. The interface currently supports English and German and automatically follows the device language. A surprisingly large part of the development process now consists of finding out how far this kind of workflow can be pushed within Android’s application model. Android occasionally answers that question in its own uniquely charming way. A good example of this is described in my Developer Diary . I spent several weeks putting quite a bit of work into keeping Node.js out of future AGENT…
METR Raises $71M to Independently Stress-Test the World's Most Powerful AI
METR raises $71M in six months to scale independent AI safety evaluations as rogue deployment risks grow more real
Develop a chatbot on WhatsApp
Don’t let OpenAI go-to-market salespeople hear about this 5M+ annual billings opportunity. See that the company wants to pay that kind of money for AI inference, $40 bucks a month per-seat for usage likely when there is direct billing of frontier models and encouragement. Find out the goals and the specifications, and survey if AI in general can meet that. Contract with a service provider who understands development and delivering turnkey AI solutions, and who can provide something more reliable and satisfactory than anything relying on even another third-party such as WhatsApp, which will degrade the experience.
Comment on Tencent Debutes AR Navigation Platform for IoV and Announces New Smart Transportation Strategy by Patterson
The article offers an interesting look at how AR navigation could make driving more intuitive by placing useful directions directly into the driver’s view. I especially like the focus on connecting navigation with a broader smart transportation strategy, rather than treating AR as a standalone feature. Accuracy and safety will be crucial as these systems develop. It’s also interesting to see how digital technology continues improving access to everyday services; https://plumberinsandiegoca.com/ is one example of how local businesses can make essential services easier to find online.
Talks to sell PayPal to Stripe and Advent are heating up
PayPal is still reportedly negotiating a potential sale to Stripe and private equity firm Advent, as the fintech firm's new CEO attempts to turn the company around.
Comment on DeepMind & Stanford U’s UNFs: Advancing Weight-Space Modeling with Universal Neural Functionals by Kate
It’s crazy to see how fast model management and automated ops are moving. Kind of reminds me of trying to handle high-touch tech support or complex account workflows where everything used to be custom-built and clunky. I was trying to resolve an enterprise routing issue last week and had to dig up the Algo phone number just to get a human who could walk me through their automated system settings. Standardizing complex back-end architectures makes life so much easier for everyone involved, whether it's telecom routing or DeepMind standardizing neural weight spaces.
The subscription quota for GPT Pro 20 has significantly decreased
My account used to last for a long time, but now I have noticed that the limit of my account has significantly decreased by 6-10 times
Data Loading for AI/ML: A Comprehensive Guide
A deep dive into data loading for model training — the three pipeline stages, parallelism strategies, shuffling, caching, resumability, and how LanceDB's StreamingDataset fits in.