From GPT-2 to gpt-oss: Analyzing the Architectural Advances
And How They Stack Up Against Qwen3
AI/ML news, top picks, and generated innovation digests.
26992 matching items
And How They Stack Up Against Qwen3
jack Morris's investigation into GPT-OSS training data https://x.com/jxmnop/status/1953899426075816164?t=3YRhVQDwQLk2gouTSACoqA&s=09
By Shaina Raza and Veronica Chatrath AI models are rapidly becoming bigger, faster, and more capable at understanding images and text together. However, while accuracy and speed are often celebrated, […] The post When AI Meets Human Matters: Evaluating Multimodal Models Through a Human-Centred Lens – Introducing HumaniBench appeared first on Vector Institute for Artificial Intelligence .
Pierrette Mahoro Mastel currently works at GIZ as a Digital Health Advisor, prior to that she was at CMU-Africa where she did her Masters in IT with a major in Machine Learning. Mastel is the IndabaX Rwanda lead and one of the 2025 General Chairs for the Indaba to be held in Kigali 17-22 August […] The post When Community Leads: Rwanda to Host the 2025 Annual Deep Learning Indaba appeared first on Deep Learning Indaba .
Putting the AI in Charge
We are delighted to welcome Nicole Ludwig at the Tübingen AI Center!
Adding attention to linear probes
Three outstanding Principal Investigators will be joining the ELLIS Institute Tübingen, co-affiliated with the MPI-IS and the Tübingen AI Center.
On 18 July 2025, the European Commission published draft Guidelines clarifying key provisions of the EU AI Act applicable to General Purpose AI (GPAI) models. The Guidelines provide interpretive guidance on the definition and scope of GPAI models, related lifecycle obligations, systemic risk criteria, and notification duties for providers. Once translated into all EU languages, […]
The Code of Practice offers a clear framework to help developers of General Purpose AI (GPAI) models meet the requirements of the EU AI Act. While providers can choose to follow the Code, they are also free to demonstrate compliance through other appropriate methods. This post provides a concise overview of each Chapter, Commitment, and […]
Lena Schlipf Honored by Students
Does process matter? We are about to find out.
Deep Learning Indaba 2025: Africa’s biggest AI Community Gathers 1000 participants in Kigali, Rwanda to Shape the Future KIGALI, RWANDA –July 14th, 2025– The Deep Learning Indaba (DLI), Africa’s premier machine learning and artificial intelligence (AI) event, proudly announces its 7th edition, set to take place in Kigali, Rwanda, under the powerful theme “Urunana – […] The post Press Release DLI 2025 appeared first on Deep Learning Indaba .
AI Singapore (AISG) and the New Zealand Ministry of Business, Innovation & Employment (MBIE) are proud to announce the awardees of the Singapore – New Zealand Joint Grant Call...
Paper: https://research.trychroma.com/context-rot Abstract: Large Language Models (LLMs) are typically presumed to process context uniformly—that is, the model should handle the 10,000th token just as reliably as the 100th. However, in practice, this assumption does not hold. We observe that model performance varies significantly as input length changes, even on simple tasks. In this report, we evaluate 18 LLMs, including the state-of-the-art GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 models. Our results reveal that models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows. Authors: Kelly Hong, Anton Troynikov, Jeff Huber Links: Homepage: https://ykilcher.com Merch: https://ykilcher.com/merch YouTube: https://www.youtube.com/c/yannickilcher Twitter: https://twitter.com/ykilcher Discord: https://ykilcher.com/discord LinkedIn: https://www.linkedin.com/in/ykilcher If you want to support me, the best thing to do is to share out the content :) If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this): SubscribeStar: https://www.subscribestar.com/yannickilcher Patreon: https://www.patreon.com/yannickilcher Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2 Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRj…
Vector’s latest annual report showcases research advancements and industry partnerships that strengthen Canada’s AI leadership Vector bridges AI research and application, translating cutting-edge science into solutions that benefit Canadians. Over […] The post Vector Institute 2024-25 annual report: Where AI research meets real-world impact appeared first on Vector Institute for Artificial Intelligence .
Paper: https://arxiv.org/abs/2507.02092 Code: https://github.com/alexiglad/EBT Website: https://energy-based-transformers.github.io/ Abstract: Inference-time computation techniques, analogous to human System 2 Thinking, have recently become popular for improving model performances. However, most existing approaches suffer from several limitations: they are modality-specific (e.g., working only in text), problem-specific (e.g., verifiable domains like math and coding), or require additional supervision/training on top of unsupervised pretraining (e.g., verifiers or verifiable rewards). In this paper, we ask the question "Is it possible to generalize these System 2 Thinking approaches, and develop models that learn to think solely from unsupervised learning?" Interestingly, we find the answer is yes, by learning to explicitly verify the compatibility between inputs and candidate-predictions, and then re-framing prediction problems as optimization with respect to this verifier. Specifically, we train Energy-Based Transformers (EBTs) -- a new class of Energy-Based Models (EBMs) -- to assign an energy value to every input and candidate-prediction pair, enabling predictions through gradient descent-based energy minimization until convergence. Across both discrete (text) and continuous (visual) modalities, we find EBTs scale faster than the dominant Transformer++ approach during training, achieving an up to 35% higher scaling rate with respect to data, batch size, parameters, FLOPs…
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design
The Future of Life Institute’s 2025 summer update to its AI Safety Index shows some companies making incremental progress, but dangerous gaps remain in key categories such as risk assessment and controlling the systems they plan to build.
Aquí puedes descargar nuestro logo oficial para aplicarlo en los distintos diseños que necesites. Logotipo en colores oficiales para fondos claros: Logotipo en colores oficiales para fondos oscuros: Logotipo monocromático para fondos claros: Logotipo monocromático para fondos oscuros:
Confronting the production-progress paradox
On July 11, the Tübingen AI Center hosted the 2025 edition of the Tübingen AI Summer Poster Event
¡Hola! Te damos la bienvenida al sitio web de Derechos Digitales, esperamos que disfrutes de los contenidos que hemos preparado para ti. Para conocer más sobre nuestra organización, quiénes somos y qué hacemos, no dudes en visitar https://www.derechosdigitales.org/quienes-somos/ A continuación encontrarás nuestra Política de Privacidad de la Información y Datos personales aplicables a tu navegación […]
تحت الياسمينة This month’s newsletter was written by two community members. Amal Nammouchi, originally from Tunisia 🇹🇳, is a co-founder of AfriClimate AI, a Pan-African non-profit using AI to tackle climate challenges across Africa. She is a PhD candidate affiliated with SOLVE and Karlstad University, Sweden, and is a long time organiser for the Deep Learning […] The post Under the Jasmine Tree appeared first on Deep Learning Indaba .
AI can help, or hurt, our thinking
I have two binary classifiers and would like to check whether there is a statistically significant difference between the area under the ROC curve (AUROC). I have reason to opt for AUROC as my evaluation metric of choice. For each classifier, I have 15 runs as I do 5-fold cross-validation and use 3 random seeds for initialisation. For evaluation, I have used unseen/independent test data. This means that for both classifiers I have 15 paired AUROC values. According to this article on Nature , DeLong test is (often) used for significance testing with AUROCs. However, as this depends on the variance and covariance I suspect that I cannot use DeLong test with these 15 AUROC values. In order to use DeLong test, I should concatenate all predictions on the test data across the 15 unique versions of each classifier. Would this be correct? Would it be a good idea to use paired t-test on these 15 AUROC pairs (assuming the differences between these pairs values are normally distributed)? Are there any arguments favouring either DeLong test or paired t-test?
A topic-organized collection of 200+ LLM research papers from 2025
× Predicting Ego-centric Video from human Actions (PEVA) . Given past video frames and an action specifying a desired change in 3D pose, PEVA predicts the next video frame. Our results show that, given the first frame and a sequence of actions, our model can generate videos of atomic actions (a), simulate counterfactuals (b), and support long video generation (c). Recent years have brought significant advances in world models that learn to simulate future outcomes for planning and control. From intuitive physics to multi-step video prediction, these models have grown increasingly powerful and expressive. But few are designed for truly embodied agents. In order to create a World Model for Embodied Agents, we need a real embodied agent that acts in the real world. A real embodied agent has a physically grounded complex action space as opposed to abstract control signals. They also must act in diverse real-life scenarios and feature an egocentric view as opposed to aesthetic scenes and stationary cameras. 💡 Tip: Click on any image to view it in full resolution. Why It’s Hard Action and vision are heavily context-dependent. The same view can lead to different movements and vice versa. This is because humans act in complex, embodied, goal-directed environments. Human control is high-dimensional and structured. Full-body motion spans 48+ degrees of freedom with hierarchical, time-dependent dynamics. Egocentric view reveals intention but hides the body. First-person vision reflects…
Tübingen AI Center is proud to be a partner in ELLlOT, a Horizon Europe-funded project aiming to develop next-generation Multimodal Generalist Foundation Models.
Johannesburg, South Africa, 30 June 2025 – Cassava Technologies, a global technology leader of African heritage, is pleased to announce that it has signed a Memorandum of Understanding (MOU) with the South African AI Association (SAAIA), an industry body focused on growing responsible AI adoption, to deliver artificial intelligence (AI) solutions and GPU-as-a-Service (GPUaas) across the […]
"Google’s rollout of tools like AI Overviews and AI Mode—chatbots that answer users’ search queries—has begun shifting online behavior from browsing links to reading AI-generated responses, drastically reducing traffic to news sites. The resulting drop in organic traffic is forcing many media organizations to rethink their sustainability models in an ecosystem increasingly dominated by tech […] The post Google’s AI search features slash traffic to news sites, deepening sustainability crisis appeared first on LatAm Journalism Review by the Knight Center .
"Google’s rollout of tools like AI Overviews and AI Mode—chatbots that answer users’ search queries—has begun shifting online behavior from browsing links to reading AI-generated responses, drastically reducing traffic to news sites. The resulting drop in organic traffic is forcing many media organizations to rethink their sustainability models in an ecosystem increasingly dominated by tech […] The post Google’s AI search features slash traffic to news sites, deepening sustainability crisis appeared first on LatAm Journalism Review by the Knight Center .
Securing the invisible paths: How cross-account event flows can become security blind spots
ByteDance introduces Astra, an innovative dual-model architecture revolutionizing robot navigation in complex indoor environments. The post ByteDance Introduces Astra: A Dual-Model Architecture for Autonomous Robot Navigation first appeared on Synced .
Which AIs to use, and how to use them
Research update on on applying local volume measurement to downstream tasks
Evaluation metrics, how to build eval datasets, eval methodology, and a review of several benchmarks.
I’m developing a tree-based model classifier (XGBoost) using some healthcare (patient visits) data. The data has a time dimension, and I want to observe if there is a longitudinal effect for the prediction of the target feature. To predict the target for the current visit (Timepoint n), it should incorporate information from the previous visit(s) (T0 to T(N-1)). The input shape is visit_time, features, and target/label. Let’s say, I have a patient with 5 visits (T1 – T5). The idea is that the first prediction (T1) will be just based on the features for this timepoint. To predict T2, I want to add information from T1. Then, for T3, it will be T1 + T2, and so on, T5 (T1+ T2 + T3 + T4). The number of timepoints (visits) vary for each patient. I read that I can add lags and rolling windows. But still couldn’t figure out the best way to do it. Any thoughts on what and how to do it for my scenario?
KV caches are one of the most critical techniques for efficient inference in LLMs in production.
MIT introduces SEAL, a framework enabling large language models to self-edit and update their weights via reinforcement learning. The post MIT Researchers Unveil “SEAL”: A New Step Towards Self-Improving AI first appeared on Synced .
"Automated failure attribution" is a crucial component in the development lifecycle of Multi-Agent systems. It has the potential to transform the challenge of identifying "what went wrong and who is to blame" from a perplexing mystery into a quantifiable and analyzable problem The post Researchers from PSU and Duke introduce “Multi-Agent Systems Automated Failure Attribution first appeared on Synced .
In this post, we will study inductive biases of the parameter-function map of random neural networks using star domain volume estimates. This builds on the ideas introduced in Estimating the Probability of Sampling a Trained Neural Network at Random and Neural Redshift: Random Networks are not Random Functions (henceforth NRS). Inductive biases To understand generalization in deep neural networks, we must understand inductive biases. Given a fixed architecture, some tasks will be easily learnable, while others can take an exponentially long time to learn (see here and here).
Our best practices for quickly identifying, resolving, and preventing issues at scale. The post How GitHub engineers tackle platform problems appeared first on The GitHub Blog .
With the conclusion of its inaugural round, the DEEP X Tübingen AI cooperation has established a successful sciencepreneurship initiative in artificial intelligence.
Announcing the Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
In many resources, I see that repeated measures analyses—such as mixed-effects models or GEE—are commonly applied when the outcome is measured multiple times within the same individuals. However, I was wondering: Can these models also be used when it’s the exposure that is repeatedly measured (e.g. at several time points), and the outcome is measured only once (e.g. at a later time point)? For example, suppose I measure maternal blood pressure at 3 time points during pregnancy (repeated exposure), and I want to study its association with birthweight (a single outcome). Can I use a mixed-effects model or GEE to account for within-subject correlation in the exposures? I’m curious which modeling approaches are most appropriate in this context, and if there are any recommended papers or examples. So for example in R Studio: model
"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry Winograd The recent successes of generative AI models have convinced some that AGI is imminent. While these models appear to capture the essence of human