AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
34834News Items
8Top Picks
202Blogs
successLast Run

NLP

122 articles tagged with this keyword, sorted by most recent first.

← All Keywords
CIO AI 2026-08-13 20:27 UTC Score 54.0 USR-0125-20260813-global-ai-ne-0b3668ac

Manual vs. AI-powered PDF redaction: protecting sensitive data in 2026

Research shows that humans play a role in 60% of breaches that expose sensitive data. That “role” often involves an employee falling for a phishing scam or using PASSWORD for their login credentials, but data exposure can also be a result of how your business redacts sensitive and personally identifiable information (PII) in your documents. Historically, manual, “black-box” redaction was considered best-practice, but this approach only obscures data, it doesn’t permanently remove it. As regulations governing data security get stricter and AI-powered redaction solutions become more accessible, organizations—especially those in highly regulated industries—are re-evaluating their PDF redaction solutions. How manual PDF redaction is different from AI-powered PDF redaction The primary difference between manual and AI-powered PDF redaction is who (or what) you rely on to do the heavy lifting. Manual redaction defined Manual redaction is a human-driven process where individuals visually scan text, select content, and apply black boxes or remove the text before sharing or storing the file. Manual redaction is only as effective as the reviewer, which makes outcomes highly variable, especially under time pressure or high document volume. AI-powered redaction defined AI-powered redaction uses machine learning and natural language processing (NLP) to automatically detect and remove sensitive information from documents. The system is trained to recognize patterns, language cues, and cont…

Synced 2026-08-13 08:52 UTC Score 57.0 AI-041-20260813-ai-specialis-1acbfefb

Comment on CMU’s Novel ‘ReStructured Pre-training’ NLP Approach Scores 40 Points Above Student Average on a Standard English Exam by mistfallhunterwiki

The idea of pretraining on restructured data is fascinating, especially how the QIN system reportedly scored 40 points above the student average on the Gaokao-English Exam while using only 1/16 of GPT-3's parameters. That efficiency gain makes the RST paradigm particularly compelling for broader NLP applications, not just exam benchmarks. I appreciate the authors sharing this direction and look forward to seeing how restructured pretraining performs across other language tasks.

ACL Anthology 2026-08-11 00:00 UTC Score 22.0 AI-079-20260811-research-pap-5af4404b

Automatic Expansion of Lexicons for Multilingual Misogyny Detection

Simona Frenda, Bilal Ghanem, Estefanía Guzmán-Falcón, Manuel Montes-y-Gómez and Luis Villaseñor-Pineda in Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2018)

ACL Anthology 2026-08-11 00:00 UTC Score 22.0 AI-079-20260811-research-pap-c39af0f6

AMI @ EVALITA2020: Automatic Misogyny Identification

Elisabetta Fersini, Debora Nozza and Paolo Rosso in Proceedings of the Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2020)

Synced 2026-08-08 09:39 UTC Score 40.0 AI-041-20260808-ai-specialis-9447d79f

Comment on Facebook Says Its ‘Blender’ Chatbot Is the Most Humanlike Ever by paramount books

Facebook’s Blender chatbot is an impressive step toward making AI conversations feel more natural and humanlike. As AI continues to improve, it could also transform how people discover educational resources, get recommendations, and interact with online stores such as Paramount Books . It will be interesting to see how this technology develops and shapes digital communication in the future.

CIO AI 2026-08-07 09:30 UTC Score 50.0 USR-0125-20260807-global-ai-ne-6fbe4864

How AI is changing the business analyst role for the better

AI’s impact has been felt across nearly every industry, and its rise has already started to alter several roles in tech, including that of the business analyst . While the rise of agentic AI may have some questioning whether AI will replace business analyst jobs entirely, as we’ve seen with most roles impacted by AI, it’s more likely that AI will augment the role and fundamentally change how BA’s conduct daily business. “As AI takes on more routine tasks, the human side of the role is becoming even more valuable. It’s becoming more of a hybrid role, where employers are often looking for candidates who can combine technical fluency with strong communication and problem-solving skills, along with sound business judgment,” says Megan Slabinski, district president of technology talent solutions at Robert Half. AI can save business analysts time in the long run, automating many of the tasks that are time consuming and repetitive around data processing, note taking, and documentation. While automation will impact the daily tasks of the role, business analysts will still be necessary for properly interpreting outputs, collaborating across teams, and maintaining compliance and AI workflows. AI-driven analysis and automated workflows With AI-driven analysis, BA’s can use machine learning models for pattern detection, determining risk, and for forecasting demand, while natural language processing (NLP) can be used for text-heavy inputs. AI tools can also assist analysts with decision-…

Apple Machine Learning Research 2026-08-07 00:00 UTC Score 49.0 AI-059-20260807-official-ai--3d79e525

Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models

Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks, including document processing and code generation. Autoregressive Language Models (ARMs), which generate tokens sequentially conditioned on all previous tokens, have been the predominant paradigm for LLMs. While these models have achieved high accuracy across a range of downstream tasks, they exhibit low arithmetic intensity due to the inherent sequential dependency in next-token prediction. Recently, Diffusion Language Models (DLMs) have emerged as a promising…

Synced 2026-08-05 02:59 UTC Score 40.0 AI-041-20260805-ai-specialis-39b7560a

Comment on Natural Language Processing In Early Education by superbearadventure.app

Great article! I really enjoyed reading this and found the information clear, practical, and easy to understand. You explained the topic in a way that's helpful for both beginners and experienced readers. Thanks for taking the time to share such valuable insights. I'm looking forward to reading more of your content. Keep up the excellent work!

Synced 2026-08-01 23:10 UTC Score 64.0 AI-041-20260801-ai-specialis-60ccc23e

Comment on Google’s ALBERT Is a Leaner BERT; Achieves SOTA on 3 NLP Benchmarks by Gonzales

What impressed me most about ALBERT was how it proved that smarter architecture can outperform simply adding more parameters. That idea of achieving better efficiency with fewer resources is valuable far beyond NLP research. The same principle applies in everyday logistics too, streamlining processes often delivers better results than adding complexity. I recently found https://moverspacker.ae/home-movers-in-ras-al-khaimah/ while looking at examples of efficient relocation planning, and it reminded me of that same optimization mindset. Great article highlighting why thoughtful design matters as much as raw scale.

ACL Anthology 2026-07-31 00:00 UTC Score 13.0 AI-079-20260731-research-pap-82652bf0

CSULoRA: Closest Safe Update Low-Rank Adaptation

Oleksandr Marchenko, Adelaide Danilov, Aria Nourbakhsh and Salima Lamsiyah in Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security

Synced 2026-07-29 14:17 UTC Score 37.0 AI-041-20260729-ai-specialis-a8632e0f

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by juego copero

Great article on FPDT! The memory hierarchy optimization is really impressive. As someone who runs a gaming site, I can appreciate efficiency improvements—smooth gameplay depends heavily on resource management. Check out https://juegocopero.com/ for more on that!

LessWrong AI 2026-07-25 01:19 UTC Score 68.0 USR-0152-20260725-community-fo-03cc5338

Linear probes tell you where quantization will hurt

Epistemic status: I have only tested one encoder family (BERT-base and its relatives) and one decoder LLM ( Qwen2.5 -3B), one seed, token-level tasks, and post-training weight quantization. I trust the results because I am not inventing anything; I am just connecting quantization with a very general idea from mech interp: the model does the easy syntax work first, the semantic work in later layers, and the prompt-relevant task work in the latest layers. That's the main idea I lean on in this project. What I'm not sure of is whether it holds when the task is diffuse, like general-knowledge QA (probably not). TL;DR. First I train a linear probe on each layer and try to see where a signal lives in a model. I only probe for classical NLP or CV tasks, like checking for depth or vision for CV or named entity recognition, parts of speech, or chunking for NLP. Then using that information to create a map of where the important work is happening and quantizing the less important layers for a task. This method is cheap; it's only one RidgeClassifier per layer. And it kind of works; a map built once on CoNLL news keeps 99–100% of full-precision accuracy at a 5-bit average on three unseen datasets, whereas Uniform kept 16–41% and opposite/anti layers kept 2–41%. This method is perfect for sharp use cases where the transformer needs specialized knowledge and can compromise on general knowledge. The itch So this all started during my 2nd semester at Northeastern, where I took Applied Progr…

Synced 2026-07-19 01:55 UTC Score 46.0 AI-041-20260719-ai-specialis-c7be6d26

Comment on Peking U & Microsoft’s Knowledge Attribution Method Enables Editing Factual Knowledge in Pretrained Transformers Without Fine-Tuning by Markdown to Doc

Impressive progress in editing factual knowledge without fine-tuning! The idea of targeting "knowledge neurons" is fascinating. For a deeper dive into transformer mechanics, check out this Markdown to Doc resource—it helps analyze such groundbreaking NLP innovations efficiently.

Data Science Stack Exchange 2026-07-16 17:56 UTC Score 32.0 AI-111-20260716-social-media-49c70074

Improving short‑text classification accuracy with overlapping classes and imbalanced data

I’m working on a multi‑class text classification problem where the input consists of very short descriptions (often only a few words) and the goal is to predict the correct category. I’m currently using an XGBoost classifier for the final prediction layer. for embeddings I am using e5 large model. The main challenges I’m facing are: The descriptions are extremely short and many classes share similar vocabulary. Because of this, the model sometimes produces high‑confidence but incorrect predictions when common words appear across multiple classes. Description is the only feature that supports the category classification. The dataset is imbalanced: some classes have significantly more training samples than others, but the distribution of the prediction data is different. This causes the model to over‑predict certain classes. I experimented with TF‑IDF features, embeddings from a generative AI model and a combination of TF‑IDF + embeddings. However, combining both actually reduced accuracy. I also tried downsampling majority class but it not really make any difference. Current model accuracy is around 30%. I’m looking for advice on: How to improve classification when different classes share highly overlapping vocabulary. How to reduce high‑confidence wrong predictions. What techniques work well for class imbalance and distribution shift between training and prediction data.

Machine Learning Mastery 2026-07-15 12:00 UTC Score 27.0 AI-039-20260715-ai-specialis-f38e483d

Scikit-Ollama for Scikit-LLM/Ollama Integration

In this article, you will learn how scikit-ollama bridges the scikit-learn interface with locally running Ollama models to perform zero-shot text classification; no cloud API...

Synced 2026-07-11 04:11 UTC Score 40.0 AI-041-20260711-ai-specialis-71f860d6

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by taichiwalk.org

Fascinating to see how Microsoft is pushing hardware efficiency for longer sequence processing—definitely a leap for AI scalability. All that computational intensity makes me think about the other side of the coin: finding calm after a deep tech session. Lately I’ve been exploring tai chi walking as a low-impact way to reset focus and improve balance, especially since sitting at a desk for hours can take a toll. There’s a beginner-friendly guide at taichiwalk.org that offers free routines and even a guided coach—no login needed, which I appreciate. Anyone else here use movement or gentle exercise to counterbalance screen time?

Synced 2026-07-07 07:10 UTC Score 51.0 AI-041-20260707-ai-specialis-3f6bf0e5

Comment on Interview with Tencent AI Lab- Four Core Research Fields:Computer Vision, Speech Recognition, Natural Language Processing and Machine Learning by Owen Parker

I've seen examples where clear tracking information greatly improves the user experience. If you're interested in how a modern tracking system presents shipment updates and delivery statuses, this guide is worth a look: https://anposttracking.org/ . It's a helpful reference for anyone building customer-focused applications

Transactions on Machine Learning Research 2026-07-06 00:00 UTC Score 49.0 AI-084-20260706-research-pap-960a167b

Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection

Attention mechanisms have revolutionized several domains of artificial intelligence, such as natural language processing and computer vision, by enabling models to selectively focus on relevant parts of the input data. While recent work has characterized the optimization dynamics of gradient descent (GD) in attention-based models and the structural properties of its preferred solutions, less is known about more general optimization algorithms such as mirror descent (MD). In this paper, we investigate the convergence properties and implicit biases of a family of MD algorithms tailored for softmax attention mechanisms, with the potential function chosen as the $p$-th power of the $\ell_p$-norm. Specifically, we show that these algorithms converge in direction to a generalized hard-margin SVM with an $\ell_p$-norm objective when applied to a classification problem using a softmax attention model. Notably, our theoretical results reveal that the convergence rate is comparable to that of traditional GD in simpler models, despite the highly nonlinear and nonconvex nature of the present problem. Additionally, we delve into the joint optimization dynamics of the key-query matrix and the decoder, establishing conditions under which this complex joint optimization converges to their respective hard-margin SVM solutions. Lastly, our numerical experiments on real data demonstrate that MD algorithms improve generalization over standard GD and excel in optimal token selection.

Transactions on Machine Learning Research 2026-07-06 00:00 UTC Score 40.0 AI-084-20260706-research-pap-bedebbba

Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective

The Transformer model is widely used in various application areas of machine learning, such as natural language processing. This paper investigates the approximation of the Hölder continuous function class $\mathcal{H}_{Q}^{\beta}\left([0,1]^{d\times n},\mathbb{R}^{d\times n}\right)$ by Transformers and constructs several Transformers that can overcome the curse of dimensionality. These Transformers consist of one self-attention layer with one head and the softmax function as the activation function, along with several feedforward layers. For example, to achieve an approximation accuracy of $\epsilon$, if the activation functions of the feedforward layers in the Transformer are ReLU and floor, only $\mathcal{O}\left(\log\frac{1}{\epsilon}\right)$ layers of feedforward layers are needed, with widths of these layers not exceeding $\mathcal{O}\left(\frac{1}{\epsilon^{2/\beta}}\log\frac{1}{\epsilon}\right)$. If other activation functions are allowed in the feedforward layers, the width of the feedforward layers can be further reduced to a constant. These results demonstrate that Transformers have a strong expressive capability. The construction in this paper is based on the Kolmogorov-Arnold Superposition Theorem and does not require the concept of contextual mapping, hence our proof is more intuitively clear compared to previous Transformer approximation works. Additionally, the translation technique proposed in this paper helps to apply the previous approximation results of fe…

Machine Learning Mastery 2026-06-16 12:00 UTC Score 27.0 AI-039-20260616-ai-specialis-e9483392

Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM

Traditional machine learning pipelines for predictive tasks like text classification usually rely on extracting structured, numerical features from raw text — for instance, TF-IDF frequencies or token embeddings — to feed into classical models such as logistic regression, ensembles, or support vector machines.

Machine Learning Mastery 2026-06-11 12:00 UTC Score 18.0 AI-039-20260611-ai-specialis-824e0fa0

Multi-Label Text Classification with Scikit-LLM

Text classification typically boils down to scenarios where a product review is "positive" or "negative", or a customer inquiry belongs to one category or another.

AI Stack Exchange 2024-05-30 06:02 UTC Score 29.0 AI-110-20240530-social-media-da3d22c6

Probability interpretation of attention mechanism in Seq2Seq

I have ready many explanations of the seq2seq model. In my opinion, however, it is really like a robot that might say something correctly, but doesn't really understand it, just as is true with an LLM generally. In my opinion, the correct way to describe Seq2Seq and similar NLP models should start from a probability view. My probability view is very simple; the output of the encoder is a representation of the probability distribution of the next word. In each step of the Decoder, it just modifies the distribution based on each word it predicted from the distribution and outputs the modified distribution. It then does this repeatedly. Assuming this probability view is correct, how could we explain the attention mechanism used in Seq2seq?

Data Science Stack Exchange 2022-08-11 00:20 UTC Score 29.0 AI-111-20220811-social-media-485f9224

Difference Between Attention and Fully Connected Layers in Deep Learning

There have been several papers in the last few years on the so-called "Attention" mechanism in deep learning (e.g. 1 2 ). The concept seems to be that we want the neural network to focus on or pay more attention to certain features, and has demonstrated some empirical success in NLP and related sequential models. When I look at some code examples such as this one , adding an Attention layer intuitively makes sense and seems to improve performance of the LSTM model. However it looks very much like a regular fully-connected layer. In that link (and with some slight change of notation), the Attention layer outputs $$ c(x) = \tanh(\mathbf{W}x + \mathbf{b} ) $$ $$ \beta(x) = \frac{e^{c(x_j)}}{\sum_{j} e^{c(x_j)}} $$ $$ f_{Attention}(x) = x\beta $$ where $W,b$ are weights/biases, $x$ is layer input, and $f(.)$ is the layer output. In contrast, a regular fully connected layer: $$ f_{Dense} = \sigma(\mathbf{W}x + \mathbf{b}) $$ for some activation function $\sigma(.)$ . My interpretation of the Attention implementation above is that it is pretty much the same thing as a standard fully connected layer, but with a $\tanh$ activation (why?), followed by a $\text{softmax}$ (okay, so that the "attention weights" $\beta$ sum to 1), followed by a linear dot product. How does this architecture allow the model to have "attention"? I do not see how it is fundamentally different or more expressive from just adding a standard fully-connected layer. Am I misunderstanding something here? Edit/My…

Stanford AI Lab Blog 2022-05-31 07:00 UTC Score 47.0 USR-0006-20220531-research-aca-a57ebba7

LinkBERT: Improving Language Model Training with Document Link

Language Model Pretraining Language models (LMs), like BERT 1 and the GPT series 2 , achieve remarkable performance on many natural language processing (NLP) tasks. They are now the foundation of today’s NLP systems. 3 These models serve important roles in products and tools that we use every day, such as search engines like Google 4 and personal assistants like Alexa 5 . These LMs are powerful because they can be pretrained via self-supervised learning on massive amounts of text data on the web without the need for labels, after which the pretrained models can be quickly adapted to a wide range of new tasks without much task-specific finetuning. For instance, BERT is pretrained to predict randomly masked words in original text (masked language modeling), e.g. predicting the masked word “dog” from “My __ is fetching the ball”. GPTs are pretrained to predict the next word given a previous sequence of text (causal language modeling), e.g. predicting the next word “ball” from “My dog is fetching the”. In either cases, through pretraining, LMs learn to encode various knowledge from a text corpus that helps to perform downstream applications involving language understanding or generation. In particular, LMs can learn world knowledge (associations between concepts like “dog”, “fetch”, “ball”) from training text where the concepts appear together, and help for knowledge-intensive applications like question answering. 6 Challenges. A challenge with most common LM pretraining strateg…

Stanford AI Lab Blog 2022-05-25 07:00 UTC Score 47.0 USR-0006-20220525-research-aca-2eecb290

Stanford AI Lab Papers and Talks at ACL 2022

The 60th Annual Meeting of the Association for Computational Linguistics (ACL) 2022 is taking place May 22nd - May 27th. We’re excited to share all the work from SAIL that’s being presented, and you’ll find links to papers, videos and blogs below. Feel free to reach out to the contact authors directly to learn more about the work that’s happening at Stanford! List of Accepted Papers LinkBERT: Pretraining Language Models with Document Links Authors : Michihiro Yasunaga, Jure Leskovec*, Percy Liang* Contact : myasu@cs.stanford.edu Links: Paper | Website Keywords : language model, pretraining, knowledge, hyperlink, bionlp When classifying grammatical role, BERT doesn’t care about word order… except when it matters Authors : Isabel Papadimitriou, Richard Futrell, Kyle Mahowald Contact : isabelvp@stanford.edu Links: Paper Keywords : large language models, analysis, word order, order invariance, grammatical role, syntax, semantics Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words Authors : Kaitlyn Zhou, Kawin Ethayarajh, Dallas Card, Dan Jurafsky Contact : katezhou@stanford.edu Keywords : cosine similarity, training data frequency, model analysis Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive Summarization Authors : Faisal Ladhak, Esin Durmus, He He, Claire Cardie, Kathleen McKeown Contact : esdurmus@stanford.edu Links: Paper Keywords : text summarization, text generation, evaluation, faithfulness Sp…

AI Stack Exchange 2022-03-15 17:05 UTC Score 29.0 AI-110-20220315-social-media-22338a14

Having the negative cases in the same batch vs. shuffling the dataset

I am working on a model for an NLP task. The model encodes the text and has a regression output layer. In this task, from each instance (positive), I create several negative cases using a specific technique and I merge them with their positive corresponding ones in a data split (training/val/test). After that, I shuffle the data split. I was thinking of the following: Isn't better to keep the negative instances with their corresponding positive ones in the same batch instead of shuffling the data? Is there an answer to this question? does it depend on the task?

Stanford AI Lab Blog 2022-01-21 08:00 UTC Score 58.0 USR-0006-20220121-research-aca-1e3c1829

Reward Isn't Free: Supervising Robot Learning with Language and Video from the Web

This work was conducted as part of SAIL and CRFM . Deep learning has enabled improvements in the capabilities of robots on a range of problems such as grasping 1 and locomotion 2 in recent years. However, building the quintessential home robot that can perform a range of interactive tasks, from cooking to cleaning, in novel environments has remained elusive. While a number of hardware and software challenges remain, a necessary component is robots that can generalize their prior knowledge to new environments, tasks, and objects in a zero or few shot manner. For example, a home robot tasked with setting the dining table cannot afford lengthy re-training for every new dish, piece of cutlery, or dining room it may need to interact with. A natural way to enable such generalization in our robots is to train them on rich data sources that contain a wide range of different environments, tasks, and objects. Indeed, this recipe of massive, diverse datasets combined with scalable offline learning algorithms (e.g. self-supervised or cheaply supervised learning) has been the backbone of the many recent successes of foundation models 3 in NLP 4 5 6 7 8 9 and vision 10 11 12 . Replicating these impressive generalization and adaptation capabilities in robot learning algorithms would certainly be a step toward robots that can be used in unstructured, real world environments. However, directly extending this recipe to robotics is nontrivial, as we neither have sufficiently large and diverse…

Stanford AI Lab Blog 2021-11-05 07:00 UTC Score 52.0 USR-0006-20211105-research-aca-02e23852

Stanford AI Lab Papers at EMNLP/CoNLL 2021

The 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021) will take place next week, colocated with CoNLL 2021. We’re excited to share all the work from SAIL that will be presented, and you’ll find links to papers, videos and blogs below. Feel free to reach out to the contact authors directly to learn more about the work that’s happening at Stanford! List of Accepted Papers Calibrate your listeners! Robust communication-based training for pragmatic speakers Authors : Rose E. Wang, Julia White, Jesse Mu, Noah D. Goodman Contact : rewang@stanford.edu Links: Paper | Video Keywords : language generation, pragmatics, communication-based training, calibration, uncertainty Cross-Domain Data Integration for Named Entity Disambiguation in Biomedical Text Authors : Maya Varma, Laurel Orr, Sen Wu, Megan Leszczynski, Xiao Ling, Christopher Ré Contact : mvarma2@stanford.edu Links: Paper | Video Keywords : named entity disambiguation, biomedical text, rare entities, data integration ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts Authors : Yuta Koreeda, Christopher D. Manning Contact : koreeda@stanford.edu Links: Paper | Website Keywords : natural language inference, contract, law, legal, dataset Venue : The Findings of EMNLP 2021 The Emergence of the Shape Bias Results from Communicative Efficiency Authors : Eva Portelance, Michael C. Frank, Dan Jurafsky, Alessandro Sordoni, Romain Laroche Contact : portelan@stanford.edu Links…

Stanford AI Lab Blog 2021-10-13 07:00 UTC Score 52.0 USR-0006-20211013-research-aca-a7eb5ef1

Selective Classification Can Magnify Disparities Across Groups

Selective classification, where models are allowed to “abstain” when they are uncertain about a prediction, is a useful approach for deploying models in settings where errors are costly. For example, in medicine, model errors can have life-or-death ramifications, but abstentions can be easily handled by backing off to a doctor, who then makes a diagnosis. Across a range of applications from vision 1 2 3 and NLP 4 5 , even simple selective classifiers, relying only on model logits, routinely and often dramatically improve accuracy by abstaining. This makes selective classification a compelling tool for ML practitioners 6 7 . However, in our recent ICLR paper, we find that despite reliably improving average accuracy, selective classification can fail to improve and even hurt the accuracy over certain subpopulations of the data . As a motivating example, consider the task of diagnosing pleural effusion, or fluid in the lungs, from chest X-rays. Pleural effusion is often treated with a chest drain, so many pleural effusion cases also have chest drains, while most cases without pleural effusion do not have chest drains 8 . While selective classification improves average accuracy for this task, we find that it does not appreciably improve accuracy on the most clinically relevant subgroup, or subpopulation, of the data: those that have pleural effusion but don’t yet have a chest drain, i.e. those that have pleural effusion but have not yet been treated for it. Practitioners, thus,…

Lilian Weng Blog 2021-03-21 00:00 UTC Score 39.0 USR-0112-20210321-ai-specialis-3e60cc8a

Reducing Toxicity in Language Models

Large pretrained language models are trained over a sizable collection of online data. They unavoidably acquire certain toxic behavior and biases from the Internet. Pretrained language models are very powerful and have shown great success in many NLP tasks. However, to safely deploy them for practical real-world applications demands a strong safety control over the model generation process.

AI Stack Exchange 2021-03-04 11:35 UTC Score 18.0 AI-110-20210304-social-media-c6ad9850

CAPTCHA based on text comprehension and random tokens

I developed a novel type of CAPTCHA based on text comprehension and random tokens . Given a task Pick the first pair of adjacent letters and a random token 8NBA596V , the user has to provide the solution NB . It offers basic protection and an attacker can solve individual tasks with specific effort. I am curious, whether contemporary AI can solve it generically ? You can access more example tasks here: https://www.topincs.com/manual/captcha There is a task database and at every attempt a new task is presented with a new random token. They always have a solution of varying length and pure guessing thus has limited chances of success. It is easy to attack an individual task by writing a small piece of code, thus a large task database is essential. What intrigues me is the question whether natural language processing or machine learning at its current state can attack the CAPTCHA generically by building a model of the meaning of the task – essentially a predicate in a tiny universe of discourse – and then applying it to the random token.

AI Stack Exchange 2021-01-13 13:34 UTC Score 23.0 AI-110-20210113-social-media-8c81b4b2

How to extract parameters from a text using AI/NLP

lets say I have three texts: "make a heading that says hello word" "make a heading of hello world" "create heading consist of hello world" How can I fetch those groups of words using AI which is referring to heading i.e hello world in this case. Which AI frameworks or libraries can do that? in all examples heading is pointing to hello world (which i am referring as group of words). so basically i want those words which will be a part of heading or in other word there is a relationship between them. another example i can give is "I am watching Breaking bad" so there is a relationship between watching and breaking bad and i want to extract what are you watching. What's the best approach? Do I have to train a model for that or there are some other techniques that can get it done?

Jay Alammar Blog 2020-12-17 00:00 UTC Score 39.0 USR-0113-20201217-ai-specialis-fb351fb3

Interfaces for Explaining Transformer Language Models

Interfaces for exploring transformer language models by looking at input saliency and neuron activation. Explorable #1: Input saliency of a list of countries generated by a language model Tap or hover over the output tokens: Explorable #2: Neuron activation analysis reveals four groups of neurons, each is associated with generating a certain type of token Tap or hover over the sparklines on the left to isolate a certain factor: The Transformer architecture has been powering a number of the recent advances in NLP. A breakdown of this architecture is provided here . Pre-trained language models based on the architecture, in both its auto-regressive (models that use their own output as input to next time-steps and that process tokens from left-to-right, like GPT2) and denoising (models trained by corrupting/masking the input and that process tokens bidirectionally, like BERT) variants continue to push the envelope in various tasks in NLP and, more recently, in computer vision. Our understanding of why these models work so well, however, still lags behind these developments. This exposition series continues the pursuit to interpret and visualize the inner-workings of transformer-based language models. We illustrate how some key interpretability methods apply to transformer-based language models. This article focuses on auto-regressive models, but these methods are applicable to other architectures and tasks as well. This is the first article in the series. In it, we present explo…

Cross Validated 2020-12-12 20:09 UTC Score 26.0 AI-113-20201212-social-media-cb7b55bc

Accuracy is 100% but model.predict is totally wrong! what could be the problem? (Autoencoder NN) [closed]

I have an Autoencoder model that receives 4 different vectors { (1,0,0,0) , (0,1,0,0) , (0,0,1,0) , (0,0,0,1) } The encoder transforms the vectors to vectors size 2, which get inside an "NLPN channel" which is a channel with nonlinear phase noise, and then goes to the decoder that reconstructs the original vector. Architecture: 4->50->2->NLPN->2->50->4 (Softmax) The loss function is mse between the input, and the input ( self.model.fit(self.x, self.x) ) The code is pretty basic: the first cell is just some simple functions defined, then it's the autoencoder class, then it's the training, and then it's the prediction which shows that for each vector 1,0,0,0 and so on, the decoder predicts (~0.25,~0.25,~0.25,~0.25) which shows that the NN has no idea what it's doing. CODE: https://colab.research.google.com/drive/1nFymmEloUSa1yjS7e5CPGsKwqP1HjS57?usp=sharing

Cross Validated 2020-06-10 14:08 UTC Score 23.0 AI-113-20200610-social-media-65d96a8d

Why do Dense layers perform better than a mix of Conv Layers, Recurrent Layers on Sentiment Analysis with BERT emebddings?

I have used BERT to make embeddings out of the imdb review dataset and I am trying out some models to check their perfomance on sentiment analysis (0 for the bad reviews and 1 for the good ones). I have seen that models with just dense Layers do work better than models which are a mixture of Convolutional with MaxPooling Layers and recurrent units such as LSTM or BiLSTM.I want to know why is this happening since BERT do captures of semantics and time dependencies so second models should perform better. I am speaking about orders from 85% accuracy with dense layers to 75% with the mix of Layers. Thanks in adavance

Cross Validated 2019-04-19 06:06 UTC Score 17.0 AI-113-20190419-social-media-01e6f0a2

How to know which statistical test to use?

I'm working on a project and was wondering how I could perform statistical tests between two correlations to prove significance. The background of my project is that I was trying to see if Twitter sentiments influence a certain market more than the other. After collecting data and calculating the correlation values using three methods (Pearson, Kendall, and Spearman) I found that the difference is very small. However, I feel that simply saying "the difference is small" is not convincing and was wondering if there were specific statistical tests for this purpose. I've done searching and have seen that people use p-value tests quite often, but I'm not sure if that would be appropriate for my case, as I'm also not familiar with statistical testing methods. I was hoping if anyone more knowledgeable would be able to give me some tips or pointers. The data that I'm specifically using looks like this: To briefly explain the data, TextBlob and NLTK Vader are two sentiment analysis tools I used. Within these two methods, I applied time lag to the data (one day forward and one day backward) and you can also see that I varied calculating correlation. X and Y are two separate markets and the correlation is between price and sentiment . The difference is the difference between these two. My original hypothesis was that market Y would have higher correlation, but apparently that is not true judging by the difference. What might be some ways that I could formally prove the significance of…

Cross Validated 2018-10-12 14:11 UTC Score 12.0 AI-113-20181012-social-media-645ca4c3

Similarity index between two texts Ask Question

I'm trying to compare two vectors in a small NLP project using Python. Code doesn't make any difference since I'm using scikit-learn, but my doubts are about my calculations. I have a query vector and some texts, all vectors are constructed using the TF-IDF algorithm and the same corpus with the same preprocessing, including Porter stemmer. The problem is that no matter the texts nor the query vector, both cosine and euclidean distances chose the same text. For instance, let $u$ be the query vector and $v_1, v_2, ..., v_n$ be the vectorized texts, then cosine similarity selects the vector $v_k$ to be more similar to $u$ , and so euclidean similarity measure. So, it's weird because I'm getting different results of similarity for those vectors, but they still chose the same vector. So my questions are: Cosine and euclidean similarities select the same vectors always? Is mandatory to normalize the euclidean distance? Why cosine distance shouldn't be normalized? Notes: To calculate the similarity coefficient of euclidean distance I use the formula ${1}\over{1+euclidean\_distance(u,v)}$ and for cosine distance $1-cosine\_distance(u,v)$ . Note that normalized means that the sum of all components of a vector is 1.