AI/ML News & Innovations Hub

AI/ML news, top picks, and generated innovation digests.

★ Visit ai-karthik.com
422Sources
60663News Items
8Top Picks
322Blogs
failedLast Run

NLP

157 articles tagged with this keyword, sorted by most recent first.

← All Keywords
Synced 2026-09-28 08:17 UTC Score 37.0 AI-041-20260928-ai-specialis-39b7560a

Comment on Natural Language Processing In Early Education by ANIME SLAYER

Topic in a simple and easy--understand way. I also recently published a related resource that your readers might find . If you have a chance, I'd appreciate it if you could take a look and consider adding it as a reference if you think it adds value. Thanks for sharing such quality content, and keep up the great work! Anime Slayer download

Synced 2026-09-28 06:38 UTC Score 42.0 AI-041-20260928-ai-specialis-b27a26cc

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by MarkItDown Fan

Great article! I recently discovered MarkItDown (markitdown.tech), an excellent tool for converting files to Markdown. Highly recommend checking out their PDF to Markdown converter at markitdown.tech/pdf-to-markdown and their online Markdown editor at markitdown.tech/markdown-online. Also worth exploring their Microsoft Word to Markdown tool at markitdown.tech/microsoft-markitdown and Markdown to PDF at markitdown.tech/markdown-to-pdf. Amazing resource for developers!

Korea AI Times 2026-09-28 05:39 UTC Score 44.0 USR-0048-20260928-global-ai-ne-3d1bf71c

셀렉트스타, EMNLP서 AI 신뢰성 연구 5편 채택

셀렉트스타(대표 김세엽)는 AI 신뢰성·안전성 관련 논문 5편이 글로벌 자연어처리(NLP) 학회 \'EMNLP 2026\'에 채택됐다고 28일 밝혔다.채택 논문은 메인 컨퍼런스 1편, 파인딩스(Findings) 1편, 워크숍 3편이다. 특히 올해 메인 컨퍼런스에 제출된 1만 8000편의 논문 중 15.4%만 채택됐다는 점에서 자체 기술력을 인정받았다는 설명이다.메인 컨퍼런스에 채택된 \'정확도 일치와 불일치 관점에서의 인간–LLM 응답 정렬 진단: 대규모 집단 수준 분석\'은 AI가 정답을 맞히는지를 넘어 틀린 경우에도 사람과 유사한 오답

Synced 2026-09-23 16:56 UTC Score 51.0 AI-041-20260923-ai-specialis-e16ea677

Comment on Say Woof? AI in Animal Language Translation by MegaTool

This is a fascinating look at how far animal language translation has come since those early novelty translators. The Zoolingua and Conservation Metrics projects really show how deep learning can bridge the gap between species, whether it's helping pet owners understand their dogs or protecting elephants from poachers. I've been following similar AI translation and NLP developments over at MegaTool , and it's exciting to see how these technologies are evolving beyond human language to decode the natural world around us. The chicken emotion detection study is particularly clever. Thanks for the great read!

Towards Data Science 2026-09-23 14:00 UTC Score 22.0 AI-036-20260923-ai-specialis-c9f59ebc

From Words to Vectors: What Happens in Between?

A Journey through TF-IDF, vector space, and text classification The post From Words to Vectors: What Happens in Between? appeared first on Towards Data Science .

Synced 2026-09-18 08:04 UTC Score 59.0 AI-041-20260918-ai-specialis-1acbfefb

Comment on CMU’s Novel ‘ReStructured Pre-training’ NLP Approach Scores 40 Points Above Student Average on a Standard English Exam by Morgan

This research on ReStructured Pre-training is fascinating—scoring 40 points above student averages on standardized tests shows how much untapped potential exists in refining NLP training approaches. It makes me wonder what other exams or real-world tasks could benefit from this method. By the way, has anyone tried combining this with tools like AISuperRemover for cleaning training data?

Synced 2026-09-14 10:54 UTC Score 65.0 AI-041-20260914-ai-specialis-01425892

Comment on Google ‘mT5’ Pretrained Text-to-Text Transformer Achieves SOTA Performance on Multilingual Benchmarks by Kate

Interesting to see how mT5 is pushing multilingual NLP forward, especially with support for so many languages. I’ve been using TextNow lately, so it’s cool to think about how better multilingual models could eventually make communication across different languages feel more seamless.

Data Science Stack Exchange 2026-09-14 00:41 UTC Score 16.0 AI-111-20260914-social-media-a89ecaea

Can the scanned fly brain do NLP?

recently a fly brain was scanned and people have took to a million creative uses of it. Is there a fundamental reason that it wouldn’t be capable of being trained to do NLP?

Data Science Stack Exchange 2026-09-13 20:13 UTC Score 39.0 AI-111-20260913-social-media-22e87e92

Is building a manual evaluation set the right approach when no reliable labeled ground truth exists for resume-job matching?

I'm building a resume screening/ranking system (matching resumes to job descriptions using pretrained sentence embeddings + cosine similarity, no fine-tuning at this stage) as a learning project aimed at becoming a market-ready NLP practitioner. Problem: I could not find a trustworthy, publicly available English dataset with genuine human-labeled resume-job match scores. I checked several Kaggle/HuggingFace options and found labels that were either AI-generated (e.g., GPT-4o) or fully synthetic with demographic columns (race/ethnicity/gender) tied to the match label, which raised bias concerns. Academic literature (ConFit, PJFNN papers) confirms that the only broadly public dataset for this exact task (Person-Job Fit) is the 2019 Alibaba matching competition dataset, which is Chinese-only; most published research instead uses private company-provided data. My current approach: Use two real (non-synthetic) datasets: a scraped resume corpus and a real LinkedIn job postings corpus, both verified for low duplication and cleaned of PII. Build a small manual evaluation set myself (~24 resume-job pairs, selected to cover clear matches, clear non-matches, and ambiguous cases), scoring them on a 0-3 relevance scale with a confidence flag. Use this manual set as ground truth to compute ranking metrics (Precision@K, MRR, NDCG) once the embedding-based matching pipeline is built. Question: Is this a sound methodology given the lack of reliable public ground truth, or is there a better-e…

Data Science Stack Exchange 2026-09-12 11:21 UTC Score 21.0 AI-111-20260912-social-media-b1276bf5

Algorithms to cluster documents with similar boiler-plate inclusions

Suppose I have a large corpus of text documents (say 100-million - in a common human language - e.g. English). I suspect that various subsets of these documents are 'similar'... by this, I mean they appear to have been created using common templates (no-longer available to me - except by inference from the corpus) containing substantial boilerplate. After applying standard NLP tokenisation to 'words' - I define the common-runs of a pair of documents as the longest sequences of words found in both documents. I define valid-common-runs as the longest sequence of common-runs that occur in the same order in both documents. Hence if "Abracadabra Systems" was (only) present at the end of one document and the beginning of another, it would be a common-run... but it would only be a valid-common-run if there were no other (distinct) common-runs in the documents. My ultimate objective is to cluster the documents using valid-common-runs (perhaps weighted by a simple function ranging over the valid-common-run text). My question is about efficient ways to compute the 'valid-common-runs' and find groups of similar documents. A naive approach to finding common-runs might be to consider every pair of documents in turn... but 100-million documents would require 10-quadrillion document comparisons... which seems prohibitively resource intensive. I would like to know (before I try to re-invent wheels) are there any well-established efficient algorithms to determine anything similar to 'valid-c…

Synced 2026-09-09 11:43 UTC Score 40.0 AI-041-20260909-ai-specialis-14ec1599

Comment on AI Analyzes Chess Commentary to Learn to Play Chess by Grace

SentiMATE presents an interesting way of combining natural language processing with chess strategy. Using human commentary to evaluate move quality shows how AI can learn from language and improve decision-making without relying entirely on traditional chess features. Regulatory awareness is important before joining any online game , since rules vary by region. Players should confirm local eligibility before depositing funds.

iAfrica 2026-09-08 08:23 UTC Score 57.0 AI-151-20260908-regional-ai--0b3f922c

Tether Releases Open Offline Translation Models for 19 African Languages, With Peer-Reviewed Benchmarks

Tether AI Research has released open-source translation models covering 19 African languages that run entirely on smartphones and laptops without an internet connection — and, unusually for a corporate AI announcement, the underlying research has been accepted for presentation at EMNLP 2026, the leading peer-reviewed conference in natural language processing. The release comprises QVAC TranslatePsy-AfriSLM, [...]

Towards Data Science 2026-08-29 13:00 UTC Score 36.0 AI-036-20260829-ai-specialis-6912748d

RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need

Enterprise Document Intelligence [Vol.1 #B00] - Retrieval answers one kind of question. Classifying a request, matching free text to a reference list, reading a table, cleaning OCR noise: each has a cheaper method that works, and the engineering is knowing which one to reach for The post RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need appeared first on Towards Data Science .

Korea AI Times 2026-08-28 07:55 UTC Score 46.0 USR-0048-20260828-global-ai-ne-c2085060

"AI 양자화 정확하고 안전하게"...옵트에이아이, EMNLP 논문 2편 채택

옵트에이아이(대표 이재호)는 논문 2편이 자연어처리(NLP) 국제학술대회 \'EMNLP 2026\'의 메인 컨퍼런스와 인더스트리 트랙에 채택됐다고 28일 밝혔다.올해 행사에는 역대 최다인 1만7669편의 논문이 제출, 메인 컨퍼런스와 파인딩스 채택률은 각각 15.4%, 14.3%를 기록했다. 옵트에이아이는 이번 채택으로 EMNLP, ACL 등 주요 NLP 학회에 3회 연속 논문이 선정됐다.채택된 논문은 ▲LLM 양자화의 신규 기준을 제시한 \'커널퀀트\' ▲온디바이스 LLM의 안전성 공백을 최초로 규명한 연구 등 2편이다.첫 번째 \'커널퀀

Entrackr AI 2026-08-27 04:39 UTC Score 35.0 USR-0212-20260827-regional-new-011d1094

InstaAstro raises $12 Mn in Series A round

Gurugram-based digital spiritual wellness platform InstaAstro has raised $12 million in a Series A funding round led by Singularity AMC and Artha Venture Fund, with participation from the company’s founders. The round comes more than two years after InstaAstro raised Rs 18.5 crore ($2.2 million) in a pre-Series A round led by Artha Venture Fund. In November 2021, the company had raised Rs 3.2 crore in a seed round. The fresh capital will be used to expand InstaAstro’s presence in regional and international markets and invest in AI-led product development. The company plans to deploy AI for sentiment analysis and multilingual support to enhance user consultations. Founded in 2021 by Nitin Verma, InstaAstro started as an astrology consultation app and has since expanded into a broader spiritual wellness platform offering astrology, tarot, numerology, vastu, Pooja Seva and spiritual commerce. It currently has 12 million registered users and more than 5,000 verified experts across 183 countries. InstaAstro’s revenue more than doubled to Rs 111.26 crore in FY26 from Rs 52 crore in FY25. Its international business, driven largely by the Indian diaspora, contributes around 25% of its revenue. The company’s user base has also grown nearly 400% over the past two years. The company claims its annualised revenue has now crossed Rs 200 crore. InstaAstro will also expand its Pooja Seva and spiritual commerce businesses. Over the longer term, it aims to build an integrated platform coveri…

AI Stack Exchange 2026-08-26 13:53 UTC Score 24.0 AI-110-20260826-social-media-ba1ece8a

Prompting McGill-NLP/AfriqueLlama-8B

McGill-NLP/AfriqueLlama-8B seems to hallucinate quite a bit. It chooses conversational outputs and goes beyond what I ask it to. I have the hyperparameters below. generated_ids = model.generate( **inputs, max_new_tokens=max_new_tokens, do_sample=False, num_beams=1, pad_token_id=tokenizer.pad_token_id) Does anyone have advice to have the model strictly follow instructions, without fine-tuning (other than including few-shot examples)?

Synced 2026-08-22 20:05 UTC Score 59.0 AI-041-20260822-ai-specialis-60ccc23e

Comment on Google’s ALBERT Is a Leaner BERT; Achieves SOTA on 3 NLP Benchmarks by John Mitchell

I found this article interesting because it shows how ALBERT improved on BERT while using parameter-sharing and other techniques to reduce model size and memory demands. The results across GLUE, SQuAD 2.0, and RACE are especially useful for understanding how NLP models can become more efficient without sacrificing performance. It also made me think about how optimization and efficiency matter in many technology-driven services today. For anyone interested in practical service providers, https://abudhabimoversandpackers.com/ is another example of a business worth exploring when reliability and organized service are important.

SiliconANGLE AI 2026-08-18 13:00 UTC Score 56.0 USR-0127-20260818-global-ai-ne-33828247

NLPatent rebrands as Clerq, launches agentic patent research workflows

Toronto patent research startup NLPatent rebranded as Clerq today and launched its first agentic workflows. The software runs a full patentability analysis in about 10 minutes. The company says every conclusion arrives with citations. That work normally takes days or weeks, and firms usually push it down to a junior associate or out to a search […] The post NLPatent rebrands as Clerq, launches agentic patent research workflows appeared first on SiliconANGLE .

Cross Validated 2026-08-17 21:16 UTC Score 47.0 AI-113-20260817-social-media-e4b59918

How can I reduce subjectivity when selecting posts for social media comment analysis?

I’m working on a computational social science project using comments from Weibo and Douyin (Chinese TikTok) for text and sentiment analysis. My study looks at how online discourse around a female public figure changed before and after a particular event. I first define several keywords related to the person and the event, search for relevant posts/videos, and then collect comments from selected posts. The problem I’m struggling with is: how should I systematically decide which posts to select? Taking the first 10 or 20 search results does not seem reliable, because search results on these platforms are algorithmically ranked and change between searches. The posts ranked at the top are also not necessarily the ones with the most discussion. Selecting posts purely by comment count is also problematic. Some videos have thousands of comments, but many are just emojis, users tagging friends, repeated/copied comments, or spam. Meanwhile, another video with fewer comments may contain much more substantive discussion relevant to my research. So in practice, I inevitably have to make a human judgment about whether a comment section contains enough meaningful, relevant text to analyze. But this creates the methodological problem I’m most worried about: Why these particular posts? If I selected another set of relevant posts, would I still get the same result? Even if I eventually collect millions of comments, I feel that a large number of comments does not solve the problem if I cannot…

CIO AI 2026-08-13 20:27 UTC Score 39.0 USR-0125-20260813-global-ai-ne-0b3668ac

Manual vs. AI-powered PDF redaction: protecting sensitive data in 2026

Research shows that humans play a role in 60% of breaches that expose sensitive data. That “role” often involves an employee falling for a phishing scam or using PASSWORD for their login credentials, but data exposure can also be a result of how your business redacts sensitive and personally identifiable information (PII) in your documents. Historically, manual, “black-box” redaction was considered best-practice, but this approach only obscures data, it doesn’t permanently remove it. As regulations governing data security get stricter and AI-powered redaction solutions become more accessible, organizations—especially those in highly regulated industries—are re-evaluating their PDF redaction solutions. How manual PDF redaction is different from AI-powered PDF redaction The primary difference between manual and AI-powered PDF redaction is who (or what) you rely on to do the heavy lifting. Manual redaction defined Manual redaction is a human-driven process where individuals visually scan text, select content, and apply black boxes or remove the text before sharing or storing the file. Manual redaction is only as effective as the reviewer, which makes outcomes highly variable, especially under time pressure or high document volume. AI-powered redaction defined AI-powered redaction uses machine learning and natural language processing (NLP) to automatically detect and remove sensitive information from documents. The system is trained to recognize patterns, language cues, and cont…

ACL Anthology 2026-08-11 00:00 UTC Score 15.0 AI-079-20260811-research-pap-c39af0f6

AMI @ EVALITA2020: Automatic Misogyny Identification

Elisabetta Fersini, Debora Nozza and Paolo Rosso in Proceedings of the Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2020)

ACL Anthology 2026-08-11 00:00 UTC Score 15.0 AI-079-20260811-research-pap-5af4404b

Automatic Expansion of Lexicons for Multilingual Misogyny Detection

Simona Frenda, Bilal Ghanem, Estefanía Guzmán-Falcón, Manuel Montes-y-Gómez and Luis Villaseñor-Pineda in Proceedings of the Sixth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA 2018)

Synced 2026-08-08 09:39 UTC Score 40.0 AI-041-20260808-ai-specialis-9447d79f

Comment on Facebook Says Its ‘Blender’ Chatbot Is the Most Humanlike Ever by paramount books

Facebook’s Blender chatbot is an impressive step toward making AI conversations feel more natural and humanlike. As AI continues to improve, it could also transform how people discover educational resources, get recommendations, and interact with online stores such as Paramount Books . It will be interesting to see how this technology develops and shapes digital communication in the future.

CIO AI 2026-08-07 09:30 UTC Score 50.0 USR-0125-20260807-global-ai-ne-6fbe4864

How AI is changing the business analyst role for the better

AI’s impact has been felt across nearly every industry, and its rise has already started to alter several roles in tech, including that of the business analyst . While the rise of agentic AI may have some questioning whether AI will replace business analyst jobs entirely, as we’ve seen with most roles impacted by AI, it’s more likely that AI will augment the role and fundamentally change how BA’s conduct daily business. “As AI takes on more routine tasks, the human side of the role is becoming even more valuable. It’s becoming more of a hybrid role, where employers are often looking for candidates who can combine technical fluency with strong communication and problem-solving skills, along with sound business judgment,” says Megan Slabinski, district president of technology talent solutions at Robert Half. AI can save business analysts time in the long run, automating many of the tasks that are time consuming and repetitive around data processing, note taking, and documentation. While automation will impact the daily tasks of the role, business analysts will still be necessary for properly interpreting outputs, collaborating across teams, and maintaining compliance and AI workflows. AI-driven analysis and automated workflows With AI-driven analysis, BA’s can use machine learning models for pattern detection, determining risk, and for forecasting demand, while natural language processing (NLP) can be used for text-heavy inputs. AI tools can also assist analysts with decision-…

Apple Machine Learning Research 2026-08-07 00:00 UTC Score 49.0 AI-059-20260807-official-ai--3d79e525

Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models

Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks, including document processing and code generation. Autoregressive Language Models (ARMs), which generate tokens sequentially conditioned on all previous tokens, have been the predominant paradigm for LLMs. While these models have achieved high accuracy across a range of downstream tasks, they exhibit low arithmetic intensity due to the inherent sequential dependency in next-token prediction. Recently, Diffusion Language Models (DLMs) have emerged as a promising…

ACL Anthology 2026-07-31 00:00 UTC Score 13.0 AI-079-20260731-research-pap-82652bf0

CSULoRA: Closest Safe Update Low-Rank Adaptation

Oleksandr Marchenko, Adelaide Danilov, Aria Nourbakhsh and Salima Lamsiyah in Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security

Synced 2026-07-29 14:17 UTC Score 37.0 AI-041-20260729-ai-specialis-a8632e0f

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by juego copero

Great article on FPDT! The memory hierarchy optimization is really impressive. As someone who runs a gaming site, I can appreciate efficiency improvements—smooth gameplay depends heavily on resource management. Check out https://juegocopero.com/ for more on that!

LessWrong AI 2026-07-25 01:19 UTC Score 68.0 USR-0152-20260725-community-fo-03cc5338

Linear probes tell you where quantization will hurt

Epistemic status: I have only tested one encoder family (BERT-base and its relatives) and one decoder LLM ( Qwen2.5 -3B), one seed, token-level tasks, and post-training weight quantization. I trust the results because I am not inventing anything; I am just connecting quantization with a very general idea from mech interp: the model does the easy syntax work first, the semantic work in later layers, and the prompt-relevant task work in the latest layers. That's the main idea I lean on in this project. What I'm not sure of is whether it holds when the task is diffuse, like general-knowledge QA (probably not). TL;DR. First I train a linear probe on each layer and try to see where a signal lives in a model. I only probe for classical NLP or CV tasks, like checking for depth or vision for CV or named entity recognition, parts of speech, or chunking for NLP. Then using that information to create a map of where the important work is happening and quantizing the less important layers for a task. This method is cheap; it's only one RidgeClassifier per layer. And it kind of works; a map built once on CoNLL news keeps 99–100% of full-precision accuracy at a 5-bit average on three unseen datasets, whereas Uniform kept 16–41% and opposite/anti layers kept 2–41%. This method is perfect for sharp use cases where the transformer needs specialized knowledge and can compromise on general knowledge. The itch So this all started during my 2nd semester at Northeastern, where I took Applied Progr…

Synced 2026-07-19 01:55 UTC Score 46.0 AI-041-20260719-ai-specialis-c7be6d26

Comment on Peking U & Microsoft’s Knowledge Attribution Method Enables Editing Factual Knowledge in Pretrained Transformers Without Fine-Tuning by Markdown to Doc

Impressive progress in editing factual knowledge without fine-tuning! The idea of targeting "knowledge neurons" is fascinating. For a deeper dive into transformer mechanics, check out this Markdown to Doc resource—it helps analyze such groundbreaking NLP innovations efficiently.

Data Science Stack Exchange 2026-07-16 17:56 UTC Score 29.0 AI-111-20260716-social-media-49c70074

Improving short‑text classification accuracy with overlapping classes and imbalanced data

I’m working on a multi‑class text classification problem where the input consists of very short descriptions (often only a few words) and the goal is to predict the correct category. I’m currently using an XGBoost classifier for the final prediction layer. for embeddings I am using e5 large model. The main challenges I’m facing are: The descriptions are extremely short and many classes share similar vocabulary. Because of this, the model sometimes produces high‑confidence but incorrect predictions when common words appear across multiple classes. Description is the only feature that supports the category classification. The dataset is imbalanced: some classes have significantly more training samples than others, but the distribution of the prediction data is different. This causes the model to over‑predict certain classes. I experimented with TF‑IDF features, embeddings from a generative AI model and a combination of TF‑IDF + embeddings. However, combining both actually reduced accuracy. I also tried downsampling majority class but it not really make any difference. Current model accuracy is around 30%. I’m looking for advice on: How to improve classification when different classes share highly overlapping vocabulary. How to reduce high‑confidence wrong predictions. What techniques work well for class imbalance and distribution shift between training and prediction data.

Machine Learning Mastery 2026-07-15 12:00 UTC Score 27.0 AI-039-20260715-ai-specialis-f38e483d

Scikit-Ollama for Scikit-LLM/Ollama Integration

In this article, you will learn how scikit-ollama bridges the scikit-learn interface with locally running Ollama models to perform zero-shot text classification; no cloud API...

Synced 2026-07-11 04:11 UTC Score 40.0 AI-041-20260711-ai-specialis-71f860d6

Comment on Microsoft’s Fully Pipelined Distributed Transformer Processes 16x Sequence Length with Extreme Hardware Efficiency by taichiwalk.org

Fascinating to see how Microsoft is pushing hardware efficiency for longer sequence processing—definitely a leap for AI scalability. All that computational intensity makes me think about the other side of the coin: finding calm after a deep tech session. Lately I’ve been exploring tai chi walking as a low-impact way to reset focus and improve balance, especially since sitting at a desk for hours can take a toll. There’s a beginner-friendly guide at taichiwalk.org that offers free routines and even a guided coach—no login needed, which I appreciate. Anyone else here use movement or gentle exercise to counterbalance screen time?

Synced 2026-07-07 07:10 UTC Score 51.0 AI-041-20260707-ai-specialis-3f6bf0e5

Comment on Interview with Tencent AI Lab- Four Core Research Fields:Computer Vision, Speech Recognition, Natural Language Processing and Machine Learning by Owen Parker

I've seen examples where clear tracking information greatly improves the user experience. If you're interested in how a modern tracking system presents shipment updates and delivery statuses, this guide is worth a look: https://anposttracking.org/ . It's a helpful reference for anyone building customer-focused applications

Transactions on Machine Learning Research 2026-07-06 00:00 UTC Score 49.0 AI-084-20260706-research-pap-960a167b

Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection

Attention mechanisms have revolutionized several domains of artificial intelligence, such as natural language processing and computer vision, by enabling models to selectively focus on relevant parts of the input data. While recent work has characterized the optimization dynamics of gradient descent (GD) in attention-based models and the structural properties of its preferred solutions, less is known about more general optimization algorithms such as mirror descent (MD). In this paper, we investigate the convergence properties and implicit biases of a family of MD algorithms tailored for softmax attention mechanisms, with the potential function chosen as the $p$-th power of the $\ell_p$-norm. Specifically, we show that these algorithms converge in direction to a generalized hard-margin SVM with an $\ell_p$-norm objective when applied to a classification problem using a softmax attention model. Notably, our theoretical results reveal that the convergence rate is comparable to that of traditional GD in simpler models, despite the highly nonlinear and nonconvex nature of the present problem. Additionally, we delve into the joint optimization dynamics of the key-query matrix and the decoder, establishing conditions under which this complex joint optimization converges to their respective hard-margin SVM solutions. Lastly, our numerical experiments on real data demonstrate that MD algorithms improve generalization over standard GD and excel in optimal token selection.

Transactions on Machine Learning Research 2026-07-06 00:00 UTC Score 40.0 AI-084-20260706-research-pap-bedebbba

Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective

The Transformer model is widely used in various application areas of machine learning, such as natural language processing. This paper investigates the approximation of the Hölder continuous function class $\mathcal{H}_{Q}^{\beta}\left([0,1]^{d\times n},\mathbb{R}^{d\times n}\right)$ by Transformers and constructs several Transformers that can overcome the curse of dimensionality. These Transformers consist of one self-attention layer with one head and the softmax function as the activation function, along with several feedforward layers. For example, to achieve an approximation accuracy of $\epsilon$, if the activation functions of the feedforward layers in the Transformer are ReLU and floor, only $\mathcal{O}\left(\log\frac{1}{\epsilon}\right)$ layers of feedforward layers are needed, with widths of these layers not exceeding $\mathcal{O}\left(\frac{1}{\epsilon^{2/\beta}}\log\frac{1}{\epsilon}\right)$. If other activation functions are allowed in the feedforward layers, the width of the feedforward layers can be further reduced to a constant. These results demonstrate that Transformers have a strong expressive capability. The construction in this paper is based on the Kolmogorov-Arnold Superposition Theorem and does not require the concept of contextual mapping, hence our proof is more intuitively clear compared to previous Transformer approximation works. Additionally, the translation technique proposed in this paper helps to apply the previous approximation results of fe…

Machine Learning Mastery 2026-06-16 12:00 UTC Score 27.0 AI-039-20260616-ai-specialis-e9483392

Building an End-to-End Sentiment Analysis Pipeline with Scikit-LLM

Traditional machine learning pipelines for predictive tasks like text classification usually rely on extracting structured, numerical features from raw text — for instance, TF-IDF frequencies or token embeddings — to feed into classical models such as logistic regression, ensembles, or support vector machines.

Machine Learning Mastery 2026-06-11 12:00 UTC Score 18.0 AI-039-20260611-ai-specialis-824e0fa0

Multi-Label Text Classification with Scikit-LLM

Text classification typically boils down to scenarios where a product review is "positive" or "negative", or a customer inquiry belongs to one category or another.

AI Stack Exchange 2026-03-13 09:45 UTC Score 26.0 AI-110-20260313-social-media-d58701e5

Why do Transformers handle long-range dependencies better than LSTMs despite lacking explicit recurrence?

Recurrent architectures such as LSTMs and GRUs were originally designed to address the vanishing gradient problem and capture long-range dependencies in sequential data. However, in recent years Transformer-based architectures have largely replaced RNN-based models in many domains such as natural language processing, time-series modeling, and even reinforcement learning. One commonly cited explanation is that Transformers allow parallel computation and avoid sequential processing, which improves training efficiency. However, this does not fully explain why they often outperform LSTMs in modeling long-range relationships. From a modeling perspective, I am trying to understand the following points: In LSTMs, the cell state is explicitly designed to propagate information across time steps. In Transformers, there is no recurrence or persistent state between tokens. Why does the self-attention mechanism still capture long-range dependencies more effectively? Is the improvement mainly due to the attention mechanism allowing direct connections between distant tokens, or are there additional factors such as representation capacity and optimization dynamics? Are there known theoretical explanations or empirical studies comparing the ability of Transformers and LSTMs to capture long-range dependencies? Are there scenarios (for example streaming data or low-resource environments) where recurrent architectures still outperform Transformers? I would appreciate references to research pape…

AI Stack Exchange 2024-06-25 00:58 UTC Score 21.0 AI-110-20240625-social-media-6dfd5a00

How do transformer models handle negation in sentiment analysis

I'm trying to understand how transformer models, such as BERT or GPT, handle negation in sentiment analysis. Specifically, I'm curious about how these models manage to correctly interpret sentences where negation changes the sentiment, such as "The movie is not good." A simple model using word embeddings + global averaging fails to handle negation properly. Intuitively, for example, if "good" has a positive sentiment score and "bad" has a negative sentiment score, a model might misinterpret "not good" by simply averaging the scores of "not" and "good". Example Without Negation Consider the following sentences with sentiment words: "The movie is good." "The movie is awesome." "The movie is terrible." Suppose we have the following word embeddings representing sentiment scores: "good" = [10] "awesome" = [12] "terrible" = [-10] Neutral words (assuming embeddings around 0): "the" = [0] "movie" = [0] "is" = [0] For these sentences, a simple global average of the sentiment scores works well: "The movie is good" = average([0, 0, 0, 10]) = 10 / 4 = 2.5 (positive sentiment) "The movie is awesome" = average([0, 0, 0, 12]) = 12 / 4 = 3 (positive sentiment) "The movie is terrible" = average([0, 0, 0, -10]) = -10 / 4 = -2.5 (negative sentiment) Example With Negation Now, consider the sentence "The movie is not good." In this case, the sentiment should be negative due to the presence of "not." However, averaging the scores naively might not handle this correctly. For example: "The movie is…

AI Stack Exchange 2024-05-30 06:02 UTC Score 29.0 AI-110-20240530-social-media-da3d22c6

Probability interpretation of attention mechanism in Seq2Seq

I have ready many explanations of the seq2seq model. In my opinion, however, it is really like a robot that might say something correctly, but doesn't really understand it, just as is true with an LLM generally. In my opinion, the correct way to describe Seq2Seq and similar NLP models should start from a probability view. My probability view is very simple; the output of the encoder is a representation of the probability distribution of the next word. In each step of the Decoder, it just modifies the distribution based on each word it predicted from the distribution and outputs the modified distribution. It then does this repeatedly. Assuming this probability view is correct, how could we explain the attention mechanism used in Seq2seq?

AI Stack Exchange 2022-11-26 20:53 UTC Score 24.0 AI-110-20221126-social-media-05252ea3

What is the difference between prompt tuning and prefix tuning?

I read prompt tuning and prefix tuning are two effective mechanisms to leverage frozen language models to perform downstream tasks. What is the difference between the two and how they work really? Prompt Tuning: https://aclanthology.org/2021.emnlp-main.243/ Prefix-Tuning: https://arxiv.org/abs/2101.00190

Data Science Stack Exchange 2022-08-11 00:20 UTC Score 29.0 AI-111-20220811-social-media-485f9224

Difference Between Attention and Fully Connected Layers in Deep Learning

There have been several papers in the last few years on the so-called "Attention" mechanism in deep learning (e.g. 1 2 ). The concept seems to be that we want the neural network to focus on or pay more attention to certain features, and has demonstrated some empirical success in NLP and related sequential models. When I look at some code examples such as this one , adding an Attention layer intuitively makes sense and seems to improve performance of the LSTM model. However it looks very much like a regular fully-connected layer. In that link (and with some slight change of notation), the Attention layer outputs $$ c(x) = \tanh(\mathbf{W}x + \mathbf{b} ) $$ $$ \beta(x) = \frac{e^{c(x_j)}}{\sum_{j} e^{c(x_j)}} $$ $$ f_{Attention}(x) = x\beta $$ where $W,b$ are weights/biases, $x$ is layer input, and $f(.)$ is the layer output. In contrast, a regular fully connected layer: $$ f_{Dense} = \sigma(\mathbf{W}x + \mathbf{b}) $$ for some activation function $\sigma(.)$ . My interpretation of the Attention implementation above is that it is pretty much the same thing as a standard fully connected layer, but with a $\tanh$ activation (why?), followed by a $\text{softmax}$ (okay, so that the "attention weights" $\beta$ sum to 1), followed by a linear dot product. How does this architecture allow the model to have "attention"? I do not see how it is fundamentally different or more expressive from just adding a standard fully-connected layer. Am I misunderstanding something here? Edit/My…

Stanford AI Lab Blog 2022-05-31 07:00 UTC Score 47.0 USR-0006-20220531-research-aca-a57ebba7

LinkBERT: Improving Language Model Training with Document Link

Language Model Pretraining Language models (LMs), like BERT 1 and the GPT series 2 , achieve remarkable performance on many natural language processing (NLP) tasks. They are now the foundation of today’s NLP systems. 3 These models serve important roles in products and tools that we use every day, such as search engines like Google 4 and personal assistants like Alexa 5 . These LMs are powerful because they can be pretrained via self-supervised learning on massive amounts of text data on the web without the need for labels, after which the pretrained models can be quickly adapted to a wide range of new tasks without much task-specific finetuning. For instance, BERT is pretrained to predict randomly masked words in original text (masked language modeling), e.g. predicting the masked word “dog” from “My __ is fetching the ball”. GPTs are pretrained to predict the next word given a previous sequence of text (causal language modeling), e.g. predicting the next word “ball” from “My dog is fetching the”. In either cases, through pretraining, LMs learn to encode various knowledge from a text corpus that helps to perform downstream applications involving language understanding or generation. In particular, LMs can learn world knowledge (associations between concepts like “dog”, “fetch”, “ball”) from training text where the concepts appear together, and help for knowledge-intensive applications like question answering. 6 Challenges. A challenge with most common LM pretraining strateg…

Stanford AI Lab Blog 2022-05-25 07:00 UTC Score 47.0 USR-0006-20220525-research-aca-2eecb290

Stanford AI Lab Papers and Talks at ACL 2022

The 60th Annual Meeting of the Association for Computational Linguistics (ACL) 2022 is taking place May 22nd - May 27th. We’re excited to share all the work from SAIL that’s being presented, and you’ll find links to papers, videos and blogs below. Feel free to reach out to the contact authors directly to learn more about the work that’s happening at Stanford! List of Accepted Papers LinkBERT: Pretraining Language Models with Document Links Authors : Michihiro Yasunaga, Jure Leskovec*, Percy Liang* Contact : myasu@cs.stanford.edu Links: Paper | Website Keywords : language model, pretraining, knowledge, hyperlink, bionlp When classifying grammatical role, BERT doesn’t care about word order… except when it matters Authors : Isabel Papadimitriou, Richard Futrell, Kyle Mahowald Contact : isabelvp@stanford.edu Links: Paper Keywords : large language models, analysis, word order, order invariance, grammatical role, syntax, semantics Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words Authors : Kaitlyn Zhou, Kawin Ethayarajh, Dallas Card, Dan Jurafsky Contact : katezhou@stanford.edu Keywords : cosine similarity, training data frequency, model analysis Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive Summarization Authors : Faisal Ladhak, Esin Durmus, He He, Claire Cardie, Kathleen McKeown Contact : esdurmus@stanford.edu Links: Paper Keywords : text summarization, text generation, evaluation, faithfulness Sp…

AI Stack Exchange 2022-04-05 17:07 UTC Score 24.0 AI-110-20220405-social-media-ec190fe9

What are examples of simple gradient based NLP models?

I am looking for a simple example of gradient-based methods for NLP. More specifically I am looking for post-hoc local explanations gradient-based methods, that is to say, which explain a single prediction by performing additional operations (after the model has emitted a prediction). Here is an example of a gradient-based method for "Finance"? (I'm not very sure of what domain it is as it is not NLP, nor Computer Vision, nor time series ...): Let us consider a minimal example, where a linear regression is used to estimate the future capital asset $y_{c}$ , based on two investments $x_{1}$ and $x_{2}$ . Let assume the assumptions above are met and the model parameters are estimated as follows: $$\mathbb{E}\left[y_{c} \mid x_{1}, x_{2}\right]=1.05 x_{1}+1.50 x_{2},$$ We can derive immediately a global interpretation of this model. Every dollar invested in fund $x_{1}$ will produce capital of $1.05 \\\$$ , while every dollar invested in $x_{2}$ will produce a capital of $1.50 \\\$$ , independently of the values $x_{1}$ and $x_{2}$ might assume in a concrete scenario. Notice that this explanation is purely based on the learned coefficient $w_{1}=1.05$ and $w_{2}=1.50$ . These, sometimes called partial regression coefficients, are themselves candidate attribution values to explain the influence of the independent variables of the target variable: $$R_{1}(x)=1.05 \quad R_{2}(x)=1.50$$ Notice also that the coefficients are the partial derivatives of the target variable with respec…

AI Stack Exchange 2022-03-15 17:05 UTC Score 29.0 AI-110-20220315-social-media-22338a14

Having the negative cases in the same batch vs. shuffling the dataset

I am working on a model for an NLP task. The model encodes the text and has a regression output layer. In this task, from each instance (positive), I create several negative cases using a specific technique and I merge them with their positive corresponding ones in a data split (training/val/test). After that, I shuffle the data split. I was thinking of the following: Isn't better to keep the negative instances with their corresponding positive ones in the same batch instead of shuffling the data? Is there an answer to this question? does it depend on the task?

Stanford AI Lab Blog 2022-01-21 08:00 UTC Score 58.0 USR-0006-20220121-research-aca-1e3c1829

Reward Isn't Free: Supervising Robot Learning with Language and Video from the Web

This work was conducted as part of SAIL and CRFM . Deep learning has enabled improvements in the capabilities of robots on a range of problems such as grasping 1 and locomotion 2 in recent years. However, building the quintessential home robot that can perform a range of interactive tasks, from cooking to cleaning, in novel environments has remained elusive. While a number of hardware and software challenges remain, a necessary component is robots that can generalize their prior knowledge to new environments, tasks, and objects in a zero or few shot manner. For example, a home robot tasked with setting the dining table cannot afford lengthy re-training for every new dish, piece of cutlery, or dining room it may need to interact with. A natural way to enable such generalization in our robots is to train them on rich data sources that contain a wide range of different environments, tasks, and objects. Indeed, this recipe of massive, diverse datasets combined with scalable offline learning algorithms (e.g. self-supervised or cheaply supervised learning) has been the backbone of the many recent successes of foundation models 3 in NLP 4 5 6 7 8 9 and vision 10 11 12 . Replicating these impressive generalization and adaptation capabilities in robot learning algorithms would certainly be a step toward robots that can be used in unstructured, real world environments. However, directly extending this recipe to robotics is nontrivial, as we neither have sufficiently large and diverse…

AI Stack Exchange 2021-11-23 15:44 UTC Score 9.0 AI-110-20211123-social-media-2131c0fb

How much labelling is required for NER with SpaCy?

I have transaction data and I would like to extract the merchant from the transaction description. I am new to this but I just came across Named Entity Recognition and SpaCy. I have hundreds of thousands of different merchants. Some questions that I have: How much labelling do I need to do given the number of merchants I need to extract? How many different instances of the same merchant I need to label to get decent results?

Stanford AI Lab Blog 2021-11-05 07:00 UTC Score 52.0 USR-0006-20211105-research-aca-02e23852

Stanford AI Lab Papers at EMNLP/CoNLL 2021

The 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021) will take place next week, colocated with CoNLL 2021. We’re excited to share all the work from SAIL that will be presented, and you’ll find links to papers, videos and blogs below. Feel free to reach out to the contact authors directly to learn more about the work that’s happening at Stanford! List of Accepted Papers Calibrate your listeners! Robust communication-based training for pragmatic speakers Authors : Rose E. Wang, Julia White, Jesse Mu, Noah D. Goodman Contact : rewang@stanford.edu Links: Paper | Video Keywords : language generation, pragmatics, communication-based training, calibration, uncertainty Cross-Domain Data Integration for Named Entity Disambiguation in Biomedical Text Authors : Maya Varma, Laurel Orr, Sen Wu, Megan Leszczynski, Xiao Ling, Christopher Ré Contact : mvarma2@stanford.edu Links: Paper | Video Keywords : named entity disambiguation, biomedical text, rare entities, data integration ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts Authors : Yuta Koreeda, Christopher D. Manning Contact : koreeda@stanford.edu Links: Paper | Website Keywords : natural language inference, contract, law, legal, dataset Venue : The Findings of EMNLP 2021 The Emergence of the Shape Bias Results from Communicative Efficiency Authors : Eva Portelance, Michael C. Frank, Dan Jurafsky, Alessandro Sordoni, Romain Laroche Contact : portelan@stanford.edu Links…

Stanford AI Lab Blog 2021-10-13 07:00 UTC Score 52.0 USR-0006-20211013-research-aca-a7eb5ef1

Selective Classification Can Magnify Disparities Across Groups

Selective classification, where models are allowed to “abstain” when they are uncertain about a prediction, is a useful approach for deploying models in settings where errors are costly. For example, in medicine, model errors can have life-or-death ramifications, but abstentions can be easily handled by backing off to a doctor, who then makes a diagnosis. Across a range of applications from vision 1 2 3 and NLP 4 5 , even simple selective classifiers, relying only on model logits, routinely and often dramatically improve accuracy by abstaining. This makes selective classification a compelling tool for ML practitioners 6 7 . However, in our recent ICLR paper, we find that despite reliably improving average accuracy, selective classification can fail to improve and even hurt the accuracy over certain subpopulations of the data . As a motivating example, consider the task of diagnosing pleural effusion, or fluid in the lungs, from chest X-rays. Pleural effusion is often treated with a chest drain, so many pleural effusion cases also have chest drains, while most cases without pleural effusion do not have chest drains 8 . While selective classification improves average accuracy for this task, we find that it does not appreciably improve accuracy on the most clinically relevant subgroup, or subpopulation, of the data: those that have pleural effusion but don’t yet have a chest drain, i.e. those that have pleural effusion but have not yet been treated for it. Practitioners, thus,…

AI Stack Exchange 2021-03-28 03:17 UTC Score 21.0 AI-110-20210328-social-media-17a262ae

Can an existing transformer model be modified to estimate the next most probable number in a sequence of numbers?

Models based on the transformer architectures (GPT, BERT, etc.) work awesome for NLP tasks including taking an input generated from words and producing probability estimates of the next word as the output. Can an existing transformer model, such as GPT-2, be modified to perform the same task on a sequence of numbers and estimate the next most probable number? If so, what modifications do we need to perform (do we still train a tokenizer to tokenize integers/floats into token IDs?)?

Lilian Weng Blog 2021-03-21 00:00 UTC Score 39.0 USR-0112-20210321-ai-specialis-3e60cc8a

Reducing Toxicity in Language Models

Large pretrained language models are trained over a sizable collection of online data. They unavoidably acquire certain toxic behavior and biases from the Internet. Pretrained language models are very powerful and have shown great success in many NLP tasks. However, to safely deploy them for practical real-world applications demands a strong safety control over the model generation process.

AI Stack Exchange 2021-03-04 11:35 UTC Score 18.0 AI-110-20210304-social-media-c6ad9850

CAPTCHA based on text comprehension and random tokens

I developed a novel type of CAPTCHA based on text comprehension and random tokens . Given a task Pick the first pair of adjacent letters and a random token 8NBA596V , the user has to provide the solution NB . It offers basic protection and an attacker can solve individual tasks with specific effort. I am curious, whether contemporary AI can solve it generically ? You can access more example tasks here: https://www.topincs.com/manual/captcha There is a task database and at every attempt a new task is presented with a new random token. They always have a solution of varying length and pure guessing thus has limited chances of success. It is easy to attack an individual task by writing a small piece of code, thus a large task database is essential. What intrigues me is the question whether natural language processing or machine learning at its current state can attack the CAPTCHA generically by building a model of the meaning of the task – essentially a predicate in a tiny universe of discourse – and then applying it to the random token.

Anyscale Blog 2021-02-10 00:00 UTC Score 39.0 USR-0085-20210210-ai-specialis-e1e07003

Retrieval Augmented Generation with Huggingface Transformers and Ray

Huggingface Transformers recently added the Retrieval Augmented Generation (RAG) model, a new NLP architecture that leverages external documents (like Wikipedia) to augment its knowledge and achieve state of the art results on knowledge-intensive tasks. In this blog post, we introduce the integration of Ray, a library for building scalable applications, into the RAG contextual document retrieval mechanism. This speeds up retrieval calls by 2x and improves the scalability of RAG distributed fine-tuning.

AI Stack Exchange 2021-01-13 13:34 UTC Score 23.0 AI-110-20210113-social-media-8c81b4b2

How to extract parameters from a text using AI/NLP

lets say I have three texts: "make a heading that says hello word" "make a heading of hello world" "create heading consist of hello world" How can I fetch those groups of words using AI which is referring to heading i.e hello world in this case. Which AI frameworks or libraries can do that? in all examples heading is pointing to hello world (which i am referring as group of words). so basically i want those words which will be a part of heading or in other word there is a relationship between them. another example i can give is "I am watching Breaking bad" so there is a relationship between watching and breaking bad and i want to extract what are you watching. What's the best approach? Do I have to train a model for that or there are some other techniques that can get it done?

Jay Alammar Blog 2020-12-17 00:00 UTC Score 39.0 USR-0113-20201217-ai-specialis-fb351fb3

Interfaces for Explaining Transformer Language Models

Interfaces for exploring transformer language models by looking at input saliency and neuron activation. Explorable #1: Input saliency of a list of countries generated by a language model Tap or hover over the output tokens: Explorable #2: Neuron activation analysis reveals four groups of neurons, each is associated with generating a certain type of token Tap or hover over the sparklines on the left to isolate a certain factor: The Transformer architecture has been powering a number of the recent advances in NLP. A breakdown of this architecture is provided here . Pre-trained language models based on the architecture, in both its auto-regressive (models that use their own output as input to next time-steps and that process tokens from left-to-right, like GPT2) and denoising (models trained by corrupting/masking the input and that process tokens bidirectionally, like BERT) variants continue to push the envelope in various tasks in NLP and, more recently, in computer vision. Our understanding of why these models work so well, however, still lags behind these developments. This exposition series continues the pursuit to interpret and visualize the inner-workings of transformer-based language models. We illustrate how some key interpretability methods apply to transformer-based language models. This article focuses on auto-regressive models, but these methods are applicable to other architectures and tasks as well. This is the first article in the series. In it, we present explo…

Cross Validated 2020-12-12 20:09 UTC Score 26.0 AI-113-20201212-social-media-cb7b55bc

Accuracy is 100% but model.predict is totally wrong! what could be the problem? (Autoencoder NN) [closed]

I have an Autoencoder model that receives 4 different vectors { (1,0,0,0) , (0,1,0,0) , (0,0,1,0) , (0,0,0,1) } The encoder transforms the vectors to vectors size 2, which get inside an "NLPN channel" which is a channel with nonlinear phase noise, and then goes to the decoder that reconstructs the original vector. Architecture: 4->50->2->NLPN->2->50->4 (Softmax) The loss function is mse between the input, and the input ( self.model.fit(self.x, self.x) ) The code is pretty basic: the first cell is just some simple functions defined, then it's the autoencoder class, then it's the training, and then it's the prediction which shows that for each vector 1,0,0,0 and so on, the decoder predicts (~0.25,~0.25,~0.25,~0.25) which shows that the NN has no idea what it's doing. CODE: https://colab.research.google.com/drive/1nFymmEloUSa1yjS7e5CPGsKwqP1HjS57?usp=sharing

Cross Validated 2020-11-11 05:16 UTC Score 15.0 AI-113-20201111-social-media-b6871e58

conceptually, how to use NLP to predict a numeric output

suppose I'm trying to use medical notes to predict the cost of medical service. For example, a patient will call in, tell the operator how they feel, their diagnosis, etc etc, and the operator will take down all the notes. Over time, as the patient visits the doctor, gets medication, treatment, they would incur the final cost of the service. My job is to predict that final cost (as close as possible) based on the initial notes. I'm familiar with how to do standard EDA on the text; word frequency, sentiment, n-grams etc etc, but I dont know how that translates to a numerical output. As an analogy, if this was a linear regression, it's intuitive to me in the sense of each variable having a beta and I can say "for each x1 increase, y will increase by whatever %..." But for text, how does it work? Should the approach to find keywords that result in a certain value? i,e, there should be a relationship between the word "cancer" and cost, etc? Trying to wrap my head around this and get a sense of where to start and what to be look for.

Cross Validated 2020-06-10 14:08 UTC Score 23.0 AI-113-20200610-social-media-65d96a8d

Why do Dense layers perform better than a mix of Conv Layers, Recurrent Layers on Sentiment Analysis with BERT emebddings?

I have used BERT to make embeddings out of the imdb review dataset and I am trying out some models to check their perfomance on sentiment analysis (0 for the bad reviews and 1 for the good ones). I have seen that models with just dense Layers do work better than models which are a mixture of Convolutional with MaxPooling Layers and recurrent units such as LSTM or BiLSTM.I want to know why is this happening since BERT do captures of semantics and time dependencies so second models should perform better. I am speaking about orders from 85% accuracy with dense layers to 75% with the mix of Layers. Thanks in adavance

Cross Validated 2019-04-19 06:06 UTC Score 17.0 AI-113-20190419-social-media-01e6f0a2

How to know which statistical test to use?

I'm working on a project and was wondering how I could perform statistical tests between two correlations to prove significance. The background of my project is that I was trying to see if Twitter sentiments influence a certain market more than the other. After collecting data and calculating the correlation values using three methods (Pearson, Kendall, and Spearman) I found that the difference is very small. However, I feel that simply saying "the difference is small" is not convincing and was wondering if there were specific statistical tests for this purpose. I've done searching and have seen that people use p-value tests quite often, but I'm not sure if that would be appropriate for my case, as I'm also not familiar with statistical testing methods. I was hoping if anyone more knowledgeable would be able to give me some tips or pointers. The data that I'm specifically using looks like this: To briefly explain the data, TextBlob and NLTK Vader are two sentiment analysis tools I used. Within these two methods, I applied time lag to the data (one day forward and one day backward) and you can also see that I varied calculating correlation. X and Y are two separate markets and the correlation is between price and sentiment . The difference is the difference between these two. My original hypothesis was that market Y would have higher correlation, but apparently that is not true judging by the difference. What might be some ways that I could formally prove the significance of…

Cross Validated 2018-10-12 14:11 UTC Score 12.0 AI-113-20181012-social-media-645ca4c3

Similarity index between two texts Ask Question

I'm trying to compare two vectors in a small NLP project using Python. Code doesn't make any difference since I'm using scikit-learn, but my doubts are about my calculations. I have a query vector and some texts, all vectors are constructed using the TF-IDF algorithm and the same corpus with the same preprocessing, including Porter stemmer. The problem is that no matter the texts nor the query vector, both cosine and euclidean distances chose the same text. For instance, let $u$ be the query vector and $v_1, v_2, ..., v_n$ be the vectorized texts, then cosine similarity selects the vector $v_k$ to be more similar to $u$ , and so euclidean similarity measure. So, it's weird because I'm getting different results of similarity for those vectors, but they still chose the same vector. So my questions are: Cosine and euclidean similarities select the same vectors always? Is mandatory to normalize the euclidean distance? Why cosine distance shouldn't be normalized? Notes: To calculate the similarity coefficient of euclidean distance I use the formula ${1}\over{1+euclidean\_distance(u,v)}$ and for cosine distance $1-cosine\_distance(u,v)$ . Note that normalized means that the sum of all components of a vector is 1.

Cross Validated 2018-04-11 12:17 UTC Score 12.0 AI-113-20180411-social-media-44a6b990

Correctness of a skewed cosine similarity graph

I am currently implementing a word2vec model that uses the cosine similarity to determine the similarity between two vectors. When plotting all the possible cosine similarities, I get the following graph: The graph is pretty skewed towards the 'most similar value'. However, I am not sure if this is an ok thing, which can just be corrected by normalizing the data, or if this skewness indicates that something is terribly wrong with the model. I am not an expert on the topic of natural language processing. Can you guys provide me with some intuition what if it is ok to normalize, or if this indicates if the model is terribly wrong? Any ideas and insights are appreciated.