Data Science Stack Exchange
2026-04-13 02:51 UTC
By Thị Mến Nguyễn
AI-111-20260413-social-media-57b8cc1a
CLV Estimation: BTYD vs. Survival Analysis
I'm working on a project using the Elo Merchant Category Recommendation dataset (Kaggle). My goal is to perform Customer Segmentation base on their transactions by combining RFM metrics with Customer Lifetime Value (CLV). I’m using the historical_transactions for this project, but I’ve hit two specific points where I’d love some expert input: Setting Identification: I am debating whether the Elo context should be treated as contractual or non-contractual. While it involves transactions across various merchants, the presence of the 'installments' variable for each transaction adds a layer of complexity. I need to clarify this to choose the most appropriate CLV estimation method: should I rely on Probabilistic Models (such as BG/NBD + Gamma-Gamma) or Survival Analysis in this specific case? Feature Engineering: Should I include 'installments' as an input feature for the clustering algorithm, or should I use it as a profiling variable to describe the clusters after they are formed? I’d appreciate any insights from those who have worked with this specific dataset or similar credit card transaction data. Thanks in advance!
I'm working on a project using the Elo Merchant Category Recommendation dataset (Kaggle). My goal is to perform Customer Segmentation base on their transactions by combining RFM metrics with Customer Lifetime Value (CLV). I’m using the historical_transactions for this project, but I’ve hit two specific points where I’d love some expert input: Setting Identification: I am debating whether the Elo context should be treated as contractual or non-contractual. While it involves transactions across various merchants, the presence of the 'installments' variable for each transaction adds a layer of complexity. I need to clarify this to choose the most appropriate CLV estimation method: should I rely on Probabilistic Models (such as BG/NBD + Gamma-Gamma) or Survival Analysis in this specific case? Feature Engineering: Should I include 'installments' as an input feature for the clustering algorithm, or should I use it as a profiling variable to describe the clusters after they are formed? I’d appreciate any insights from those who have worked with this specific dataset or similar credit card transaction data. Thanks in advance!
Full article content could not be extracted automatically. Read the original below.
Source:
Data Science Stack Exchange
· datascience.stackexchange.com