Golden sunlight falls across Ateneo’s empty Red Brick Road beneath a canopy of rain trees.

CH 00 · GABRIEL LABARIENTO / MANILA

Welcome, thanks for stopping by!

Product engineering · AI · Data

GABE TV

CH 01 / Project story

Data / 2026

Data Analyst

Telecom Churn Prediction

Churn modeling for 100K customers. The highest-risk group had 78% churn, a 1.6× lift over the baseline.

Telecom Churn Prediction project preview

The problem

Company A, a telecom operator, needed a way to identify customers at risk of leaving. Churn was close to 50% in the test data, so contacting everyone would spread the retention budget thin. I needed to identify whom to contact first and explain the reasoning behind that choice.

  • RoleData analyst
  • ContextEnd-to-end churn analysis
  • Scale100,000 customer records
  • Result78% churn in the highest-risk decile

What I built

I joined the client and records tables one-to-one on Customer_ID. The resulting dataset had 100,000 rows and about 100 columns covering demographics, equipment, revenue, usage, tenure, and churn. Customers who left had median monthly revenue of 47.49, close to 48.88 for those who stayed. The clearest risk appeared around renewal at months 11 to 12. I created retention segments, including 'Renewal + silent-switching risk', trained an XGBoost classifier, and used its risk tiers to propose a targeted 'Renewal Rescue System'. I engineered features for equipment age, usage decline, and renewal timing. The 11-to-12-month renewal cohort had a 64% churn rate; the proposed retention rules map high-risk customers to specific actions.

BUILT WITH
PythonpandasNumPyscikit-learnXGBoostMatplotlibSeaborn

Technical decisions

Keeping preprocessing inside the training pipeline

I put median and most-frequent imputation, along with one-hot encoding, inside an sklearn Pipeline and ColumnTransformer. Each transform fits only on the training folds. This made preprocessing repeatable across roughly 100 numeric and categorical columns and kept test data out of training.

python
def build_churn_pipeline(X):
    numeric_features = X.select_dtypes(include=["number", "bool"]).columns.tolist()
    categorical_features = X.select_dtypes(include=["object", "category"]).columns.tolist()

    preprocessor = ColumnTransformer(transformers=[
        ("num", SimpleImputer(strategy="median"), numeric_features),
        ("cat", Pipeline([
            ("imputer", SimpleImputer(strategy="most_frequent")),
            ("onehot", OneHotEncoder(handle_unknown="ignore")),
        ]), categorical_features),
    ])

    model = XGBClassifier(
        n_estimators=250, max_depth=4, learning_rate=0.05,
        subsample=0.9, colsample_bytree=0.9,
        eval_metric="logloss", random_state=42, n_jobs=-1,
    )

    return Pipeline([("preprocess", preprocessor), ("model", model)])

Evaluating who to contact first

The model's ROC-AUC was about 0.69. I also checked whether its ranking could help a team prioritize outreach: 78% of customers in the highest-risk decile had churned, and the 'Very high risk' tier had a 73% churn rate against a roughly 50% baseline. Those comparisons make the model's use clearer when the retention budget is fixed.

Tradeoffs

Prioritization rather than precise probabilities

I evaluated the model around risk tiers and the highest-risk decile because a retention team has a fixed budget. The model can help order a contact list, but its scores should not be treated as precise probabilities for individual customers.

Keeping missing data visible

Missing values were concentrated in demographic and equipment fields. I kept that data-quality problem in the analysis rather than implying that imputation had resolved it. Some features remained weak predictors.

Results

Customers analyzed
100K
Model ROC-AUC
0.69
Top-decile churn rate
78%
  • I grouped 100,000 customer records into risk tiers to help a retention team prioritize its budget.
  • The highest-risk decile had a 78% churn rate, a 1.6× lift over the baseline. This is the churn rate within the group, not the share of all churners captured.

Command Palette

Search the portfolio. Use arrow keys, Enter to select.