Skip to content
Automotive / RetailPseudonymized

ML Customer Segmentation for an Automotive Dealer Group

CorrDyn built an ML clustering pipeline on Databricks with LLM-powered profiling to segment customers across vehicle brands for a dealer group.

Editorial photograph evoking ml customer segmentation for an automotive dealer group

300+ features from vehicle sales, service history, and customer survey

Feature integration

8 clustering algorithms with Bayesian hyperparameter optimization

Algorithms evaluated

Interactive Power BI dashboard for segment exploration across brands and models

Analytical tool

The situation

A regional automotive dealer group operating across multiple vehicle brands wanted to move from intuition-based marketing toward data-driven customer understanding. The group had three data assets that had never been analyzed together: vehicle sales history spanning years of transactions, service repair order records capturing maintenance and repair behavior, and a customer survey with over 8,700 responses covering purchase motivations, lifestyle indicators, media habits, demographics, and brand attitudes.

Leadership wanted answers to specific questions. What distinguishes the customers who buy one brand versus another? What do owners of a specific model care about? Which customer segments respond to which marketing channels? A prior segmentation study by a global consulting firm had provided a high-level framework, but the dealer group wanted something more quantitative, more granular, and operationalized for ongoing use.

What we built

CorrDyn built an automated machine learning segmentation system on the client’s Databricks platform, integrating all three data sources through a shared customer identity key.

Data engineering. We built a four-stage pipeline that produced a unified customer dataset. Stage one extracted approximately 90 features from vehicle deal records: purchase counts, new-versus-used ratios, financial patterns, brand loyalty indicators, and trade-in behavior. Stage two processed service repair orders into roughly 40 features covering visit frequency, service types, spend patterns, and dealership preferences. Stage three transformed survey responses into approximately 200 binary features using one-hot encoding. Stage four joined all sources by customer ID and added derived features like total lifetime spend and customer age.

ML pipeline. The clustering engine evaluates eight algorithms across both Spark ML and scikit-learn frameworks. Optuna provides Bayesian hyperparameter optimization. MLflow tracks every experiment. The pipeline runs at three levels: all customers together, by vehicle brand, and by individual vehicle model. For each level, it tests cluster counts from two through five and selects the optimal configuration based on Silhouette Score, Davies-Bouldin Index, Calinski-Harabasz Index, and density-based validation.

LLM-powered profiling. After optimal clusters are identified, a statistical profiler generates comparison tables across all features. Claude then interprets those statistics to produce segment names, distinctive characteristics, behavioral profiles, and business-relevant descriptions. The combined outputs are stored in a Databricks table for downstream use.

Exploratory dashboard. When statistical validation proved inconclusive for many vehicle-level subsets due to sample size limitations, CorrDyn pivoted the primary deliverable to a self-hosted Power BI dashboard. The dashboard lets marketing stakeholders filter by any combination of vehicle brand, model, demographic attributes, and survey responses. Exploration pages display distributions for every survey question. Cluster pages show segment names, descriptions, characteristics, and customer counts with the same filtering capability.

What changed

The dealer group’s marketing team gained an interactive tool for understanding their customer base at a level of detail that was not previously accessible. Rather than relying on broad demographic assumptions or prior consulting frameworks, they can now explore how purchase motivations, lifestyle indicators, and brand attitudes differ across specific vehicle lines. The automated ML pipeline and its reference outputs remain available in Databricks for reuse when additional survey waves or expanded transaction data improve statistical power. The infrastructure is reusable and positions the company for a future Customer 360 initiative that would extend segmentation beyond survey respondents to the full customer base.

Frequently Asked
Questions

We have customer data across sales, service, and surveys but no unified view — where do we start?
Start by integrating your data sources around a shared customer identity key. CorrDyn built a pipeline on Databricks that unified 300+ features from vehicle sales, service history, and 8,700+ survey responses for an automotive dealer group. That unified dataset is the foundation for segmentation, marketing targeting, and eventually a full Customer 360.
How does ML-driven customer segmentation work when data is spread across multiple systems?
The pipeline evaluates multiple clustering algorithms with Bayesian hyperparameter optimization via Optuna, validated against statistical quality metrics. For this automotive client, eight algorithms were tested across 300+ features from sales, service, and survey data. When sample sizes are too small for robust clusters, CorrDyn pivots to interactive dashboards that let marketing teams explore the data directly rather than delivering segments that cannot be statistically defended.
What happens if the data does not support the original project scope?
CorrDyn adjusts the deliverable to match what the data can actually support. For this client, vehicle-model-level segments lacked statistical power, so we pivoted to an exploratory dashboard while preserving the ML pipeline for future use as data density improves. You get honest analysis, not a polished deliverable built on shaky assumptions.

Get your free
proposal.

Tell us about your challenges. We will be honest about whether we can help.

  • No pitch decks. We start by listening.
  • Discovery calls are free.
  • We respond within one business day.

Or email us directly at [email protected]

No sales scripts. No commitments.