ML Customer Segmentation for an Automotive Dealer Group
CorrDyn built an ML clustering pipeline on Databricks with LLM-powered profiling to segment customers across vehicle brands for a dealer group.

300+ features from vehicle sales, service history, and customer survey
Feature integration
8 clustering algorithms with Bayesian hyperparameter optimization
Algorithms evaluated
Interactive Power BI dashboard for segment exploration across brands and models
Analytical tool
The situation
A regional automotive dealer group operating across multiple vehicle brands wanted to move from intuition-based marketing toward data-driven customer understanding. The group had three data assets that had never been analyzed together: vehicle sales history spanning years of transactions, service repair order records capturing maintenance and repair behavior, and a customer survey with over 8,700 responses covering purchase motivations, lifestyle indicators, media habits, demographics, and brand attitudes.
Leadership wanted answers to specific questions. What distinguishes the customers who buy one brand versus another? What do owners of a specific model care about? Which customer segments respond to which marketing channels? A prior segmentation study by a global consulting firm had provided a high-level framework, but the dealer group wanted something more quantitative, more granular, and operationalized for ongoing use.
What we built
CorrDyn built an automated machine learning segmentation system on the client’s Databricks platform, integrating all three data sources through a shared customer identity key.
Data engineering. We built a four-stage pipeline that produced a unified customer dataset. Stage one extracted approximately 90 features from vehicle deal records: purchase counts, new-versus-used ratios, financial patterns, brand loyalty indicators, and trade-in behavior. Stage two processed service repair orders into roughly 40 features covering visit frequency, service types, spend patterns, and dealership preferences. Stage three transformed survey responses into approximately 200 binary features using one-hot encoding. Stage four joined all sources by customer ID and added derived features like total lifetime spend and customer age.
ML pipeline. The clustering engine evaluates eight algorithms across both Spark ML and scikit-learn frameworks. Optuna provides Bayesian hyperparameter optimization. MLflow tracks every experiment. The pipeline runs at three levels: all customers together, by vehicle brand, and by individual vehicle model. For each level, it tests cluster counts from two through five and selects the optimal configuration based on Silhouette Score, Davies-Bouldin Index, Calinski-Harabasz Index, and density-based validation.
LLM-powered profiling. After optimal clusters are identified, a statistical profiler generates comparison tables across all features. Claude then interprets those statistics to produce segment names, distinctive characteristics, behavioral profiles, and business-relevant descriptions. The combined outputs are stored in a Databricks table for downstream use.
Exploratory dashboard. When statistical validation proved inconclusive for many vehicle-level subsets due to sample size limitations, CorrDyn pivoted the primary deliverable to a self-hosted Power BI dashboard. The dashboard lets marketing stakeholders filter by any combination of vehicle brand, model, demographic attributes, and survey responses. Exploration pages display distributions for every survey question. Cluster pages show segment names, descriptions, characteristics, and customer counts with the same filtering capability.
What changed
The dealer group’s marketing team gained an interactive tool for understanding their customer base at a level of detail that was not previously accessible. Rather than relying on broad demographic assumptions or prior consulting frameworks, they can now explore how purchase motivations, lifestyle indicators, and brand attitudes differ across specific vehicle lines. The automated ML pipeline and its reference outputs remain available in Databricks for reuse when additional survey waves or expanded transaction data improve statistical power. The infrastructure is reusable and positions the company for a future Customer 360 initiative that would extend segmentation beyond survey respondents to the full customer base.
Frequently Asked
Questions
We have customer data across sales, service, and surveys but no unified view — where do we start?
How does ML-driven customer segmentation work when data is spread across multiple systems?
What happens if the data does not support the original project scope?
Related Work
Similar engagements across our portfolio.

Parking Revenue Optimization for a Major Entertainment Venue
CorrDyn built demand prediction and dynamic pricing models for a multi-venue entertainment complex, optimizing parking revenue across nine lots.

Manufacturing Throughput Optimization for a Biotech Producer
CorrDyn built process time analytics and a batch production simulation for a biotech manufacturer, enabling data-driven throughput optimization.

IoT Analytics for a Biotech Manufacturing Operation
CorrDyn built machine monitoring pipelines, ML failure detection, and real-time Grafana dashboards for a global biotech DNA/RNA synthesis manufacturer.
Related Conversations
Podcast episodes that cover the same ground.
Unlocking the Power of Semantic BI with Hashboard
with Carlos Aguilar, Hashboard
Carlos Aguilar explains how Hashboard eliminates multiple sources of truth in analytics through semantic BI, version control, and governance.
Physics, Free Energy, and Computational Drug Discovery
with Robert Abel, Schrödinger
Robert Abel of Schrödinger on why ML alone fails in 10^60 chemical space and how physics-based simulation reaches near-experimental accuracy.
Applying ML/AI to Drug Development with Anil Kane
with Dr. Anil Kane, Thermo Fisher Scientific
Dr. Anil Kane of Thermo Fisher Scientific discusses how AI, machine learning, and digital tools are reshaping drug development and manufacturing efficiency.