What is Latent Class Analysis in Choice Modelling?

What is Latent Class Analysis in Choice Modelling?

Latent Class (LC) analysis is a mixture-model segmentation technique that discovers distinct patterns of preference in choice data. It assumes the population is made up of $K$ underlying segments, each with its own choice model, and each respondent belongs probabilistically to one segment.

How it works

The population is represented as a mixture of $K$ classes. Each class $k$ has its own vector of part-worth utilities $\boldsymbol{\beta}_k$ and a mixing weight $\pi_k$ (the proportion of the population in that class). Class membership is unobserved and marginalised out:

$$\log p(y_n) = \log \sum_{k=1}^{K} \pi_k \prod_t P(y_{nt} \mid \boldsymbol{\beta}_k)$$

Fitting is done by Expectation-Maximisation in frequentist settings, or MCMC in Bayesian. Once fit, each respondent’s posterior probability of belonging to each class follows from Bayes’ rule.

Choosing the number of classes

$K$ is chosen by information criterion rather than visual inspection. The Bayesian Information Criterion is the standard:

$$\mathrm{BIC} = -2 \log \hat{\mathcal{L}} + p \log(N T)$$

where $p$ is the parameter count and $NT$ the total number of observed choices. BIC’s log-$N$ penalty is stricter than AIC’s constant penalty, which matters because overfitting here means manufacturing spurious micro-segments.

When to use Latent Class

  • You need named, interpretable segments for a Marketing or Strategy audience

  • You want to know what share of the population is in each segment

  • You want to run share-of-preference simulations by segment and sum back to the aggregate

  • Your sample is big enough to support reliable estimates per class (100–150 per class as a floor)

When Hierarchical Bayes beats Latent Class

HB with a continuous normal prior over individual utilities is better when:

  • You want individual-level targeting rather than segment-level

  • You care about the full tail of the preference distribution

  • Your sample is small — HB’s partial pooling degrades more gracefully with few respondents

The best practice: run both

Run both k-means-on-HB-utilities AND Latent Class on the same study, then compare. If both methodologies find the same story, segmentation is robust. If they diverge, the story is methodology-dependent — itself an important finding.

Related terms

Conjoint analysis · Hierarchical Bayes · Full methodology

Subscribe to this product here and start using it today!

ordeen

At Ordeen, our mission is to make AI useful, not just impressive. Because the future shouldn’t just be AI-powered. It should be data-backed and reality-checked.

Verified data first. AI last.

© 2026

All Rights Reserved