What is Hierarchical Bayesian Estimation in Choice Modelling?

What is Hierarchical Bayesian Estimation in Choice Modelling?

Hierarchical Bayesian (HB) estimation is the gold-standard statistical method for extracting individual-level utilities from conjoint-style choice data. It treats respondents as draws from a population distribution, then uses Bayesian inference to estimate each person’s preferences individually — even on the sparse 10–20 choice tasks each respondent typically answers.

Why HB is needed

A single respondent in a conjoint study answers only a handful of choice tasks — far too few to pin down a full utility vector on their own. The naive alternatives both fail:

  • Aggregate logit — fits one average utility to the whole sample. Ignores heterogeneity. Produces the classic IIA failures (red-bus / blue-bus problem) when used for market simulation.

  • Per-respondent fit — estimates each person’s utilities from their own data alone. Massively overfits on the sparse ~15 choices per person. Collapses on respondents who made ambiguous choices.

HB is the principled middle. It partially pools information across respondents: each person’s estimate is shrunk toward the population mean in proportion to how noisy their own data is. Respondents with clear patterns get estimates close to their observed behaviour; respondents with ambiguous answers get estimates closer to the population average.

How it works mathematically

Each respondent $n$ has a utility vector $\boldsymbol{\beta}_n$. A multivariate normal prior is placed over respondents:

$$\boldsymbol{\beta}_n \sim \mathcal{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})$$

where $\boldsymbol{\mu}$ is the population mean and $\boldsymbol{\Sigma}$ the population covariance. Hyper-priors are placed on $\boldsymbol{\mu}$ and $\boldsymbol{\Sigma}$ (the covariance factored via a Cholesky decomposition with an LKJ prior on the correlation matrix for numerical stability).

The full posterior over everything is

$$p(\{\boldsymbol{\beta}_n\}, \boldsymbol{\mu}, \boldsymbol{\Sigma} \mid \mathbf{y}) \propto \prod_n \Big[\prod_t P(y_{nt} \mid \boldsymbol{\beta}_n) \cdot \mathcal{N}(\boldsymbol{\beta}_n \mid \boldsymbol{\mu}, \boldsymbol{\Sigma})\Big] p(\boldsymbol{\mu}) p(\boldsymbol{\Sigma})$$

Posterior is drawn by Markov Chain Monte Carlo — specifically Hamiltonian Monte Carlo with the No-U-Turn Sampler (NUTS), which is dramatically more efficient than random-walk methods on this correlated high-dimensional geometry.

Why individual-level utilities matter

Aggregate utilities tell you what the “average” customer wants — which is often a customer who does not exist. Individual-level utilities let you:

  • Build a market simulator that preserves genuine heterogeneity (not IIA)

  • Discover latent segments via clustering or latent-class on the utilities

  • Compute WTP distributions, not just averages

  • Target different customers with different products or messages

Diagnostics

Modern HB implementations enforce convergence via the Gelman-Rubin diagnostic ($\hat{R} < 1.1$), effective sample size thresholds, and divergent transition monitoring. A run that fails these checks is rejected rather than shipped.

Related terms

Conjoint analysis · Latent Class · Krinsky-Robb · Full methodology

Subscribe to this product here and start using it today!

ordeen

At Ordeen, our mission is to make AI useful, not just impressive. Because the future shouldn’t just be AI-powered. It should be data-backed and reality-checked.

Verified data first. AI last.

© 2026

All Rights Reserved