Hierarchical Bayesian (HB) estimation is the gold-standard statistical method for extracting individual-level utilities from conjoint-style choice data. It treats respondents as draws from a population distribution, then uses Bayesian inference to estimate each person’s preferences individually — even on the sparse 10–20 choice tasks each respondent typically answers.
Why HB is needed
A single respondent in a conjoint study answers only a handful of choice tasks — far too few to pin down a full utility vector on their own. The naive alternatives both fail:
Aggregate logit — fits one average utility to the whole sample. Ignores heterogeneity. Produces the classic IIA failures (red-bus / blue-bus problem) when used for market simulation.
Per-respondent fit — estimates each person’s utilities from their own data alone. Massively overfits on the sparse ~15 choices per person. Collapses on respondents who made ambiguous choices.
HB is the principled middle. It partially pools information across respondents: each person’s estimate is shrunk toward the population mean in proportion to how noisy their own data is. Respondents with clear patterns get estimates close to their observed behaviour; respondents with ambiguous answers get estimates closer to the population average.
How it works mathematically
Each respondent $n$ has a utility vector $\boldsymbol{\beta}_n$. A multivariate normal prior is placed over respondents:
$$\boldsymbol{\beta}_n \sim \mathcal{N}(\boldsymbol{\mu}, \boldsymbol{\Sigma})$$
where $\boldsymbol{\mu}$ is the population mean and $\boldsymbol{\Sigma}$ the population covariance. Hyper-priors are placed on $\boldsymbol{\mu}$ and $\boldsymbol{\Sigma}$ (the covariance factored via a Cholesky decomposition with an LKJ prior on the correlation matrix for numerical stability).
The full posterior over everything is
$$p(\{\boldsymbol{\beta}_n\}, \boldsymbol{\mu}, \boldsymbol{\Sigma} \mid \mathbf{y}) \propto \prod_n \Big[\prod_t P(y_{nt} \mid \boldsymbol{\beta}_n) \cdot \mathcal{N}(\boldsymbol{\beta}_n \mid \boldsymbol{\mu}, \boldsymbol{\Sigma})\Big] p(\boldsymbol{\mu}) p(\boldsymbol{\Sigma})$$
Posterior is drawn by Markov Chain Monte Carlo — specifically Hamiltonian Monte Carlo with the No-U-Turn Sampler (NUTS), which is dramatically more efficient than random-walk methods on this correlated high-dimensional geometry.
Why individual-level utilities matter
Aggregate utilities tell you what the “average” customer wants — which is often a customer who does not exist. Individual-level utilities let you:
Build a market simulator that preserves genuine heterogeneity (not IIA)
Discover latent segments via clustering or latent-class on the utilities
Compute WTP distributions, not just averages
Target different customers with different products or messages
Diagnostics
Modern HB implementations enforce convergence via the Gelman-Rubin diagnostic ($\hat{R} < 1.1$), effective sample size thresholds, and divergent transition monitoring. A run that fails these checks is rejected rather than shipped.
Related terms
Conjoint analysis · Latent Class · Krinsky-Robb · Full methodology