This document is the full statistical specification of the Ordeen Choice Modelling estimation stack. It is written for quantitative researchers who want to know exactly what is being computed, not an executive summary. Each method is described with its model, its estimator and the correctness property it guarantees — so you can match any claim on the product against the reference literature.
1. Random utility theory and the multinomial logit kernel
Discrete choice modelling rests on random utility theory: a respondent assigns each alternative a latent utility and chooses the one with the highest realised value. Utility splits into a systematic part explained by observed attributes and an independent random shock.
For respondent $n$ facing task $t$ with alternatives $j = 1 \ldots J$, the utility of alternative $j$ is
$$U_{ntj} = \mathbf{x}_{ntj}^{\top}\,\boldsymbol{\beta}_{n} + \varepsilon_{ntj}$$
where $\mathbf{x}_{ntj}$ is the attribute vector (effects-coded categorical attributes plus numeric terms), $\boldsymbol{\beta}_{n}$ is that respondent's vector of part-worth utilities, and the shocks $\varepsilon_{ntj}$ are i.i.d. Type-I extreme value. That distributional assumption yields the closed-form multinomial logit (MNL) choice probability:
$$P(y_{nt}=j \mid \boldsymbol{\beta}_{n}) = \frac{\exp\!\big(\mathbf{x}_{ntj}^{\top}\boldsymbol{\beta}_{n}\big)}{\sum_{k=1}^{J}\exp\!\big(\mathbf{x}_{ntk}^{\top}\boldsymbol{\beta}_{n}\big)}$$
A "none of these" option is modelled explicitly as an extra alternative carrying its own per-respondent threshold utility — a respondent who keeps rejecting the menu informs the model instead of biasing it.
2. Hierarchical Bayesian estimation
A single respondent answers only a handful of tasks — far too few to pin down a full utility vector on their own. The fix is partial pooling: respondents are treated as draws from a population distribution, so each person's estimate is shrunk toward the population mean in proportion to how noisy their own data is.
A multivariate normal prior is placed over individual vectors:
$$\boldsymbol{\beta}_{n} \sim \mathcal{N}\!\big(\boldsymbol{\mu},\, \boldsymbol{\Sigma}\big), \qquad n = 1 \ldots N$$
with hyper-priors on the population mean $\boldsymbol{\mu}$ and covariance $\boldsymbol{\Sigma}$. The covariance is factored through its Cholesky decomposition with an LKJ prior on the correlation matrix, which keeps the geometry well-behaved and lets the sampler explore correlations between part-worths without producing degenerate covariance draws.
The full posterior over everything is
$$p\big(\{\boldsymbol{\beta}_n\}, \boldsymbol{\mu}, \boldsymbol{\Sigma} \mid \mathbf{y}\big) \;\propto\; \prod_{n=1}^{N}\Bigg[\,\prod_{t} P(y_{nt}\mid\boldsymbol{\beta}_n)\;\mathcal{N}(\boldsymbol{\beta}_n\mid\boldsymbol{\mu},\boldsymbol{\Sigma})\Bigg]\,p(\boldsymbol{\mu})\,p(\boldsymbol{\Sigma})$$
Why this is the right method. Aggregate logit assumes everyone is identical and produces the classic IIA pathologies in the simulator; fitting each respondent separately overfits wildly on the handful of tasks that respondent answered. Hierarchical Bayes is the principled middle — it recovers genuine heterogeneity while remaining stable on sparse per-person data.
3. Sampling the posterior with HMC and NUTS
The posterior has no closed form, so it is drawn by Markov Chain Monte Carlo. The sampler is Hamiltonian Monte Carlo with the No-U-Turn Sampler (NUTS), which augments the parameters with momentum and simulates Hamiltonian dynamics to propose distant, high-acceptance moves. HMC/NUTS is dramatically more efficient than random-walk Metropolis on the correlated, high-dimensional geometry of a hierarchical choice model.
Models are reparameterised into a non-centred form,
$$\boldsymbol{\beta}_n = \boldsymbol{\mu} + \mathbf{L}\,\mathbf{z}_n, \qquad \mathbf{z}_n \sim \mathcal{N}(\mathbf{0}, \mathbf{I})$$
with $\mathbf{L}$ the Cholesky factor of $\boldsymbol{\Sigma}$. This decouples the scale of the population spread from the individual draws and removes the funnel geometry that otherwise causes divergent transitions.
Several independent chains run in parallel from dispersed starting points, with convergence enforced via the Gelman-Rubin diagnostic — every parameter must satisfy $\hat{R} < 1.1$ (typically $< 1.01$) with adequate effective sample size. Divergent transitions are monitored and a run that fails these checks is rejected rather than shipped.4. Adaptive CBC — the respondent-specific design
ACBC's adaptive front end — the Build-Your-Own concept, the must-have / unacceptable screener and the per-respondent tournament — means two respondents are literally shown different sets of alternatives, built from their own earlier answers. The right thing to estimate against is therefore the exact design each respondent actually saw, not an aggregate design or a replayed facsimile.
For respondent $n$, the likelihood contribution uses the alternatives $\mathcal{A}_{nt}$ actually presented in task $t$:
$$P(y_{nt}=j \mid \boldsymbol{\beta}_n, \mathcal{A}_{nt}) = \frac{\exp\!\big(\mathbf{x}_{ntj}^{\top}\boldsymbol{\beta}_n\big)}{\sum_{k \in \mathcal{A}_{nt}}\exp\!\big(\mathbf{x}_{ntk}^{\top}\boldsymbol{\beta}_n\big)}$$
Estimating an adaptive study as if everyone saw the same choice sets biases the resulting utilities toward the aggregate screener logic rather than the respondent's revealed trade-offs.
5. Latent-class segmentation
Where interpretable segments are preferred to a continuous distribution, the latent-class model represents the population as a mixture of $K$ classes, each with its own utility vector $\boldsymbol{\beta}_k$ and mixing weight $\pi_k$. Class membership is unobserved and marginalised out, so each respondent's likelihood is a weighted sum over classes:
$$\log p(y_n) = \log \sum_{k=1}^{K} \pi_k \prod_{t} P\big(y_{nt}\mid \boldsymbol{\beta}_k\big)$$
The number of classes is chosen by information criterion rather than by eye. Candidate solutions are scored with the Bayesian Information Criterion:
$$\mathrm{BIC} = -2\,\log \hat{\mathcal{L}} + p \,\log(N T)$$
where $p$ is the parameter count and $NT$ the number of observed choices. BIC's heavier penalty resists the tendency to manufacture spurious micro-segments.
6. MaxDiff — best-worst scaling
For pure prioritisation — ranking a long list of features, messages or needs — MaxDiff shows small sets and asks for the best and worst item. A best-worst answer is decomposed into a rank-ordered logit:
$$P(\text{best}=b,\,\text{worst}=w \mid S) = \frac{e^{u_b}}{\sum_{i\in S} e^{u_i}} \cdot \frac{e^{-u_w}}{\sum_{i\in S\setminus\{b\}} e^{-u_i}}$$
Item utilities are estimated hierarchically, then re-centred so the population mean per item is zero. Anchored MaxDiff adds a per-respondent threshold so the scale becomes "important vs not", not merely relative — which is what makes money-denominated results possible.
7. Menu-Based Choice — independent take decisions
When respondents assemble a basket rather than pick one product, each menu item is modelled as its own binary decision. The probability of taking item $i$ is a logistic function of its utility relative to a per-respondent "take-it" threshold, with an optional price term:
$$P(\text{take } i \mid n) = \frac{1}{1 + \exp\!\big(-(u_{ni} - \theta_n + \gamma_n\, p_i)\big)}$$
where $u_{ni}$ is respondent $n$'s utility for item $i$, $\theta_n$ their threshold, $\gamma_n$ their price sensitivity and $p_i$ the item's price. Item utilities, thresholds and price coefficients are all estimated hierarchically.8. Willingness-to-pay with honest uncertainty
WTP converts a utility gain into money by dividing it by the marginal utility of price:
$$\mathrm{WTP} = \frac{\beta_{\text{target}} - \beta_{\text{ref}}}{\lvert \beta_{\text{price}} \rvert}$$
A point estimate alone is dangerous — the price coefficient sits in the denominator, so when it is small or uncertain the ratio explodes. The ratio is instead evaluated at every posterior draw and the credible interval of the resulting distribution is reported (a Krinsky-Robb style construction). Degenerate cases — where the price coefficient straddles zero, or where a segment's price slope has the wrong sign relative to the population — are detected and refused rather than reported as a confident number.
The reference level is anchored to the population baseline rather than silently defaulting to the first alphabetical level, and the reference is surfaced on every WTP table. A market-calibrated WTP variant is also computed by searching for the price change that holds a product's simulated share constant after an upgrade.
9. Attribute importance
Once utilities are estimated, each attribute's importance is the spread it commands — the range between its best and worst level — expressed as a share of the total range across all attributes:
$$I_a = \frac{\max_l \beta_{a,l} - \min_l \beta_{a,l}}{\sum_{a'} \big(\max_l \beta_{a',l} - \min_l \beta_{a',l}\big)} \times 100$$
Importance is computed per respondent and averaged, so an attribute that strongly divides the audience is not washed out by averaging utilities first. Importance is reported with a posterior credible interval.
10. The market simulator and price step response
The simulator computes individual-level share-of-preference and averages across respondents, preserving heterogeneity:
$$\widehat{\text{share}}(P) = \frac{1}{N}\sum_{n=1}^{N} \frac{\exp\!\big(\mathbf{x}_{P}^{\top}\boldsymbol{\beta}_n\big)}{\sum_{P' \in \mathcal{M}} \exp\!\big(\mathbf{x}_{P'}^{\top}\boldsymbol{\beta}_n\big)}$$
for product $P$ in market $\mathcal{M}$. Because the full set of posterior draws is retained, each share carries a credible interval and the posterior probability that a given product is the most preferred is reported — not just a point estimate.
Sweeping a product's price traces an empirical demand curve. The simulator reports the full price step response rather than collapsing to a single optimal price: at each evaluated step, share, revenue and (if margin is supplied) contribution are returned, each with a credible interval. Reporting the curve rather than an argmax exposes plateaus and cliffs — where revenue is nearly flat (and there is pricing headroom), where share falls off sharply, and which competitor absorbs the walk-away as price rises.
11. Interactions, controlled for multiplicity
A main-effects utility model treats each attribute level as acting on choice independently. In practice, levels sometimes pull together or push apart beyond what main effects predict. The interaction grid tests each attribute-pair cell for a departure from the additive baseline — but testing many cells at once requires multiplicity control.
Each $p$-value is corrected with the Benjamini-Hochberg procedure, which controls the expected proportion of false discoveries at a stated level $q$:
$$p_{(i)} \le \frac{i}{m}\,q \;\;\Rightarrow\;\; \text{reject } H_{0,(i)}$$
where $p_{(1)} \le \cdots \le p_{(m)}$ are the sorted $m$ cell $p$-values. Surviving pairs are labelled Complement or Substitute by the sign of the interaction coefficient and reported in plain language. A small-cell Haldane correction prevents a near-empty cell from manufacturing an interaction all on its own.
12. TURF — honest reach for item portfolios
Given a long list of items, TURF (Total Unduplicated Reach and Frequency) asks for the smallest portfolio of size $k$ that reaches the most people, where "reach" means at least one item clears the respondent's take-threshold:
$$\mathrm{Reach}(\mathcal{S}) = \frac{1}{N}\sum_{n=1}^{N} \mathbb{1}\!\left[\max_{i \in \mathcal{S}} u_{ni} \ge \theta_n\right]$$
Exhaustive search over all $\binom{M}{k}$ subsets grows faster than any usable design can honestly resolve. The search is bounded by considering only top-ranked candidates by marginal reach, and the reported curve is truncated at the honest coverage limit implied by what each respondent could realistically evaluate. A benchmark line (random pick of the same size) is plotted alongside so a nominally "high" reach that merely matches random is flagged visually.
13. Experimental design efficiency
The quality of every downstream estimate is capped by the design that generated the data. The generator builds balanced, near-orthogonal designs that respect prohibited level combinations and the chosen overlap rule, then reports a transparent efficiency score combining:
Level balance — how evenly each level of each attribute appears across the design, so no level is under-sampled
Orthogonality — how independent the attributes are from one another, so their effects can be told apart
The score is surfaced before fieldwork begins — a weak design is caught at the design stage, not after the data is in.
For practitioners: this methodology is designed to be independently reproducible. The method names, model equations, estimator choices and diagnostic thresholds above are the specification; any Hierarchical Bayes implementation that satisfies them (Stan, NumPyro, PyMC) should produce posterior means within Monte Carlo noise of each other.
Questions on methodology, model choice or validation: im@ordeen.com.
