## Identification and Estimation {#sec-estimation}

Search models are identified from patterns that ordinary choice data cannot
produce: *which* alternatives a respondent opened, *in what order*, *how many*,
and *which* of the opened alternatives was ultimately bought. This section sets
out what each of those margins identifies, what SIFT's designed variation adds
to arguments that field studies must make on observational grounds, and what
remains normalized rather than estimated.

Throughout, the object of interest is the pair $(\bfbeta, \bfgamma)$ —
preferences and search costs — together with the post-search standard deviation
$\tilde\sigma$, which most of the literature fixes to one and which we argue
must be estimated.

### What the Designed Variation Buys {#sec-est-variation}

#### Position enters search costs, not utility

The identifying assumption underlying every position-based search-cost shifter
is that a listing's position affects what it costs to inspect an alternative,
and nothing else. @Ursu_2018 §4.3 takes this seriously rather than asserting it,
enumerating the ways position could instead enter the model — through pre-search
utility, through the post-search spread, or through combinations — and ruling
them out with observational tests. Because SIFT randomizes position within every
task and records the full click sequence, we can run those tests as experimental
contrasts rather than conditional comparisons.

Two statistics suffice, and neither requires fitting a model.

**T1. Among clicked alternatives, does position predict purchase?**
**T2. Is the first click near the top of the page?**

The sign of T1 warrants care, because it runs opposite to the natural
intuition. One might reason that if position acts only on costs, it should bear
on whether an alternative is opened but not on what is bought once it has been;
a flat T1 would then indicate position-in-cost. Simulation shows the reverse,
and the mechanism is selection. When position raises the search cost, an
alternative in a poor position must have unusually high pre-search utility to be
worth opening at all. Conditional on having been opened, poorly placed
alternatives are therefore *better* on average than well-placed ones, and T1 is
strongly **positive**. When position instead enters pre-search utility, the
direct and selection effects approximately cancel and T1 is **flat**.

In simulations of 400 respondents completing 8 tasks each, T1 has slope $+0.39$
($p \approx 5\times10^{-15}$) when position acts on costs, and $-0.05$
($p = 0.06$) when it acts on pre-search utility. Interpreting a flat T1 as
evidence for position-in-cost would invert the conclusion.

T2 separates the remaining case. If position widens the post-search spread, the
reservation value *rises* with position, so the worst-placed alternatives would
be opened first. That is what we see: mean first-click position 4.96 against a
uniform benchmark of 3.5, versus 2.21 and 2.57 in the cost and utility arms.

The two tests together form a decision rule:

| T2 | T1 | position acts on |
|---|---|---|
| first click near top | strongly positive | **search cost** |
| first click near top | flat | pre-search utility |
| first click near bottom | — | post-search spread |

We regard this as a design check to be run and reported in every SIFT study
rather than assumed. Details, code, and the full three-arm simulation are in
[the position comparison](../comparisons/compare-ursu-position.qmd).

#### Separating preferences from search costs

The sharper problem is not whether position shifts costs but whether a high
search cost can be distinguished from a low taste. @Morozov_2021 give the
cleanest statement of the separation (their §4.3.2): stopping decisions depend on
the *realized* utilities of alternatives already searched, but not on their
search costs, which are sunk by the time the stopping decision is made. Trade
the utility intercept against the search cost so as to hold the mean reservation
value fixed, and the search *order* is unchanged — but the higher intercept
raises realized utility and so makes stopping more likely.

Their Appendix D sharpens this into an observable moment: the **recall rate**, the
probability of purchasing an alternative encountered earlier in the sequence
rather than the most recent one, responds to the utility intercept but not to
the search cost. SIFT observes recall directly, because respondents have free
recall and we record the entire click sequence. This is our cleanest separation
moment, and it is available in the raw data without any modelling.

#### What the panel buys, and what it does not

@Morozov_2021 prove nonparametric identification of individual-level
preferences *and* search costs from panel search data. It is worth being precise
about the strength of that result, because it is easy to overclaim. The argument
assumes the number of sessions per consumer grows without bound; with between 2
and 35 sessions per consumer and a mean of 5, @Morozov_2021 estimate
hyperparameters rather than individual parameters (their §5.1).

At the task counts SIFT can realistically field — call it $T=10$ — we are in the
same position. What the panel delivers is *hierarchical posteriors in the
conjoint tradition* rather than nonparametric individual identification. That is the standard the conjoint audience already applies to
part-worths, and it is the right standard here.

The panel also does something less obvious. @Ursu_2018 reports baseline search
costs above \$102 and attributes the magnitude to observing a single click per
search impression with no way to link a consumer's impressions; restricting
attention to impressions with two or more clicks cuts mean estimated search
costs by up to 83%. Single-spell data mechanically inflates search costs. The
panel is therefore a *levels* argument as much as a heterogeneity argument.

#### What neither the panel nor the shifter buys alone

The two ingredients address different problems, and the literature illustrates
this by having one or the other but never both. @Morozov_2021 have the panel and
still normalize the post-search standard deviation, reporting that little
variation in their data speaks to it (their Appendix C). @Yavorsky_2021 have an
exogenous search-cost shifter and no panel, and do estimate it — finding
$\tilde\sigma = 8.16$ against the conventional normalization of one, with a
confidence interval that excludes one and consequences large enough to reverse
the sign of one search-cost coefficient and to understate a counterfactual by a
factor of three to four.

SIFT has both, and it is important not to conflate what each contributes.
Simulations holding the total number of tasks fixed while varying tasks per
respondent show that the panel does not improve recovery of the post-search
spread; the randomized shifter does that work on its own. What the panel buys is
individual-level parameters, which is a separate claim resting on separate
variation.

SIFT also has something neither has: because we design the product page, part of
the post-search component is *fixed by the design* rather than inferred.

This last point is stronger than it first appears. In a field setting the analyst never observes what inspection
revealed; the post-search shock is a residual, and all that can be recovered is
its variance. In SIFT the analyst *wrote* the product page. So for every clicked
alternative the revealed attributes $\bfL_{ijt}$ are known, and they enter
realized utility as $\bfL_{ijt}'\bfkappa$ — a mean shift, not a variance.
Preferences over attributes that are only visible *after* a click are therefore
identified in exactly the way conjoint part-worths are: from which alternative
gets bought given what was shown.

The distinction is easy to lose in implementation, with consequences that are
silent. The ex ante spread $\tilde\sigma$ belongs in the reservation value,
because that is computed before the click; the realized $\bfL$ belongs in the
utility, because that is evaluated after it. Use $\tilde\sigma$ for both and
$\bfkappa$ enters the likelihood only through
$\tilde\sigma^2 = \bfkappa'\mathrm{Var}(\bfL)\bfkappa + \sigma_\xi^2$, where it
trades off exactly against $\sigma_\xi$ and is not identified at all — while
every other parameter continues to recover normally, so nothing looks wrong.
This is the sharpest version of the identification claim in the paper and it
belongs in the introduction.

### Normalization {#sec-est-normalization}

Setting the standard deviation of the pre-search shock to one fixes the scale of
utility, as in @Ursu_Seiler_Honka_2024 §4.4.2 and @Morozov_2021 Appendix C. That
normalization is innocuous and we adopt it.

The post-search standard deviation $\tilde\sigma$ is a different matter. It is
identified in principle by functional form — $\tilde\sigma$ appears on both
sides of $g(\zeta) = \cost/\tilde\sigma$ — but @Yavorsky_2021 shows by
simulation that functional-form identification alone is not enough to estimate
it in practice, and that an exogenous, alternative-specific cost shifter is what
makes estimation work.

The consequence of fixing it deserves to be stated without hedging. With both
variances normalized, search costs cannot be monetized and search-cost
counterfactuals are not interpretable (@Ursu_Seiler_Honka_2024 §4.4.4). That
would rule out every friction counterfactual in @sec-managerial. Estimating
$\tilde\sigma$ is therefore load-bearing for this paper rather than a
refinement, and it is the single clearest reason SIFT needs a randomized
position rather than merely a rich attribute design.

### Testing the Zero-Search-Cost Restriction {#sec-est-nesting}

SIFT nests conjoint: as the search cost goes to zero every alternative is
inspected, choice is made over fully realized utilities, and the model collapses
to a standard choice-based conjoint model. Whether search costs are
distinguishable from zero is therefore a testable question rather than a
maintained assumption — but the test requires care.

The natural test — a Wald test on the estimated search cost — is not valid
here. Zero is a limit rather than an interior point of the parameter space:
$\cost \to 0$ corresponds to $g(\cost) \to +\infty$ and to
$\log \cost \to -\infty$. The null therefore sits at a boundary under every
parameterization, and the standard asymptotics do not apply.

We construct the test from observable implications instead, in three ways. The
first is holdout fit: SIFT against the nested conjoint model on tasks withheld
from estimation. The second is the observable implication itself — whether any
respondent leaves alternatives unopened, which under zero search costs no one
would. The third is a likelihood-ratio comparison against a *simulated* rather
than $\chi^2$ reference distribution.

This also bears on a parameterization choice. @Ursu_Seiler_Honka_2024 §5.2 notes
that estimating $g$ directly avoids a root-find inside the optimization loop.
Convenient — but it puts the quantity we most want to test on a scale where the
null is at infinity. We estimate $\log \cost$ and pay the root-find, which in
our implementation is a spline evaluation rather than an iterative solve (see
[`R/README.md`](../R/README.md)).

### Individual-Level Estimation from a Task Panel {#sec-est-panel}

There are two routes to the individual-level deliverable that makes SIFT a
conjoint substitute rather than a search paper.

**Full hierarchical Bayes** is what the conjoint audience expects and what makes
individual part-worths and individual search costs directly reportable.

**Two-stage empirical Bayes** estimates the hyperparameters by simulated maximum
likelihood and then draws individual posteriors by Metropolis–Hastings
conditional on those estimates. This is what @Morozov_2021 do to compute
personalized prices (their §7.2 and Appendix H), and it is considerably cheaper.

Our runtime work reverses the ordering we originally expected. Because HB never
integrates over the heterogeneity distribution — each respondent sits at their
own $\bftheta_i$ — the mixing-draw multiplier that makes random-coefficients
SMLE expensive simply does not arise. We therefore treat HB as the primary
estimator and empirical Bayes as the robustness check.

Runtime should be stated precisely, since the sampler's cost is easy to
underestimate from per-evaluation timings alone. The implemented sampler measures
**1.8 ms per task per iteration** — linear in the number of tasks, and only
weakly increasing in the number of simulation draws, because per-task overhead
rather than draw count dominates. Each iteration makes two passes over the data
and carries per-respondent overhead in addition. For 800 respondents completing
10 tasks each, at 20,000 iterations, this amounts to approximately **80 hours
single-threaded**.

The $\bftheta_i$ step is, however, embarrassingly parallel: each respondent's
Metropolis update touches only that respondent's tasks. Parallelised across 12
cores we measure a 7.5-fold speedup, which is conservative in the sense that
communication accounts for a material share of each iteration at that problem
size. A commercial-scale fit then completes in approximately **11 hours**.

The claim is therefore conditional. Hierarchical Bayes is the cheaper route to
heterogeneity and completes within an overnight budget, but only when the
respondent step is parallelised; a single-threaded implementation does not.

### Estimation {#sec-est-estimation}

Runtime is a binding practical constraint rather than an implementation detail:
a fit has to complete overnight, not over a week. The published timings make the
point. @Ursu_Seiler_Honka_2024 Table 2, for 1,000 consumers with heterogeneity,
report 5h20m for importance sampling, 29h48m for the kernel-smoothed simulator
plus roughly 32 further hours to calibrate its scaling vector, and 198h24m for
GHK. @Chung_2025 Table 1, for a cross-sectional design of comparable size,
reports 0.48 minutes for their probability-mapping simulator against 5–7 minutes
kernel-smoothed and 15–127 minutes for crude frequency — one to two orders of
magnitude.

We have ported and validated the @Chung_2025 estimator and reproduced their
Table 1; the port, its five validation checks, and two bugs we found in their
shipped code are documented in [`R/README.md`](../R/README.md).

For SIFT itself we exploit a structural feature of the design. Conditional on
the observed purchase, every Weitzman condition is a conjunction of linear
inequalities — the stopping rule is a statement about a maximum and therefore a
union in general, but the choice rule names which alternative attains that
maximum. Continuation then collapses to the single binding constraint
$\resv_{s_K}$, because reservation values decrease along the search order. And
conditional on $\bfeta$ and on the post-search shock to the chosen alternative,
the remaining shocks are independent one-sided constraints that integrate out in
closed form, leaving a one-dimensional integral.

This yields a hybrid estimator: GHK on the pre-search block, quadrature on the
post-search block. Measured against the kernel-smoothed accept–reject simulator
of @Yavorsky_2021 at equal draw counts, its simulation variance is lower by a
factor of 21 to 100 depending on the design. Because runtime is exactly linear
in the number of draws, the relevant summary is time to a given precision, on
which the hybrid is 3 to 21 times cheaper; crude frequency is not competitive on
either margin. The kernel-smoothed simulator also shows the upward bias in the
negative log-likelihood that @Yavorsky_2021 notes is intrinsic to it. The
construction and its validation are in
[`notebooks/sift_likelihood.qmd`](../notebooks/sift_likelihood.qmd).

### Software {#sec-est-software}

All estimation code is in R and is released with the paper. The reservation-value
inverse $g^{-1}$ is evaluated by a cubic spline built once in the standardized
argument $\cost/\tilde\sigma$, so a single spline serves every $\tilde\sigma$;
it covers costs down to $3.9\times10^{-16}$, well below the range any optimizer
will propose.
