Algorithm Appreciation: The Asymmetric Counterpart to Algorithm Aversion
Algorithm appreciation is the empirical finding that people sometimes weight algorithm-labeled advice more than human-labeled advice, particularly in numeric or estimation tasks under conditions of unfamiliarity, before observed errors, and where users retain some interpretive agency. It is best understood not as the inverse of Algorithm Aversion but as the same underlying problem — calibrating reliance to evidence — viewed at a different moment in the user-system interaction. This article treats appreciation and aversion as context-triggered reliance states within a single theory of Calibrated Reliance, anchored in Logg, Minson, and Moore (2019) and Dietvorst, Simmons, and Massey (2015), and traces the design implications for systems that involve AI advice.
Coverage note: verified through May 2026. Generative-AI-specific evidence remains thinner and less settled than the predictive-advice literature; claims about LLM trust dynamics are flagged where the evidence base shifts.
Why appreciation is not the opposite of aversion
For most of the 2000s and 2010s, the modal view in the judgment-and-decision-making literature was that humans were broadly averse to algorithmic advice. The intuition was old — Meehl had documented decades earlier that clinicians often outperformed statistical models in self-assessment but underperformed them in actual prediction — and the finding had become folk wisdom: people resist algorithmic judgment even when it is demonstrably better than their own. Dietvorst, Simmons, and Massey (2015) gave this folk wisdom one of its sharpest formulations. After participants observed an algorithm err, they were less willing to use it on subsequent forecasts than they were to use a human forecaster who had erred equivalently or worse. The asymmetric penalty for algorithmic imperfection became one of the most-cited results in the field.
Logg, Minson, and Moore (2019) showed that this story was incomplete. Across six preregistered studies — using numeric estimation tasks ranging from weight-from-photo judgments to song popularity to romantic attractiveness — participants who received identical advice gave it greater weight when it was labeled as coming from an algorithm than when it was labeled as coming from another person. Forecasters surveyed beforehand expected the opposite pattern, which is itself an important methodological signal: the experts in the field had absorbed the aversion finding as a general theorem when the underlying evidence supported only a context-specific claim.
The natural rhetorical move is to call appreciation the "counterpart" or "mirror image" of aversion. That framing is useful as a headline but misleading as a theory. Appreciation and aversion are not equal-and-opposite attitudes toward algorithms; they are reliance patterns that emerge under different evidence conditions, at different points in the interaction, and through partially distinct mechanisms. The appreciation finding is strongest:
- in numeric or estimation tasks where the algorithm's output reads as a neutral statistical instrument,
- among users who lack strong domain identity or whose unaided performance they have reason to distrust,
- before the user has observed the algorithm err on a salient case,
- and when the experimental cost of advice-taking is low.
The aversion finding is strongest:
- after observed algorithmic error, especially error on cases where the user's intuition suggested an obvious answer,
- in tasks the user perceives as subjective, moral, or identity-laden,
- when the user is accountable to others for the decision,
- and when accepting the algorithm's output feels like surrendering agency rather than incorporating advice.
Treating these as two sides of the same coin flattens the asymmetry that makes either finding interesting. A user can appreciate algorithmic advice on a calorie estimate at 9:00, reject it on a medical decision at 9:15, and over-rely on it for a portfolio rebalance at 9:30 — without any of these moves contradicting the others. The conceptual heart of this article is that reliance changes when the evidence environment changes, and the appreciation-aversion pair is best read as a snapshot of that dynamic at two illustrative moments.
Definitions and neighboring concepts
The vocabulary in this area is treacherous. The same word ("trust," "reliance," "acceptance") often denotes different things in different papers, and the failure to separate them produces a great deal of muddled secondary commentary. The article uses the following definitions throughout.
Algorithm appreciation (Logg et al. 2019): the behavioral pattern of weighting algorithm-labeled advice more than identical human-labeled advice, typically measured by advice-taking weights or final estimates in judge-advisor experimental paradigms.
Algorithm aversion (Dietvorst et al. 2015): the behavioral pattern of reducing reliance on an algorithm after observing it err, often disproportionately compared to reductions in reliance on similarly-erring human advisors. See Algorithm Aversion for the dedicated entry.
Automation bias (Parasuraman and Riley 1997; Parasuraman and Manzey 2010): the tendency to over-rely on automated cues, including by failing to monitor for automation errors, omitting independent verification, or accepting automated outputs that disagree with available contradictory evidence. See Automation Bias.
Trust in automation (Lee and See 2004; Hoff and Bashir 2015): the attitudinal disposition to rely on an automated agent under conditions of uncertainty or vulnerability. Trust is a state, not a behavior; users can trust without relying and rely without trusting.
Reliance: the behavioral act of using an automated agent's output in a decision, with or without modification.
Calibrated reliance: a reliance pattern in which the user relies more on the system in conditions where it outperforms unaided judgment and less where it does not, with the difference detectable in decision quality. See Calibrated Reliance.
Acceptance: a coarser construct, often measured by self-report or by binary use/don't-use, that conflates trust, reliance, and convenience.
These distinctions matter because the deployment failure modes look different at each level. A user who trusts an algorithm but does not rely on it produces no behavioral consequence; a user who relies without trusting produces compliance behavior that may collapse under stress; a user who exhibits appreciation in self-report but aversion in behavior is doing what humans often do in surveys; a system that maximizes acceptance may have no effect on calibrated reliance.
Two anchor studies
The appreciation and aversion literatures rest, to a first approximation, on two seminal papers. They use overlapping methods but ask different questions, and reading them together is the cleanest way to see why the two findings are not in tension.
Logg, Minson, and Moore (2019)
The Logg et al. paper reports six studies. The core design uses the judge-advisor paradigm: participants make an initial estimate, see advice from another source, and produce a final estimate. The dependent variable is the weight given to the advice, computed as the proportional move from the initial estimate toward the advisor's. The manipulation is the source label: in the algorithm condition, participants are told the advice was produced by an algorithm; in the human condition, by another person. The advice itself is identical in content.
Across all six studies, participants weighted the algorithm-labeled advice more heavily. The effect appears in:
- estimating a person's weight from a photograph,
- predicting song popularity,
- predicting romantic attractiveness,
- and forecasting other quantitative outcomes.
The paper also reports that lay forecasters predicted the opposite result, which is the cleanest piece of evidence that the aversion intuition had outrun the evidence base. Notably, the appreciation effect attenuated when participants had reason to believe their own expertise was superior and when they observed the algorithm err — bringing the result into contact with the aversion literature rather than contradicting it.
The study's scope is bounded in ways that matter for deployment generalization. The tasks are quantitative, the stakes are low, the advice is presented as a number rather than a recommendation with justification, and accountability is minimal. Logg et al. did not show that humans broadly prefer algorithms; they showed that the prior literature had over-stated aversion in a class of advice-taking tasks where, under controlled conditions, appreciation actually dominated. The result is real and important; its policy implications are narrower than headlines often suggested.
Dietvorst, Simmons, and Massey (2015)
Dietvorst et al. ran a complementary set of studies in which participants forecasted outcomes (such as student academic performance) and were given access to a statistical model that, on average, made smaller errors than the participants. The critical manipulation was whether participants observed the model's forecasts on a held-out set of cases before they were asked to choose between using the model or their own judgment on subsequent cases.
Participants who observed the algorithm err became less willing to use it, even when its observed performance still substantially exceeded their own. Participants who observed humans err did not show the same disproportionate abandonment. The authors interpret this as evidence that algorithm errors are read as diagnostic of system unreliability, while human errors are read as ordinary noise.
Dietvorst, Simmons, and Massey (2018) followed up with a key refinement. When participants were given the ability to modify the algorithm's forecast — even by small amounts — their willingness to use it increased substantially, and overall forecast accuracy improved. Agency, in other words, reduced aversion without requiring full delegation. This finding is one of the most actionable in the deployment literature and is what the design implications below build on.
The two findings, side by side
| Dimension | Logg et al. 2019 (appreciation) | Dietvorst et al. 2015 (aversion) |
|---|---|---|
| Paradigm | Judge-advisor; advice-weighting | Model-vs-self choice; reliance over trials |
| Task type | Numeric estimation across multiple domains | Numeric forecasting (academic performance) |
| When measured | Single advice-taking moment | After observed algorithm performance |
| Comparator | Identical human-labeled advice | Self-judgment with known-better model available |
| User has seen errors? | Typically no | Yes, by design |
| Outcome | Algorithm-labeled advice weighted more | Algorithm abandoned despite superior accuracy |
| Mechanism implicated | Perceived objectivity, source authority, low-confidence in self | Disproportionate penalty for algorithm error; loss of agency |
| Generalization scope | Bounded to numeric, low-stakes advice-taking | Bounded to post-error reliance decisions |
The table makes the asymmetry concrete: the two findings answer different questions, and a coherent theory of human-AI reliance must accommodate both.
A coexistence and transition model
The synthesis that fits the data best is that appreciation and aversion are reliance states that the same person can move between as the evidence environment shifts. The transitions are not random; they are driven by a small set of inputs that the deployment design controls or fails to control.
A schematic of the typical trajectory in an AI advice system:
-
First contact. The user has no local performance evidence. Source cues dominate: framing as "algorithmic," "AI-powered," or "data-driven" can elicit appreciation, especially if the task reads as numeric, unfamiliar, or computational. This is the Logg et al. regime.
-
Initial reliance. The user accepts the system's output, possibly with modification. If the system's outputs are reasonable on the user's first few cases — or if the user has no easy way to detect error — appreciation persists.
-
First salient error. The user observes a case where the algorithm is clearly wrong. The character of the error matters more than its size: an algorithm that errs in a way the user finds obvious (recommending a clearly inappropriate route, missing a domain-evident exception, hallucinating a citation) produces a sharper trust collapse than an equivalent numeric error on an ambiguous case. This is the Dietvorst et al. regime.
-
Post-error reliance. The user's reliance drops, often disproportionately. If the system provides no narrative for the error (no uncertainty signal, no out-of-distribution warning, no contestability), aversion can become durable. If the system provides agency to modify the output — Dietvorst et al. (2018) — aversion is partially repaired.
-
Stable miscalibration. In many real deployments, users settle into a reliance pattern that is not well-calibrated to actual system performance. Some users continue to defer to the system on cases where it is unreliable (automation bias). Others reject it on cases where it would help (residual aversion). Both failure modes can coexist in the same workflow.
The states are not stages in a fixed sequence; users skip steps, regress, and partition reliance across subtasks. A clinician may appreciate the algorithm's literature-summarization output, distrust its differential-diagnosis suggestions, and over-rely on its dosing calculations — all on the same patient. The model's value is not in predicting which state a user occupies but in naming the inputs that move users between them: observed performance, error visibility, agency, accountability, perceived task type, and the local evidence environment.
This is the conceptual move that separates a useful synthesis from a misleading one. "Appreciation versus aversion" treated as a personality dimension produces bad design recommendations ("identify the appreciators and target them"). Appreciation and aversion treated as reliance states under shifting evidence produces a tractable design problem: surface the evidence the user needs to update reliance appropriately.
Moderators
The literature identifies a set of moderators that govern when appreciation or aversion dominates. Treating them as independent psychological variables — the way some review papers do — misses the design point. Each moderator can be read as a question about what evidence the user has, or thinks they have, about the system's competence and their own role.
Task objectivity
Castelo, Bos, and Lehmann (2019) show that algorithm aversion is stronger in tasks perceived as subjective and weaker in tasks perceived as objective. The same algorithm receives more reliance in a numeric forecasting task than in a partner-recommendation task, even when stripped of irrelevant differences. The mechanism is that algorithms read as appropriate instruments for problems with right answers and as inappropriate substitutes for problems requiring judgment.
The deployment implication is that systems should not present themselves as authoritative on tasks users perceive as subjective unless the system can demonstrate domain-appropriate competence. Forcing appreciation in a subjective task by interface polish is precisely the failure mode the deployment-risk literature warns about.
User expertise
Logg et al. found that the appreciation effect attenuates among participants with high domain expertise. Experts give less weight to algorithmic advice in their domain, partly because they have stronger priors and partly because they have more reason to suspect the algorithm has missed context. This is consistent with a long line of advice-taking research showing that domain confidence reduces advice-weighting in general, not only from algorithmic sources.
The implication is that "build trust in the AI" reads differently for novices and experts. A novice can be moved toward over-reliance with light-touch design; an expert often requires evidence of comparable-to-superior performance on cases the expert finds difficult. Treating all users as a single trust population invites both miscalibrations.
Observed performance and error exposure
Dietvorst et al. 2015 is the canonical study here, but the broader pattern is robust across the literature: observed algorithm performance is the single largest mover of reliance, and the asymmetric penalty for algorithm error is one of the most replicable findings in the field. The penalty depends on:
- the visibility of the error,
- whether the error is interpretable as a "stupid" mistake versus a hard-case mistake,
- the user's comparative belief about their own performance,
- and whether the system signals uncertainty that might have predicted the error.
A system that errs without warning is more punished than a system that errs after flagging uncertainty. This is one of the practical levers most under-used in real deployment design.
Perceived AI capability
Distinct from observed performance, the user's prior beliefs about AI capability act as a strong moderator. Users who believe AI is generally competent give algorithm-labeled advice more weight; users who believe AI is unreliable in the user's domain give it less. This is harder to manipulate at deployment time, but framing, branding, and comparative communication (showing the system's track record against named baselines) influences it.
The risk is that this lever can be used to inflate perceived capability beyond what observed performance supports. Authority laundering — using "AI-powered" framing to legitimize decisions the underlying system cannot justify — is one of the more serious deployment failure modes, and it operates precisely on this moderator.
Accountability
When users are accountable to others for a decision, reliance dynamics change in ways that the laboratory advice-taking literature only partly captures. A doctor signing a treatment plan, a loan officer approving a mortgage, an engineer authorizing a deployment, and a parent deciding on a school placement all face accountability burdens that the typical Mechanical Turk participant does not. The literature on accountability and judgment, going back to Tetlock, suggests that anticipated accountability often increases information-seeking and shifts users toward defensible rather than accurate decisions.
For AI advice, accountability cuts in multiple directions. Users may reject algorithmic advice they would have accepted privately, because the decision needs to be justified to a third party who would find algorithmic justification insufficient. Or users may accept algorithmic advice they would have rejected privately, because the algorithm provides post-hoc cover ("the system recommended it"). Both patterns are documented in field studies of clinical and legal AI deployment.
Perceived agency and control
Dietvorst, Simmons, and Massey (2018) is the cleanest evidence on this dimension. When users can modify the algorithm's recommendation — even slightly — their willingness to use it increases substantially, and accuracy improves. Agency moves reliance toward appreciation even when observed performance is unchanged.
The mechanism likely involves both psychological ownership (users feel responsible for a modified recommendation in a way they do not for a raw algorithmic output) and informational benefit (users sometimes have local context the algorithm lacks). The deployment implication is that systems should preserve meaningful override paths, but the override has to be substantive: a UI button that lets the user nudge the output by 5% to "feel involved" is design theater unless that nudge is doing real informational work.
Stakes
Stakes interact with everything above. In low-stakes tasks, appreciation is easy to elicit and aversion is rare. In high-stakes tasks, both effects intensify: users become more attentive to evidence of capability, more sensitive to error, and more concerned about accountability. The literature on high-stakes AI advice — clinical decision support, autonomous driving, criminal risk assessment — finds that the same user can show different reliance patterns on the same algorithm at different stakes levels.
Explanation and uncertainty communication
Bansal et al. (2019) and Buçinca et al. (2021), among others, find that explanations and confidence displays have complicated effects on reliance. Long explanations can produce an illusion of understanding without improving calibration. Confidence scores can be ignored if the workflow rewards speed. Explanations that focus on the system's process can backfire when users find the process unconvincing. The crude finding that "explanations build trust" has not survived contact with calibration-sensitive measurement.
The more careful finding is that uncertainty and explanation are useful when they help users form an accurate mental model of when the system is and is not reliable. The same explanation that improves calibration in one task can degrade it in another. This is one of the strongest places to apply the calibrated-reliance frame: the question is not whether explanations increase reliance, but whether they shift reliance in the direction of accuracy.
Whether the AI advises or decides
Finally — and underappreciated in much of the literature — the architectural question of whether the AI advises a human decision-maker or makes the decision itself changes the entire reliance dynamic. Advice systems leave the user in the loop and produce the advice-taking patterns that Logg et al. and Dietvorst et al. studied. Decision systems remove the user from the loop and produce questions about appeal, recourse, and institutional legitimacy that the advice-taking literature was not designed to answer. The deployment design should be explicit about which architecture is in play, because the user-facing affordances differ.
Moderator summary
| Moderator | Direction of effect | Primary evidence |
|---|---|---|
| Task objectivity | More objective → more appreciation | Castelo et al. 2019 |
| User domain expertise | More expert → less appreciation | Logg et al. 2019 |
| Observed error | Visible error → sharper aversion | Dietvorst et al. 2015 |
| Perceived AI capability | Higher → more appreciation | Glikson and Woolley 2020 (review) |
| Accountability | Bidirectional; context-dependent | Field studies; Tetlock-tradition work |
| Perceived agency / control | More control → less aversion | Dietvorst et al. 2018 |
| Stakes | Higher → both effects intensify | Mixed evidence; deployment studies |
| Explanation / uncertainty | Calibration-dependent | Bansal et al. 2019; Buçinca et al. 2021 |
| AI advises vs decides | Decides → different dynamics entirely | Lee and See 2004; recent governance work |
The right way to read this table is not as a list of independent dials but as a description of the evidence environment the user is operating in. Reliance follows from what users believe about the system, themselves, the task, and the stakes — and design choices control what evidence they have access to.
The generative AI question
The literatures synthesized above were built almost entirely on predictive and actuarial systems: statistical forecasts, classification models, decision-support tools that take inputs and produce numeric or categorical outputs. Generative AI is a different empirical case, and the article's most explicit hedge is that the appreciation-aversion balance under generative AI is not yet settled.
The reasons to expect generative AI to amplify appreciation include:
- Fluency and social legibility. LLMs produce text that reads as competent, helpful, and well-organized. Source cues that previously required deliberate signaling ("this is an algorithm") are now built into the medium.
- Broad domain coverage. A user can ask the same system about cooking, code, contract drafting, and emotional advice. The illusion of general competence is structural to the interface.
- Convenience. Advice-taking from an LLM has near-zero friction. The Logg et al. effect — appreciation in low-cost advice-taking — operates here at scale.
- Anthropomorphic interaction. Conversational turn-taking, first-person framing, and apologetic recovery from errors all reduce the perceived "algorithm-ness" of the interaction in ways that may shift users out of the aversion-triggering frame entirely.
The reasons to expect generative AI to amplify aversion include:
- Hallucination. Fabricated citations, fictional case law, and confidently wrong answers are highly salient errors. They activate the Dietvorst et al. mechanism — algorithm error read as diagnostic of unreliability — particularly strongly.
- Opacity. Users have little insight into why a generative system produced one output and not another. This frustrates the user-expertise moderator: users with relevant domain knowledge often find the system's "reasoning" hollow on inspection.
- Sycophancy and persuasive wrongness. Generative systems can produce confidently wrong outputs in registers that human experts would not use. The mismatch between confidence and correctness can trigger sharp trust collapse.
- Brittle failure modes. A small prompt change can produce qualitatively different outputs. Users who encounter this brittleness develop domain-specific aversion that the older predictive literature did not anticipate.
The most defensible interpretation of the current evidence is that generative AI is likely to produce more rapid oscillation between appreciation and aversion than predictive systems did, with both states more intense and both more brittle. The implication for deployment is that the article's central design recommendation — calibrate reliance, do not maximize trust — applies even more strongly under generative AI than under the systems Logg et al. and Dietvorst et al. studied. The field has not yet produced large, preregistered, field-based studies of long-run reliance on generative AI under realistic accountability conditions, and until it does, deployment claims about generative AI trust dynamics should be treated as design hypotheses rather than settled findings.
A specific warning is that the predictive-AI trust literature should not be back-ported wholesale onto generative AI. The same word — "trust" — refers to different mental objects when applied to a forecasting model with a documented error rate on a held-out test set versus a conversational system that is fluent across domains and never produces the same output twice. Wiki entries that cite Logg et al. 2019 as evidence that "users trust AI" without distinguishing the systems are committing a category error.
Deployment design: calibrated reliance, not trust maximization
The temptation in applied work is to take the appreciation finding as license for "build trust in AI" as a design objective. This is the central failure mode the article exists to caution against. The right design objective is calibrated reliance: users should rely on the system more on cases where it improves outcomes and less on cases where it does not, with the difference measurable in decision quality.
The reframing matters because the failure modes of trust maximization are not benign. Polished interfaces, confident framings, institutional branding, and fluent explanations can produce premature reliance that the system's actual performance does not warrant. The classic automation-bias literature (Parasuraman and Riley 1997; Parasuraman and Manzey 2010) and a generation of subsequent work on automation in cockpits, control rooms, medical decision support, and driver-assist systems have documented the costs of trust-maximizing design: complacency, omission errors, skill decay, and accountability laundering.
Concrete design levers that move systems toward calibrated reliance:
Surface local performance evidence
Users update reliance most strongly on observed performance in their domain. Systems should show comparative performance on the user's task — not generic benchmark scores, but error rates on cases the user recognizes as similar to their own. A medical AI showing its sensitivity and specificity on the user's patient population is doing real work; showing "trained on 10 million cases" is doing brand work.
Communicate uncertainty without false precision
Uncertainty displays are useful when they help users decide when to trust an output. They are not useful when they are decorative ("85% confident") or when their meaning is unclear. The design challenge is to produce uncertainty signals that vary meaningfully across cases and that users can validate against outcomes. A confidence score that is always 0.8 is worse than no confidence score, because it teaches the user to ignore it.
Preserve meaningful agency
Dietvorst et al. (2018) is the strongest evidence here: even modest user control over the algorithm's recommendation increases willingness to use it and improves accuracy. The agency has to be substantive. Override mechanisms that exist only to satisfy a regulatory checkbox produce liability theater, not calibrated reliance.
Separate recommendation from authority
A system that recommends a treatment is doing something different from a system that approves a treatment. Users — and their organizations — should know which architecture is in play. Recommendation systems require the user to accept the decision and the responsibility; decision systems require institutional governance, appeal mechanisms, and audit trails. Hybrid designs that obscure the boundary are a frequent source of accountability gaps.
Train users on characteristic failure modes
Users develop calibrated reliance faster when they know how the system fails, not only how it succeeds. Showing failure cases during onboarding — and continuing to surface them during use — improves the user's mental model and reduces both over-reliance and reflexive rejection. Systems that hide failure modes for marketing reasons produce users whose reliance is brittle to the first surprise.
Measure reliance behavior, not just trust attitudes
Self-reported trust is a weak proxy for behavioral reliance. Deployment metrics should track:
- acceptance rates of system recommendations, segmented by case type,
- override frequency and the user's stated reasons,
- post-error reliance change,
- decision quality outcomes where measurable,
- and the rate of silent acceptance of wrong outputs (the hardest and most important to measure).
A system with rising acceptance and unchanging decision quality is not winning trust; it is producing complacency.
Provide contestability and recourse
In high-stakes deployments, users — and the people affected by decisions — need ways to challenge outputs and trigger review. Contestability is not the same as override; it is the institutional infrastructure for catching errors that the user did not catch in the moment. Systems without contestability convert algorithmic errors into permanent decisions.
Watch for second-order failures
Fixing aversion by smoothing the interface can produce over-reliance. Fixing over-reliance by adding friction can produce reflexive rejection. The design problem is genuinely two-sided, and design choices should be evaluated against both failure modes simultaneously. Deployment monitoring that tracks only one direction is undermeasuring.
Failure modes the appreciation framing invites
If the article were left at "users sometimes prefer algorithms; design to harness this," it would invite the failure modes the deployment-risk literature spends most of its time documenting. Naming them explicitly is part of the article's purpose.
Automation bias. Users defer to the system on cases where their own judgment would have been better. The deployment signal is silent acceptance of wrong outputs without any user-side flag. Buçinca et al. (2021) document this directly in AI-assisted decision-making: explanations that increase reliance often do not improve accuracy and sometimes degrade it.
Deskilling. Users who consistently defer to a system lose the skills they would have used to override it. This is most documented in aviation and medicine but applies anywhere expert judgment is partially automated. Once skill has eroded, the user is incapable of meaningful override, and the system's nominal "advisory" architecture has effectively become a decision architecture.
Authority laundering. Organizations deploy AI to legitimize decisions that would not be defensible on their own merits. "Users appreciate the algorithm" becomes the user-research justification for a deployment whose underlying performance does not warrant it. The appreciation finding is particularly susceptible to this misuse because Logg et al. tested user preference in low-stakes advice-taking, not institutional appropriateness in high-stakes deployment.
Liability diffusion. When an AI recommendation contributes to a bad decision, the locus of responsibility becomes diffuse. The user blames the system; the organization blames the user; the vendor blames the deployment context. Without explicit governance, diffuse liability produces under-corrected errors and chilling effects on user override.
Brittle trust collapse. A system that has cultivated appreciation through interface design and framing — rather than through demonstrated performance — is vulnerable to sharp reliance collapse after the first visible failure. The Dietvorst mechanism applies most strongly when users had higher expectations than the performance evidence supported.
Heterogeneity-flattening. Treating "users" as a single trust population produces design that works for the average and fails for the tails. Experts, novices, occasional users, and high-frequency users have different reliance dynamics, and a system that optimizes for one segment can degrade performance for others.
Surveillance and discipline. AI advice systems often function as discipline mechanisms — users who deviate from the algorithm's recommendation must justify their deviation in writing, while users who comply do not. This produces a structural pressure toward compliance that is invisible in the user-preference data and that the appreciation finding does not address.
The article's risk-management posture is not anti-automation. Many of the appreciation findings reflect calibrated responses to algorithmic systems that are genuinely better than unaided human judgment, and resistance to such systems is often the more dangerous error. The point is that the appreciation finding should not become ammunition for shallow deployment arguments. The same evidence that warrants taking algorithmic advice seriously in numeric forecasting warrants taking accountability, contestability, and calibrated communication of uncertainty seriously in the surrounding design.
What the broader trust literature adds
The advice-taking work anchored by Logg et al. and Dietvorst et al. is a small slice of the human-AI trust literature. Reading them in isolation produces a narrower view than the topic warrants.
Lee and See (2004) provide the foundational framework: trust in automation is a multidimensional attitude shaped by perceived performance, process (understanding of how the system works), and purpose (alignment between system and user goals). The framework predates the modern appreciation-aversion debate and survives it intact — both findings can be located within the Lee and See dimensions, and the deployment design recommendations above are essentially applications of the framework.
Hoff and Bashir (2015) review the trust-in-automation literature and emphasize the dynamic, calibrating nature of trust over the course of an interaction. Dispositional trust (the user's prior trust in automation generally), situational trust (trust given the context), and learned trust (trust updated by experience) interact in ways that the lab-paradigm advice-taking studies only partially capture.
Glikson and Woolley (2020) review human-AI trust specifically and identify a number of moderators relevant to AI deployment: tangibility (is the AI embodied or virtual), task interdependence, transparency, immediacy of behavior, anthropomorphism, and reliability. Their review is one of the more useful syntheses for designers because it organizes the evidence around manipulable system features.
Bansal et al. (2019) and Buçinca et al. (2021) provide more recent and methodologically careful work on AI-assisted decision-making, focusing on the question of when AI advice actually improves human decisions versus when it produces appearance of help without substance. Their core finding — that increased reliance and increased accuracy are not the same thing — is the empirical backbone of the calibrated-reliance framing.
Castelo, Bos, and Lehmann (2019) extend the aversion literature to task subjectivity and produce one of the better moderator findings: the same system is appreciated more in objective tasks and resisted more in subjective tasks. Longoni, Bonezzi, and Morewedge (2019) extend it to medical AI specifically, showing that resistance to algorithmic medical care is partly driven by users' belief that their case is unique and that algorithms cannot accommodate uniqueness.
The classic clinical-versus-statistical-prediction literature — Meehl's original 1954 review, Dawes' subsequent work, and the Grove et al. meta-analysis — provides the base-rate context: in structured prediction tasks, mechanical models often outperform unaided human judgment, and resistance to mechanical prediction has historically been costly to outcomes. This base rate is sometimes used to argue that aversion is irrational and appreciation is the rational corrective. That conclusion is too strong. The clinical-versus-statistical literature is about a narrow class of prediction problems with specific data conditions; generalizing it to all AI deployment, including generative AI in subjective domains, is not warranted by the underlying evidence.
The overall posture of the broader literature is that human-AI reliance is dynamic, context-specific, and shaped by both system properties and user properties in ways that resist simple summaries. The appreciation finding fits inside this larger story; it does not replace it.
Open empirical questions
Several questions are not currently answered well enough by the literature to support strong deployment claims. Naming them is more honest than hedging the claims that depend on them.
Does the appreciation-aversion balance change when AI capability becomes overwhelming evidence? The clinical-versus-statistical literature suggests that even large performance gaps do not always produce reliance, but most of that evidence predates the current generation of AI systems. As AI capability in specific domains substantially exceeds unaided human performance — and as that excess becomes visible in field outcomes — does reliance follow? The early evidence from radiology, protein structure prediction, and code completion is mixed: reliance has increased, but calibration to actual case-by-case performance remains imperfect.
Does generative AI change the pattern compared to predictive AI? The hedge above is the honest answer: probably yes, but the direction and magnitude are not settled. The most useful framing is that generative AI likely produces faster and larger swings between appreciation and aversion than predictive systems did, but durable reliance patterns after sustained use under realistic accountability conditions have not yet been measured at scale.
How does long-run reliance evolve under repeated error and recovery? Almost all of the experimental evidence is short-horizon. Real deployments involve months or years of interaction with errors of varying severity. Whether users develop calibrated reliance over time, or whether they oscillate between over-reliance and reflexive rejection, is poorly measured.
What is the effect of organizational and institutional context on individual reliance? A user's reliance on AI advice in a clinical workflow is shaped by hospital policy, billing structure, malpractice exposure, peer norms, and electronic health record design — none of which appear in the lab paradigm. The translation from advice-taking findings to deployed reliance dynamics is an open question.
Are there durable individual differences in reliance disposition? Some research suggests stable individual differences in trust in automation, but the predictive validity of these measures for actual reliance behavior is weak. Whether "appreciators" and "aversives" are useful user-segmentation categories, or whether reliance is determined almost entirely by context, is empirically unclear.
Does the appreciation finding replicate in non-WEIRD samples? Most of the advice-taking literature is built on Western, educated, industrialized, rich, democratic samples. Whether the appreciation effect holds in other cultural contexts — and whether the moderators behave similarly — is undermeasured.
The cheapest decisive experiment for the deployment-relevant questions is a preregistered, within-subjects advice-taking study crossing four factors: task type (numeric forecast versus subjective or value-laden advice), source label (human expert versus predictive model versus generative AI), performance feedback (none versus visible error with comparative human error), and agency (accept-only versus adjustable recommendation). The primary outcomes should be behavioral advice weight, final decision accuracy, willingness to use again, and post-error reliance change. Self-reported trust should be a secondary outcome. For a deployed system, the practical version is an A/B test comparing a calibrated-reliance interface (showing model track record, uncertainty, known failure modes) against a generic "AI-powered" interface, with downstream decision quality as the outcome.
Practical takeaways
The article's deployment-relevant claims, distilled:
-
Treat reliance, not trust, as the design target. Trust is an attitude; reliance is a behavior; calibrated reliance is the goal. Measuring trust without measuring reliance produces a misleading picture of system effectiveness.
-
Appreciation is real but bounded. Users will often weight algorithmic advice more than identical human advice in numeric, unfamiliar, low-stakes tasks. This is a real finding, not a marketing claim, but it does not generalize automatically to high-stakes, subjective, or generative-AI contexts.
-
Aversion is real and asymmetric. The penalty for observed algorithm error is sharper than the penalty for observed human error. Systems that hide failure modes for marketing reasons set themselves up for sharper trust collapse later.
-
Coexistence is the default state. The same user can appreciate the system in one subtask and resist it in another, and the resistance pattern can flip after a single salient error. The design implication is that subtask-level reliance dynamics matter more than user-level dispositions.
-
Agency reduces aversion without sacrificing accuracy. Dietvorst et al. (2018) is one of the most actionable findings in the literature. Preserve meaningful override, and reliance improves on both sides — users accept more and accept better.
-
Communicate uncertainty in ways that vary meaningfully across cases. Uniform confidence scores are decorative. The calibration test is whether the user's reliance after seeing the uncertainty signal is better aligned with actual case-by-case performance.
-
Show local performance evidence. Generic benchmarks do not change reliance behavior in well-calibrated ways. Domain-specific, comparator-anchored performance evidence does.
-
Watch both directions of failure. Reducing aversion is not the same as improving the system. Over-reliance can be the more dangerous failure mode, especially under generative AI.
-
Treat generative AI as a distinct empirical case. Until field evidence catches up, claims about LLM trust dynamics should be marked as design hypotheses rather than settled findings.
-
Don't let "users appreciate the algorithm" launder authority. User preference is not evidence of system reliability. Institutional deployment requires governance, contestability, and accountability that user-preference data cannot substitute for.
The deepest practical takeaway is that the appreciation finding is most valuable not as a justification for AI deployment but as a corrective to the prior overreach of the aversion narrative. The previous story — humans dislike algorithms, and design must overcome this — was too simple. The story replacing it should not be that humans now like algorithms and design should harness this. The story should be that reliance is a moving target shaped by evidence, and the design problem is to produce evidence that moves reliance in the direction of better decisions.
Companion entries
Core theory: Algorithm Aversion Calibrated Reliance Automation Bias Trust in Automation Human-AI Trust Clinical Versus Statistical Prediction
Mechanisms and moderators: Advice-Taking Paradigm Source Labeling Effects Anthropomorphism in AI Interaction Uncertainty Communication Explanation Effects on Reliance
Practice: Human-in-the-Loop Architectures Override Design Confidence Display Design Calibration Curves Post-Deployment Monitoring of Reliance AI Decision-Support Systems Contestability and Recourse
Counterarguments and risks: Authority Laundering Liability Diffusion in AI Deployment Deskilling Under Automation Surveillance Effects of Decision Support Brittle Trust Collapse
Generative AI extensions: Hallucination and Trust Sycophancy in Language Models Fluency-Accuracy Mismatch LLM Reliance Dynamics